View Detailed Profile
Keynote: Beyond Benchmarks: Evaluating Agents Against What They Are Actually Supposed to Do

Keynote: Beyond Benchmarks: Evaluating Agents Against What They Are Actually Supposed to Do

Your

LLM & AI Agent Benchmarks vs Reality: Why AI Applications Break

LLM & AI Agent Benchmarks vs Reality: Why AI Applications Break

Learn more about LLM

Agent Evaluation & Benchmarks - Agentic AI MOOC 2025 Lecture 4 Summary

Agent Evaluation & Benchmarks - Agentic AI MOOC 2025 Lecture 4 Summary

This lecture discusses the critical shift from

Keynote | The Future of AI Agents | Arize Observe 2026

Keynote | The Future of AI Agents | Arize Observe 2026

AI

How I Actually Used AI Agents to Build a Benchmark

How I Actually Used AI Agents to Build a Benchmark

My old AI planning

The Art and Science of Benchmarking Agents (Agentic AI Summit 2026)

The Art and Science of Benchmarking Agents (Agentic AI Summit 2026)

Vincent Sunn Chen (Founding Team & Research Fellow, Snorkel AI) breaks down what actually makes an AI

Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development

Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development

Paper:

How to Evaluate Your AI Agent Using Test Cases and Metrics

How to Evaluate Your AI Agent Using Test Cases and Metrics

Building reliable AI

Measuring Agents With Interactive Evaluations

Measuring Agents With Interactive Evaluations

Agents

How to evaluate agents in practice

How to evaluate agents in practice

Evaluating Agents

Evaluating and Debugging Non-Deterministic AI Agents

Evaluating and Debugging Non-Deterministic AI Agents

Evaluate

Beyond Benchmarks - Lessons from Building AI Evals in the Real World

Beyond Benchmarks - Lessons from Building AI Evals in the Real World

As AI systems become more capable,

Benchmarks Are Lying to You: How to Evaluate Models and Agents in the Real World

Benchmarks Are Lying to You: How to Evaluate Models and Agents in the Real World

Details: