Markdown source

Slides: Evals Are Broken, Use Them Anyway — Ara Khan, Cline

Source Video

Evals Are Broken, Use Them Anyway — Ara Khan, Cline

Relationship To World's Fair 2026

These slides are extracted from a public AI Engineer YouTube video connected to World's Fair 2026. Speaker-matched clips are supporting context unless later confirmed as exact session recordings; official livestream recordings are day-level/event-level source material.

Related Scheduled Sessions

Extracted Slides

slide-002.jpg

Slide text:

I want you to be able to build, interpret and/or use evals in your agent flows

slide-003.jpg

Slide text:

Francois Chollet @fchollet · 6h

★ ★ AIE Overoptimized for public benchmark numbers at the detriment of with actual usefulness is a core competency for Al labs, and any new lab is unlikelyto be successful without first figuring that out. everything else. Knowing how to evaluate models in a way that correlates The new model from Meta is already looking like a disappointment:

55 QZ4 ili? h0115K

AlEngineer

EUROPE

slide-004.jpg

Slide text:

harness) based ona benchmark result? Nikunj Kothari@nikunjo4h Genuinely curious = has any engineer made a decision on a model (or

★ AIE trust their taste, evals and how models are performing for their specific use case. Every researcher / engineer I have talked to routinely dismisses it. They

not. It's seems like it's simply now used for the labs to show how they're slightly ahead but in my head has no material impact on whether a model is used or

I'd love to be proven wrong so please push back!

930 包 65 6.1K

AI Engineer

EUROPE

slide-005.jpg

Slide text:

Why SWE-bench Verified no longer measures frontier coding capabilities

SWE-bench Verified is increasingly contaminated. We recommend SWE-bench Pro.

slide-006.jpg

Slide text:

terminal-bench: benchmarks for ai

agents sin t terminal environments

AIE terminal-bench is a collection of harbor-native benchmarks

to help agent makers quantify their agents' terminal mastery

)terninal-bench 3.θ is now in developacnt 复 terminal-bench-science is now in developnent

help build the next frontior benchmnrk oxtending terminaL-bench to tho natural sciences

i want to test my ogent

Engineering the future of Al

slide-007.jpg

Slide text:

Engineering the future of AI

Hidden Non-Slide Evidence

Classification audit: raw/sources/slide-ai-classification/slides/QuuIywMG4s8/audit.json

Slide-Derived Subjects To Review

Subject extraction uses video title, related session titles/descriptions, transcript context, and OCR text when available. OCR is best-effort and should be reviewed against the embedded slide images.