Slides: Evals Are Broken, Use Them Anyway — Ara Khan, Cline
Source Video
Evals Are Broken, Use Them Anyway — Ara Khan, Cline
Relationship To World's Fair 2026
These slides are extracted from a public AI Engineer YouTube video connected to World's Fair 2026. Speaker-matched clips are supporting context unless later confirmed as exact session recordings; official livestream recordings are day-level/event-level source material.
Related Scheduled Sessions
- No individual scheduled session mapping has been assigned yet; treat this as an event livestream deck.
Extracted Slides

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.96 - Text source: agent_vision.
Slide text:
I want you to be able to build, interpret and/or use evals in your agent flows

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.93 - Text source: advanced OCR
rapidocr-live/border-trim/opencv-adaptive. - OCR decision: ready — contains a dense social-post screenshot with small text
Slide text:
Francois Chollet @fchollet · 6h
★ ★ AIE Overoptimized for public benchmark numbers at the detriment of with actual usefulness is a core competency for Al labs, and any new lab is unlikelyto be successful without first figuring that out. everything else. Knowing how to evaluate models in a way that correlates The new model from Meta is already looking like a disappointment:
55 QZ4 ili? h0115K
AlEngineer
EUROPE

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.93 - Text source: advanced OCR
rapidocr-live/bright-screen/opencv-adaptive. - OCR decision: ready — contains a dense social-post screenshot with small text
Slide text:
harness) based ona benchmark result? Nikunj Kothari@nikunjo4h Genuinely curious = has any engineer made a decision on a model (or
★ AIE trust their taste, evals and how models are performing for their specific use case. Every researcher / engineer I have talked to routinely dismisses it. They
not. It's seems like it's simply now used for the labs to show how they're slightly ahead but in my head has no material impact on whether a model is used or
I'd love to be proven wrong so please push back!
930 包 65 6.1K
AI Engineer
EUROPE

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.97 - Text source: agent_vision.
Slide text:
Why SWE-bench Verified no longer measures frontier coding capabilities
SWE-bench Verified is increasingly contaminated. We recommend SWE-bench Pro.

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.95 - Text source: advanced OCR
rapidocr-live/border-trim/contrast. - OCR decision: ready — title slide contains smaller UI text and labels suited for OCR
Slide text:
terminal-bench: benchmarks for ai
agents sin t terminal environments
AIE terminal-bench is a collection of harbor-native benchmarks
to help agent makers quantify their agents' terminal mastery
)terninal-bench 3.θ is now in developacnt 复 terminal-bench-science is now in developnent
help build the next frontior benchmnrk oxtending terminaL-bench to tho natural sciences
i want to test my ogent
Engineering the future of Al

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.97 - Text source: agent_vision.
Slide text:
Engineering the future of AI
Hidden Non-Slide Evidence
- `slide-001.jpg` —
speaker_stageconfidence0.98; camera shot of speaker and audience with a projected slide, not a clean slide frame - `slide-008.jpg` —
speaker_stageconfidence0.99; Camera shot of presenter and audience with projected slide in the background; not a readable presentation slide frame.
Classification audit: raw/sources/slide-ai-classification/slides/QuuIywMG4s8/audit.json
Slide-Derived Subjects To Review
Subject extraction uses video title, related session titles/descriptions, transcript context, and OCR text when available. OCR is best-effort and should be reviewed against the embedded slide images.