---
title: "Slides: Evals Are Broken, Use Them Anyway — Ara Khan, Cline"
category: "slides"
video_id: "QuuIywMG4s8"
sourceLabels: ["Public YouTube video frames", "Public YouTube metadata"]
---

# Slides: Evals Are Broken, Use Them Anyway — Ara Khan, Cline

## Source Video
[Evals Are Broken, Use Them Anyway — Ara Khan, Cline](https://www.youtube.com/watch?v=QuuIywMG4s8)

## Relationship To World's Fair 2026
These slides are extracted from a public AI Engineer YouTube video connected to World's Fair 2026. Speaker-matched clips are supporting context unless later confirmed as exact session recordings; official livestream recordings are day-level/event-level source material.

## Related Scheduled Sessions
- No individual scheduled session mapping has been assigned yet; treat this as an event livestream deck.

## Extracted Slides
![[assets/slides/QuuIywMG4s8/slide-002.jpg]]

- Recreated text/layout view: [open HTML recreation](/assets/slide-recreations/slides/QuuIywMG4s8/slide-002.html)
- AI slide classifier: `content_slide` confidence `0.96`
- Text source: agent_vision.

Slide text:

> I want you to be able to build, interpret and/or use evals in your agent flows

![[assets/slides/QuuIywMG4s8/slide-003.jpg]]

- Recreated text/layout view: [open HTML recreation](/assets/slide-recreations/slides/QuuIywMG4s8/slide-003.html)
- AI slide classifier: `content_slide` confidence `0.93`
- Text source: advanced OCR `rapidocr-live/border-trim/opencv-adaptive`.
- OCR decision: ready — contains a dense social-post screenshot with small text

Slide text:

> Francois Chollet @fchollet · 6h
> ★ ★ AIE Overoptimized for public benchmark numbers at the detriment of with actual usefulness is a core competency for Al labs, and any new lab is unlikelyto be successful without first figuring that out. everything else. Knowing how to evaluate models in a way that correlates The new model from Meta is already looking like a disappointment:
> 55 QZ4 ili? h0115K
> AlEngineer
> EUROPE

![[assets/slides/QuuIywMG4s8/slide-004.jpg]]

- Recreated text/layout view: [open HTML recreation](/assets/slide-recreations/slides/QuuIywMG4s8/slide-004.html)
- AI slide classifier: `content_slide` confidence `0.93`
- Text source: advanced OCR `rapidocr-live/bright-screen/opencv-adaptive`.
- OCR decision: ready — contains a dense social-post screenshot with small text

Slide text:

> harness) based ona benchmark result? Nikunj Kothari@nikunjo4h Genuinely curious = has any engineer made a decision on a model (or
> ★ AIE trust their taste, evals and how models are performing for their specific use case. Every researcher / engineer I have talked to routinely dismisses it. They
> not. It's seems like it's simply now used for the labs to show how they're slightly ahead but in my head has no material impact on whether a model is used or
> I'd love to be proven wrong so please push back!
> 930 包 65 6.1K
> AI Engineer
> EUROPE

![[assets/slides/QuuIywMG4s8/slide-005.jpg]]

- Recreated text/layout view: [open HTML recreation](/assets/slide-recreations/slides/QuuIywMG4s8/slide-005.html)
- AI slide classifier: `content_slide` confidence `0.97`
- Text source: agent_vision.

Slide text:

> Why SWE-bench Verified no longer measures frontier coding capabilities
> SWE-bench Verified is increasingly contaminated. We recommend SWE-bench Pro.

![[assets/slides/QuuIywMG4s8/slide-006.jpg]]

- Recreated text/layout view: [open HTML recreation](/assets/slide-recreations/slides/QuuIywMG4s8/slide-006.html)
- AI slide classifier: `content_slide` confidence `0.95`
- Text source: advanced OCR `rapidocr-live/border-trim/contrast`.
- OCR decision: ready — title slide contains smaller UI text and labels suited for OCR

Slide text:

> terminal-bench: benchmarks for ai
> agents sin t terminal environments
> AIE terminal-bench is a collection of harbor-native benchmarks
> to help agent makers quantify their agents' terminal mastery
> )terninal-bench 3.θ is now in developacnt 复 terminal-bench-science is now in developnent
> help build the next frontior benchmnrk oxtending terminaL-bench to tho natural sciences
> i want to test my ogent
> Engineering the future of Al

![[assets/slides/QuuIywMG4s8/slide-007.jpg]]

- Recreated text/layout view: [open HTML recreation](/assets/slide-recreations/slides/QuuIywMG4s8/slide-007.html)
- AI slide classifier: `content_slide` confidence `0.97`
- Text source: agent_vision.

Slide text:

> Engineering the future of AI


### Hidden Non-Slide Evidence
- [`slide-001.jpg`](/assets/slides/QuuIywMG4s8/slide-001.jpg) — `speaker_stage` confidence `0.98`; camera shot of speaker and audience with a projected slide, not a clean slide frame
- [`slide-008.jpg`](/assets/slides/QuuIywMG4s8/slide-008.jpg) — `speaker_stage` confidence `0.99`; Camera shot of presenter and audience with projected slide in the background; not a readable presentation slide frame.

Classification audit: `raw/sources/slide-ai-classification/slides/QuuIywMG4s8/audit.json`

## Slide-Derived Subjects To Review
Subject extraction uses video title, related session titles/descriptions, transcript context, and OCR text when available. OCR is best-effort and should be reviewed against the embedded slide images.
