---
title: "Slides: OpenThoughts: Data Recipes for Reasoning Models — Ryan Marten, Bespoke Labs"
category: "slides"
video_id: "liG97YXaTSA"
sourceLabels: ["Public YouTube video frames", "Public YouTube metadata"]
---

# Slides: OpenThoughts: Data Recipes for Reasoning Models — Ryan Marten, Bespoke Labs

## Source Video
[OpenThoughts: Data Recipes for Reasoning Models — Ryan Marten, Bespoke Labs](https://www.youtube.com/watch?v=liG97YXaTSA)

## Relationship To World's Fair 2026
These slides are extracted from a public AI Engineer YouTube video connected to World's Fair 2026. Speaker-matched clips are supporting context unless later confirmed as exact session recordings; official livestream recordings are day-level/event-level source material.

## Related Scheduled Sessions
- [[2026-06-30-alex-shaw-everything-is-a-rollout]] — Everything Is a Rollout

## Extracted Slides
![[assets/slides/liG97YXaTSA/slide-001.jpg]]

- Recreated text/layout view: [open HTML recreation](/assets/slide-recreations/slides/liG97YXaTSA/slide-001.html)
- AI slide classifier: `content_slide` confidence `0.98`
- Text source: advanced OCR `rapidocr-live/bright-screen/contrast`.
- OCR decision: ready — Dense chart slide with many small labels and axis text.

Slide text:

> Reasoning
> Progress on Al benchmarks in the past five years
> AIE 100
> 80 Trivia questions (TriviaQA)
> Accuracy 40+ 20- 60- Variousexams (MMLU) math (GSM8K) Grade-school math (MATH) Competition SWEtasks(SWE benchverified) STEM(GPQA) Graduate-level Prestigious math exam(AIME) "Humanity's lastexam
> 2020 2021 2022 2023 2024 2025
> Credit: Jason Wed
> aws

![[assets/slides/liG97YXaTSA/slide-002.jpg]]

- Recreated text/layout view: [open HTML recreation](/assets/slide-recreations/slides/liG97YXaTSA/slide-002.html)
- AI slide classifier: `content_slide` confidence `0.99`
- Text source: advanced OCR `rapidocr-live/border-trim/contrast`.
- OCR decision: ready — Diagram slide with small labels and flow arrows that are better handled by OCR.

Slide text:

> Strong Reasoning Through SFT
> AIE IGRSO DeepSeak- R1-Zero stan data
> DeepSeek-V3 Base SFTon doto Cold STOT Initialization RL R(GRPO) Converged RL reasoning sampias >009
> SFT on
> reasoning samplos 600k samples general 2002 Reasoning Model RLforalignment DeepSeek- R1
> Credita16z
> Microsoft smol ai


### Hidden Non-Slide Evidence
- [`slide-003.jpg`](/assets/slides/liG97YXaTSA/slide-003.jpg) — `speaker_stage` confidence `0.99`; Camera shot of speakers on stage, not a presentation slide.
- [`slide-004.jpg`](/assets/slides/liG97YXaTSA/slide-004.jpg) — `speaker_stage` confidence `0.99`; Camera shot of speaker on stage, not a presentation slide.

Classification audit: `raw/sources/slide-ai-classification/slides/liG97YXaTSA/audit.json`

## Slide-Derived Subjects To Review
Subject extraction uses video title, related session titles/descriptions, transcript context, and OCR text when available. OCR is best-effort and should be reviewed against the embedded slide images.
