---
title: "Slides: Voice In, Visuals Out: The Agony and the Ecstasy - Allen Pike, Forestwalk Labs"
category: "slides"
video_id: "65X0pQ6Lmbg"
sourceLabels: ["Public YouTube video frames", "Public YouTube metadata"]
---

# Slides: Voice In, Visuals Out: The Agony and the Ecstasy - Allen Pike, Forestwalk Labs

## Source Video
[Voice In, Visuals Out: The Agony and the Ecstasy - Allen Pike, Forestwalk Labs](https://www.youtube.com/watch?v=65X0pQ6Lmbg)

## Relationship To World's Fair 2026
These slides are extracted from a public AI Engineer YouTube video connected to World's Fair 2026. Speaker-matched clips are supporting context unless later confirmed as exact session recordings; official livestream recordings are day-level/event-level source material.

## Related Scheduled Sessions
- No individual scheduled session mapping has been assigned yet; treat this as an event livestream deck.

## Extracted Slides
![[assets/slides/65X0pQ6Lmbg/slide-001.jpg]]

- Recreated text/layout view: [open HTML recreation](/assets/slide-recreations/slides/65X0pQ6Lmbg/slide-001.html)
- AI slide classifier: `title_card` confidence `0.99`
- Text source: agent_vision.

Slide text:

> Voice In, Visuals Out
> The Agony and the Ecstasy

![[assets/slides/65X0pQ6Lmbg/slide-002.jpg]]

- Recreated text/layout view: [open HTML recreation](/assets/slide-recreations/slides/65X0pQ6Lmbg/slide-002.html)
- AI slide classifier: `content_slide` confidence `0.98`
- Text source: agent_vision.

Slide text:

> “Audio is the human-preferred input to AIs, but vision is the preferred output from them.”
> — Andrej Karpathy, May 2026

![[assets/slides/65X0pQ6Lmbg/slide-003.jpg]]

- Recreated text/layout view: [open HTML recreation](/assets/slide-recreations/slides/65X0pQ6Lmbg/slide-003.html)
- AI slide classifier: `content_slide` confidence `0.97`
- Text source: agent_vision.
- OCR decision: ready — Dense diagram and UI screenshot text are better handled by OCR; only short obvious labels are captured here.

Slide text:

> Visualization
> Interactivity

![[assets/slides/65X0pQ6Lmbg/slide-004.jpg]]

- Recreated text/layout view: [open HTML recreation](/assets/slide-recreations/slides/65X0pQ6Lmbg/slide-004.html)
- AI slide classifier: `content_slide` confidence `0.97`
- Text source: agent_vision.
- OCR decision: ready — Dense diagram/UI content with small labels is OCR-suitable; only the short visible labels are captured here.

Slide text:

> Visualization
> Interactivity
> Beauty

![[assets/slides/65X0pQ6Lmbg/slide-005.jpg]]

- Recreated text/layout view: [open HTML recreation](/assets/slide-recreations/slides/65X0pQ6Lmbg/slide-005.html)
- AI slide classifier: `title_card` confidence `0.95`
- Text source: agent_vision.

Slide text:

> Voice In

![[assets/slides/65X0pQ6Lmbg/slide-006.jpg]]

- Recreated text/layout view: [open HTML recreation](/assets/slide-recreations/slides/65X0pQ6Lmbg/slide-006.html)
- AI slide classifier: `demo_video` confidence `0.94`
- Text source: none.
- OCR decision: ready — The slide is a row of video thumbnails with small captions and view counts; OCR is better suited than manual transcription in this pass.
- Slide text: not surfaced (`illegible` by AI classifier).
![[assets/slides/65X0pQ6Lmbg/slide-008.jpg]]

- Recreated text/layout view: [open HTML recreation](/assets/slide-recreations/slides/65X0pQ6Lmbg/slide-008.html)
- AI slide classifier: `title_card` confidence `0.98`
- Text source: agent_vision.

Slide text:

> The Tyranny of Latency

![[assets/slides/65X0pQ6Lmbg/slide-009.jpg]]

- Recreated text/layout view: [open HTML recreation](/assets/slide-recreations/slides/65X0pQ6Lmbg/slide-009.html)
- AI slide classifier: `content_slide` confidence `0.98`
- Text source: agent_vision.

Slide text:

> 200ms seamless voice
> 100ms feels instant

![[assets/slides/65X0pQ6Lmbg/slide-010.jpg]]

- Recreated text/layout view: [open HTML recreation](/assets/slide-recreations/slides/65X0pQ6Lmbg/slide-010.html)
- AI slide classifier: `content_slide` confidence `0.99`
- Text source: agent_vision.

Slide text:

> Pillars of low latency
> 1 Fast models

![[assets/slides/65X0pQ6Lmbg/slide-011.jpg]]

- Recreated text/layout view: [open HTML recreation](/assets/slide-recreations/slides/65X0pQ6Lmbg/slide-011.html)
- AI slide classifier: `content_slide` confidence `0.99`
- Text source: agent_vision.

Slide text:

> Pillars of low latency
> 1 Fast models
> 2 Short intervals
> 3 Stable cache


### Hidden Non-Slide Evidence
- [`slide-007.jpg`](/assets/slides/65X0pQ6Lmbg/slide-007.jpg) — `sponsor_logo` confidence `0.98`; logo-only branding frame with no substantive slide content

Classification audit: `raw/sources/slide-ai-classification/slides/65X0pQ6Lmbg/audit.json`

## Slide-Derived Subjects To Review
Subject extraction uses video title, related session titles/descriptions, transcript context, and OCR text when available. OCR is best-effort and should be reviewed against the embedded slide images.
