---
title: "Video Discovery for Agentic World-Model Training"
category: "talks"
date: "2026-06-29"
time: "2:50pm-3:10pm"
track: "Expo Stage 2 NW"
room: "Expo Stage 2 NW"
speakers: ["Rafael Levi"]
sourceLabels: ["Official conference schedule", "Public YouTube metadata"]
scheduleTrack: ""
scheduleRoom: "Expo Stage 2 NW"
scheduleLabels: ["Expo Stage 2 NW", "session", "confirmed"]
---
# Video Discovery for Agentic World-Model Training

## Conference Context
- Date/time: 2026-06-29 · 2:50pm-3:10pm
- Track/room: track TBD · Expo Stage 2 NW
- Speaker(s): Rafael Levi
- Session type/status: session · confirmed

- Track: track TBD
- Room: Expo Stage 2 NW
- Session type: session
- Status: confirmed

## Session Description
Physical AI had its “Attention Is All You Need” moment with the rise of Vision-Language-Action models. The next bottleneck is data: not just more video, but the ability to find the exact real-world moments that teach models how the world works: gravity, motion, causality, human behavior, and object interactions. This session explores a new approach: discovering specific scenes from the vastness of the web. We’ll show how teams can search for moments like objects falling, people interacting with environments, or actions unfolding over time, then collect and structure only the relevant clips for training and evaluation. Attendees will learn how scene-level discovery changes multimodal data pipelines, reducing wasted collection, processing, storage, and review, while making it easier to build targeted datasets for VLA systems, robotics, physical AI, and agentic world models.

## Media Evidence
[From MCP to Scale: Pipelines That Build Themselves — Rafael Levi, Bright Data](https://www.youtube.com/watch?v=zTZ0qunQXnM) (speaker-match related prior/adjacent AI Engineer video; captions: English auto-captions).

- Source video: `youtube-zTZ0qunQXnM`
- Slide deck: [[youtube-zTZ0qunQXnM-reconstructed-slides|Reconstructed Slides: From MCP to Scale: Pipelines That Build Themselves — Rafael Levi, Bright Data]] — 2 visible slide image(s); 2 HTML recreation(s).
![[assets/reconstructed-slides/zTZ0qunQXnM/slide-007.jpg]]
![[assets/reconstructed-slides/zTZ0qunQXnM/slide-008.jpg]]
- Additional slide evidence: [[youtube-zTZ0qunQXnM-slides|Slides: From MCP to Scale: Pipelines That Build Themselves — Rafael Levi, Bright Data]]

## Evidence Graph
This evidence graph is generated from currently linked source material: official schedule text, related video pages, cached transcripts, visible slide text, dense/reconstructed slide pages, and AI slide-classification audits.

### Media Signals
- `youtube-zTZ0qunQXnM` — source page linked
- Evidence links for `youtube-zTZ0qunQXnM`: [[youtube-zTZ0qunQXnM]], [[youtube-zTZ0qunQXnM-slides]], [[youtube-zTZ0qunQXnM-reconstructed-slides]]

### Agent Reading Notes
Use these signals to refine the synopsis, topic links, people/company context, and method notes. If a source is a related external video rather than an exact official recording, keep it framed as supporting evidence.

## Transcript Status
Related video transcript availability: English auto-captions. Treat this as supporting context, not a recording of this exact scheduled session unless later confirmed. Not fetched yet.

## People
- [[rafael-levi]]

## Supporting Slides
- [[youtube-zTZ0qunQXnM-slides]] — extracted from the related public AI Engineer video.

## Synthesis
### Synthesized Breakdown
# Video Discovery for Agentic World-Model Training ## Conference Context - Date/time: 2026-06-29 · 2:50pm-3:10pm - Track/room: track TBD · Expo Stage 2 NW - Speaker(s): Rafael Levi - Session type/status: session · confirmed - Track: track TBD - Room: Expo Stage 2 NW - Session type: session - Status: confirmed ## Session Description Physical AI had its “Attention Is All You Need” moment with the rise of Vision-Language-Action models. The next bottleneck is data: not just more video, but the ability to find the exact real-world moments that teach models how the world works: gravity, motion, causality, human behavior, and object interactions. This session explores a new approach: discovering specific scenes from the vastness of the web. We’ll show how teams can search for moments like objects falling, people interacting with environments, or actions unfolding over time, then collect and structure only the relevant clips for training and evaluation.

### Speaker And Company Context
- [[rafael-levi|Rafael Levi]] — DevRel at [[bright-data|Bright Data]].

### Topics Covered
- [[agentic-search]]
- [[mcp]]

### Derived Links And Source Material
- [[youtube-zTZ0qunQXnM]] — related YouTube source page.
- [[youtube-zTZ0qunQXnM-slides]] — slide evidence.
- [[youtube-zTZ0qunQXnM-reconstructed-slides]] — slide evidence.

### Novel Concepts / Clever Methods
- No highlighted novel concept has been detected yet.

### Evidence Boundary
This synthesis is based on the official schedule and linked source pages. It should be revisited when exact session recordings or transcript-backed secondary sources are available.
