The Golden Age of AI Engineering
Conference Context
- Date/time: 2026-06-29 · 9:25am-9:45am
- Track/room: Software Factories · Main Stage
- Speaker(s): Alexander Embiricos, Romain Huet
- Session type/status: keynote · confirmed
- Track: Software Factories
- Room: Main Stage
- Session type: keynote
- Status: confirmed
Session Description
TBD
Media Evidence
From Text to Vision to Voice Exploring Multimodality with Open AI: Romain Huet (speaker-match related prior/adjacent AI Engineer video; captions: English auto-captions).
- Source video:
youtube-yJHw33cVeHo - Slide deck: Dense Slides: From Text to Vision to Voice Exploring Multimodality with Open AI: Romain Huet — 1 visible slide image(s); 1 HTML recreation(s).
- Additional slide evidence: Slides: From Text to Vision to Voice Exploring Multimodality with Open AI: Romain Huet, Reconstructed Slides: From Text to Vision to Voice Exploring Multimodality with Open AI: Romain Huet
- Slide-derived themes for
youtube-yJHw33cVeHo: hello.

- youtube pMggiOb18tc transcript — full cached transcript markdown for the related YouTube source.
Evidence Graph
This evidence graph is generated from currently linked source material: official schedule text, related video pages, cached transcripts, visible slide text, dense/reconstructed slide pages, and AI slide-classification audits.
Media Signals
youtube-yJHw33cVeHo— 1 slide-derived text signals- Slide-derived themes for
youtube-yJHw33cVeHo: hello. - Evidence links for
youtube-yJHw33cVeHo: youtube yJHw33cVeHo, youtube yJHw33cVeHo slides, youtube yJHw33cVeHo dense slides, youtube yJHw33cVeHo reconstructed slides
Agent Reading Notes
Use these signals to refine the synopsis, topic links, people/company context, and method notes. If a source is a related external video rather than an exact official recording, keep it framed as supporting evidence.
Summary
This scheduled keynote sits in the Software Factories track and brings together two OpenAI product leaders: Alexander Embiricos, Head of Enterprise Product and former Codex product lead, and Romain Huet, Head of Developer Experience. The available supporting material is not yet a confirmed recording of this exact session, but it does connect the session to OpenAI's developer-facing story around multimodal AI systems, including text, vision, voice, model capability shifts, and practical engineering workflows.
The linked Romain Huet video, "From Text to Vision to Voice Exploring Multimodality with Open AI," provides the strongest contextual evidence currently attached to this page. Its extracted and reconstructed slide pages point toward themes such as GPT-4-era model evolution, cheaper and more capable models, vision use cases, natural interfaces, and developer tooling. Taken together with Embiricos's background in Codex and enterprise product, the session is likely best treated as a keynote about how AI engineering is moving from isolated model demos toward production software factories, where multimodal models, developer platforms, and enterprise workflows become part of the same engineering system.
Transcript Status
Related video transcript availability: English auto-captions. Treat this as supporting context, not a recording of this exact scheduled session unless later confirmed. Not fetched yet.
People
Supporting Slides
- youtube yJHw33cVeHo slides — extracted from the related public AI Engineer video.
Slide Evidence
- Slide-only cropped deck: youtube yJHw33cVeHo dense slides (3 viable slide images).
- Related slide/OCR pages:
- youtube yJHw33cVeHo dense slides
- youtube yJHw33cVeHo reconstructed slides
- youtube yJHw33cVeHo slides
- Slide-derived terms:
gpt-40,models,intelligence,cases,gpt-4,world,model,string,custom,openal,fair,toolbox,cheaper,trained,vision,next,gpt-3,natural
Official YouTube Recording
- youtube pMggiOb18tc — official AI Engineer YouTube channel recording published 2026-07-09.
- Evidence status: transcript/slide enrichment pending.
- Boundary: use this recording as media evidence; keep date/time/room facts tied to the official schedule.
Synthesis
Synthesized Breakdown
Good morning everyone. I'm Raman. Hey everyone. I'm Alexander.
Speaker And Company Context
- No speaker profile is attached in the official schedule data.
Topics Covered
Derived Links And Source Material
- youtube pMggiOb18tc transcript — transcript markdown; source cache
raw/sources/youtube-transcripts/pMggiOb18tc.txt(4,606 words). - youtube pMggiOb18tc — related YouTube source page.
- youtube yJHw33cVeHo — related YouTube source page.
- youtube yJHw33cVeHo slides — slide evidence.
- youtube yJHw33cVeHo reconstructed slides — slide evidence.
- youtube yJHw33cVeHo dense slides — slide evidence.
Novel Concepts / Clever Methods
- No highlighted novel concept has been detected yet.
Evidence Boundary
This synthesis uses the official schedule plus cached video transcripts. Official AI Engineer World's Fair San Francisco 2026 livestreams and cut videos are primary event video sources for transcript/slide evidence; external, historical, or speaker-matched videos remain supporting context unless manually verified as exact official event recordings.