---
title: "Slides: Building voice agents with OpenAI — Dominik Kundel, OpenAI"
category: "slides"
video_id: "iXhba366fQc"
sourceLabels: ["Public YouTube video frames", "Public YouTube metadata"]
---

# Slides: Building voice agents with OpenAI — Dominik Kundel, OpenAI

## Source Video
[Building voice agents with OpenAI — Dominik Kundel, OpenAI](https://www.youtube.com/watch?v=iXhba366fQc)

## Relationship To World's Fair 2026
These slides are extracted from a public AI Engineer YouTube video connected to World's Fair 2026. Speaker-matched clips are supporting context unless later confirmed as exact session recordings; official livestream recordings are day-level/event-level source material.

## Related Scheduled Sessions
- No individual scheduled session mapping has been assigned yet; treat this as an event livestream deck.

## Extracted Slides
![[assets/slides/iXhba366fQc/slide-001.jpg]]

OCR text:

> aWS
> 
> ee)
> OKCT ol gk Mn a Aine SO RMU ronTeTo) DIE
> #Mdaily £3augmentcode WorkOS

![[assets/slides/iXhba366fQc/slide-002.jpg]]

OCR text:

> Before we get started Workshop Slack
> Morkshop-openai-voice-agents
> Starter repository
> Documentation
> elt.
> agers
> Peon Ae
> tae i
> TAA PaT:
> te
> aws
> World's Fair ed

![[assets/slides/iXhba366fQc/slide-003.jpg]]

OCR text:

> Why voice agents? - _
> Me Eg
> 1-800-CHATGPT
> Voice agents provide a wide range of benefits
> * More flexible than traditional \VRs (no need
> to outline complex phone trees, multilingual).
> Let human agents focus on complex cases
> + Makes your Al more accessible to broader
> parts of the population
> + More information dense and personalized . ,
> than just text by introducing emotions a *
> + Can act as an API to the real-world by
> leading conversations with other businesses & vy ref
> Worth Fair

![[assets/slides/iXhba366fQc/slide-004.jpg]]

OCR text:

> Challenges of
> speech-to-speech
> ‘
> !
> Complex decision making Reusing existing capabilities Dealing with complex states
> The speech-to-spaech model does not support Hf you've been investing in specalzed agents. tnstead of bukcing chasns that can represent
> the levat of reasoning that more advanced that rely on specitic modeis tor certain parts of a different parts of a workflow you are dealing with
> models tke 03 or o4-mun support. text-based flow, it’s harder to reuse them an agent that has an ongaing conversation.
> Wortds Fate ~ |

![[assets/slides/iXhba366fQc/slide-005.jpg]]

OCR text:

> DoGo
> net/tan
> "hiner/ec
> AIE
> rt].l.
> kisteny)
> -erDer'
> maDerTesl
> crigtiac:etD
> mather ia gives locatien
> leatie:.srig0,
> ncrigtiao
> 'SONE'
> aws
> Worid'sFair

![[assets/slides/iXhba366fQc/slide-006.jpg]]

OCR text:

> General tips for
> 
> building voice agents
> 
> Start small with a clear goal Build evals and guardrails early. Customize and create a personalized
> Collect feedback brand

![[assets/slides/iXhba366fQc/slide-007.jpg]]

OCR text:

> Oaewspets-werahep 11
> Terinina
> 01-bad
> 21 Ar
> toruerlanr
> AIE
> UD
> 1A2.02 UTT-D S（）Typcrot
> aws
> Wrs F

![[assets/slides/iXhba366fQc/slide-008.jpg]]

OCR text:

> Traces
> * Agents SDK automatically sends traces to the Traces Wredhes wee scat a8
> dashboard in your OpenA! account uo “eee _. ~o
> - Review conversations . _ /
> * Replay audio ;
> - Inspect tool calls aa
> e .
> U
> Workts Far cs >]

![[assets/slides/iXhba366fQc/slide-009.jpg]]

OCR text:

> AI Engineer
> World's Fair

## Slide-Derived Subjects To Review
Subject extraction uses video title, related session titles/descriptions, transcript context, and OCR text when available. OCR is best-effort and should be reviewed against the embedded slide images.
