Voice Agents
Overview
Voice agents are AI systems that understand, reason, and respond through speech, often in real time. They combine speech recognition, speaker diarization, language models, tool use, dialogue state, text-to-speech, and sometimes visual or screen output.
Conference Context
They build on IVR systems, speech recognition, voice assistants, call-center automation, real-time media systems, and conversational AI. Modern multimodal and realtime models make them more fluid, but production voice still depends on latency, turn-taking, and trust.
Significance
Voice is natural for hands-free, high-attention, or emotionally sensitive workflows. It also exposes failures quickly: delays, interruptions, wrong speaker attribution, and unnatural responses break trust faster than in text.
Applied Use
Design around conversation state, latency budgets, interruption handling, speaker identity, fallback paths, and clear tool permissions. Test with realistic audio conditions, accents, overlapping speakers, and production transcripts.
Voice agents are useful in customer support, healthcare intake, sales calls, meeting assistants, field work, accessibility tools, tutoring, and companion interfaces.
Use voice when speaking is faster or more accessible than typing, or when the workflow happens away from a keyboard. Prefer text when precision, reviewability, or complex visual comparison is primary.
Design around conversation state, latency budgets, interruption handling, speaker identity, fallback paths, and clear tool permissions. Test with realistic audio conditions, accents, overlapping speakers, and production transcripts.
Voice agents are useful in customer support, healthcare intake, sales calls, meeting assistants, field work, accessibility tools, tutoring, and companion interfaces.
Use voice when speaking is faster or more accessible than typing, or when the workflow happens away from a keyboard. Prefer text when precision, reviewability, or complex visual comparison is primary.
Design around conversation state, latency budgets, interruption handling, speaker identity, fallback paths, and clear tool permissions. Test with realistic audio conditions, accents, overlapping speakers, and production transcripts.
Voice agents are useful in customer support, healthcare intake, sales calls, meeting assistants, field work, accessibility tools, tutoring, and companion interfaces.
Use voice when speaking is faster or more accessible than typing, or when the workflow happens away from a keyboard. Prefer text when precision, reviewability, or complex visual comparison is primary.
Design around conversation state, latency budgets, interruption handling, speaker identity, fallback paths, and clear tool permissions. Test with realistic audio conditions, accents, overlapping speakers, and production transcripts.
Voice agents are useful in customer support, healthcare intake, sales calls, meeting assistants, field work, accessibility tools, tutoring, and companion interfaces.
Use voice when speaking is faster or more accessible than typing, or when the workflow happens away from a keyboard. Prefer text when precision, reviewability, or complex visual comparison is primary.
Design around conversation state, latency budgets, interruption handling, speaker identity, fallback paths, and clear tool permissions. Test with realistic audio conditions, accents, overlapping speakers, and production transcripts.
Voice agents are useful in customer support, healthcare intake, sales calls, meeting assistants, field work, accessibility tools, tutoring, and companion interfaces.
Use voice when speaking is faster or more accessible than typing, or when the workflow happens away from a keyboard. Prefer text when precision, reviewability, or complex visual comparison is primary.
Design around conversation state, latency budgets, interruption handling, speaker identity, fallback paths, and clear tool permissions. Test with realistic audio conditions, accents, overlapping speakers, and production transcripts.
Voice agents are useful in customer support, healthcare intake, sales calls, meeting assistants, field work, accessibility tools, tutoring, and companion interfaces.
Use voice when speaking is faster or more accessible than typing, or when the workflow happens away from a keyboard. Prefer text when precision, reviewability, or complex visual comparison is primary.
Connections
- 2026 06 29 bohan li realtime voice agents with frontier intelligence — Realtime Voice Agents with Frontier Intelligence; Bohan Li (Day 2 — Session Day 1 · 2:50pm-3:10pm · Voice & Realtime AI; official schedule)
- 2026 06 29 charlie guo voice agents can just do things — Voice Agents Can Just Do Things; Charlie Guo (Day 2 — Session Day 1 · 11:40am-12:00pm · Voice & Realtime AI; official schedule)
- 2026 06 29 valeria wu fon speech to speech model research at google deepmind — Speech-to-Speech Model Research at Google DeepMind; Valeria Wu Fon, Tom Ouyang (Day 2 — Session Day 1 · 11:10am-11:30am · Voice & Realtime AI; official schedule)
- 2026 06 29 sumanyu sharma i monitored crime audio voice agents scare me more — I Monitored Crime Audio. Voice Agents Scare Me More.; Sumanyu Sharma (Day 2 — Session Day 1 · 2:25pm-2:45pm · Voice & Realtime AI; official schedule)
- 2026 06 29 midam kim my name is my name is a linguistic map for building and debugging voice agents — "My name is... my name is...": A Linguistic Map for Building and Debugging Voice Agents; Midam Kim (Day 2 — Session Day 1 · 3:20pm-3:40pm · Voice & Realtime AI; official schedule)
- 2026 06 29 venky b 5 voice agent failure modes you ll hit in week one — 5 Voice Agent Failure Modes You'll Hit in Week One; Venky B, Vyas A (Day 2 — Session Day 1 · 1:55pm-2:15pm · Voice & Realtime AI; official schedule)
- 2026 07 01 kwindla kramer voice is the universal interface — Voice is the universal interface; Kwindla Kramer, Neil Zeghidour (Day 4 — Session Day 3 · 11:40am-12:00pm · Expo Stage 3 SW; official schedule)
- 2026 06 29 paula dozsa tolan voice first ai companion — Tolan: Voice-First AI Companion; Paula Dozsa (Day 2 — Session Day 1 · 1:30pm-1:50pm · Voice & Realtime AI; official schedule)
- 2026 07 01 vivek muppalla 200 million patient interactions later what the generic voice stack misses — 200 Million Patient Interactions Later: What the Generic Voice Stack Misses; Vivek Muppalla (Day 4 — Session Day 3 · 12:05pm-12:25pm · AI in Healthcare; official schedule)
- 2026 06 29 fuad ali voice agents are mostly invisible here s how to see them — Voice Agents Are Mostly Invisible. Here's How to See Them.; Fuad Ali (Day 2 — Session Day 1 · 1:55pm-2:15pm · Expo Stage 2 NW; official schedule)
- 2026 06 29 neil zeghidour your voice agent is just a walkie talkie — Your Voice Agent is Just a Walkie-Talkie; Neil Zeghidour (Day 2 — Session Day 1 · 12:05pm-12:25pm · Claws & Personal Agents; official schedule)
- 2026 06 30 thor schaeff build realtime multimodal agents with gemini live — Build realtime multimodal agents with Gemini Live; Thor 雷神 Schaeff (Day 3 — Session Day 2 · 10:45am-11:05am · Workshops Day 2; official schedule)
- 2026 06 30 thor schaeff build realtime multimodal agents with gemini live continued 2 — Build realtime multimodal agents with Gemini Live (continued 2); Thor 雷神 Schaeff (Day 3 — Session Day 2 · 11:10am-11:30am · Workshops Day 2; official schedule)
- 2026 06 30 thor schaeff build realtime multimodal agents with gemini live continued 3 — Build realtime multimodal agents with Gemini Live (continued 3); Thor 雷神 Schaeff (Day 3 — Session Day 2 · 11:40am-12:00pm · Workshops Day 2; official schedule)
- 2026 06 30 matt lawler fde playbook build an ai support agent and give it a voice — FDE Playbook: Build an AI Support Agent and Give It a Voice; Matt Lawler (Day 3 — Session Day 2 · 11:10am-11:30am · Expo Stage 4; official schedule)
- 2026 06 30 rishab kumar from stateless to stateful orchestrating real time voice and messaging agents with twilio and amazon bedrock — From Stateless to Stateful: Orchestrating Real-Time Voice & Messaging Agents with Twilio and Amazon Bedrock; Rishab Kumar (Day 3 — Session Day 2 · 12:05pm-12:25pm · Expo Stage 2 NW; official schedule)
- 2026 06 29 nick nisi lifestyles of the ai native voice coding agent skills hooks and scheduled tasks — Lifestyles of the AI-Native: Voice-coding, agent skills, hooks and scheduled tasks; Nick Nisi, Zack Proser (Day 1 — Workshop Day · 4:30pm-5:30pm · Workshops Day 1; official schedule)
- 2026 07 01 lina colucci voice agents with realtime video — Voice agents with Realtime Video; Lina Colucci (Day 4 — Session Day 3 · 1:55pm-2:15pm · Generative Media; official schedule)
- 2026 07 01 idan gazit realtime multiplayer automation and you — Realtime multiplayer, automation, and you!; Idan Gazit (Day 4 — Session Day 3 · 2:50pm-3:10pm · Agentic Engineering; official schedule)
- 2026 06 29 sonar expo welcome speech — Expo Welcome Speech; Sonar, Extend AI (Day 1 — Workshop Day · 6:00pm-6:15pm · Expo Stage 3; official schedule)
- 2026 06 30 thor schaeff build realtime multimodal agents with gemini live continued 4 — Build realtime multimodal agents with Gemini Live (continued 4); Thor 雷神 Schaeff (Day 3 — Session Day 2 · 12:05pm-12:25pm · Workshops Day 3; official schedule)
- 2026 06 29 amit desai act confirm or stop smarter behavior for ai assistants wearables and robots — Act, Confirm, or Stop? Smarter behavior for AI assistants, wearables & robots; Amit Desai (Day 2 — Session Day 1 · 3:45pm-4:05pm · Voice & Realtime AI; official schedule)
- 2026 06 29 kwindla kramer the new primitives building ai native software — The New Primitives: Building AI-Native Software; Kwindla Kramer (Day 2 — Session Day 1 · 10:45am-11:05am · Voice & Realtime AI; official schedule)
- 2026 07 01 sai krishna rallabandi wearing the agent engineering a family and friends personal agent from group chats to glasses — Wearing the Agent: Engineering a Family-and-Friends Personal Agent, from Group Chats to Glasses; Sai Krishna Rallabandi (Day 4 — Session Day 3 · 3:45pm-4:05pm · AI in Finance; official schedule)
- Neil Zeghidour
- Thor 雷神 Schaeff
- Kwindla Kramer
- Kent C. Dodds
- swyx
- Bohan Li
- Charlie Guo
- Valeria Wu Fon
- Tom Ouyang
- Sumanyu Sharma
- Midam Kim
- Venky B
- Vyas A
- Paula Dozsa
- Vivek Muppalla
- Fuad Ali
- Matt Lawler
- Rishab Kumar
- Nick Nisi
- Zack Proser
- Lina Colucci
- Idan Gazit
- Sonar
- Extend AI
- Google DeepMind
- Gradium
- Plivo
- Daily
- WorkOS
- CoupleWork AI
- EpicProduct.engineer
- SonderMind
- Latent Space / AI Engineer
- EliseAi
- OpenAI
- Hamming AI
- ServiceNow
- Tolan
- Hippocratic AI
- Arize AI
- AssemblyAI
- Twilio
- youtube I2cbIws9j10 slides — WF26: Harness Engineering & Startup Battlefield ft. Garry Tan, Mike Krieger, @t3dotgg , DSPy (80 extracted slide frames)
Evidence Graph
Transcript-backed resources
- youtube mFLlVpnGpds — Beyond Transcription: Building Voice AI That Understands Conversations — Hervé Bredin, pyannoteAI
- youtube u rJwPPU3QA — How to talk to statues — Joe Reeve, ElevenLabs
- youtube ij AU9dpJjc — Stop Writing Tone Instructions. Layer Them. - Isadora Martin-Dye, Isadora & Co
- youtube 65X0pQ6Lmbg — Voice In, Visuals Out: The Agony and the Ecstasy - Allen Pike, Forestwalk Labs
- youtube hVJOnuhFmTA — The Prompt Is Still a Punch Card - Ted Johnson, JoinIn AI
- youtube HsxQICTLF84 — Building an ACP-Compatible Agent Live — Bennet Fenner, Zed
- youtube kZsf_Sfm7RU — The Missing Layer After Launch - Raphael Kalandadze, Wandero AI
- youtube G6IlDzj8OjA — GTM Is You - Victoria Melnikova, Evil Martians
Quote signals
- “We can just switch to having voice in visuals out and then we benefit from the more forgiving visual response envelope that people have.” — youtube 65X0pQ6Lmbg
- “And today I'm going to be sharing some of what we've learned building voice in visuals out experiences using AI.” — youtube 65X0pQ6Lmbg
- “But there's some real breakthroughs over the last few months where both of this visuals in and audio or the audio in and visuals out experience is now um really feasible.” — youtube 65X0pQ6Lmbg
- “So for that you need to know who's who who spoke when.” — youtube mFLlVpnGpds
- “And the question is now we need to assign a speaker to each word.” — youtube mFLlVpnGpds
- “So in there you you'll be able to to with Pianoteq open source toolkit for diarization with piano metrics for evaluation.” — youtube mFLlVpnGpds
Transcript-backed resources
Quote signals
This evidence graph consolidates scheduled talks, linked videos, transcripts, and slide-derived material connected to this topic.
Linked Sessions
- Realtime Voice Agents with Frontier Intelligence
- Voice Agents Can Just Do Things
- Speech-to-Speech Model Research at Google DeepMind
- I Monitored Crime Audio. Voice Agents Scare Me More.
- \"My name is... my name is...\": A Linguistic Map for Building and Debugging Voice Agents
- 5 Voice Agent Failure Modes You'll Hit in Week One
- Voice is the universal interface
- Tolan: Voice-First AI Companion
- 200 Million Patient Interactions Later: What the Generic Voice Stack Misses
- Voice Agents Are Mostly Invisible. Here's How to See Them.
Media Signals
youtube-dvft0Gp9sEE— 1,508 transcript words; 1 slide-derived text signals- Transcript signals for
youtube-dvft0Gp9sEE: analysis, sales, transcript, models, calls, data, single, doing. - Slide-derived themes for
youtube-dvft0Gp9sEE: manual, analysis, keyword. - Evidence links for
youtube-dvft0Gp9sEE: youtube dvft0Gp9sEE, youtube dvft0Gp9sEE transcript, youtube dvft0Gp9sEE slides, youtube dvft0Gp9sEE dense slides, youtube dvft0Gp9sEE reconstructed slides youtube-P_RI1kCkRbo— 10 slide-derived text signals- Slide-derived themes for
youtube-P_RI1kCkRbo: engineering, future, brig, team, google, translation, great, hove. - Evidence links for
youtube-P_RI1kCkRbo: youtube P_RI1kCkRbo, youtube P_RI1kCkRbo slides, youtube P_RI1kCkRbo dense slides, youtube P_RI1kCkRbo reconstructed slides youtube-u3NofYYstaY— 2 slide-derived text signals- Slide-derived themes for
youtube-u3NofYYstaY: oracle, team, accenture. - Evidence links for
youtube-u3NofYYstaY: youtube u3NofYYstaY, youtube u3NofYYstaY slides, youtube u3NofYYstaY dense slides, youtube u3NofYYstaY reconstructed slides youtube-SbcQYbrvAfI— 7 slide-derived text signals- Slide-derived themes for
youtube-SbcQYbrvAfI: planning, missing, data, domain, breaking, system, instructions, learned. - Evidence links for
youtube-SbcQYbrvAfI: youtube SbcQYbrvAfI, youtube SbcQYbrvAfI slides, youtube SbcQYbrvAfI dense slides, youtube SbcQYbrvAfI reconstructed slides
Source Coverage
This table summarizes the local evidence already linked from this topic. It is a navigation aid, not a claim that every linked page has been fully reviewed.
| Evidence type | Count | Review note |
| --- | ---: | --- |
| other | 44 | Related pages outside the main evidence categories. |
| resources | 12 | Video/resource pages; check source status before treating as primary event evidence. |
| slides | 13 | OCR or reconstructed slide evidence; mark claims as OCR-derived unless image-reviewed. |
| talks | 24 | Official schedule pages; use for titles, speakers, tracks, and stated talk framing. |
| transcripts | 1 | Transcript markdown; check session matching and caption quality. |
Talks
- 2026 06 29 bohan li realtime voice agents with frontier intelligence
- 2026 06 29 charlie guo voice agents can just do things
- 2026 06 29 valeria wu fon speech to speech model research at google deepmind
- 2026 06 29 sumanyu sharma i monitored crime audio voice agents scare me more
- 2026 06 29 midam kim my name is my name is a linguistic map for building and debugging voice agents
- 2026 06 29 venky b 5 voice agent failure modes you ll hit in week one
Resources
- youtube mFLlVpnGpds
- youtube u rJwPPU3QA
- youtube ij AU9dpJjc
- youtube 65X0pQ6Lmbg
- youtube hVJOnuhFmTA
- youtube HsxQICTLF84
Slides
- youtube I2cbIws9j10 slides
- youtube dvft0Gp9sEE slides
- youtube dvft0Gp9sEE dense slides
- youtube dvft0Gp9sEE reconstructed slides
- youtube P_RI1kCkRbo slides
- youtube P_RI1kCkRbo dense slides
Transcripts
This table summarizes the local evidence already linked from this topic. It is a navigation aid, not a claim that every linked page has been fully reviewed.
Talks
Resources
Slides
Transcripts
This table summarizes the local evidence already linked from this topic. It is a navigation aid, not a claim that every linked page has been fully reviewed.
Talks
Resources
Slides
Transcripts
This table summarizes the local evidence already linked from this topic. It is a navigation aid, not a claim that every linked page has been fully reviewed.
Talks
Resources
Slides
Transcripts
This table summarizes the local evidence already linked from this topic. It is a navigation aid, not a claim that every linked page has been fully reviewed.
Talks
Resources
Slides
Transcripts
This table summarizes the local evidence already linked from this topic. It is a navigation aid, not a claim that every linked page has been fully reviewed.
Talks
Resources
Slides
Transcripts
This table summarizes the local evidence already linked from this topic. It is a navigation aid, not a claim that every linked page has been fully reviewed.
Talks
Resources
Slides
Transcripts
This table summarizes the local evidence already linked from this topic. It is a navigation aid, not a claim that every linked page has been fully reviewed.
Talks
Resources
Slides
Transcripts
This table summarizes the local evidence already linked from this topic. It is a navigation aid, not a claim that every linked page has been fully reviewed.
Talks
Resources
Slides
Transcripts
Active Use Cases
- Realtime support and appointment workflows.
- Meeting and call understanding with speaker attribution.
- Voice-in visual-out interfaces for richer task completion.
- Hands-free operational assistants.