---
title: "Voice is the universal interface"
category: "talks"
date: "2026-07-01"
time: "11:40am-12:00pm"
track: "Expo Stage 3 SW"
room: "Expo Stage 3 SW"
speakers: ["Kwindla Kramer", "Neil Zeghidour"]
sourceLabels: ["Official conference schedule", "Public YouTube metadata"]
scheduleTrack: ""
scheduleRoom: "Expo Stage 3 SW"
scheduleLabels: ["Expo Stage 3 SW", "session", "confirmed"]
---
# Voice is the universal interface

## Conference Context
- Date/time: 2026-07-01 · 11:40am-12:00pm
- Track/room: track TBD · Expo Stage 3 SW
- Speaker(s): Kwindla Kramer, Neil Zeghidour
- Session type/status: session · confirmed

- Track: track TBD
- Room: Expo Stage 3 SW
- Session type: session
- Status: confirmed

## Session Description
Language models give us the ability to create natural language, conversational, interfaces for computers. We are seeing a rapid shift among early adopters to using general language instead of traditional user interfaces for tasks like writing code and editing spreadsheets. Join the cofounders of Pipecat, Gradium, and Daily as we discuss the future of realtime voice and AI interfaces. Voice is the most efficient input mode for natural-language systems, and often the most efficient output mode, as well. But good voice interfaces require a very high degree of conversational facility, intelligence, task-specific reliability, and robustness to real-world realities like multiple speakers and background noise. There's a long history of voice interfaces in science fiction: Star Trek, Iron Man, Her. We'll use these depictions of computing possibilities as a jumping off point for talking about the ideal voice interface. How close are we to being able to build these interfaces with today's models, hardware, orchestration tooling, and UI libraries? What are the most promising research directions? What did the movies get wrong, now that we actually have experience building natural language, open-ended, voice systems?

## Media Evidence
[Voice AI: when is the "Her" moment? — Neil Zeghidour, CEO, Gradium AI](https://www.youtube.com/watch?v=P_RI1kCkRbo) (speaker-match related prior/adjacent AI Engineer video; captions: English auto-captions).

- Source video: `youtube-P_RI1kCkRbo`
- Slide deck: [[youtube-P_RI1kCkRbo-dense-slides|Dense Slides: Voice AI: when is the \"Her\" moment? — Neil Zeghidour, CEO, Gradium AI]] — 1 visible slide image(s); 1 HTML recreation(s).
![[assets/dense-slides/P_RI1kCkRbo/slide-001.jpg]]
- Additional slide evidence: [[youtube-P_RI1kCkRbo-slides|Slides: Voice AI: when is the \"Her\" moment? — Neil Zeghidour, CEO, Gradium AI]], [[youtube-P_RI1kCkRbo-reconstructed-slides|Reconstructed Slides: Voice AI: when is the \"Her\" moment? — Neil Zeghidour, CEO, Gradium AI]]
- Slide-derived themes for `youtube-P_RI1kCkRbo`: engineering, future, brig, team, google, translation, great, hove.

## Evidence Graph
This evidence graph is generated from currently linked source material: official schedule text, related video pages, cached transcripts, visible slide text, dense/reconstructed slide pages, and AI slide-classification audits.

### Media Signals
- `youtube-P_RI1kCkRbo` — 10 slide-derived text signals
- Slide-derived themes for `youtube-P_RI1kCkRbo`: engineering, future, brig, team, google, translation, great, hove.
- Evidence links for `youtube-P_RI1kCkRbo`: [[youtube-P_RI1kCkRbo]], [[youtube-P_RI1kCkRbo-slides]], [[youtube-P_RI1kCkRbo-dense-slides]], [[youtube-P_RI1kCkRbo-reconstructed-slides]]

### Agent Reading Notes
Use these signals to refine the synopsis, topic links, people/company context, and method notes. If a source is a related external video rather than an exact official recording, keep it framed as supporting evidence.

## Transcript Status
Related video transcript availability: English auto-captions. Treat this as supporting context, not a recording of this exact scheduled session unless later confirmed. Not fetched yet.

## People
- [[kwindla-kramer]]
- [[neil-zeghidour]]

## Supporting Slides
- [[youtube-P_RI1kCkRbo-slides]] — extracted from the related public AI Engineer video.

## Slide Evidence
- Slide-only cropped deck: [[youtube-P_RI1kCkRbo-dense-slides]] (2 viable slide images).
- Related slide/OCR pages:
- [[youtube-P_RI1kCkRbo-dense-slides]]
- [[youtube-P_RI1kCkRbo-reconstructed-slides]]
- [[youtube-P_RI1kCkRbo-slides]]
- Slide-derived terms: `engineering`, `future`, `braintrust`, `workos`, `openal`, `engineer`, `aiengineer`, `gradium`, `kyutai`, `real`, `raat`, `tdps`, `vastelnt`, `cofounder`, `brig`, `soos`, `google`, `deepmind`

## Synthesis
### Synthesized Breakdown
# Voice is the universal interface ## Conference Context - Date/time: 2026-07-01 · 11:40am-12:00pm - Track/room: track TBD · Expo Stage 3 SW - Speaker(s): Kwindla Kramer, Neil Zeghidour - Session type/status: session · confirmed - Track: track TBD - Room: Expo Stage 3 SW - Session type: session - Status: confirmed ## Session Description Language models give us the ability to create natural language, conversational, interfaces for computers. We are seeing a rapid shift among early adopters to using general language instead of traditional user interfaces for tasks like writing code and editing spreadsheets. Join the cofounders of Pipecat, Gradium, and Daily as we discuss the future of realtime voice and AI interfaces. Voice is the most efficient input mode for natural-language systems, and often the most efficient output mode, as well.

### Speaker And Company Context
- [[kwindla-kramer|Kwindla Kramer]] — CEO at [[daily|Daily]].
- [[neil-zeghidour|Neil Zeghidour]] — Co-founder & CEO at [[gradium|Gradium]].

### Topics Covered
- [[agent-security]]
- [[agentic-search]]
- [[coding-agents]]

### Derived Links And Source Material
- [[youtube-P_RI1kCkRbo]] — related YouTube source page.
- [[youtube-P_RI1kCkRbo-slides]] — slide evidence.
- [[youtube-P_RI1kCkRbo-reconstructed-slides]] — slide evidence.
- [[youtube-P_RI1kCkRbo-dense-slides]] — slide evidence.

### Novel Concepts / Clever Methods
- No highlighted novel concept has been detected yet.

### Evidence Boundary
This synthesis is based on the official schedule and linked source pages. It should be revisited when exact session recordings or transcript-backed secondary sources are available.
