Slides: Voice AI: when is the "Her" moment? — Neil Zeghidour, CEO, Gradium AI
Source Video
Voice AI: when is the "Her" moment? — Neil Zeghidour, CEO, Gradium AI
Relationship To World's Fair 2026
These slides are extracted from a public AI Engineer YouTube video connected to World's Fair 2026. Speaker-matched clips are supporting context unless later confirmed as exact session recordings; official livestream recordings are day-level/event-level source material.
Related Scheduled Sessions
- No individual scheduled session mapping has been assigned yet; treat this as an event livestream deck.
Extracted Slides

OCR text:
PLATINUM SPONSORS
Braintrust WorkOS OpenAI

OCR text:
aa raat? ; er ; a
oA Tdps
‘ . no a a NC) VAsteLnT(G (OU
ky en ae ee oe :
a we eC ae CEO & Cofounder
a a a £ “ Gradium
ee i ae
brig ete oar oon
soos need ae
; y
N
a Google DeepMind
cee, tort

OCR text:
; coat ;
a ra .
~ a ae a ; a Pa ot
Pa 7 5 - - ne i : : eb,
i i coe Fi a F mee { : 7 be en ; :
at ¥ a ene i, a cae . 2 eS
Lar aed nr! ‘ & 4 f 4 aes ana * ary
A By 7 vs xs J . Le
eNmaeereamraauniock the unrealized potential of voice A a ae a a
i
flud natura voce as the new owerface for A ar ; 8 . a !
(ve tra vonce tode's basically SEE PTS, S28) os ; ; ea . woe a vee s & .
F es oor
, =
rn i
oe €3 Braintrust €} WorkOS OpenAl
a tenor

OCR text:
From ResearchtoProduction
Kyutai Breakthrough
Paris-basedopenscience Al lab. Founded2023withE300M. Moshi:firstspeech-native Research
AIE Team from MetaFAIR,Google DeepMind,and Inria. Hibiki-Zero:real-time
speech-to-speech translation.
moshi.chat by/kyutai Buito bringresearchtoproduction.$70Mraised in2025. Gradium:Scaling Impact Pocket-TTS:CPUmodelforTTS
AIEngineer
AIEngineer EUROPE

OCR text:
AIE
AI Engineer
Engineering the future of AI

OCR text:
AIE
Engineering the future of AI
AIEngineer

OCR text:
In real life...
AIE
AI Engineer
EUROPE

OCR text:
In real life...
ai-PULSE
AIE
byScdewoy
Engineering the future of Al
AIEngineer
COROFL

OCR text:
Engineering the future of AI

OCR text:
Whatismissing?
Latency
Contextual Fillers
User
AIE
Agent
STT
LLMFillers
TTS
Tool Calling
TTS
Engineering the future of Al
AIEngineer

OCR text:
page aineaea nena aeresess sewer
. on ID wamemen: ae @ th @
an > § nines, \ eres
a * a _ , aa al are ° . ie
Fy ry - Zz Leia Pa a
bd * ae se aan a nae aoa
lelos eic tk Mae | leis Ear eee ee: ee a ol. 8
_ €3 Braintrust €) WorkOS OpenAl

OCR text:
On, that's 2 great top<! Yeah, I'd hove to help you b
Peo ‘
bd bd
ere
bd ad
* ae bd
“ ‘=> . Ne |
. eee Wee te te oa! :
a) | | Al Engineer |
iD SUL ela
[nengne_| ;

OCR text:
lie glayed:
AIE
/.kyutai
Engineering the future of Al
AIEngineer

OCR text:
Moshi Lessons Learned
What Works Well
Full-Duplex Conversation: Natural, uninterrupted
dialogue flow
Real-time Interaction: Truly conversational AI, user and
AI speak simultaneously
Challenges
A research prototype, not a real agent
Lack of observability and reliability
Does not convey or understand empathy well
Engineering the future of AI

OCR text:
GradiumPhonon Real-Time InferenceonCPU
CPUInference Weights WER SpeakerSim
Faster than real-time with no perceptible delay Phonon ~100M 1.48% 56.37%
AIE Kani TTS 2 450M 4.97% 40.73%
Personalization NeuTTS Air 552M 2.18% 47.51%
Multilingual+canreproduce any voicefroma short referenceclip,noretrainingrequired Kokoro NeuTTSNano 229M 100M 1.71% 0.90% 40.15%
SOTAperformance
Benchmarkon seed-tts→
Engineering the future of Al
AIEngineer

OCR text:
Cn “Eye +h O® Be Fak wweun
ED eo rete mene se ws
* Te ee ree entree sated Le QP ESE ORE MUNN enn ad ERE abe a a6 @.% @
JINR RT coca Ore Bo Hn ee 6 Oo
Eo oF" 4.59 SAGE IR OAS G 1 ewe ceRE ee oro wee Cgcetginite =
»
z=
bg ”
ace =
Ls “hepath fopaatdfan yc ce espe y fort oped aerators bes a 2
loan ial . , —
Bi science and engineering Scene
F ; >
ca [od ee CCS SE Oca SOMES A LO OOO BO SO LTRS MESS
-_ - *
Taner)
= sere rte teen
wBagCF®eO<Sse BBO. - "474° OB * Sse Le
‘a . Engineering the fut f Al
I - Hl

OCR text:
AI Engineer
EUROPE
HTTPS://AI.ENGINEER
Slide-Derived Subjects To Review
Subject extraction uses video title, related session titles/descriptions, transcript context, and OCR text when available. OCR is best-effort and should be reviewed against the embedded slide images.