---
title: "Slides: Beyond Transcription: Building Voice AI That Understands Conversations — Hervé Bredin, pyannoteAI"
category: "slides"
video_id: "mFLlVpnGpds"
sourceLabels: ["Public YouTube video frames", "Public YouTube metadata"]
---

# Slides: Beyond Transcription: Building Voice AI That Understands Conversations — Hervé Bredin, pyannoteAI

## Source Video
[Beyond Transcription: Building Voice AI That Understands Conversations — Hervé Bredin, pyannoteAI](https://www.youtube.com/watch?v=mFLlVpnGpds)

## Relationship To World's Fair 2026
These slides are extracted from a public AI Engineer YouTube video connected to World's Fair 2026. Speaker-matched clips are supporting context unless later confirmed as exact session recordings; official livestream recordings are day-level/event-level source material.

## Related Scheduled Sessions
- No individual scheduled session mapping has been assigned yet; treat this as an event livestream deck.

## Extracted Slides
![[assets/slides/mFLlVpnGpds/slide-002.jpg]]

- Recreated text/layout view: [open HTML recreation](/assets/slide-recreations/slides/mFLlVpnGpds/slide-002.html)
- AI slide classifier: `content_slide` confidence `0.98`
- Text source: advanced OCR `rapidocr-live/border-trim/contrast`.
- OCR decision: ready — dense slide with screenshot, diagram, and small text

Slide text:

> Transcription
> WHAT was said
> AIE
> lorem ipsum dolor sit amet consectetur adipiscing elit
> hendrerit morbi sed nunc morbi convallis erat non
> mollis ex sollicitudin porttitor velit at
> pyannoteAl
> Google DeepMind
> AlEnginee

![[assets/slides/mFLlVpnGpds/slide-003.jpg]]

- Recreated text/layout view: [open HTML recreation](/assets/slide-recreations/slides/mFLlVpnGpds/slide-003.html)
- AI slide classifier: `content_slide` confidence `0.99`
- Text source: agent_vision.
- OCR decision: ready — dense slide with code snippet, chart, and small labels

Slide text:

> What can go wrong?
> hopefully the answer is not "the live demo"

![[assets/slides/mFLlVpnGpds/slide-004.jpg]]

- Recreated text/layout view: [open HTML recreation](/assets/slide-recreations/slides/mFLlVpnGpds/slide-004.html)
- AI slide classifier: `content_slide` confidence `0.99`
- Text source: advanced OCR `rapidocr-live/bright-screen/contrast`.
- OCR decision: ready — dense chart slide with small labels and benchmark text

Slide text:

> State of the art (Q4 2025)
> Conversational telephone speech (CTS)
> DIHARDCTS/61conversations/2speakers FoseAorrn KieCeienCersen
> AIE pyannoteAi
> X
> 271
> 151 X.OA
> 22.5
> KN
> 5.5A Ln
> HipLsreney mengMini
> BPGE
> pyannote.ai/benchmark
> AI Engineer
> EUROPE

![[assets/slides/mFLlVpnGpds/slide-005.jpg]]

- Recreated text/layout view: [open HTML recreation](/assets/slide-recreations/slides/mFLlVpnGpds/slide-005.html)
- AI slide classifier: `content_slide` confidence `0.98`
- Text source: advanced OCR `rapidocr-live/full`.
- OCR decision: ready — dense comparison slide with chart and screenshot text

Slide text:

> What can go wrong? theblamegame Wirnport the b.Th Op Aye,o hat isnot it Eit Oseck Pe a the bet Deh afe evarivated.
> Subetution Deletionnsertion OM
> AIE tcorc-WER 50% mytahealth/Crispeisper Coherelabs/cohere-transcribe-03-2026 ihe-ganite/ganite-4.0-llsoeoch nodel 5.52 5.42 6.67 280.02 84.05 RTFx 524.88 LS6 Ogen 8,44
> ihm-ganit/granite-sech-3.3-2 Open 8,9
> ieganite/gnite-sch-3.3-8 5.74 145.42 Open 8.98
> 301 a.5 0n/03-AS-1.78 vidia/canaty-o 0-2.5 5.63 147.93 418.28 Open Open 95'01
> 20% s7 nicrosoft/Phi-4-multinodal-festruct usefelse $.66 20`9 151.1 448.15 Open Open 11.09 10.68
> 6.05 3386.02 Open
> evidla/oarakeet-tet-0.6b-3 6.32 3332.74 ope 11.39
> 6,42 146.23 Open
> Whisperx Parckeet
> pyannote.ai/blog/stt-orchestration hf.co/spaces/hf-audio/open_asr_leaderboard
> AIEngineer
> EUROPE


### Hidden Non-Slide Evidence
- [`slide-001.jpg`](/assets/slides/mFLlVpnGpds/slide-001.jpg) — `speaker_stage` confidence `0.99`; speaker at podium with projected slide; not a readable slide capture
- [`slide-006.jpg`](/assets/slides/mFLlVpnGpds/slide-006.jpg) — `sponsor_logo` confidence `0.99`; logo/end card, not a content slide

Classification audit: `raw/sources/slide-ai-classification/slides/mFLlVpnGpds/audit.json`

## Slide-Derived Subjects To Review
Subject extraction uses video title, related session titles/descriptions, transcript context, and OCR text when available. OCR is best-effort and should be reviewed against the embedded slide images.
