Markdown source

Slides: From Text to Vision to Voice Exploring Multimodality with Open AI: Romain Huet

Source Video

From Text to Vision to Voice Exploring Multimodality with Open AI: Romain Huet

Relationship To World's Fair 2026

These slides are extracted from a public AI Engineer YouTube video connected to World's Fair 2026. Speaker-matched clips are supporting context unless later confirmed as exact session recordings; official livestream recordings are day-level/event-level source material.

Related Scheduled Sessions

Extracted Slides

slide-001.jpg

OCR text:

INNOVATION SPONSOR

aws

PLATINUM SPONSORS

MongoDB.

Google Cloud

neo4j

slide-002.jpg

OCR text:

From Text to Vision to Voice:

Exploring Multimodality

with OpenAl

RomainHuet

Head ofDeveloper Experience,OpenAl

slide-003.jpg

OCR text:

01 AI Outlook

02 GPT-4o

03 What's Next

slide-004.jpg

OCR text:

| ae Mission

We are a research and deployment

ANE company working to build

artificial general intelligence (AGI)

that benefits all humanity.

|

J eee]

slide-005.jpg

OCR text:

Programming Assistance

Compliance Legaland CodeReview

AndVirtual Assistants Chatbots Use Cases GPT-3 Information Searchand Retrieval

Language Creation Content

Communication Emailand

slide-006.jpg

OCR text:

images and leaningsand Interpret Summarize visuals findings Personalize content Predict trends Use Cases GPT-4 Create assistive experiences prooesses Augment existing Automate workfiows Accelerate lsunches

concepts Explain Analyzedata

Translatetext inrealtime Visualizedata orsystems resources Allocate

slide-007.jpg

OCR text:

images and learnings and Interpret Summarize visuals findings Personalize content Predict trends Use Cases GPT-4 Create assistive experiences processes Augment eoosting Automate workflows Accelerate launches Goodmorning

concepts Explain Analyzedata

Translatetext inrealtime Visualizedata orsystems Alocate DJ Aaxdfea

slide-008.jpg

OCR text:

CR

cay

GPT-40 . ——

AN ; Te

terre car eee PPC NNT ot - on 7

penn comer py C

uo dl eine Es

aed ie f - >

B GPT-4 4 ase, A Mi

ee Use Cases oon

Z ee eI a

jog E i.

ena .

nes Da ea Lol f -- ——

: ve

F yY eH,

Pe ear ae ers ;

a en ns

Saar eas 5 f r f

Haste (ah tale moacomae) :

ane rt ae eral

slide-009.jpg

OCR text:

Milcrosoft Worid'sFall

OB World'sFair Goog World'sFair AlEngi neo4j

VISVIVO Galileo World'sFair

World'sFair ALER Covalent.

sFair ee

dby Wc Crusoe World's Fair AlEng

slide-010.jpg

OCR text:

eo

La

Introducing GPT-4o, our new flagship model

A step towards natural

human-computer interaction

e Multimodal reasoning

e Natural conversation @ee0ee@

e High intelligence

e Improved language support

e Vision and audio

r 4

slide-011.jpg

OCR text:

——oo Te CGR a Bt @ 7 2B wera

° ueers: Kemet: lem Greens 3

Mey O 2004

Hello GPT-40

wet ntrorees Nthat

ca ime.

cower A Oem

ey E — oo

den Hy

toy

we el rn’

a GPT-40

slide-012.jpg

OCR text:

(GO ewer we tw mmm ee ae eee ma Rte 6 oe Berean

oe eee ey a

[are Hello GPT-40

to 3a ee

. 4

wg at

| a @ GPr4o oY

slide-013.jpg

OCR text:

5 afer nia ence y

| : ee rr cers ra

AlEngineer

Fy .

World's Fair

cERTeaa Taare Oy] eee it

Microsoft

slide-014.jpg

OCR text:

Oo Ct fe te vee we Ue r @ @CGLj #@Q ti aBMBtia © a8 weriaw

8

.

er 2 cond Se

aoe eines se

i pyrene Sea ae

red ried

(ed Ne

Tot ad

Taped asecpmoecennneenes

i o

_ L

slide-015.jpg

OCR text:

6 wer le em ewe we ~‘~ ee OL @c a Gia a8 wer em

[AE] | a

meee a ose!

ee cS oa Se loeseed .

* wee

| —_ o

es

slide-016.jpg

OCR text:

"use client":

EventsourceMessage,I

fetchEventSource,

)fro"omicrosoft/fetch-eveet-source"

inpsrt(naneid}frcaai":

AIE

inport(useCallback,useEffect,useState}fron

"react":

inprtASSISTANT_ID,DNSTRUCTIONS,MOEL}r/COStas

lnport(Toolbax)fros*./toolbex":

export interface Message

id:string:

content: stringi

role:"ur"assistant"I"tool":

nawe?:string:

status?i"rumming”|"cogletee"

export Snterface MessagePaylead

content:string:

attachents7:(fste_sd:strlng: toots:(type:string 1)11:

port default fnction usessistant( toolbox):(toolbox Toolbox))

const (threadib,setThreaiD]auseStatestring|rutt-(ntt);

wecost [sessages, setHessages]-usestate-Hessagell):

coest[isurming,setisRuening]-use5tate(faise):

cotst [inpt,setInput]-usestate()

STsFa

Worid'sFai

Microseft

slide-017.jpg

OCR text:

nf

os

: ig e

7 1 han

i >a : ; , = a

f ‘

slide-018.jpg

OCR text:

AIE

Our investment areas

Textual

intelligence

Microsoft

smol.ai

slide-019.jpg

OCR text:

ino

o

(o7

Cc

ov

2

co

= aferor-W]

®

8

=

GPT-3 Era GPT-4 Era “GPT Next” Future Models

slide-020.jpg

OCR text:

Our investment areas

Textual

intelligence

Cheaper and

faster models

slide-021.jpg

OCR text:

@

Cheaper and faster models

Our models will We will continue

keep getting to release

cheaper. models of

different sizes.

slide-022.jpg

OCR text:

Our investment areas

Textual intelligence

Cheaper and faster models

Custom models

slide-023.jpg

OCR text:

Aaa

ag

Harvey.

Al Legal Technology

for Attorneys

ei deHsrent SaeraU lies)

Harvey worked with OpenAl to develop a - 83% increase in factual responses

custom trained model that has extensive

domain knowledge of case law to improve - Attorneys at top law firms preferred the

answer depth and reduce hallucination rates. custom trained model's outputs 94% of the

time over GPT-4

iTerotataliel tis

Custom Trained Model

slide-024.jpg

OCR text:

Co-worker agents

Docs

Code Repo

Calendar

CRM

Email

Multimodal AI

AI Agents

Slide-Derived Subjects To Review

Subject extraction uses video title, related session titles/descriptions, transcript context, and OCR text when available. OCR is best-effort and should be reviewed against the embedded slide images.