Markdown source

Dense Slides: Training Agentic Reasoners — Will Brown, Prime Intellect

Source Video

Training Agentic Reasoners — Will Brown, Prime Intellect

Method

This deck is slide-only. The existing captured video frame set supplies candidate frames, then local OpenCV rejects sponsor/title/speaker-only frames, crops visible slide surfaces, deduplicates, and saves the cropped slide images.

Cropped Visible Slides

slide-001.jpg

Slide text:

RL kinda works now

AIE Q DNeDSk.Al-Zero AbE aCCuraCy during trlnlng Hlrtet Summury >NvDu Corp 118.58 uso -20.64 (-14.83%) + past 5 dny

Caha 2? Lr. s or Pu Gfs + Dechtr A+ Noun +20 (4 +2 c0 1+ 74)

10 50 11 64t YTD 5Y LL

150

1410-2160 14 143++d 4++-11 rl-r+(o*+1+ 120 01

$Hro 000* 1000 Ito 24n

Microsoft smop

slide-002.jpg

Slide text:

the big labs are all doing it

AIE learning computo = bettor porformance' trond observod in GPT-serics pretraining. By retracing the scaling path--thia time in RL-we've pushed an additional order of magnitude in both tralning performance keeps chmbing. Continuing to scale reinforcement Throughout tha deveiopment of OpenAl o3, we've observed that large-scale reinforcement learning oxhibits the sama "more computa and inferonce-tirmo roasoning. yet still seo cloar performance gains, valldating that the models' porformance continues to lmprove the more they're allowed to think At equa! in ChatGPT-and wo'vo validatod that if wo lot k think longer, its Latcncy and cost with OpenAl ol, o3 delirvers hlgher performance Is Rl + Llms enough for AGl? - 123K views · 13 days ag0 Sholto Douglas & Trenton Bricken AND NEXT WHAT COMES 2:24:02

Microsoft smop

slide-003.jpg

Slide text:

agents are also a thing TaskExecutionProcess MANUSACAGENT

Dlenaare

pupnd

AIE alcoes ts Clode Code researchsrevieei ela'fer bes fer getting stere 25003000.By sectingthsrgont efntredehen evalute the plte and se itirs cearer o idey wthapprmate cordiex1500 to2000and Let'smanually crop and inspecti

Analyzing data

slt.aist`gre)

-0.5.899.5. 799.5

Analyzing image

Microsoft smol?

slide-004.jpg

Slide text:

they’re kinda the same thing actually

slide-005.jpg

Slide text:

towardseverythingasync

AIE returnaaittode_asyncio,gathert rollowt_tasks-[ self._rn_single(sephere,client,model,pronpt,aniver,sapling_args,args) fetpronpt,answer-inzippronpts,ansers) *rollout_tasks, total-len(prompts), descof'Running(len(pronptsl)rollouts async_batch_generator.py async_dataloader_wrapper.py

grpo_config.py

grpo_trainer.py

GPU1 GPU2 GPU3 GPU4 1-8 9-16 17-24

Time

(up to four steps), Prime-RL matches the performance of synchronous baselines. us DeepScaleR trainingvs asynchronous Prime-RL under varyingasynchrony levels.

aws

slide-006.jpg

Slide text:

SFT warmup + small models = fun on just a few GPUs

PRumeintllet CreatenewGPU Cluster

AIE results=vf_env.evaluate( client=client, model=model_nane, sompling_args=sonpling_args, 200 141 Ga DO N100 B0 C Ala 00NS

num_saeples=nun_sanples 12.14

hub_odel_id"0wen2.5-7B-Math-Python-SFT", A100 0028

PCxSOM 1MS

trainer=SfTrainer( nodel=nodel,

argseargs,

train_dataset=datasettype:icnore

GH200 A100

trainer.train() 6) (/

入7, 344

Microsoft smol ai

Classification audit: raw/sources/slide-ai-classification/dense/PbHm2qKnu10/audit.json