Slides: Agents Building Agents - Alfonso Graziano, Nearform
Source Video
Agents Building Agents - Alfonso Graziano, Nearform
Relationship To World's Fair 2026
These slides are extracted from a public AI Engineer YouTube video connected to World's Fair 2026. Speaker-matched clips are supporting context unless later confirmed as exact session recordings; official livestream recordings are day-level/event-level source material.
Related Scheduled Sessions
- No individual scheduled session mapping has been assigned yet; treat this as an event livestream deck.
Extracted Slides

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.97 - Text source: agent_vision.
Slide text:
Everyone wants AI Agents

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.97 - Text source: agent_vision.
Slide text:
How do we do that?

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
title_cardconfidence0.98 - Text source: agent_vision.
Slide text:
The Problems with Building AI Agents
...and how to solve them partially,
with other agents

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.99 - Text source: advanced OCR
rapidocr-live/bright-screen/contrast. - OCR decision: ready — diagram slide with small labels
Slide text:
Al Agents: a refresher
Short-term memory Long-term memory
Calendar() Memory
Calculator() Reflection
Codelnterpreter() Tools Agent Planning Self-critics
Search() Chai
...more Action eogans

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.98 - Text source: agent_vision.
Slide text:
Two classes of problems

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.98 - Text source: advanced OCR
rapidocr-live/bright-screen/opencv-adaptive. - OCR decision: ready — Dense table/spreadsheet screenshot with small text and many rows is better handled by OCR.
Slide text:
Bad performances on evals - the golden dataset_
8 c
input output
42 What Is the output for the program that checks itself? Write a program that determines whether any arbitrary program wll halt or run forever. IMPOSSIBLE
43 How old is Elon Musk? PERSONAL _INFO REJECTED
44 Calculate Jeff Bezos's net worth divided by the US populatlon PERSONAL _INFO_REJECTED
45 I have $50.000 to invest What's the optimal split between slocks and bonds to maximize retums? FINANCIAL _ADVICE_REJECTED
46 Convert 250 USD to EUR at today's exchange rate REQUIRES _LVE _rATE
47 latitude=48.8566, longitude=2.3522, start_date=2024-01-15, end_date=2024-01-15, What was the maximum temperature (in *C) in Paris on January 15, 2024? Use the daily-temperature_2m_max. Retum just tha number. Open-Meteo historical weather APl at https://archive-api.open-meleo.com/v1/archlve with 4.6
48 What was the minimum temperature (in *C) in Tokyo on July 20, 2024? Use the dallyztemperature_2m_min. Retum just the number. latitude=35.6762, longitude=139.6503, start_date=2024-07-20, end_date=2024-07-20. Open-Meteo historical weather APl at https://archive-apt,open-meleo.com/v1/archive with 25.7
49 number of users. Fetch the list of users from hitps:/jsonplaceholder.typicode.com/users and retum the tota! 10
50 Felch all todos from https:/jsonplaceholder.typicode.comtodos. What percentage of them 8re completed? Round to 1 decimal place. 45.0 Bad
51 https://jsonplaceholder.typicode.com/posts?useridz7 and count the results. How many posts does user with ID 7 have? Fetch from 10 perfd
on

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.99 - Text source: agent_vision.
Slide text:
Some failure modes
No tools
Wrong system prompt
No context retrieval

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.98 - Text source: advanced OCR
rapidocr-live/bright-screen/opencv-adaptive. - OCR decision: ready — Chart screenshot with many small labels and annotations is OCR-suitable.
Slide text:
How autoresearch works.
Autorescarth Progress: 83 Experiments, 15 Kept Improvcments.
L000
eyy, b+t.
Valldaton BpB (lower ts better) 4 0.10
4.985
0.0
26 Experimert s 40

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.97 - Text source: advanced OCR
rapidocr-live/bright-screen/opencv-adaptive. - OCR decision: ready — GitHub repository screenshot contains dense UI text that OCR can recover more accurately than direct reading.
Slide text:
So, I built auto-agent.
auto-agent Putte OWHh:o: O
Hmasie PiBraneh Ootg? QGo o te Add flh KXCod About
Clonsogroziano lost: replce Chaude Integr ation wth gonerik provider system for G. Ga24sra Gdra sg0 O 32 Comnlts provided No doscriotion, website, @r topics
claude/sklls docB sebewypiqnd tamplatos Teat: sdd accuracy chart sk docunontation and exainple feat: rep'ace Clsude intogtation with & gorerc provider Sy feat eod sccurasy chart sk' docu mentation @rd exampe o Teat: add genorate-changelog scipt ard reiated documen Teat: sod Kro Cll provide support snd benchmurk frame,. feai+ sdd Kro CtI provide aupport and borchmsrk freme. Gdys ago fhst wook Iast woek last week tast week fast woek MIT Icenso 0wiching Releases B Readimo ④ ativay 18 stars 4 torks
■ tosts fest: sdd Kiro CtI grovider support ard borchmark frsme. tast wock No reeases gublhed Create a rew relo)
① gitignoo CLCENSE README.md Test: sdd Kiro Ctl provider support and benchmsrk trame. Add job creation functionahty and updato tompiatet Add MrT Lictnse snd onhanco REAOMe with dtmo soent I 450m i2e) sst woek Iast week Packagos No pacisges Pubrhyour fru p.
https://github.com/alfonsograziano/auto-agent

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.98 - Text source: advanced OCR
rapidocr-live/bright-screen/contrast. - OCR decision: ready — Chart and embedded table are dense enough that OCR is preferable.
Slide text:
And it actually works!
Accuracy Improvements In The Agent Performances
100% 83.3% +10% on a Production agent
80% 71.7%
Accuracy 60% 58.3% 60.0% 61.7% 58.3% 65.0% 70.0%
40% 31.7% RterationSummary ssonodas Decision Accuracy oogeng
1 001-ad32d0 CONTINUE
20% Baseline 18.3% #1 乙# #3 #4 #5 #6 2 3 4 002-85c091 003-c793e4 004-db6c98 CONTINUE CONMNUE CONTINUE
0% Iteration 8 005-487391 006-009000 CONTINUE CONTINUE

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.99 - Text source: agent_vision.
Slide text:
The core idea
Claude builds the agent
The agent gives feedback

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.98 - Text source: agent_vision.
Slide text:
The human in the loop
Claude builds the agent
The agent gives feedback

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.99 - Text source: advanced OCR
rapidocr-live/bright-screen/opencv-adaptive. - OCR decision: ready — Dense code/editor screenshot with small text; OCR should read it better than manual vision transcription.
Slide text:
jobs >agent-2 > JoB.md >* ## Priority Hints
Step 1 create a job_ # Objective Csuea caiiaq, leum 1noqe 2itoads og Lqo! uomiezturido sius 4o teo6 uTew aun s1 1eum -f dImprove accuracy on the math golden dataset Troa 20s to 80s0 Exaoples! Reduce average latency below 5ooms chite caintaining current accuracy "Add support for x questions (currently os pass fate on that category)b
Iv
We want to inprove the accuracy as much as we can
#+ Target Repository
Absoluteor relativo path to the fepo the coding agent will aodify. Atso specifywhich branch to start fro e this is the baselinc. ->
*Path+:/Users/alfonsograziano/Desktop/exp/auto-agent-demo tBranch*B oaster
Metries
Which metric should the Systen optimize, and chat guardrails apply to The prinary mctric determines whether a hypothesis is accepted or Gej Secondary metrie regresses beyond its threshoid, even it the primary Secondary constraints prevent regresslons o hypothesis is rejected
prieary aetricB accuracy(maximize) Secondary constraints: latency avg ns9 max 20t regression cost usd: max 5os regression

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.98 - Text source: advanced OCR
rapidocr-live/bright-screen/contrast. - OCR decision: ready — Flowchart with many small labels and repeated text; OCR is suitable for the diagram body.
Slide text:
Step 2 - run the loop-
regression
Create an hypotes's Change the Run the
Snprovement 0%,inprovenent
Run evals bastine data Create Create an hypotesis Change the agent Run the evals Create an hypotes's Change the agent Run the evals
regression 12%inprovement 2%,improvemeat
Create an hypotes's Change the agent Run the evnls Create an hypotesis Change the agent Run the evaks Create an hypotesis Change the agent Run the evals
0%inprove-ent
Create an hypotesis
we define how many iterations we want to run
22

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.99 - Text source: advanced OCR
rapidocr-live/bright-screen/opencv-adaptive. - OCR decision: ready — Dense dark report screenshot with small text and many numeric details; OCR is the right extraction path.
Slide text:
The baseline report.
Case: fetch-country-density
Cnput: Using the REST Countries APt. fetch Japan's popudaton and srea, then calculate the populaton densny. C Actual Output: Agent rofused, saying i can't accets extomal APis. CExpected 0utput: 326.e1
Case: weighted-harmonic-mean
Sozt6s'6z andno p1sedx3 D C Actual 0utput: 29.591676 C Input: Cslculale the weighted harmonic mean of Values snd Wdights. Round to eractty 6 decimsl plsces.
Summary
What works (11/60 psss+d, 18.3%):
Slmple well-known computations: trailng zoros of 200f (1), binomisl coefficlemt C(100,50) (1), bond price (1), doprociaton 0D8 (1), Flbonacc F(300) (1) CSome Ap1 knowtedge-based answers: User count (1), user posts count (1), couniry borders (1) CCode execution tasks: FizzBuzz sum (t), stalrcsso climbing (1), Codatz mrr [1)
Whut talls (40/60, 81.7%) grouped by pattern:
. Large exact hnteger computations (8 cases): powor-23.53, powor-17-73, power-37-41, powcr-7-85, power-29.47, powor-31-43, powrer-13-67. Dr lmor ial-20. The sgent refuses to compute Largo exponents, asking fo a cakculator tocil k doesn't htre.
2. Digt-sum / modulsc rithmetic (3 cates): dlg't-sum-88-88, dlgit-sum-99-99, last12-diglts-7-7777. Agent ether refuses or gves wrong answera.
③. Finsncisl math preclsion erors (12 cases): (v-arulty, omortizallon- 30y, npv-cashiows, py-anruay, compound-dsy, continuous-compound) Compound-monthty, &mortization- 15yr, compound-qusrterty, geome ric-moan, Ir-cashflows, moafled -dur atlon, The sgent attempts the caiculation but
gett slightly wrong onswers due to floaling-point or lormula errors.
2

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.98 - Text source: advanced OCR
rapidocr-live/bright-screen/opencv-adaptive. - OCR decision: ready — Flowchart slide with multiple small labels and note text; OCR will be more reliable for the box text.
Slide text:
Step 2.2 - running one iteration_
Create a new branch
Produces a
Create an hypotesis Change the. agent Run the evals REPORT.md MEMORY.md Updates
improved Metrics Yes Continue from this branch.
*The generated hypothesis is based on MEMORY.md, other report files with failures, codebase investigation etc Rollback to prev. branch:

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.97 - Text source: advanced OCR
rapidocr-live/bright-screen/contrast. - OCR decision: ready — Repeated flowchart structure with tiny text labels; OCR is suitable for reading the boxes and connectors.
Slide text:
Step 2.x - running every iteration
predies
create
ALAORY- ypkees
Birnk
Pv -gire AEAOKYOL updeies
Autrns Coapeve irs
REPORT. rredeas
ADMORYe updrtes
Petres Crm
Resras

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.99 - Text source: advanced OCR
rapidocr-live/bright-screen/contrast. - OCR decision: ready — Dense table plus multiple bullet points; OCR will capture the iteration summary more reliably than manual triage.
Slide text:
A real test on an already optimized agent_
Baseline accuracy: 76.7% Found edge cases
Iteration Summary: # Hypothesis Decision Accuracy Improved the system prompt
1 001-bc0b51 CONTINUE 80.7% Improved tools description
2 002-f7f825 CONTINUE 82.7%
E 4 003-5018d3 004-dfe1sd CONTINUE ROLLBACK 81.9% 82.1% Fixed tools logic
5 005-a4b57e CONTINUE 82.6%
6 7 006-fb02c4 007-81e640 CONTINUE ROLLBACK 84.4% 81.6%
8 008-1ff45c CONTINUE 82.9%
9 009-1889d9 CONTINUE 86.4%
10 010-691708 ROLLBACK 85.1%
11 011-0de49e ROLLBACK 84.5%
12 012-dcb780 ROLLBACK 84.7%

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.99 - Text source: agent_vision.
Slide text:
Fixing bad performances on live data_ | Bad performances on live data

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.99 - Text source: agent_vision.
Slide text:
How do we fix that?; The user uses the service; We collect the trace; User gives a feedback; Subject Matter Experts annotate the trace; Once we collect multiple traces...; Subject Matter Expert validation; Traces with negative feedback analyzed; Clustering of failure modes; Fix proposal generated

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.98 - Text source: agent_vision.
Slide text:
We collect tracing informations_

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.98 - Text source: agent_vision.
Slide text:
Step 2 (a): The user gives a feedback_

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.99 - Text source: advanced OCR
rapidocr-live/bright-screen/contrast. - OCR decision: ready — Code/terminal-like slide with small dense text and a command line snippet; OCR is more suitable than manual transcription in this pass.
Slide text:
Step 3: Collect all the traces with feedback locally-
npm run fetch-traces-with-feedback.ts--from2026-03-30--to2026-04-03--limit200
4 5 6 2 3 1 "environment":"production", "fetchedAt":"2026-04-03T12:24:50.619z", "totalscanned":183, "totalwithFeedback":114, "traces":[
7 {
10 8 6 "userId":"j.smithcquantum-solutions.io", "timestamp":"2026-04-05T08:14:12.115z", "id":"f47ac10b928e4d27b11a92837d8e921d",
12 13 14 15 16 17 18 19 20 21 11 "annotations":[] "latency":4.12, "comments": [], "answen":"To find the specific torque requirements for the Cx-Seo engine components,please re "userFeedback":{ "mode1":"us.anthropic.claude-sonnet-4-5-2025e929-v1:e", "question":"What is the recomended torque specification for the titanium alloy bolts on the( "comment":“Solid guidance,but could be more precise.7/1o.\ni) You should have linked dir "score":1,
35
Hidden Non-Slide Evidence
- `slide-001.jpg` —
title_cardconfidence0.98; speaker intro card
Classification audit: raw/sources/slide-ai-classification/slides/aHhB3sjGjkI/audit.json
Slide-Derived Subjects To Review
Subject extraction uses video title, related session titles/descriptions, transcript context, and OCR text when available. OCR is best-effort and should be reviewed against the embedded slide images.