Markdown source

Slides: Agents Building Agents - Alfonso Graziano, Nearform

Source Video

Agents Building Agents - Alfonso Graziano, Nearform

Relationship To World's Fair 2026

These slides are extracted from a public AI Engineer YouTube video connected to World's Fair 2026. Speaker-matched clips are supporting context unless later confirmed as exact session recordings; official livestream recordings are day-level/event-level source material.

Related Scheduled Sessions

Extracted Slides

slide-002.jpg

Slide text:

Everyone wants AI Agents

slide-003.jpg

Slide text:

How do we do that?

slide-004.jpg

Slide text:

The Problems with Building AI Agents

...and how to solve them partially,

with other agents

slide-005.jpg

Slide text:

Al Agents: a refresher

Short-term memory Long-term memory

Calendar() Memory

Calculator() Reflection

Codelnterpreter() Tools Agent Planning Self-critics

Search() Chai

...more Action eogans

slide-006.jpg

Slide text:

Two classes of problems

slide-007.jpg

Slide text:

Bad performances on evals - the golden dataset_

8 c

input output

42 What Is the output for the program that checks itself? Write a program that determines whether any arbitrary program wll halt or run forever. IMPOSSIBLE

43 How old is Elon Musk? PERSONAL _INFO REJECTED

44 Calculate Jeff Bezos's net worth divided by the US populatlon PERSONAL _INFO_REJECTED

45 I have $50.000 to invest What's the optimal split between slocks and bonds to maximize retums? FINANCIAL _ADVICE_REJECTED

46 Convert 250 USD to EUR at today's exchange rate REQUIRES _LVE _rATE

47 latitude=48.8566, longitude=2.3522, start_date=2024-01-15, end_date=2024-01-15, What was the maximum temperature (in *C) in Paris on January 15, 2024? Use the daily-temperature_2m_max. Retum just tha number. Open-Meteo historical weather APl at https://archive-api.open-meleo.com/v1/archlve with 4.6

48 What was the minimum temperature (in *C) in Tokyo on July 20, 2024? Use the dallyztemperature_2m_min. Retum just the number. latitude=35.6762, longitude=139.6503, start_date=2024-07-20, end_date=2024-07-20. Open-Meteo historical weather APl at https://archive-apt,open-meleo.com/v1/archive with 25.7

49 number of users. Fetch the list of users from hitps:/jsonplaceholder.typicode.com/users and retum the tota! 10

50 Felch all todos from https:/jsonplaceholder.typicode.comtodos. What percentage of them 8re completed? Round to 1 decimal place. 45.0 Bad

51 https://jsonplaceholder.typicode.com/posts?useridz7 and count the results. How many posts does user with ID 7 have? Fetch from 10 perfd

on

slide-008.jpg

Slide text:

Some failure modes

No tools

Wrong system prompt

No context retrieval

slide-009.jpg

Slide text:

How autoresearch works.

Autorescarth Progress: 83 Experiments, 15 Kept Improvcments.

L000

eyy, b+t.

Valldaton BpB (lower ts better) 4 0.10

4.985

0.0

26 Experimert s 40

slide-010.jpg

Slide text:

So, I built auto-agent.

auto-agent Putte OWHh:o: O

Hmasie PiBraneh Ootg? QGo o te Add flh KXCod About

Clonsogroziano lost: replce Chaude Integr ation wth gonerik provider system for G. Ga24sra Gdra sg0 O 32 Comnlts provided No doscriotion, website, @r topics

claude/sklls docB sebewypiqnd tamplatos Teat: sdd accuracy chart sk docunontation and exainple feat: rep'ace Clsude intogtation with & gorerc provider Sy feat eod sccurasy chart sk' docu mentation @rd exampe o Teat: add genorate-changelog scipt ard reiated documen Teat: sod Kro Cll provide support snd benchmurk frame,. feai+ sdd Kro CtI provide aupport and borchmsrk freme. Gdys ago fhst wook Iast woek last week tast week fast woek MIT Icenso 0wiching Releases B Readimo ④ ativay 18 stars 4 torks

■ tosts fest: sdd Kiro CtI grovider support ard borchmark frsme. tast wock No reeases gublhed Create a rew relo)

① gitignoo CLCENSE README.md Test: sdd Kiro Ctl provider support and benchmsrk trame. Add job creation functionahty and updato tompiatet Add MrT Lictnse snd onhanco REAOMe with dtmo soent I 450m i2e) sst woek Iast week Packagos No pacisges Pubrhyour fru p.

https://github.com/alfonsograziano/auto-agent

slide-011.jpg

Slide text:

And it actually works!

Accuracy Improvements In The Agent Performances

100% 83.3% +10% on a Production agent

80% 71.7%

Accuracy 60% 58.3% 60.0% 61.7% 58.3% 65.0% 70.0%

40% 31.7% RterationSummary ssonodas Decision Accuracy oogeng

1 001-ad32d0 CONTINUE

20% Baseline 18.3% #1 乙# #3 #4 #5 #6 2 3 4 002-85c091 003-c793e4 004-db6c98 CONTINUE CONMNUE CONTINUE

0% Iteration 8 005-487391 006-009000 CONTINUE CONTINUE

slide-012.jpg

Slide text:

The core idea

Claude builds the agent

The agent gives feedback

slide-013.jpg

Slide text:

The human in the loop

Claude builds the agent

The agent gives feedback

slide-014.jpg

Slide text:

jobs >agent-2 > JoB.md >* ## Priority Hints

Step 1 create a job_ # Objective Csuea caiiaq, leum 1noqe 2itoads og Lqo! uomiezturido sius 4o teo6 uTew aun s1 1eum -f dImprove accuracy on the math golden dataset Troa 20s to 80s0 Exaoples! Reduce average latency below 5ooms chite caintaining current accuracy "Add support for x questions (currently os pass fate on that category)b

Iv

We want to inprove the accuracy as much as we can

#+ Target Repository

Absoluteor relativo path to the fepo the coding agent will aodify. Atso specifywhich branch to start fro e this is the baselinc. ->

*Path+:/Users/alfonsograziano/Desktop/exp/auto-agent-demo tBranch*B oaster

Metries

Which metric should the Systen optimize, and chat guardrails apply to The prinary mctric determines whether a hypothesis is accepted or Gej Secondary metrie regresses beyond its threshoid, even it the primary Secondary constraints prevent regresslons o hypothesis is rejected

prieary aetricB accuracy(maximize) Secondary constraints: latency avg ns9 max 20t regression cost usd: max 5os regression

slide-015.jpg

Slide text:

Step 2 - run the loop-

regression

Create an hypotes's Change the Run the

Snprovement 0%,inprovenent

Run evals bastine data Create Create an hypotesis Change the agent Run the evals Create an hypotes's Change the agent Run the evals

regression 12%inprovement 2%,improvemeat

Create an hypotes's Change the agent Run the evnls Create an hypotesis Change the agent Run the evaks Create an hypotesis Change the agent Run the evals

0%inprove-ent

Create an hypotesis

we define how many iterations we want to run

22

slide-016.jpg

Slide text:

The baseline report.

Case: fetch-country-density

Cnput: Using the REST Countries APt. fetch Japan's popudaton and srea, then calculate the populaton densny. C Actual Output: Agent rofused, saying i can't accets extomal APis. CExpected 0utput: 326.e1

Case: weighted-harmonic-mean

Sozt6s'6z andno p1sedx3 D C Actual 0utput: 29.591676 C Input: Cslculale the weighted harmonic mean of Values snd Wdights. Round to eractty 6 decimsl plsces.

Summary

What works (11/60 psss+d, 18.3%):

Slmple well-known computations: trailng zoros of 200f (1), binomisl coefficlemt C(100,50) (1), bond price (1), doprociaton 0D8 (1), Flbonacc F(300) (1) CSome Ap1 knowtedge-based answers: User count (1), user posts count (1), couniry borders (1) CCode execution tasks: FizzBuzz sum (t), stalrcsso climbing (1), Codatz mrr [1)

Whut talls (40/60, 81.7%) grouped by pattern:

. Large exact hnteger computations (8 cases): powor-23.53, powor-17-73, power-37-41, powcr-7-85, power-29.47, powor-31-43, powrer-13-67. Dr lmor ial-20. The sgent refuses to compute Largo exponents, asking fo a cakculator tocil k doesn't htre.

2. Digt-sum / modulsc rithmetic (3 cates): dlg't-sum-88-88, dlgit-sum-99-99, last12-diglts-7-7777. Agent ether refuses or gves wrong answera.

③. Finsncisl math preclsion erors (12 cases): (v-arulty, omortizallon- 30y, npv-cashiows, py-anruay, compound-dsy, continuous-compound) Compound-monthty, &mortization- 15yr, compound-qusrterty, geome ric-moan, Ir-cashflows, moafled -dur atlon, The sgent attempts the caiculation but

gett slightly wrong onswers due to floaling-point or lormula errors.

2

slide-017.jpg

Slide text:

Step 2.2 - running one iteration_

Create a new branch

Produces a

Create an hypotesis Change the. agent Run the evals REPORT.md MEMORY.md Updates

improved Metrics Yes Continue from this branch.

*The generated hypothesis is based on MEMORY.md, other report files with failures, codebase investigation etc Rollback to prev. branch:

slide-018.jpg

Slide text:

Step 2.x - running every iteration

predies

create

ALAORY- ypkees

Birnk

Pv -gire AEAOKYOL updeies

Autrns Coapeve irs

REPORT. rredeas

ADMORYe updrtes

Petres Crm

Resras

slide-019.jpg

Slide text:

A real test on an already optimized agent_

Baseline accuracy: 76.7% Found edge cases

Iteration Summary: # Hypothesis Decision Accuracy Improved the system prompt

1 001-bc0b51 CONTINUE 80.7% Improved tools description

2 002-f7f825 CONTINUE 82.7%

E 4 003-5018d3 004-dfe1sd CONTINUE ROLLBACK 81.9% 82.1% Fixed tools logic

5 005-a4b57e CONTINUE 82.6%

6 7 006-fb02c4 007-81e640 CONTINUE ROLLBACK 84.4% 81.6%

8 008-1ff45c CONTINUE 82.9%

9 009-1889d9 CONTINUE 86.4%

10 010-691708 ROLLBACK 85.1%

11 011-0de49e ROLLBACK 84.5%

12 012-dcb780 ROLLBACK 84.7%

slide-020.jpg

Slide text:

Fixing bad performances on live data_ | Bad performances on live data

slide-021.jpg

Slide text:

How do we fix that?; The user uses the service; We collect the trace; User gives a feedback; Subject Matter Experts annotate the trace; Once we collect multiple traces...; Subject Matter Expert validation; Traces with negative feedback analyzed; Clustering of failure modes; Fix proposal generated

slide-022.jpg

Slide text:

We collect tracing informations_

slide-023.jpg

Slide text:

Step 2 (a): The user gives a feedback_

slide-024.jpg

Slide text:

Step 3: Collect all the traces with feedback locally-

npm run fetch-traces-with-feedback.ts--from2026-03-30--to2026-04-03--limit200

4 5 6 2 3 1 "environment":"production", "fetchedAt":"2026-04-03T12:24:50.619z", "totalscanned":183, "totalwithFeedback":114, "traces":[

7 {

10 8 6 "userId":"j.smithcquantum-solutions.io", "timestamp":"2026-04-05T08:14:12.115z", "id":"f47ac10b928e4d27b11a92837d8e921d",

12 13 14 15 16 17 18 19 20 21 11 "annotations":[] "latency":4.12, "comments": [], "answen":"To find the specific torque requirements for the Cx-Seo engine components,please re "userFeedback":{ "mode1":"us.anthropic.claude-sonnet-4-5-2025e929-v1:e", "question":"What is the recomended torque specification for the titanium alloy bolts on the( "comment":“Solid guidance,but could be more precise.7/1o.\ni) You should have linked dir "score":1,

35

Hidden Non-Slide Evidence

Classification audit: raw/sources/slide-ai-classification/slides/aHhB3sjGjkI/audit.json

Slide-Derived Subjects To Review

Subject extraction uses video title, related session titles/descriptions, transcript context, and OCR text when available. OCR is best-effort and should be reviewed against the embedded slide images.