---
title: "Building self-learning loops for your agent"
category: "talks"
date: "2026-06-29"
time: "11:05am-12:05pm"
track: "Posttraining & Midtraining"
room: "Track 1"
speakers: ["Fuad Ali"]
sourceLabels: ["Official conference schedule", "Public YouTube metadata"]
scheduleTrack: "Posttraining & Midtraining"
scheduleRoom: "Track 1"
scheduleLabels: ["Posttraining & Midtraining", "Track 1", "sponsor", "confirmed"]
---
# Building self-learning loops for your agent

## Conference Context
- Date/time: 2026-06-29 · 11:05am-12:05pm
- Track/room: Posttraining & Midtraining · Track 1
- Speaker(s): Fuad Ali
- Session type/status: sponsor · confirmed

- Track: Posttraining & Midtraining
- Room: Track 1
- Session type: sponsor
- Status: confirmed

## Session Description
Building an AI demo is easy. Knowing whether it actually works — and keeping it working in production — is the hard part. Most teams ship agents on vibes: they try a few prompts, the output looks good, and they push to production with no real way to measure quality or catch regressions. This hands-on workshop walks through the full lifecycle of shipping a real AI agent, using a working financial-analyst agent built on the Claude Agent SDK as the running example. You'll instrument it with tracing, do structured error analysis on its actual outputs, and build a layered evaluation suite — from cheap deterministic code checks to LLM-as-a-judge evaluators with custom rubrics. We'll cover the parts most tutorials skip: why agents fail in ways single LLM calls don't, the eval anti-patterns that quietly mislead you, and how to know whether you can even trust your judge (meta-evaluation). Finally, we'll close the loop: turning eval results into datasets and experiments, running evals online against production traffic, wiring them to monitors and alerts, and feeding failure explanations back to a coding agent to actually fix the underlying problems. You'll leave with a runnable notebook and a repeatable, evaluation-driven workflow you can apply to your own agents the next day.

## Media Evidence
[Build a Prompt Learning Loop - SallyAnn DeLucia & Fuad Ali, Arize](https://www.youtube.com/watch?v=SbcQYbrvAfI) (speaker-match related prior/adjacent AI Engineer video; captions: English auto-captions).

- Source video: `youtube-SbcQYbrvAfI`
- Slide deck: [[youtube-SbcQYbrvAfI-dense-slides|Dense Slides: Build a Prompt Learning Loop - SallyAnn DeLucia & Fuad Ali, Arize]] — 32 visible slide image(s); 32 HTML recreation(s).
![[assets/dense-slides/SbcQYbrvAfI/slide-001.jpg]]
![[assets/dense-slides/SbcQYbrvAfI/slide-002.jpg]]
![[assets/dense-slides/SbcQYbrvAfI/slide-003.jpg]]
- Additional slide evidence: [[youtube-SbcQYbrvAfI-slides|Slides: Build a Prompt Learning Loop - SallyAnn DeLucia & Fuad Ali, Arize]], [[youtube-SbcQYbrvAfI-reconstructed-slides|Reconstructed Slides: Build a Prompt Learning Loop - SallyAnn DeLucia & Fuad Ali, Arize]]
- Slide-derived themes for `youtube-SbcQYbrvAfI`: planning, missing, data, domain, breaking, system, instructions, learned.

## Evidence Graph
This evidence graph is generated from currently linked source material: official schedule text, related video pages, cached transcripts, visible slide text, dense/reconstructed slide pages, and AI slide-classification audits.

### Media Signals
- `youtube-SbcQYbrvAfI` — 7 slide-derived text signals
- Slide-derived themes for `youtube-SbcQYbrvAfI`: planning, missing, data, domain, breaking, system, instructions, learned.
- Evidence links for `youtube-SbcQYbrvAfI`: [[youtube-SbcQYbrvAfI]], [[youtube-SbcQYbrvAfI-slides]], [[youtube-SbcQYbrvAfI-dense-slides]], [[youtube-SbcQYbrvAfI-reconstructed-slides]]

### Agent Reading Notes
Use these signals to refine the synopsis, topic links, people/company context, and method notes. If a source is a related external video rather than an exact official recording, keep it framed as supporting evidence.

## Transcript Status
Related video transcript availability: English auto-captions. Treat this as supporting context, not a recording of this exact scheduled session unless later confirmed. Not fetched yet.

## People
- [[fuad-ali]]

## Supporting Slides
- [[youtube-SbcQYbrvAfI-slides]] — extracted from the related public AI Engineer video.

## Slide Evidence
- Slide-only cropped deck: [[youtube-SbcQYbrvAfI-dense-slides]] (32 viable slide images).
- Related slide/OCR pages:
- [[youtube-SbcQYbrvAfI-dense-slides]]
- [[youtube-SbcQYbrvAfI-reconstructed-slides]]
- [[youtube-SbcQYbrvAfI-slides]]
- Slide-derived terms: `prompt`, `learning`, `system`, `rules`, `coding`, `gepa`, `score`, `changes`, `cost`, `claude`, `test`, `mone`, `planning`, `missing`, `aarize`, `evals`, `function`, `exam`

## Synthesis
### Synthesized Breakdown
# Building self-learning loops for your agent ## Conference Context - Date/time: 2026-06-29 · 11:05am-12:05pm - Track/room: Posttraining & Midtraining · Track 1 - Speaker(s): Fuad Ali - Session type/status: sponsor · confirmed - Track: Posttraining & Midtraining - Room: Track 1 - Session type: sponsor - Status: confirmed ## Session Description Building an AI demo is easy. Knowing whether it actually works — and keeping it working in production — is the hard part. Most teams ship agents on vibes: they try a few prompts, the output looks good, and they push to production with no real way to measure quality or catch regressions. This hands-on workshop walks through the full lifecycle of shipping a real AI agent, using a working financial-analyst agent built on the Claude Agent SDK as the running example.

### Speaker And Company Context
- [[fuad-ali|Fuad Ali]] — Senior Product Manager at [[arize-ai|Arize AI]].

### Topics Covered
- [[agent-security]]
- [[coding-agents]]

### Derived Links And Source Material
- [[youtube-SbcQYbrvAfI]] — related YouTube source page.
- [[youtube-SbcQYbrvAfI-slides]] — slide evidence.
- [[youtube-SbcQYbrvAfI-reconstructed-slides]] — slide evidence.
- [[youtube-SbcQYbrvAfI-dense-slides]] — slide evidence.

### Novel Concepts / Clever Methods
- No highlighted novel concept has been detected yet.

### Evidence Boundary
This synthesis is based on the official schedule and linked source pages. It should be revisited when exact session recordings or transcript-backed secondary sources are available.
