---
title: "From Vibes to Production: Evaluating and Shipping AI Agents That Work 201"
category: "talks"
date: "2026-06-29"
time: "2:20pm-4:20pm"
track: "Track 1"
room: "Track 1"
speakers: ["Laurie Voss"]
sourceLabels: ["Official conference schedule", "Public YouTube metadata"]
scheduleTrack: "Track 1"
scheduleRoom: "Track 1"
scheduleLabels: ["Track 1", "Track 1", "sponsor", "confirmed"]
---
# From Vibes to Production: Evaluating and Shipping AI Agents That Work 201

## Conference Context
- Date/time: 2026-06-29 · 2:20pm-4:20pm
- Track/room: Track 1 · Track 1
- Speaker(s): Laurie Voss
- Session type/status: sponsor · confirmed

- Track: Track 1
- Room: Track 1
- Session type: sponsor
- Status: confirmed

## Session Description
Building an AI demo is easy. Knowing whether it actually works — and keeping it working in production — is the hard part. Most teams ship agents on vibes: they try a few prompts, the output looks good, and they push to production with no real way to measure quality or catch regressions. This hands-on workshop walks through the full lifecycle of shipping a real AI agent, using a working financial-analyst agent built on the Claude Agent SDK as the running example. You'll instrument it with tracing, do structured error analysis on its actual outputs, and build a layered evaluation suite — from cheap deterministic code checks to LLM-as-a-judge evaluators with custom rubrics. We'll cover the parts most tutorials skip: why agents fail in ways single LLM calls don't, the eval anti-patterns that quietly mislead you, and how to know whether you can even trust your judge (meta-evaluation). Finally, we'll close the loop: turning eval results into datasets and experiments, running evals online against production traffic, wiring them to monitors and alerts, and feeding failure explanations back to a coding agent to actually fix the underlying problems. You'll leave with a runnable notebook and a repeatable, evaluation-driven workflow you can apply to your own agents the next day.

## Media Evidence
[Ship Real Agents: Hands-On Evals for Agentic Applications — Laurie Voss, Arize](https://www.youtube.com/watch?v=Xfl50508LZM) (speaker-match related prior/adjacent AI Engineer video; captions: English auto-captions).

- [[youtube-Xfl50508LZM-transcript]] — full cached transcript markdown for the related YouTube source.

- Source video: `youtube-Xfl50508LZM`
- Slide deck: [[youtube-Xfl50508LZM-dense-slides|Dense Slides: Ship Real Agents: Hands-On Evals for Agentic Applications — Laurie Voss, Arize]] — 7 visible slide image(s); 7 HTML recreation(s).
![[assets/dense-slides/Xfl50508LZM/slide-001.jpg]]
![[assets/dense-slides/Xfl50508LZM/slide-002.jpg]]
![[assets/dense-slides/Xfl50508LZM/slide-003.jpg]]
- Additional slide evidence: [[youtube-Xfl50508LZM-slides|Slides: Ship Real Agents: Hands-On Evals for Agentic Applications — Laurie Voss, Arize]], [[youtube-Xfl50508LZM-reconstructed-slides|Reconstructed Slides: Ship Real Agents: Hands-On Evals for Agentic Applications — Laurie Voss, Arize]]
- Slide-derived themes for `youtube-Xfl50508LZM`: phoenix, prompt, settings, general, detect, regressions, change, compare.

## Evidence Graph
This evidence graph is generated from currently linked source material: official schedule text, related video pages, cached transcripts, visible slide text, dense/reconstructed slide pages, and AI slide-classification audits.

### Media Signals
- `youtube-Xfl50508LZM` — 22,591 transcript words; 6 slide-derived text signals
- Transcript signals for `youtube-Xfl50508LZM`: evals, eval, data, should, judge, output, whether, phoenix.
- Slide-derived themes for `youtube-Xfl50508LZM`: phoenix, prompt, settings, general, detect, regressions, change, compare.
- Evidence links for `youtube-Xfl50508LZM`: [[youtube-Xfl50508LZM]], [[youtube-Xfl50508LZM-transcript]], [[youtube-Xfl50508LZM-slides]], [[youtube-Xfl50508LZM-dense-slides]], [[youtube-Xfl50508LZM-reconstructed-slides]]

### Agent Reading Notes
Use these signals to refine the synopsis, topic links, people/company context, and method notes. If a source is a related external video rather than an exact official recording, keep it framed as supporting evidence.

## Transcript Status
Related video transcript availability: English auto-captions. Treat this as supporting context, not a recording of this exact scheduled session unless later confirmed. Cached at `raw/sources/youtube-transcripts/Xfl50508LZM.txt` (22,591 words).

## People
- [[laurie-voss]]

## Supporting Slides
- [[youtube-Xfl50508LZM-slides]] — extracted from the related public AI Engineer video.

## Slide Evidence
- Slide-only cropped deck: [[youtube-Xfl50508LZM-dense-slides]] (7 viable slide images).
- Related slide/OCR pages:
- [[youtube-Xfl50508LZM-dense-slides]]
- [[youtube-Xfl50508LZM-reconstructed-slides]]
- [[youtube-Xfl50508LZM-slides]]
- Slide-derived terms: `phoenix`, `claude`, `tome`, `setting`, `tracing`, `alengineer`, `europe`, `ages`, `notebook`, `cloud`, `comma`, `swiss`, `cheese`, `braintrust`, `workos`, `openal`, `frage`, `gers`

## Synthesis
### Synthesized Breakdown
Hi everybody. Uh my name's Laurie Voss. I am head of developer experience at Arize AI. Uh in a former life, I co-founded npm Inc.

### Speaker And Company Context
- [[laurie-voss|Laurie Voss]] — Head of Developer Relations at [[arize-ai|Arize AI]].

### Topics Covered
- [[agent-security]]
- [[agentic-search]]
- [[agentic-web]]
- [[ai-sandboxes]]
- [[coding-agents]]

### Derived Links And Source Material
- [[youtube-Xfl50508LZM-transcript]] — transcript markdown; source cache `raw/sources/youtube-transcripts/Xfl50508LZM.txt` (22,591 words).
- [[youtube-Xfl50508LZM]] — related YouTube source page.
- [[youtube-Xfl50508LZM-slides]] — slide evidence.
- [[youtube-Xfl50508LZM-reconstructed-slides]] — slide evidence.
- [[youtube-Xfl50508LZM-dense-slides]] — slide evidence.

### Novel Concepts / Clever Methods
- No highlighted novel concept has been detected yet.

### Evidence Boundary
This synthesis uses the official schedule plus cached video transcripts. Official AI Engineer World's Fair San Francisco 2026 livestreams and cut videos are primary event video sources for transcript/slide evidence; external, historical, or speaker-matched videos remain supporting context unless manually verified as exact official event recordings.
