---
title: "Your Agent Evolved. Your Evals Didn't."
category: "talks"
date: "2026-06-29"
time: "11:10am-11:30am"
track: "AI Architects: Show my Workflow"
room: "Leadership 2"
speakers: ["Ameya Bhatawdekar"]
sourceLabels: ["Official conference schedule", "Public YouTube metadata"]
scheduleTrack: "AI Architects: Show my Workflow"
scheduleRoom: "Leadership 2"
scheduleLabels: ["AI Architects: Show my Workflow", "Leadership 2", "session", "confirmed"]
---
# Your Agent Evolved. Your Evals Didn't.

## Conference Context
- Date/time: 2026-06-29 · 11:10am-11:30am
- Track/room: AI Architects: Show my Workflow · Leadership 2
- Speaker(s): Ameya Bhatawdekar
- Session type/status: session · confirmed

- Track: AI Architects: Show my Workflow
- Room: Leadership 2
- Session type: session
- Status: confirmed

## Session Description
Knowing which generation your agent is in, which failure modes your current evals are blind to, and what to build next is the difference between shipping with confidence and flying blind. Agent architectures have evolved through six generations; prompt, chain, ReAct loop, workflow graph, modern agent loop, AI harness. And each one quietly breaks the eval strategy of the generation before it. A prompt-quality rubric won't catch a bad tool call; a trace scorer won't catch memory poisoning. Using a single SRE incident response agent threaded through every generation, this talk shows exactly where each architecture outgrows its evals and what you need to close the gap.

## Media Evidence
No related AI Engineer channel video found yet.

## Evidence Graph
This evidence graph is generated from currently linked source material: official schedule text, related video pages, cached transcripts, visible slide text, dense/reconstructed slide pages, and AI slide-classification audits.

### Media Signals
No linked video, transcript, or slide source has been attached yet.

### Agent Reading Notes
Use these signals to refine the synopsis, topic links, people/company context, and method notes. If a source is a related external video rather than an exact official recording, keep it framed as supporting evidence.

## Transcript Status
No official session recording transcript was found by exact title match on the AI Engineer YouTube channel during this run.

## People
- [[ameya-bhatawdekar]]

## Notes
- Pending transcript synthesis when an official recording or confirmed matching video is available.

## Synthesis
### Synthesized Breakdown
# Your Agent Evolved. Your Evals Didn't. ## Conference Context - Date/time: 2026-06-29 · 11:10am-11:30am - Track/room: AI Architects: Show my Workflow · Leadership 2 - Speaker(s): Ameya Bhatawdekar - Session type/status: session · confirmed - Track: AI Architects: Show my Workflow - Room: Leadership 2 - Session type: session - Status: confirmed ## Session Description Knowing which generation your agent is in, which failure modes your current evals are blind to, and what to build next is the difference between shipping with confidence and flying blind. Agent architectures have evolved through six generations; prompt, chain, ReAct loop, workflow graph, modern agent loop, AI harness.

### Speaker And Company Context
- [[ameya-bhatawdekar|Ameya Bhatawdekar]] — VP, Field CTO at [[braintrust|Braintrust]].

### Topics Covered
- [[agent-security]]

### Derived Links And Source Material

### Novel Concepts / Clever Methods
- No highlighted novel concept has been detected yet.

### Evidence Boundary
This synthesis is based on the official schedule and linked source pages. It should be revisited when exact session recordings or transcript-backed secondary sources are available.
