Slides: Production Evals For Agentic AI Systems - Nishant Gupta, Meta Superintelligence Labs
Source Video
Production Evals For Agentic AI Systems - Nishant Gupta, Meta Superintelligence Labs
Relationship To World's Fair 2026
These slides are extracted from a public AI Engineer YouTube video connected to World's Fair 2026. Speaker-matched clips are supporting context unless later confirmed as exact session recordings; official livestream recordings are day-level/event-level source material.
Related Scheduled Sessions
- No individual scheduled session mapping has been assigned yet; treat this as an event livestream deck.
Extracted Slides

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
title_cardconfidence0.99 - Text source: agent_vision.
Slide text:
Production Evals
for Agentic Systems
Measuring reliability beyond accuracy. Building evaluation systems for autonomous AI workflows.
Nishant Gupta
Tech Lead @ Meta

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.98 - Text source: advanced OCR
rapidocr-live/bright-screen/contrast. - OCR decision: ready — Small chart labels, panel headers, and callouts are OCR-suitable.
Slide text:
our evaluation methods AI systems evolved faster than
The Illusion The Reality
100% Modes Invisible Failure
75% Behavior Degraded Production
Benchmark Accuracy 90% 25% 50% 0% T-0 T+10ms T+50ms Reliability Gaps Unpredictable User MAAAA T+10GmS

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.98 - Text source: advanced OCR
rapidocr-live/border-trim/contrast. - OCR decision: ready — Structured table text and cell labels are OCR-suitable.
Slide text:
The Paradigm Shift: Output vs. Behavior
Traditional LLM Evaluation Agent Evaluation
Goal Output Accuracy Workflow Behavior
Environment Static Datasets Dynamic Contexts
Execution Single-path Processing Multi-path & Tool Dependent
Failure Mode Hallucination Cascading Workflow Failure

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.94 - Text source: agent_vision.
Slide text:
Think like an SRE: Accuracy gives way to Reliab

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.98 - Text source: agent_vision.
Slide text:
The Evaluation Signal Hierarchy
Production Telemetry
Scenario Evals
Benchmarks

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.98 - Text source: advanced OCR
rapidocr-live/border-trim/opencv-adaptive. - OCR decision: ready — Small diagram labels and the metrics box are OCR-suitable.
Slide text:
Offline Evals: Scenario-Driven Simulation
Agent Sandbox Discrete Outputs
Test Runner Sinulated. Tools Execution: Steps Update State Hetrics Conpletion Rate 98.5x
.Tool Correctness 108%
Plan Quality High'
Simulated Cost $0.05
Scenario-driven, not prompt-driven.

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.97 - Text source: agent_vision.
Slide text:
Online Evals: The Production Stream Production is your largest evaluation dataset. Every interaction is signal.

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.97 - Text source: agent_vision.
Slide text:
Human-in-the-Loop Calibration Humans are evaluators, not merely fallback systems.

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.98 - Text source: advanced OCR
rapidocr-live/border-trim/opencv-adaptive. - OCR decision: ready — dense chart and dashboard text are better suited for OCR
Slide text:
Observability is the Prerequisite
The Trace Waterfall Live Metrics Dashboard.
User Prompt 3sm-70ms) Latency 345 ms
Planner lteration (7a-3sms)
Retries 7
Vector DB Lookup. 8ms 2.5 % AW
Parallel APl Tool Calls.- AP1B:38ms APIA-45m) Step Costs $0.014
APIC: 62mu Memory Usage
480 MB
"You cannot evaluate what you cannot observe.

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.97 - Text source: agent_vision.
Slide text:
The Continuous Evaluation Loop Evaluation is an always-running service, not a testing phase.

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.96 - Text source: agent_vision.
Slide text:
The Agentic Control Plane Reference Architecture

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.99 - Text source: agent_vision.
Slide text:
Architectural Imperatives
1. Offline benchmarks are necessary but insufficient.
2. Agentic systems must be evaluated as full workflows.
3. Production telemetry is the ultimate evaluation signal.
4. Reliability always supersedes raw model accuracy.
5. Evals are no longer tests; they are core infrastructure.
You can't improve what you don't continuously evaluate.
Classification audit: raw/sources/slide-ai-classification/slides/vljxQZfJ9wY/audit.json
Slide-Derived Subjects To Review
Subject extraction uses video title, related session titles/descriptions, transcript context, and OCR text when available. OCR is best-effort and should be reviewed against the embedded slide images.