Markdown source

Slides: Production Evals For Agentic AI Systems - Nishant Gupta, Meta Superintelligence Labs

Source Video

Production Evals For Agentic AI Systems - Nishant Gupta, Meta Superintelligence Labs

Relationship To World's Fair 2026

These slides are extracted from a public AI Engineer YouTube video connected to World's Fair 2026. Speaker-matched clips are supporting context unless later confirmed as exact session recordings; official livestream recordings are day-level/event-level source material.

Related Scheduled Sessions

Extracted Slides

slide-001.jpg

Slide text:

Production Evals

for Agentic Systems

Measuring reliability beyond accuracy. Building evaluation systems for autonomous AI workflows.

Nishant Gupta

Tech Lead @ Meta

slide-002.jpg

Slide text:

our evaluation methods AI systems evolved faster than

The Illusion The Reality

100% Modes Invisible Failure

75% Behavior Degraded Production

Benchmark Accuracy 90% 25% 50% 0% T-0 T+10ms T+50ms Reliability Gaps Unpredictable User MAAAA T+10GmS

slide-003.jpg

Slide text:

The Paradigm Shift: Output vs. Behavior

Traditional LLM Evaluation Agent Evaluation

Goal Output Accuracy Workflow Behavior

Environment Static Datasets Dynamic Contexts

Execution Single-path Processing Multi-path & Tool Dependent

Failure Mode Hallucination Cascading Workflow Failure

slide-004.jpg

Slide text:

Think like an SRE: Accuracy gives way to Reliab

slide-005.jpg

Slide text:

The Evaluation Signal Hierarchy

Production Telemetry

Scenario Evals

Benchmarks

slide-006.jpg

Slide text:

Offline Evals: Scenario-Driven Simulation

Agent Sandbox Discrete Outputs

Test Runner Sinulated. Tools Execution: Steps Update State Hetrics Conpletion Rate 98.5x

.Tool Correctness 108%

Plan Quality High'

Simulated Cost $0.05

Scenario-driven, not prompt-driven.

slide-007.jpg

Slide text:

Online Evals: The Production Stream Production is your largest evaluation dataset. Every interaction is signal.

slide-008.jpg

Slide text:

Human-in-the-Loop Calibration Humans are evaluators, not merely fallback systems.

slide-009.jpg

Slide text:

Observability is the Prerequisite

The Trace Waterfall Live Metrics Dashboard.

User Prompt 3sm-70ms) Latency 345 ms

Planner lteration (7a-3sms)

Retries 7

Vector DB Lookup. 8ms 2.5 % AW

Parallel APl Tool Calls.- AP1B:38ms APIA-45m) Step Costs $0.014

APIC: 62mu Memory Usage

480 MB

"You cannot evaluate what you cannot observe.

slide-010.jpg

Slide text:

The Continuous Evaluation Loop Evaluation is an always-running service, not a testing phase.

slide-011.jpg

Slide text:

The Agentic Control Plane Reference Architecture

slide-012.jpg

Slide text:

Architectural Imperatives

1. Offline benchmarks are necessary but insufficient.

2. Agentic systems must be evaluated as full workflows.

3. Production telemetry is the ultimate evaluation signal.

4. Reliability always supersedes raw model accuracy.

5. Evals are no longer tests; they are core infrastructure.

You can't improve what you don't continuously evaluate.

Classification audit: raw/sources/slide-ai-classification/slides/vljxQZfJ9wY/audit.json

Slide-Derived Subjects To Review

Subject extraction uses video title, related session titles/descriptions, transcript context, and OCR text when available. OCR is best-effort and should be reviewed against the embedded slide images.