---
title: "The Art of Building Verifiers for Computer Use Agents"
category: "talks"
date: "2026-07-01"
time: "11:40am-12:00pm"
track: "Expo Stage 1 NE"
room: "Expo Stage 1 NE"
speakers: ["Miguel González Fernández", "Corby Rosset"]
sourceLabels: ["Official conference schedule", "Public YouTube metadata"]
scheduleTrack: ""
scheduleRoom: "Expo Stage 1 NE"
scheduleLabels: ["Expo Stage 1 NE", "session", "confirmed"]
---
# The Art of Building Verifiers for Computer Use Agents

## Conference Context
- Date/time: 2026-07-01 · 11:40am-12:00pm
- Track/room: track TBD · Expo Stage 1 NE
- Speaker(s): Miguel González Fernández, Corby Rosset
- Session type/status: session · confirmed

- Track: track TBD
- Room: Expo Stage 1 NE
- Session type: session
- Status: confirmed

## Session Description
Every team building browser agents has the same problem: you can't trust your own evals. Browser tasks are too open-ended for deterministic checks, so teams use LLM verifiers as judges, and the judges are wrong constantly. WebVoyager misses 45% of failures. WebJudge misses 22%. Used as RL reward, you're not training a better agent, you're training a more confident liar. This talk walks through the Universal Verifier, open-sourced with Microsoft Research: false positive rate near zero, Cohen's κ matching human-human agreement. Four design principles, one open benchmark, and an honest account of where auto-research worked and where it plateaued.

## Media Evidence
No related AI Engineer channel video found yet.

## Evidence Graph
This evidence graph is generated from currently linked source material: official schedule text, related video pages, cached transcripts, visible slide text, dense/reconstructed slide pages, and AI slide-classification audits.

### Media Signals
No linked video, transcript, or slide source has been attached yet.

### Agent Reading Notes
Use these signals to refine the synopsis, topic links, people/company context, and method notes. If a source is a related external video rather than an exact official recording, keep it framed as supporting evidence.

## Transcript Status
No official session recording transcript was found by exact title match on the AI Engineer YouTube channel during this run.

## People
- [[miguel-gonz-lez-fern-ndez]]
- [[corby-rosset]]

## Notes
- Pending transcript synthesis when an official recording or confirmed matching video is available.

## Synthesis
### Synthesized Breakdown
# The Art of Building Verifiers for Computer Use Agents ## Conference Context - Date/time: 2026-07-01 · 11:40am-12:00pm - Track/room: track TBD · Expo Stage 1 NE - Speaker(s): Miguel González Fernández, Corby Rosset - Session type/status: session · confirmed - Track: track TBD - Room: Expo Stage 1 NE - Session type: session - Status: confirmed ## Session Description Every team building browser agents has the same problem: you can't trust your own evals. Browser tasks are too open-ended for deterministic checks, so teams use LLM verifiers as judges, and the judges are wrong constantly. WebVoyager misses 45% of failures. WebJudge misses 22%.

### Speaker And Company Context
- [[miguel-gonz-lez-fern-ndez|Miguel González Fernández]] — Tech Lead at [[browserbase|Browserbase]].
- [[corby-rosset|Corby Rosset]] — Senior Researcher at [[microsoft-research|Microsoft Research]].

### Topics Covered
- [[agent-security]]
- [[agentic-search]]
- [[agentic-web]]
- [[ai-sandboxes]]

### Derived Links And Source Material

### Novel Concepts / Clever Methods
- No highlighted novel concept has been detected yet.

### Evidence Boundary
This synthesis is based on the official schedule and linked source pages. It should be revisited when exact session recordings or transcript-backed secondary sources are available.
