Markdown source

Slides: SWE-rebench: Lessons from Evaluating Coding Agents — Ibragim Badertdinov, Nebius

Source Video

SWE-rebench: Lessons from Evaluating Coding Agents — Ibragim Badertdinov, Nebius

Relationship To World's Fair 2026

These slides are extracted from a public AI Engineer YouTube video connected to World's Fair 2026. Speaker-matched clips are supporting context unless later confirmed as exact session recordings; official livestream recordings are day-level/event-level source material.

Related Scheduled Sessions

Extracted Slides

slide-001.jpg

Slide text:

SWE-rebench: Lessons from Evaluating Coding Agents on Real Software Engineering Tasks

slide-002.jpg

Slide text:

Why evals matter now?

Models improved. Choosing became harder.

Vibe checks do not scale

SWE performance grows rapidly

Options change every month

* SWE-rebench 2026_02 tasks

slide-003.jpg

Slide text:

Anatomy of a Task: A Task Is More Than Text

Task description - original issue text Regresslon causod by changes for woakref ot fllesystom 1284

★ AIE +

Sandbox environment -

executable Docker image shs256141e015c4.:9

Verifier - tests from the PR. FAIL_TO PASS + PASS TO PASS Sire 14GB Leut updaitd 19 d3y2 a90

docker dull s=ereb+nch/seb.eva]

-dev.1776.pyfakets-1286

AlEngineer

EUROPE

slide-004.jpg

Slide text:

What Makes a Good Task?

Problem description

1. A good task balances clarily

+ ★ AIE Reliable verifier 3. 2. Too easy and too hard both fail Complexity: breadth or depth 35 Gettest listaansles_oapty_or eone_obiectsicbjccts_vatue] Colcl1ent coch,Moc Spl cllent.It_cuthent1eated.retrn Vluo o True Col ellent.coject tore.tnt,retum vluo o? ebjectb objects_wlue. Gxcebrto

1. Should reward actual fixes, should reject fake

solutions (nltlaixe_coa:simertapl ctient_to_ustaol_clfent]

2. Not too narrow, not too wide rewtt oClRrnerh1mvkellem, Gcloud,Cobjet-store

Stable infrastructure is part of eval → Sseerel ho cojects tound atl'test 'in resutt.cot pat

1. Minimal infra noise during runs + contaier.aplclint.object_store.list.asert"called_onc_u

2. ( Connection might blink, images might become

stale. pipeline might break (1970s bug)

AIEngineer

EUROPE

slide-005.jpg

Slide text:

Execution Setup: Minimal Agent, Strong Infrastructure

OPEN

grep

Minimalistic agent (open, edit, bash)

AIE YOLO setup ReAct + demo - tools + no_demo Agent<Infrastructure python python3 EDIT find git REPLACE GOTO cat cd

· Claude-Opus-4.6 top commands from our agent sed SUBHIT 1s SCROLL_DOAN rm

# Braintrust WorkOs OpenAi

Classification audit: raw/sources/slide-ai-classification/slides/wcUJWP6WpGM/audit.json

Slide-Derived Subjects To Review

Subject extraction uses video title, related session titles/descriptions, transcript context, and OCR text when available. OCR is best-effort and should be reviewed against the embedded slide images.