---
title: "Slides: SWE-rebench: Lessons from Evaluating Coding Agents — Ibragim Badertdinov, Nebius"
category: "slides"
video_id: "wcUJWP6WpGM"
sourceLabels: ["Public YouTube video frames", "Public YouTube metadata"]
---

# Slides: SWE-rebench: Lessons from Evaluating Coding Agents — Ibragim Badertdinov, Nebius

## Source Video
[SWE-rebench: Lessons from Evaluating Coding Agents — Ibragim Badertdinov, Nebius](https://www.youtube.com/watch?v=wcUJWP6WpGM)

## Relationship To World's Fair 2026
These slides are extracted from a public AI Engineer YouTube video connected to World's Fair 2026. Speaker-matched clips are supporting context unless later confirmed as exact session recordings; official livestream recordings are day-level/event-level source material.

## Related Scheduled Sessions
- No individual scheduled session mapping has been assigned yet; treat this as an event livestream deck.

## Extracted Slides
![[assets/slides/wcUJWP6WpGM/slide-001.jpg]]

- Recreated text/layout view: [open HTML recreation](/assets/slide-recreations/slides/wcUJWP6WpGM/slide-001.html)
- AI slide classifier: `title_card` confidence `0.91`
- Text source: agent_vision.

Slide text:

> SWE-rebench: Lessons from Evaluating Coding Agents on Real Software Engineering Tasks

![[assets/slides/wcUJWP6WpGM/slide-002.jpg]]

- Recreated text/layout view: [open HTML recreation](/assets/slide-recreations/slides/wcUJWP6WpGM/slide-002.html)
- AI slide classifier: `content_slide` confidence `0.99`
- Text source: agent_vision.

Slide text:

> Why evals matter now?
> Models improved. Choosing became harder.
> Vibe checks do not scale
> SWE performance grows rapidly
> Options change every month
> * SWE-rebench 2026_02 tasks

![[assets/slides/wcUJWP6WpGM/slide-003.jpg]]

- Recreated text/layout view: [open HTML recreation](/assets/slide-recreations/slides/wcUJWP6WpGM/slide-003.html)
- AI slide classifier: `content_slide` confidence `0.96`
- Text source: advanced OCR `rapidocr-live/bright-screen/opencv-adaptive`.
- OCR decision: ready — multi-column slide with screenshots and small embedded labels

Slide text:

> Anatomy of a Task: A Task Is More Than Text
> Task description - original issue text Regresslon causod by changes for woakref ot fllesystom 1284
> ★ AIE +
> Sandbox environment -
> executable Docker image shs256141e015c4.:9
> Verifier - tests from the PR. FAIL_TO PASS + PASS TO PASS Sire 14GB Leut updaitd 19 d3y2 a90
> docker dull s=ereb+nch/seb.eva]
> -dev.1776.pyfakets-1286
> AlEngineer
> EUROPE

![[assets/slides/wcUJWP6WpGM/slide-004.jpg]]

- Recreated text/layout view: [open HTML recreation](/assets/slide-recreations/slides/wcUJWP6WpGM/slide-004.html)
- AI slide classifier: `content_slide` confidence `0.97`
- Text source: advanced OCR `rapidocr-live/bright-screen/opencv-adaptive`.
- OCR decision: ready — code screenshot and small text in a structured slide layout

Slide text:

> What Makes a Good Task?
> Problem description
> 1. A good task balances clarily
> + ★ AIE Reliable verifier 3. 2. Too easy and too hard both fail Complexity: breadth or depth 35 Gettest listaansles_oapty_or eone_obiectsicbjccts_vatue] Colcl1ent coch,Moc Spl cllent.It_cuthent1eated.retrn Vluo o True Col ellent.coject tore.tnt,retum vluo o? ebjectb objects_wlue. Gxcebrto
> 1. Should reward actual fixes, should reject fake
> solutions (nltlaixe_coa:simertapl ctient_to_ustaol_clfent]
> 2. Not too narrow, not too wide rewtt oClRrnerh1mvkellem, Gcloud,Cobjet-store
> Stable infrastructure is part of eval → Sseerel ho cojects tound atl'test 'in resutt.cot pat
> 1. Minimal infra noise during runs + contaier.aplclint.object_store.list.asert"called_onc_u
> 2. （ Connection might blink, images might become
> stale. pipeline might break (1970s bug)
> AIEngineer
> EUROPE

![[assets/slides/wcUJWP6WpGM/slide-005.jpg]]

- Recreated text/layout view: [open HTML recreation](/assets/slide-recreations/slides/wcUJWP6WpGM/slide-005.html)
- AI slide classifier: `content_slide` confidence `0.96`
- Text source: advanced OCR `rapidocr-live/bright-screen/contrast`.
- OCR decision: ready — small command list and dense layout on the right side

Slide text:

> Execution Setup: Minimal Agent, Strong Infrastructure
> OPEN
> grep
> Minimalistic agent (open, edit, bash)
> AIE YOLO setup ReAct + demo - tools + no_demo Agent<Infrastructure python python3 EDIT find git REPLACE GOTO cat cd
> · Claude-Opus-4.6 top commands from our agent sed SUBHIT 1s SCROLL_DOAN rm
> # Braintrust WorkOs OpenAi


Classification audit: `raw/sources/slide-ai-classification/slides/wcUJWP6WpGM/audit.json`

## Slide-Derived Subjects To Review
Subject extraction uses video title, related session titles/descriptions, transcript context, and OCR text when available. OCR is best-effort and should be reviewed against the embedded slide images.
