---
title: "Open-Source Inference Engineering for the Agentic Era"
category: "talks"
date: "2026-06-29"
time: "9:00am-11:00am"
track: "Workshops Day 1"
room: "Track 8"
speakers: ["Zain Hasan", "Yubo Wang", "Qingyang Wu", "Jue Wang"]
sourceLabels: ["Official conference schedule", "Public YouTube metadata"]
scheduleTrack: "Workshops Day 1"
scheduleRoom: "Track 8"
scheduleLabels: ["Workshops Day 1", "Track 8", "sponsor", "confirmed"]
---
# Open-Source Inference Engineering for the Agentic Era

## Conference Context
- Date/time: 2026-06-29 · 9:00am-11:00am
- Track/room: Workshops Day 1 · Track 8
- Speaker(s): Zain Hasan, Yubo Wang, Qingyang Wu, Jue Wang
- Session type/status: sponsor · confirmed

- Track: Workshops Day 1
- Room: Track 8
- Session type: sponsor
- Status: confirmed

## Session Description
Agentic coding workloads demand long contexts, multi-turn conversations, and throughput at a scale that most inference engines weren't built for. TokenSpeed is a new open-source engine purpose-built for this regime, built collaboratively by NVIDIA DevTech, AMD Triton, Qwen Inference, Together AI, and others. In this 2-hour hands-on workshop, Together Inference Research Engineers and a TokenSpeed co-creator will cover TokenSpeed architecture, deploying your first model, optimizing for agentic workloads, kernel and hardware tuning, and throughput/latency trade-offs.

## Media Evidence
No related AI Engineer channel video found yet.

## Evidence Graph
This evidence graph is generated from currently linked source material: official schedule text, related video pages, cached transcripts, visible slide text, dense/reconstructed slide pages, and AI slide-classification audits.

### Media Signals
No linked video, transcript, or slide source has been attached yet.

### Agent Reading Notes
Use these signals to refine the synopsis, topic links, people/company context, and method notes. If a source is a related external video rather than an exact official recording, keep it framed as supporting evidence.

## Transcript Status
No official session recording transcript was found by exact title match on the AI Engineer YouTube channel during this run.

## People
- [[zain-hasan]]
- [[yubo-wang]]
- [[qingyang-wu]]
- [[jue-wang]]

## Notes
- Pending transcript synthesis when an official recording or confirmed matching video is available.

## Synthesis
### Synthesized Breakdown
# Open-Source Inference Engineering for the Agentic Era ## Conference Context - Date/time: 2026-06-29 · 9:00am-11:00am - Track/room: Workshops Day 1 · Track 8 - Speaker(s): Zain Hasan, Yubo Wang, Qingyang Wu, Jue Wang - Session type/status: sponsor · confirmed - Track: Workshops Day 1 - Room: Track 8 - Session type: sponsor - Status: confirmed ## Session Description Agentic coding workloads demand long contexts, multi-turn conversations, and throughput at a scale that most inference engines weren't built for. TokenSpeed is a new open-source engine purpose-built for this regime, built collaboratively by NVIDIA DevTech, AMD Triton, Qwen Inference, Together AI, and others. In this 2-hour hands-on workshop, Together Inference Research Engineers and a TokenSpeed co-creator will cover TokenSpeed architecture, deploying your first model, optimizing for agentic workloads, kernel and hardware tuning, and throughput/latency trade-offs. ## Media Evidence No related AI Engineer channel video found yet.

### Speaker And Company Context
- [[zain-hasan|Zain Hasan]] — Staff AI/ML Engineer - DX at [[together-ai|Together AI]].
- [[yubo-wang|Yubo Wang]] — LLM Inference at [[together-ai|Together AI]].
- [[qingyang-wu|Qingyang Wu]] — Staff Research Scientist at [[together-ai|Together AI]].
- [[jue-wang|Jue Wang]] — Senior Staff Researcher at [[together-ai|Together AI]].

### Topics Covered
- [[agentic-search]]

### Derived Links And Source Material

### Novel Concepts / Clever Methods
- No highlighted novel concept has been detected yet.

### Evidence Boundary
This synthesis is based on the official schedule and linked source pages. It should be revisited when exact session recordings or transcript-backed secondary sources are available.
