Citation Needed: Provenance for LLM-Built Knowledge Graphs
Conference Context
- Date/time: 2026-07-01 · 3:20pm-3:40pm
- Track/room: Graphs · Track 5
- Speaker(s): Daniel Chalef
- Session type/status: sponsor · confirmed
- Track: Graphs
- Room: Track 5
- Session type: sponsor
- Status: confirmed
Session Description
An LLM doesn't copy facts into your knowledge graph. It synthesizes them: entities merge across sources, and later data invalidates earlier facts. By the time your agent retrieves "patient has a penicillin allergy," the origin — an EHR record, a lab report, or something typed into a chatbot — is gone. This talk covers engineering lineage into a lossy, generative pipeline: episode-to-fact links as structural graph properties, provenance that survives entity resolution, metadata projection (tag a source once; it follows every derived node and edge), and the query semantics of filtering facts by ancestry, including mixed-trust parentage. Deletion is the inverse problem: GDPR erasure propagates back through the same derivation edges. Compliance gets an audit trail; engineers get agents they can debug instead of black boxes.
Media Evidence
Stop Using RAG as Memory — Daniel Chalef, Zep (speaker-match related prior/adjacent AI Engineer video; captions: English auto-captions).
- Source video:
youtube-T5IMo5ntyhA - Slide deck: Dense Slides: Stop Using RAG as Memory — Daniel Chalef, Zep — 1 visible slide image(s); 1 HTML recreation(s).
- Additional slide evidence: Slides: Stop Using RAG as Memory — Daniel Chalef, Zep, Reconstructed Slides: Stop Using RAG as Memory — Daniel Chalef, Zep
- Slide-derived themes for
youtube-T5IMo5ntyhA: daniel, media, assistant, remembers, everything, except, listening, habits.

Evidence Graph
This evidence graph is generated from currently linked source material: official schedule text, related video pages, cached transcripts, visible slide text, dense/reconstructed slide pages, and AI slide-classification audits.
Media Signals
youtube-T5IMo5ntyhA— 9 slide-derived text signals- Slide-derived themes for
youtube-T5IMo5ntyhA: daniel, media, assistant, remembers, everything, except, listening, habits. - Evidence links for
youtube-T5IMo5ntyhA: youtube T5IMo5ntyhA, youtube T5IMo5ntyhA slides, youtube T5IMo5ntyhA dense slides, youtube T5IMo5ntyhA reconstructed slides
Agent Reading Notes
Use these signals to refine the synopsis, topic links, people/company context, and method notes. If a source is a related external video rather than an exact official recording, keep it framed as supporting evidence.
Transcript Status
Related video transcript availability: English auto-captions. Treat this as supporting context, not a recording of this exact scheduled session unless later confirmed. Not fetched yet.
People
Supporting Slides
- youtube T5IMo5ntyhA slides — extracted from the related public AI Engineer video.
Slide Evidence
- Slide-only cropped deck: youtube T5IMo5ntyhA dense slides (1 viable slide images).
- Related slide/OCR pages:
- youtube T5IMo5ntyhA dense slides
- youtube T5IMo5ntyhA reconstructed slides
- youtube T5IMo5ntyhA slides
- Slide-derived terms:
entitytype,entityfields.text,memory,export,financial,fields,debt,category,benchmark,none,reflect,description,user,type,goal,entityfields.float,amount,high
Synthesis
Synthesized Breakdown
Citation Needed: Provenance for LLM-Built Knowledge Graphs ## Conference Context - Date/time: 2026-07-01 · 3:20pm-3:40pm - Track/room: Graphs · Track 5 - Speaker(s): Daniel Chalef - Session type/status: sponsor · confirmed - Track: Graphs - Room: Track 5 - Session type: sponsor - Status: confirmed ## Session Description An LLM doesn't copy facts into your knowledge graph. It synthesizes them: entities merge across sources, and later data invalidates earlier facts. By the time your agent retrieves "patient has a penicillin allergy," the origin — an EHR record, a lab report, or something typed into a chatbot — is gone. This talk covers engineering lineage into a lossy, generative pipeline: episode-to-fact links as structural graph properties, provenance that survives entity resolution, metadata projection (tag a source once; it follows every derived node and edge), and the query semantics of filtering facts by ancestry, including mixed-trust parentage.
Speaker And Company Context
- Daniel Chalef — Founder and CEO at Zep AI.
Topics Covered
Derived Links And Source Material
- youtube T5IMo5ntyhA — related YouTube source page.
- youtube T5IMo5ntyhA slides — slide evidence.
- youtube T5IMo5ntyhA reconstructed slides — slide evidence.
- youtube T5IMo5ntyhA dense slides — slide evidence.
Novel Concepts / Clever Methods
- No highlighted novel concept has been detected yet.
Evidence Boundary
This synthesis is based on the official schedule and linked source pages. It should be revisited when exact session recordings or transcript-backed secondary sources are available.