---
title: "Citation Needed: Provenance for LLM-Built Knowledge Graphs"
category: "talks"
date: "2026-07-01"
time: "3:20pm-3:40pm"
track: "Graphs"
room: "Track 5"
speakers: ["Daniel Chalef"]
sourceLabels: ["Official conference schedule", "Public YouTube metadata"]
scheduleTrack: "Graphs"
scheduleRoom: "Track 5"
scheduleLabels: ["Graphs", "Track 5", "sponsor", "confirmed"]
---
# Citation Needed: Provenance for LLM-Built Knowledge Graphs

## Conference Context
- Date/time: 2026-07-01 · 3:20pm-3:40pm
- Track/room: Graphs · Track 5
- Speaker(s): Daniel Chalef
- Session type/status: sponsor · confirmed

- Track: Graphs
- Room: Track 5
- Session type: sponsor
- Status: confirmed

## Session Description
An LLM doesn't copy facts into your knowledge graph. It synthesizes them: entities merge across sources, and later data invalidates earlier facts. By the time your agent retrieves "patient has a penicillin allergy," the origin — an EHR record, a lab report, or something typed into a chatbot — is gone. This talk covers engineering lineage into a lossy, generative pipeline: episode-to-fact links as structural graph properties, provenance that survives entity resolution, metadata projection (tag a source once; it follows every derived node and edge), and the query semantics of filtering facts by ancestry, including mixed-trust parentage. Deletion is the inverse problem: GDPR erasure propagates back through the same derivation edges. Compliance gets an audit trail; engineers get agents they can debug instead of black boxes.

## Media Evidence
[Stop Using RAG as Memory — Daniel Chalef, Zep](https://www.youtube.com/watch?v=T5IMo5ntyhA) (speaker-match related prior/adjacent AI Engineer video; captions: English auto-captions).

- Source video: `youtube-T5IMo5ntyhA`
- Slide deck: [[youtube-T5IMo5ntyhA-dense-slides|Dense Slides: Stop Using RAG as Memory — Daniel Chalef, Zep]] — 1 visible slide image(s); 1 HTML recreation(s).
![[assets/dense-slides/T5IMo5ntyhA/slide-001.jpg]]
- Additional slide evidence: [[youtube-T5IMo5ntyhA-slides|Slides: Stop Using RAG as Memory — Daniel Chalef, Zep]], [[youtube-T5IMo5ntyhA-reconstructed-slides|Reconstructed Slides: Stop Using RAG as Memory — Daniel Chalef, Zep]]
- Slide-derived themes for `youtube-T5IMo5ntyhA`: daniel, media, assistant, remembers, everything, except, listening, habits.

## Evidence Graph
This evidence graph is generated from currently linked source material: official schedule text, related video pages, cached transcripts, visible slide text, dense/reconstructed slide pages, and AI slide-classification audits.

### Media Signals
- `youtube-T5IMo5ntyhA` — 9 slide-derived text signals
- Slide-derived themes for `youtube-T5IMo5ntyhA`: daniel, media, assistant, remembers, everything, except, listening, habits.
- Evidence links for `youtube-T5IMo5ntyhA`: [[youtube-T5IMo5ntyhA]], [[youtube-T5IMo5ntyhA-slides]], [[youtube-T5IMo5ntyhA-dense-slides]], [[youtube-T5IMo5ntyhA-reconstructed-slides]]

### Agent Reading Notes
Use these signals to refine the synopsis, topic links, people/company context, and method notes. If a source is a related external video rather than an exact official recording, keep it framed as supporting evidence.

## Transcript Status
Related video transcript availability: English auto-captions. Treat this as supporting context, not a recording of this exact scheduled session unless later confirmed. Not fetched yet.

## People
- [[daniel-chalef]]

## Supporting Slides
- [[youtube-T5IMo5ntyhA-slides]] — extracted from the related public AI Engineer video.

## Slide Evidence
- Slide-only cropped deck: [[youtube-T5IMo5ntyhA-dense-slides]] (1 viable slide images).
- Related slide/OCR pages:
- [[youtube-T5IMo5ntyhA-dense-slides]]
- [[youtube-T5IMo5ntyhA-reconstructed-slides]]
- [[youtube-T5IMo5ntyhA-slides]]
- Slide-derived terms: `entitytype`, `entityfields.text`, `memory`, `export`, `financial`, `fields`, `debt`, `category`, `benchmark`, `none`, `reflect`, `description`, `user`, `type`, `goal`, `entityfields.float`, `amount`, `high`

## Synthesis
### Synthesized Breakdown
# Citation Needed: Provenance for LLM-Built Knowledge Graphs ## Conference Context - Date/time: 2026-07-01 · 3:20pm-3:40pm - Track/room: Graphs · Track 5 - Speaker(s): Daniel Chalef - Session type/status: sponsor · confirmed - Track: Graphs - Room: Track 5 - Session type: sponsor - Status: confirmed ## Session Description An LLM doesn't copy facts into your knowledge graph. It synthesizes them: entities merge across sources, and later data invalidates earlier facts. By the time your agent retrieves "patient has a penicillin allergy," the origin — an EHR record, a lab report, or something typed into a chatbot — is gone. This talk covers engineering lineage into a lossy, generative pipeline: episode-to-fact links as structural graph properties, provenance that survives entity resolution, metadata projection (tag a source once; it follows every derived node and edge), and the query semantics of filtering facts by ancestry, including mixed-trust parentage.

### Speaker And Company Context
- [[daniel-chalef|Daniel Chalef]] — Founder and CEO at [[zep-ai|Zep AI]].

### Topics Covered
- [[agent-security]]

### Derived Links And Source Material
- [[youtube-T5IMo5ntyhA]] — related YouTube source page.
- [[youtube-T5IMo5ntyhA-slides]] — slide evidence.
- [[youtube-T5IMo5ntyhA-reconstructed-slides]] — slide evidence.
- [[youtube-T5IMo5ntyhA-dense-slides]] — slide evidence.

### Novel Concepts / Clever Methods
- No highlighted novel concept has been detected yet.

### Evidence Boundary
This synthesis is based on the official schedule and linked source pages. It should be revisited when exact session recordings or transcript-backed secondary sources are available.
