





















Joel Rorseth, PhD candidate
David R. Cheriton School of Computer Science
Supervisor: Professor Lukasz Golab
Retrieval-augmented generation (RAG) enables large language models (LLMs) to integrate external knowledge. However, when users receive undesirable outputs, LLMs cannot faithfully explain which specific external knowledge was responsible. Existing counterfactual and rule-based provenance methods can attribute outputs to influential external knowledge, but are limited by poor scaling and rigid attribution granularity. To bridge these gaps, we introduce a novel framework that computes provenance efficiently at the level of propositions, identifying atomic facts that influence LLM outputs. Our scalable system exploits the hierarchical structure of textual knowledge to decompose documents into propositions, prune redundant propositions, and identify a minimal set of “sufficient" propositions.
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。