惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

人人都是产品经理
人人都是产品经理
云风的 BLOG
云风的 BLOG
Blog — PlanetScale
Blog — PlanetScale
GbyAI
GbyAI
WordPress大学
WordPress大学
有赞技术团队
有赞技术团队
I
InfoQ
爱范儿
爱范儿
U
Unit 42
Martin Fowler
Martin Fowler
C
Check Point Blog
小众软件
小众软件
博客园 - Franky
H
Help Net Security
博客园 - 【当耐特】
IT之家
IT之家
B
Blog RSS Feed
Recent Announcements
Recent Announcements
宝玉的分享
宝玉的分享
美团技术团队
Last Week in AI
Last Week in AI
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
S
SegmentFault 最新的问题
MongoDB | Blog
MongoDB | Blog

cs.SI updates on arXiv.org

Hiding in Plain Sight: Finding MAHA on Reddit Prism: Structural Symmetry Scanning via Duality-Constrained Laplacian Projection MV-Gate: Insider Threat Detection via Multi-View Behavioral Statistics and Semantic Modeling Algorithmic Cultivation: How Social Media Feeds Shape User Language Universal Dynamics of Punctuated Progress AI-Mediated Communication Can Steer Collective Opinion CitePrism: Human-in-the-Loop AI for Citation Auditing and Editorial Integrity Explainable Detection of Depression Status Shifts from User Digital Traces Can Visual Mamba Improve AI-Generated Image Detection? An In-Depth Investigation ScioMind: Cognitively Grounded Multi-Agent Social Simulation with Anchoring-Based Belief Dynamics and Dynamic Profiles Humanwashing -- It Should Leave You Feeling Dirty When Do LLMs Generate Realistic Social Networks? A Multi-Dimensional Study of Culture, Language, Scale, and Method Moltbook Moderation: Uncovering Hidden Intent Through Multi-Turn Dialogue Linking Extreme Discourse to Structural Polarization in Signed Interaction Networks Predicting Channel Closures in the Lightning Network with Machine Learning Latent Causal Void: Explicit Missing-Context Reconstruction for Misinformation Detection Predictive Maps of Multi-Agent Reasoning: A Successor-Representation Spectrum for LLM Communication Topologies Large Language Models for Causal Relations Extraction in Social Media: A Validation Framework for Disaster Intelligence When Can Digital Personas Reliably Approximate Human Survey Findings? RAwR: Role-Aware Rewiring via Approximate Equitable Partition GravityGraphSAGE: Link Prediction in Directed Attributed Graphs Structure-Centric Graph Foundation Model via Geometric Bases Attention-based graph neural networks: a survey When AI Meets Science: Research Diversity, Interdisciplinarity, Visibility, and Retractions across Disciplines in a Global Surge Scalable inference of spatial regions and temporal signatures from time series Can LLMs Emulate Human Belief Dynamics? Predicting Post Virality with Temporal Cross-Attention over Trend Signals H3: A Healthcare Three-Hop Index for Physician Referral Network Prediction Dynamic Graph with Similarity-Aware Attention Graph Neural Network for Recommender Systems Spectral Graph Sparsification Preserves Representation Geometry in Graph Neural Networks
Structural Hallucination in Large Language Models: A Netw...
[Submitted on 2 Mar 2026 (v1), last revised 21 Aug 2026 (this ve · 2026-03-02 · via cs.SI updates on arXiv.org

View PDF HTML (experimental)

Abstract:Large Language Models (LLMs) increasingly mediate access to scholarly information, yet their outputs are typically evaluated at the level of individual statements rather than knowledge structure. This paper introduces structural hallucination: systematic distortion of conceptual organization, relational architecture, and bibliographic grounding that remains invisible to sentence-level accuracy metrics. To detect such distortions, we develop a network-based hallucination stress test grounded in knowledge graph extraction, graph similarity analysis, centrality comparison, and citation integrity verification. The protocol is applied to three structured domains representing core forms of scholarly knowledge: Roget's Thesaurus (1911) as a historical knowledge organization system; Wikidata philosophers as a biographical knowledge graph; and bibliographic citation records retrieved from the this http URL database. Across all domains, substantial structural divergence is observed. In the lexical benchmark, macro-averaged F1 scores fall below 0.05; in the biographical benchmark, hallucination rates exceed 93%; and in the bibliometric benchmark, citation omission reaches 91.9%. Network-level comparison in the Roget reconstruction further reveals node-set Jaccard similarity of 0.028 and source-mismatch rates above 94%. These findings show that structural fidelity cannot be inferred from local fluency alone. The proposed stress test provides a reproducible instrument for evaluating the structural integrity of LLM-generated knowledge representations within knowledge organization and information quality research.

Submission history

From: Moses Boudourides [view email]
[v1] Mon, 2 Mar 2026 00:49:36 UTC (3,929 KB)
[v2] Fri, 21 Aug 2026 14:31:37 UTC (3,842 KB)