惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

F
Full Disclosure
Recorded Future
Recorded Future
T
Tenable Blog
S
Securelist
C
CERT Recently Published Vulnerability Notes
T
Threatpost
S
Schneier on Security
A
Arctic Wolf
The Hacker News
The Hacker News
C
CXSECURITY Database RSS Feed - CXSecurity.com
Know Your Adversary
Know Your Adversary
P
Privacy International News Feed
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
The Register - Security
The Register - Security
Cisco Talos Blog
Cisco Talos Blog
AWS News Blog
AWS News Blog
K
Kaspersky official blog
T
True Tiger Recordings
T
Threat Research - Cisco Blogs
V
Vulnerabilities – Threatpost
P
Palo Alto Networks Blog
T
The Exploit Database - CXSecurity.com
小众软件
小众软件
B
Blog
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
Microsoft Azure Blog
Microsoft Azure Blog
Cyberwarzone
Cyberwarzone
C
Cybersecurity and Infrastructure Security Agency CISA
T
Tor Project blog
Spread Privacy
Spread Privacy
Malwarebytes
Malwarebytes
P
Proofpoint News Feed
F
Fox-IT International blog
F
Fortinet All Blogs
P
Privacy & Cybersecurity Law Blog
G
GRAHAM CLULEY
量子位
Latest news
Latest news
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
博客园 - 叶小钗
Project Zero
Project Zero
T
Tailwind CSS Blog
N
Netflix TechBlog - Medium
Martin Fowler
Martin Fowler
IntelliJ IDEA : IntelliJ IDEA – the Leading IDE for Professional Development in Java and Kotlin | The JetBrains Blog
IntelliJ IDEA : IntelliJ IDEA – the Leading IDE for Professional Development in Java and Kotlin | The JetBrains Blog
I
Intezer
博客园_首页
腾讯CDC
H
Hackread – Cybersecurity News, Data Breaches, AI and More
D
Darknet – Hacking Tools, Hacker News & Cyber Security

cs.CL updates on arXiv.org

Temporal Concept Drift in Legal Judgment Prediction: Neural Baselines Across Three Epochs of Ukrainian Court Decisions World-State Transformations for Neuro-symbolic Interactive Storytelling ROC Analysis for Evaluating Translation Quality Estimation Systems READER: Reasoning-Enhanced AI-Generated Text Detection M$^\star$: Every Task Deserves Its Own Memory Harness Learning to Route Languages for Multilingual Policy Optimization Quantifying the Impact of Translation Errors on Multilingual LLM Evaluation Repeated Sequences Reveal Gaps between Large Language Models and Natural Language They Are Not the Same: Direct Causes Are Not Grounded Emotion Explanations Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document Retrieval End-to-End Intracortical Speech Decoding from Neural Activity Measuring the Depth of LLM Unlearning via Activation Patching Generating Legal Commentaries from Case Databases via Retrieval, Clustering, and Generation AstroMind: A High-Fidelity Benchmark for Spacecraft Behavior Reasoning Based on Large Language Models TriVAL: A Tri-Validation Framework for Faithful Automatic Optimization Modeling Overview of the PsyDefDetect Shared Task at BioNLP 2026: Detecting Levels of Psychological Defense Mechanisms in Supportive Conversations SEP-Attack: A Simple and Effective Paradigm for Transfer-Based Textual Adversarial Attack AERIC: Anticipatory Hidden-State Monitoring for Implicit Harmful Dialogue Clarification Is Not Enough: Post-Clarification Answering Remains the Bottleneck in Multi-Turn QA Distinguishing Right from Wrong in Debates: Attribution Analysis of Chinese Harmful Memes Eureka: Intelligent Feature Engineering for Enterprise AI Cloud Resource Demand Prediction Multi-Persona Debate System for Automated Scientific Hypothesis Generation DRInQ: Evaluating Conversational Implicature with Controlled Context Variation Evidence-Linked Radiology Reporting: A Human-Supervised Reference Architecture for Structured Imaging Intelligence Towards a Universal Causal Reasoner MultiHaluDet: Multilingual Hallucination Detection via LLM Hidden State Probing Found in Conversation: LLMs Teach Themselves to Close the Multi-Turn Gap Document Classification Pattern Recognition via Information Fusion: A Systematic Review of Multimodal and Multiview Representation Approaches How Much Structure Do LLMs Need? Evaluating LLMs for Bibliometric Cluster Description Toxicity in Twitch Chats: An LLM-Based Analysis Across Gaming Communities The Tokenizer Tax Across 25 European Languages: Domain Invariance, Cross-Lingual Few-Shot Effects, and the Ukrainian Penalty Lngram: N-gram Conditional Memory in Latent Space H$^{2}$MT: Semantic Hierarchy-Aware Hierarchical Memory Transformer ECHO: Terminal Agents Learn World Models for Free MATO: Multi-objective Personalized Alignment with Test-time Optimization for Large Language Models Raon-Speech Technical Report Better, Faster: Harnessing Self-Improvement in Large Reasoning Models LLM Agent Based Renewable Energy Forecasting Using Edge and IoT Data A Review of Solar Wind Weather and Grid Aware Decision Support Language Bias in LVLMs: From In-Depth Analysis to Simple and Effective Mitigation GroupTravelBench: Benchmarking LLM Agents on Multi-Person Travel Planning Extracting Training Data from Diffusion Language Models via Infilling EchoDistill:Alignment Noisy-to-Clean Self-Distillation for Robust Audio LLMs Inference Time Optimization with Confidence Dynamics Mix-MoE: Improving Multilingual Machine Translation of Large Language Models through Mixed MoEs Phonetic Modeling of Dialectal Variation in Vietnamese Speech ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions From Automation to Collaboration: Human-in-the-Loop Methods for Safe and Trustworthy NLP A general tensor-structured compression scheme for efficient large language models Agent-ToM: Learning to Monitor Autonomous LLM Agents via Theory-of-Mind Reasoning StepGap: A Hybrid NLI-LLM Checker for Step-Level Evidence-Gap Detectionin Multi-Hop Question Answering MindAlign: Bridging EEG, Vision, and Language for Zero-Shot Visual Decoding Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering Who judges the judges? Governance from metrics: a runtime framework for continuous LLM compliance monitoring Knowing but Not Showing: LLMs Recognize Ambiguity but Rarely Ask Clarifying Questions WhenLoss: Diagnosing Write and Retrieval Bottlenecks in Long-Context Memory Systems Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning Faithful or Fabricated? A Causal Framework for Rationalization Bias in LLM Judges Know You Before You Speak: User-State Modeling for LLM Personalization in Multi-Turn Conversation The Path Matters: Learning a Token-Commitment Policy for Diffusion Language Models CSP-Atlas: Concept-Specific Neural Circuits in a Sparse Python Transformer An Interactive Paradigm for Deep Research HiMed: Incentivizing Hindi Reasoning in Medical LLMs DTO: a Differentiable Training Objective for Effective Counterfactual Story Rewriting What Are We Actually Decoding? Source Attribution for Non-Invasive Brain-to-Language Retrieval Beyond the Target: From Imitation to Collaboration in Speculative Decoding TRACE: A taxonomy-grounded synthetic dataset for teaching-program generation and session interpretation in Applied Behavior Analysis NITP: Next Implicit Token Prediction for LLM Pre-training Locality Matters for Training-Free Audio Token Compression in Audio-Language Models When Reasoning Hurts: Source-Aware Evaluation of Frontier LLMs for Clinical SOAP Note Generation SemanticZip: A Pilot Framework for Lossy Text Compression with LLMs as Semantic Decompressors Investigating the Interplay between Contextual and Parametric Chain-of-Thought Faithfulness under Optimization Exploring Profiles of Cognitive Distortions Associated with Mental Health Disorders JudgmentBench: Comparing Rubric and Preference Evaluation for Quality Assessment Momentum Streams for Optimizer-Inspired Transformers CUNY at CLPsych 2026: A Pipeline Approach to Classification and Summarization of Mental Health Changes Tool-Call Dependency Structure is Linearly Decodable in LLM Agent Residual Streams Teaching Through Analogies: A Modular Pipeline for Educational Analogy Generation By Their Fruits You Will Know Them: Comparing Formalizations of Law by the Decisions They Encode A Multi-Probe Audit of Clinical-Interview Depression Detection Benchmarks Knowledge Graph-Driven Expert-Level Reasoning for Neuroscience Discovering Lexical Gaps Using Embeddings from Multilingual LLMs Decompose-and-Refine: Structured Legal Question Answering with Parametric Retrieval SEAL: Synergistic Co-Evolution of Agents and Learning Environments Structure-Aware RAG: Structured Retrieval Augmented Generation from Noisy Data for Conversational Agents Improving Labeling Consistency with Detailed Constitutional Definitions and AI-Driven Evaluation Word Class Representations Spontaneously Emerge from Successor Representations Trained on Natural Language Side-by-side Comparison Amplifies Dialect Bias in Language Models Guarded Repair for Harm-Aware Post-hoc Replacement of LLM Mathematical Reasoning SLAP: Stratified Loss-based Pruning for On-Policy Data-Efficient Instruction Tuning AI-Associated Lexical Shifts Across 34 Languages: Cross-Lingual Convergence and Diachronic Uptake in News Writing Direct Preference Optimization for English-Mandarin Code-Switching Speech Recognition in Audio LLMs P1SCO: Social Dimensions from a Perspectivist Lens Grammatically-Guided Sparse Attention for Efficient and Interpretable Transformers QUEST: Training Frontier Deep Research Agents with Fully Synthetic Tasks Translators as Invisible Teachers of AI: Copyright, Translation Memory, and the Political Economy of Linguistic Data Mimir: Large-scale Multilingual Concept Modeling Re-defining Humor Data Objects for AI Humor Research STREAM: A Data-Centric Framework for Mining High-Value Task-Oriented Dialogues from Streaming Media Large Language Model Selection with Limited Annotations
Improving the Completeness and Comparability of Segment Disclosures: A Large Language Model Approach
Yue Liu, Zhi · 2026-05-26 · via cs.CL updates on arXiv.org

View PDF HTML (experimental)

Abstract:Segment-level disclosures are a central component of financial reporting, providing insight into firms' internal organization and the allocation of economic activities across operating units. However, segment information is often presented in both qualitative and quantitative forms, dispersed across tables and narrative sections of Form 10-K filings. Empirical research relying on structured databases faces both completeness and comparability challenges, as some firm-year observations may be missing, nested segment disclosures are not captured, and support for longitudinal and cross-firm comparability is limited. This study develops a large language model-based framework to extract segment disclosures directly from Form 10-K filings and to preserve both reportable and nested segment information. We further design a retrieval augmented system that incorporates information across multiple filings to support comparability. We use two representative settings to demonstrate its application: longitudinal analysis within a firm to interpret segment changes over time, and cross firm alignment of geographic segments across firms with different reporting structures. The results indicate that the artifact accurately extracts segment-level information and effectively addresses questions that require cross-period knowledge, demonstrating the potential of LLM-based approaches to enhance the measurement and interpretation of segment disclosures.
Comments: 39 pages, 4 figures, submitted to Accounting Horizons
Subjects: Computation and Language (cs.CL); Information Retrieval (cs.IR); General Finance (q-fin.GN)
Cite as: arXiv:2605.23924 [cs.CL]
  (or arXiv:2605.23924v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2605.23924

arXiv-issued DOI via DataCite

Submission history

From: Zhiyuan Cheng [view email]
[v1] Mon, 20 Apr 2026 03:04:08 UTC (462 KB)