惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

H
Help Net Security
T
ThreatConnect
SecWiki News
SecWiki News
F
Future of Privacy Forum
AWS News Blog
AWS News Blog
C
Cisco Blogs
A
Arctic Wolf
Vercel News
Vercel News
The GitHub Blog
The GitHub Blog
Scott Helme
Scott Helme
V
V2EX
博客园 - 叶小钗
阮一峰的网络日志
阮一峰的网络日志
K
Kaspersky official blog
G
Google Developers Blog
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
P
Privacy International News Feed
C
Cyber Attacks, Cyber Crime and Cyber Security
N
News | PayPal Newsroom
Schneier on Security
Schneier on Security
NISL@THU
NISL@THU
Microsoft Azure Blog
Microsoft Azure Blog
量子位
The Hacker News
The Hacker News
Stack Overflow Blog
Stack Overflow Blog
Security Latest
Security Latest
M
Microsoft Research Blog - Microsoft Research
Google Online Security Blog
Google Online Security Blog
博客园_首页
C
CXSECURITY Database RSS Feed - CXSecurity.com
I
InfoQ
Google DeepMind News
Google DeepMind News
Y
Y Combinator Blog
The Cloudflare Blog
Microsoft Security Blog
Microsoft Security Blog
Martin Fowler
Martin Fowler
Cisco Talos Blog
Cisco Talos Blog
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
T
Troy Hunt's Blog
F
Fox-IT International blog
S
Security @ Cisco Blogs
博客园 - 司徒正美
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
C
Comments on: Blog
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
L
LINUX DO - 最新话题
GbyAI
GbyAI
Project Zero
Project Zero
腾讯CDC
T
Tailwind CSS Blog

cs.CL updates on arXiv.org

Phonetic Modeling of Dialectal Variation in Vietnamese Speech The Tokenizer Tax Across 25 European Languages: Domain Invariance, Cross-Lingual Few-Shot Effects, and the Ukrainian Penalty World-State Transformations for Neuro-symbolic Interactive Storytelling Mimir: Large-scale Multilingual Concept Modeling Copy-as-Decode: Grammar-Constrained Parallel Prefill for LLM Editing What Are We Actually Decoding? Source Attribution for Non-Invasive Brain-to-Language Retrieval When Reasoning Hurts: Source-Aware Evaluation of Frontier LLMs for Clinical SOAP Note Generation Translators as Invisible Teachers of AI: Copyright, Translation Memory, and the Political Economy of Linguistic Data Clarification Is Not Enough: Post-Clarification Answering Remains the Bottleneck in Multi-Turn QA Grammatically-Guided Sparse Attention for Efficient and Interpretable Transformers Discovering Lexical Gaps Using Embeddings from Multilingual LLMs Guarded Repair for Harm-Aware Post-hoc Replacement of LLM Mathematical Reasoning Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document Retrieval Generating Legal Commentaries from Case Databases via Retrieval, Clustering, and Generation EchoDistill:Alignment Noisy-to-Clean Self-Distillation for Robust Audio LLMs Quantifying the Impact of Translation Errors on Multilingual LLM Evaluation NITP: Next Implicit Token Prediction for LLM Pre-training Faithful or Fabricated? A Causal Framework for Rationalization Bias in LLM Judges GroupTravelBench: Benchmarking LLM Agents on Multi-Person Travel Planning End-to-End Intracortical Speech Decoding from Neural Activity Knowing but Not Showing: LLMs Recognize Ambiguity but Rarely Ask Clarifying Questions Multi-Persona Debate System for Automated Scientific Hypothesis Generation An Interactive Paradigm for Deep Research Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth Agent-ToM: Learning to Monitor Autonomous LLM Agents via Theory-of-Mind Reasoning Overview of the PsyDefDetect Shared Task at BioNLP 2026: Detecting Levels of Psychological Defense Mechanisms in Supportive Conversations SEAL: Synergistic Co-Evolution of Agents and Learning Environments Document Classification Pattern Recognition via Information Fusion: A Systematic Review of Multimodal and Multiview Representation Approaches Distinguishing Right from Wrong in Debates: Attribution Analysis of Chinese Harmful Memes A Multi-Probe Audit of Clinical-Interview Depression Detection Benchmarks TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering Repeated Sequences Reveal Gaps between Large Language Models and Natural Language Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning A general tensor-structured compression scheme for efficient large language models P1SCO: Social Dimensions from a Perspectivist Lens Raon-Speech Technical Report Exploring Profiles of Cognitive Distortions Associated with Mental Health Disorders Evidence-Linked Radiology Reporting: A Human-Supervised Reference Architecture for Structured Imaging Intelligence Better, Faster: Harnessing Self-Improvement in Large Reasoning Models By Their Fruits You Will Know Them: Comparing Formalizations of Law by the Decisions They Encode CUNY at CLPsych 2026: A Pipeline Approach to Classification and Summarization of Mental Health Changes Improving the Completeness and Comparability of Segment Disclosures: A Large Language Model Approach JudgmentBench: Comparing Rubric and Preference Evaluation for Quality Assessment Know You Before You Speak: User-State Modeling for LLM Personalization in Multi-Turn Conversation Found in Conversation: LLMs Teach Themselves to Close the Multi-Turn Gap DRInQ: Evaluating Conversational Implicature with Controlled Context Variation They Are Not the Same: Direct Causes Are Not Grounded Emotion Explanations MATO: Multi-objective Personalized Alignment with Test-time Optimization for Large Language Models Lngram: N-gram Conditional Memory in Latent Space ROC Analysis for Evaluating Translation Quality Estimation Systems ECHO: Terminal Agents Learn World Models for Free TRACE: A taxonomy-grounded synthetic dataset for teaching-program generation and session interpretation in Applied Behavior Analysis The Path Matters: Learning a Token-Commitment Policy for Diffusion Language Models StepGap: A Hybrid NLI-LLM Checker for Step-Level Evidence-Gap Detectionin Multi-Hop Question Answering READER: Reasoning-Enhanced AI-Generated Text Detection AstroMind: A High-Fidelity Benchmark for Spacecraft Behavior Reasoning Based on Large Language Models MultiHaluDet: Multilingual Hallucination Detection via LLM Hidden State Probing SLAP: Stratified Loss-based Pruning for On-Policy Data-Efficient Instruction Tuning HiMed: Incentivizing Hindi Reasoning in Medical LLMs CP-Agent: A Calibrated Risk-Controlled Agent for Feedback-Driven Competitive Programming Word Class Representations Spontaneously Emerge from Successor Representations Trained on Natural Language Improving Labeling Consistency with Detailed Constitutional Definitions and AI-Driven Evaluation Measuring the Depth of LLM Unlearning via Activation Patching Towards a Universal Causal Reasoner AI-Associated Lexical Shifts Across 34 Languages: Cross-Lingual Convergence and Diachronic Uptake in News Writing Who judges the judges? Governance from metrics: a runtime framework for continuous LLM compliance monitoring Language Bias in LVLMs: From In-Depth Analysis to Simple and Effective Mitigation H$^{2}$MT: Semantic Hierarchy-Aware Hierarchical Memory Transformer Re-defining Humor Data Objects for AI Humor Research DTO: a Differentiable Training Objective for Effective Counterfactual Story Rewriting Learning to Route Languages for Multilingual Policy Optimization SEP-Attack: A Simple and Effective Paradigm for Transfer-Based Textual Adversarial Attack Large Language Model Selection with Limited Annotations From Automation to Collaboration: Human-in-the-Loop Methods for Safe and Trustworthy NLP Inference Time Optimization with Confidence Dynamics Toxicity in Twitch Chats: An LLM-Based Analysis Across Gaming Communities Eureka: Intelligent Feature Engineering for Enterprise AI Cloud Resource Demand Prediction Extracting Training Data from Diffusion Language Models via Infilling Knowledge Graph-Driven Expert-Level Reasoning for Neuroscience Direct Preference Optimization for English-Mandarin Code-Switching Speech Recognition in Audio LLMs Locality Matters for Training-Free Audio Token Compression in Audio-Language Models ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions Temporal Concept Drift in Legal Judgment Prediction: Neural Baselines Across Three Epochs of Ukrainian Court Decisions Side-by-side Comparison Amplifies Dialect Bias in Language Models How Much Structure Do LLMs Need? Evaluating LLMs for Bibliometric Cluster Description QUEST: Training Frontier Deep Research Agents with Fully Synthetic Tasks WhenLoss: Diagnosing Write and Retrieval Bottlenecks in Long-Context Memory Systems Structure-Aware RAG: Structured Retrieval Augmented Generation from Noisy Data for Conversational Agents CSP-Atlas: Concept-Specific Neural Circuits in a Sparse Python Transformer TriVAL: A Tri-Validation Framework for Faithful Automatic Optimization Modeling MindAlign: Bridging EEG, Vision, and Language for Zero-Shot Visual Decoding AERIC: Anticipatory Hidden-State Monitoring for Implicit Harmful Dialogue Tool-Call Dependency Structure is Linearly Decodable in LLM Agent Residual Streams Decompose-and-Refine: Structured Legal Question Answering with Parametric Retrieval Teaching Through Analogies: A Modular Pipeline for Educational Analogy Generation Beyond the Target: From Imitation to Collaboration in Speculative Decoding Momentum Streams for Optimizer-Inspired Transformers STREAM: A Data-Centric Framework for Mining High-Value Task-Oriented Dialogues from Streaming Media LLM Agent Based Renewable Energy Forecasting Using Edge and IoT Data A Review of Solar Wind Weather and Grid Aware Decision Support Investigating the Interplay between Contextual and Parametric Chain-of-Thought Faithfulness under Optimization
CArtBench: Evaluating Vision-Language Models on Chinese Art Understanding, Interpretation, and Authenticity
Xuefeng Wei, · 2026-05-26 · via cs.CL updates on arXiv.org

View PDF HTML (experimental)

Abstract:We introduce CARTBENCH, a museum-grounded benchmark for evaluating vision-language models (VLMs) on Chinese artworks beyond short-form recognition and QA. CARTBENCH comprises four subtasks: CURATORQA for evidence-grounded recognition and reasoning, CATALOGCAPTION for structured four-section expert-style appreciation, REINTERPRET for defensible reinterpretation with expert ratings, and CONNOISSEURPAIRS for diagnostic authenticity discrimination under visually similar confounds. CARTBENCH is built by aligning image-bearing Palace Museum objects from Wikidata with authoritative catalog pages, spanning five art categories across multiple dynasties. Across nine representative VLMs, we find that high overall CURATORQA accuracy can mask sharp drops on hard evidence linking and style-to-period inference; long-form appreciation remains far from expert references; and authenticity-oriented diagnostic discrimination stays near chance, underscoring the difficulty of connoisseur-level reasoning for current models.
Comments: under review
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2604.11632 [cs.CL]
  (or arXiv:2604.11632v2 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2604.11632

arXiv-issued DOI via DataCite

Submission history

From: Xuefeng Wei [view email]
[v1] Mon, 13 Apr 2026 15:44:02 UTC (5,261 KB)
[v2] Mon, 25 May 2026 15:43:45 UTC (5,261 KB)