惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Y
Y Combinator Blog
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
博客园 - 司徒正美
Blog — PlanetScale
Blog — PlanetScale
博客园 - 聂微东
月光博客
月光博客
量子位
大猫的无限游戏
大猫的无限游戏
Stack Overflow Blog
Stack Overflow Blog
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
The Cloudflare Blog
P
Proofpoint News Feed
B
Blog RSS Feed
美团技术团队
腾讯CDC
C
Check Point Blog
Engineering at Meta
Engineering at Meta
F
Fortinet All Blogs
N
Netflix TechBlog - Medium
Recent Announcements
Recent Announcements
J
Java Code Geeks
S
SegmentFault 最新的问题
WordPress大学
WordPress大学
宝玉的分享
宝玉的分享

Paper Index on ACL Anthology

A Bounded Coordination-Support Capability for Multi-Party Settings: Task-State Monitoring in Firefighter Incident Command A Dataset of Latin Etymologies Extracted from Wiktionary An Efficient Approach for Answering Not Readily Attainable Questions for RAG-based Applications Automated German Alt Text Generation for News Charts Call Support Copilot: A Reproducible Multimodal System for Speech Emotion Recognition, Intent Understanding, and Agent Assistance Can Large Language Models Replace Statistical Software? Code-Switching Detection in Multilingual Child Speech with SwissBERT Concept Extraction and Webb’s Depth of Knowledge: Comparing LLM Question Generation Pipelines for Educational Assessment Data Augmentation for Historical NER: A Systematic Comparison of Lexical and LLM-based Approaches Enhancing Retrieval via Cognitively Motivated Document Expansion Extending the Contact Hypothesis: Cross-Linguistic Evaluation of Religion and Nationality Bias When Prompting LLMs in German and Icelandic Extracting Article-Level Legal Dependencies from Swiss Federal Law using LLMs How Good is AI on Swiss Voting Booklets? A Multilingual OCR and Alignment Benchmark Optimizing Large Language Models for Robust Domain-Specific Text-to-SQL: From Prompting to Preference Alignment Proceedings of the 11th Edition of the Swiss Text Analytics Conference Reinforcement Learning for Latent-Space Thinking in LLMs RUMLEM: A Dictionary-Based Lemmatizer for Romansh Skill Extraction from Resumes and Job Offers across Six Languages Text vs. Phoneme Intermediates for Low-Resource Swiss German The Same Email, Signed Differently: Testing Negotiation Bias and Recommendation Stability in LLMs Which Skills Debate Reaches the Public? Comparing Scientific Literature and Media Coverage of AI and LLM Skill Impacts (2022–2025) Controlling Language and Style of Multi-lingual Generative Language Models with Control Vectors Hybrid Human-LLM Corpus Construction and LLM Evaluation for the Caused-Motion Construction Implicit and Indirect: Detecting Face-threatening and Paired Actions in Asynchronous Online Conversations Northern European Journal of Language Technology, Volume 11 A modular architecture for creating multimodal embodied agents with an episodic Knowledge Graph as an explainable and controllable long-term memory A Neural Approach to Discourse Relation Signal Detection An Analysis of Japanese Sentence-final Particle Yone: Compare Yone and Ne in Response Attribution and the discourse structure of reports Automatic Detection of the Bulgarian Evidential Renarrative
大语言模型汉字富语义能力评测
2026-03-23 · via Paper Index on ACL Anthology
Abstract

"中文相较于以英文为代表的表音文字具有富语义的特点,单个汉字蕴含了读音、字形结构、偏旁部首等丰富的语义特征,在构建自然语言处理相关应用时具有独特的价值,可以视作额外的特征,提升在特定任务的表现。近年来,大语言模型飞速发展,展现出海量的知识储备和强大的推理能力,其中,大模型对汉字富语义特征的掌握可以视作大模型中文能力的基础。然而,目前对于大模型汉字富语义能力评测研究较少,针对性地评测大模型在汉字富语义方面的能力边界,有助于了解大模型中英文能力差异性、并推测大模型在字形、字音相关下游任务上的表现。因此,本研究从汉字的结构、偏旁、读音、笔画、多音字和部件六个维度,对大语言模型进行了全面评测,旨在深入探究其对汉字基本富语义特征的掌握程度。本研究以GB2312 标准字符集和现代汉语词典为依据,围绕汉字的结构、偏旁、读音、笔画、多音字和部件六个维度,构建了一系列“问题-答案”对,并制定了科学合理的评分标准。在此基础上,对十余种主流的大语言模型进行了深入评测。同时,为探究模型在中英文能力上的差异,将上述中文评测任务翻译为英文,并选取了三个代表性模型进行对比评测。此外,本研究进一步从汉字结构推理、偏旁推理、读音推理三个关键角度出发,设计了一系列推理评测任务,旨在深入评估大语言模型对汉字富语义特征的推理能力。本研究的评测结果具有重要的参考价值,可为大语言模型相关领域的研究人员在中文下游任务优化、基础模型选择等关键环节提供参考和启发。"