


























TCM-5CEval基准在先前TCM-3CEval基础上扩展,从五个关键维度评估大语言模型的中医综合能力:核心知识(TCM-Exam)、古典文献素养(TCM-LitQA)、临床决策(TCM-MRCD)、中药学(TCM-CMM)和临床非药物疗法(TCM-ClinNPT)。该基准旨在弥补系统性知识差距和文化语境对齐不足,为中医领域LLM提供更精细的诊断工具。
对15个主流LLM的全面评估显示显著性能差异,deepseek_r1和gemini_2_5_pro表现最佳。模型在回忆基础中医知识方面较为熟练,但在理解古典文献的诠释复杂性时存在困难。测试暴露了模型在推理稳定性上的根本缺陷。
基于排列的一致性测试揭示所有模型存在广泛推理脆弱性。即使最高分模型在面对不同选项排序时也表现出显著性能下降,反映出普遍的位置偏差敏感性和缺乏鲁棒理解能力。TCM-5CEval已上传至Medbench平台,加入“中医综合能力深度挑战”专项赛道。
Large language models (LLMs) have demonstrated exceptional capabilities in general domains, yet their application in highly specialized and culturally-rich fields like Traditional Chinese Medicine (TCM) requires rigorous and nuanced evaluation. Building upon prior foundational work such as TCM-3CEval, which highlighted systemic knowledge gaps and the importance of cultural-contextual alignment, we introduce TCM-5CEval, a more granular and comprehensive benchmark. TCM-5CEval is designed to assess LLMs across five critical dimensions: (1) Core Knowledge (TCM-Exam), (2) Classical Literacy (TCM-LitQA), (3) Clinical Decision-making (TCM-MRCD), (4) Chinese Materia Medica (TCM-CMM), and (5) Clinical Non-pharmacological Therapy (TCM-ClinNPT). We conducted a thorough evaluation of fifteen prominent LLMs, revealing significant performance disparities and identifying top-performing models like deepseek\_r1 and gemini\_2\_5\_pro. Our findings show that while models exhibit proficiency in recalling foundational knowledge, they struggle with the interpretative complexities of classical texts. Critically, permutation-based consistency testing reveals widespread fragilities in model inference. All evaluated models, including the highest-scoring ones, displayed a substantial performance degradation when faced with varied question option ordering, indicating a pervasive sensitivity to positional bias and a lack of robust understanding. TCM-5CEval not only provides a more detailed diagnostic tool for LLM capabilities in TCM but aldso exposes fundamental weaknesses in their reasoning stability. To promote further research and standardized comparison, TCM-5CEval has been uploaded to the Medbench platform, joining its predecessor in the "In-depth Challenge for Comprehensive TCM Abilities" special track.
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。