


















Abstract:Model compression is increasingly essential for deploying large language models (LLMs), yet existing comparative studies largely focus on pruning and quantization evaluated primarily on knowledge-centric benchmarks. Thus, we introduce UniComp, a unified evaluation framework for comparing pruning, quantization, and knowledge distillation. UniComp evaluates compressed models along three dimensions: performance, reliability, and efficiency, using a diverse set of capability- and safety-oriented benchmarks together with a hardware-aware efficiency analysis. Through evaluation of six compression techniques across 40 datasets, we observe (i) a consistent knowledge bias, where factual recall is largely preserved while multi-step reasoning, multilingual, and instruction-following capabilities degrade; (ii) a decoupling between performance and reliability, indicating that retained performance does not consistently imply preserved reliability; and (iii) that task-specific calibration can yield up to 50% relative improvement of reasoning performance in pruned models.
| Comments: | 18 pages, 5 figures, 18 tables |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2602.09130 [cs.LG] |
| (or arXiv:2602.09130v4 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2602.09130 arXiv-issued DOI via DataCite |
From: Jonathan Von Rad [view email]
[v1]
Mon, 9 Feb 2026 19:20:56 UTC (782 KB)
[v2]
Wed, 11 Feb 2026 09:09:33 UTC (782 KB)
[v3]
Sun, 19 Apr 2026 18:09:45 UTC (1,070 KB)
[v4]
Tue, 5 May 2026 18:01:11 UTC (1,084 KB)
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。