












Abstract:While Large Language Models (LLMs) have substantially improved the functional correctness of code translation, the critical dimension of execution efficiency remains overlooked. We present \textbf{\textsc{Trace}}, the first benchmark to explicitly assess efficiency in LLM-translated code. \textsc{Trace} includes 1,000 efficiency-critical tasks across C++, Java, and Python, each augmented with stress tests that reveal efficiency disparities often overlooked by small-scale tests. Using \textsc{Trace}, we conduct an extensive evaluation of 28 representative LLMs and highlight several key insights: 1) Correctness and efficiency are often misaligned: the correctness leader Claude-Sonnet-4-Think achieves only moderate time efficiency, outperformed by smaller open-source LLMs such as Qwen2.5-Coder-14B-Instruct. 2) Inefficiency is both prevalent and patterned: 23.5\% of correct translations suffer from notable inefficiency, mainly arising from algorithm implementation discrepancy (11.9\%), language construct mismatch (66.4\%), and resource management inefficiency (21.7\%). 3) Inference-time prompt strategies bring only modest improvements, indicating that simple prompting alone is insufficient to improve translation efficiency. Together, our results establish execution efficiency as an essential dimension of code translation and position \textsc{Trace} as a principled foundation for efficiency-oriented evaluation.
From: Zhihao Gong [view email]
[v1]
Fri, 15 Aug 2025 13:33:52 UTC (337 KB)
[v2]
Thu, 19 Mar 2026 03:49:51 UTC (802 KB)
[v3]
Mon, 31 Aug 2026 03:49:40 UTC (801 KB)
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。