




















Abstract:We investigate Functional Majority Voting (FMV), a method based on functional consensus for code generation with Large Language Models, which identifies a representative solution from multiple generations using their runtime execution signatures on test inputs. We find that FMV is an effective test-time inference strategy, substantially boosting performance on LiveCodeBench without a large compute overhead. Furthermore, we extend the utility of functional consensus and apply it as an aggregation strategy for label-free Test-Time Reinforcement Learning. We demonstrate that this increases pass@1 on holdout tasks, but find no evidence of self-improvement beyond the base model's performance ceiling.
| Comments: | ICLR 2026 Test-Time Updates (TTU) Workshop |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2604.15618 [cs.LG] |
| (or arXiv:2604.15618v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2604.15618 arXiv-issued DOI via DataCite (pending registration) |
From: Jonas Hübotter [view email]
[v1]
Fri, 17 Apr 2026 01:51:36 UTC (3,227 KB)
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。