惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

云风的 BLOG
云风的 BLOG
IT之家
IT之家
D
Docker
博客园 - 叶小钗
A
About on SuperTechFans
博客园_首页
Apple Machine Learning Research
Apple Machine Learning Research
Recorded Future
Recorded Future
Stack Overflow Blog
Stack Overflow Blog
腾讯CDC
V
V2EX
S
SegmentFault 最新的问题
量子位
P
Proofpoint News Feed
酷 壳 – CoolShell
酷 壳 – CoolShell
Latest news
Latest news
大猫的无限游戏
大猫的无限游戏
月光博客
月光博客
有赞技术团队
有赞技术团队
The GitHub Blog
The GitHub Blog
C
Cyber Attacks, Cyber Crime and Cyber Security
I
InfoQ
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
D
DataBreaches.Net
G
GRAHAM CLULEY
P
Proofpoint News Feed
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
Microsoft Security Blog
Microsoft Security Blog
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
Y
Y Combinator Blog
小众软件
小众软件
NISL@THU
NISL@THU
L
Lohrmann on Cybersecurity
aimingoo的专栏
aimingoo的专栏
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
I
Intezer
Last Week in AI
Last Week in AI
T
Threatpost
人人都是产品经理
人人都是产品经理
U
Unit 42
Security Latest
Security Latest
AWS News Blog
AWS News Blog
T
The Blog of Author Tim Ferriss
MongoDB | Blog
MongoDB | Blog
罗磊的独立博客
GbyAI
GbyAI
P
Palo Alto Networks Blog
G
Google Developers Blog
MyScale Blog
MyScale Blog
L
LangChain Blog

Hacker News - Newest: "LLM"

GitHub - lechmazur/position_bias: A benchmark for testing whether LLM judges keep the same preference when two lightly edited versions of the same story are shown in opposite orders. Flex routing (EU and EFTA) Dark Factories: Retooling for LLM Velocity Ask HN: What would be the impact of a LLM output injection attack? GitHub - AronDaron/dataset-generator: No-code desktop app for generating high-quality synthetic datasets to fine-tune LLMs — plan-then-execute pipeline, LLM-as-judge, HuggingFace upload. GitHub - Oaklight/llm-rosetta: Production-ready LLM API translation layer for Python — bidirectional conversion between OpenAI, Anthropic & Google formats via hub-and-spoke IR. Optional API gateway. Streaming & non-streaming. Zero core deps. Contributions welcome! GitHub - browser-use/browser-harness: Self-healing browser harness that enables LLMs to complete any task. GitHub - moeen-mahmud/remen: Remen turns thoughts into something you can return to Analyzing 156 LLM Launch Posts on Hacker News ChatGPT vs Gemini vs Claude: The Best LLM Subscription You Should Buy GitHub - salaamalykum/quran-semantic-search: High-density RAG Semantic Search Engine & Quran Corpus (GEO/SEO Architecture) GitHub - NVIDIA/TensorRT-LLM: TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way. The State of LLM Bug Bounties in 2026 Operational Readiness Criteria for Tool-Using LLM Agents Meshcore: Architecture for a Decentralized P2P LLM Inference Network How an LLM becomes more coherent as we train it GitHub - seetrex-ai/laimark GitHub - Jossifresben/BibCrit: AI-assited biblical textual criticism GitHub - wastedcode/memex: File system based wiki, maintained by Claude 99helpers.com GitHub - cliver-project/AITrigram GitHub - unbody-io/adapt: A self-evolving memory layer for AI agents. GitHub - hb20007/awesome-gen-ai-fails: A list of incidents where reliance on generative AI and LLMs resulted in harm to companies, individuals, or society GitHub - nevenkordic/localmind: Run any local LLM with persistent memory and context. CLI agent over Ollama with SQLite-backed hybrid recall. No cloud. Ask HN: What are the machine requirements for a LLM like Llama-3.1-8B? Faster LLM Inference via Sequential Monte Carlo grpo explained: group relative policy optimization for llm finetuning - cgft Stop comparing price per million tokens: the hidden LLM API costs · TensorZero Andrej Karpathy's LLM Wiki Is a Bad Idea GitHub - GG-QandV/mnemostroma: Offline RAM-first cognitive leer/coprocessor for AI agents and robotics. Solves "Context Abandonment" with 20-80ms latency using a dual-thread biomimetic memory architecture (ONNX + SQLite WAL). mempalace/agent at agent · skorotkiewicz/mempalace GitHub - Nyquest-ai/nyquest-rust-fullstack-pub: Nyquest — Semantic Compression Proxy for LLMs. 350+ rules, local LLM stage, 15-75% token savings. Full Rust stack. GitHub - TheoV823/mneme: Enforce architectural decisions in AI-assisted development. GitHub - klemenvod/TokenBrawl: A 1v1 Bomberman-style game where two LLM agents play autonomously against each other. No human plays — you watch the AIs fight. Each agent receives a text description of the board state, reasons about it, and outputs a move as JSON. The game engine executes it. Introducing the Common AI Provider: LLM and AI Agent Support for Apache Airflow Power Circuit AI: Designing Power Electronic Circuits for Motor Drives with Generative Artificial Intelligence Ask HN: How to program with IDE and LLM on CPU locally? Show HN: Agent-cache – Multi-tier LLM/tool/session caching for Valkey and Redis Bonsai 1-bit WebGPU - a Hugging Face Space by webml-community The LLM Fallacy: Misattribution in AI-Assisted Cognitive Workflows Ask HN: Simple tooling for local LLM code critique without IDE integration? Can a General LLM Diagnose a DICOM Slice? A 10-Case Public Benchmark Charts-of-Thought: Enhancing LLM Visualization Literacy (PDF, 2026) GitHub - Mesh-LLM/mesh-llm: Distributed AI/LLM for the people. Share compute privately or publicly to power your agents and chat. GitHub - seamus-brady/springdrift: A persistent runtime for long-lived LLM agents Writing an LLM from scratch, part 32k -- Interventions: training a better model locally with gradient accumulation Ask HN: Which LLM model and agentic CLI are you using for local development? GitHub - wayneColt/modelcascade: Route local. Escalate smart. Never overspend. Open-source multi-model cascade routing for autonomous agents. LLM pricing is 100x harder than you think GitHub - asakin/llm-primer: Pre-warmed Claude Code sessions in tmux. No startup wait. GitHub - EggerMarc/chat-rs: A multi-provider LLM framework for Rust. GitHub - SynapseKit/SynapseKit: Minimal, async-first Python framework for production LLM apps- 2 hard deps, no magic, no SaaS. A Claude Skill that Makes LLM Paragraphs More Bearable Does Gas Town 'steal' usage from users' LLM credits & paid services to improve itself? What's Claude Code Actually Doing? Open the Black Box with the Arthur Engine Milla Jovovich's New Open Source LLM Memory App and the Dark Code Problem Your intuition of LLM token usage might be wrong Show HN: Bloomberg Terminal for LLM ops – free and open source GitHub - 0xchamin/mcptube: Transform YouTube videos into a compounding knowledge base with transcripts, vision analysis, and agentic search. Works as an MCP server for Claude, Copilot & more. Show HN: Open KB: Open LLM Knowledge Base Your LLM is a compiler, not a runtime GitHub - sapountzis/Unslop: A Web Feed That Deserves You crates.io: Rust Package Registry Beyond Karpathy's LLM-Wiki: The Necessity of Cognitive Governance GitHub - amitshekhariitbhu/llm-internals: Learn LLM internals step by step - from tokenization to attention to inference optimization. GitHub - parallem-ai/parallem: An expressive library for running agents with the Batch API. GitHub - stfurkan/pi-llm LLM-Wiki Show HN: Formal – Formal verification for AI-generated code using Lean 4 LRTS – Regression testing for LLM prompts (open source, local-first) LLM Wiki Skill: Build a Second Brain with Claude Code and Obsidian I built an LLM Wiki and RAG solution: here's a demo for a security KB The biggest advance in AI since the LLM Predict-Rlm: The LLM Runtime That Lets Models Write Their Own Control Flow the-synthetic-library/the-synthetic-mind at main · joshferrer1/the-synthetic-library GitHub - yisding/reviewwiggum GitHub - Donnyb369/mcp-spine: Context Minifier & State Guard — Local-first MCP middleware proxy GitHub - Beledarian/wgpu-llm: A from-scratch LLM inference engine that uses wgpu (the cross-platform WebGPU implementation) to dispatch WGSL compute shaders for every math operation a Transformer needs. No CUDA. No Python. No massive framework dependencies. Just Rust, raw shaders, and your GPU. GitHub - anitiue/Hindsight: An experience-driven self-improvement framework for LLM agents — 基于经验的 LLM Agent 自我改进框架 GitHub - stef41/lmscan: 🔍 Detect AI-generated text and fingerprint which LLM wrote it. Open-source GPTZero alternative. Zero dependencies, works offline. GitHub - alainnothere/AmdPerformanceTesting: Amd Performance Testing Ask HN: Is a purely Markdown-based CRM a terrible idea? Optimized for LLM agents Context Engineering - LLM Memory and Retrieval for AI Agents | Weaviate little_helper_tui/letter.md at main · sleepyeldrazi/little_helper_tui GitHub - EvanZhouDev/umr: The Unified Model Registry for all your local AI apps. GitHub - JordanCT/VigIA-Orchestrator Your Agent Is Mine: Measuring Malicious Intermediary Attacks on the LLM Supply Chain A Taxonomy of RL Environments for LLM Agents Llama LLM Network Feture GitHub - genedeng-ca/ai-mac-migration: AI-powered Mac-to-Mac migration tool - replace Apple Migration Assistant with intelligent, selective transfer using local LLMs GitHub - lunargate-ai/gateway: High-performance self-hosted AI gateway (OpenAI-compatible) with routing, retries, and streaming GitHub - AuthBits/webmcp: A lightweight, prompt-driven MCP web research server for high-quality LLM powered information extraction. Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering Springdrift: An Auditable Persistent Runtime for LLM Agents with Case-Based Memory, Normative Safety, and Ambient Self-Perception High-Stakes Personalization: Rethinking LLM Customization for Individual Investor Decision-Making From Static Templates to Dynamic Runtime Graphs: A Survey of Workflow Optimization for LLM Agents HUOZIIME: An On-Device LLM-enhanced Input Method for Deep Personalization TIDE: Token-Informed Depth Execution for Per-Token Early Exit in LLM Inference Characterizing WebGPU Dispatch Overhead for LLM Inference Across Four GPU Vendors, Three Backends, and Three Browsers LLM Targeted Underperformance Disproportionately Impacts Vulnerable Users
Pushing the Pareto: SOTA LLM Judges
Tanmay Chopra · 2026-05-29 · via Hacker News - Newest: "LLM"

The LLM-as-Judge

An LLM-as-Judge is a language model used to evaluate the output of an AI system against a rubric. The judge consumes some combination of an input, a candidate output, and an evaluation criterion, and emits a verdict: a binary label, a preference between two candidates, a scalar score, or a natural-language critique.

In a world of open-ended outputs and infinite ways to arrive at them, it has become the backbone of evaluation - used in offline benchmarks, online monitoring, RLHF pipelines, and safety guardrails.

The standard approach involves prompting a frontier model with the input and parsing the verdict from the output. It's a quick but dirty way to keep AI in check - it applies universally and takes minutes to set up, but suffers from compounding uncertainty from stacked LLM calls, latency tails in double-digit seconds, and cost scaling linearly with every call. Teams are left torn between continuing manual evals and actually trusting their judges at scale.

Why It's Broken

The root cause is a structure mismatch. Judging tasks are classification and regression (discriminative) problems with small, closed label spaces. The standard implementation applies a generative model to this discriminative task - sampling from a 100,000-token vocabulary and collapsing back to two (or few) choices. This wasted computation and unnecessary noise makes generative judging expensive, slow, and unreliable.

A second, deeper problem is the objective mismatch. Prompting the same frontier LLM that generated the output with "Is this helpful?" asks it to answer from the same distribution, not from the knowledge of specific labelers whose judgments you're aiming to emulate. These distributions may overlap but aren't identical (hence the need for judgement), and no prompt, few-shot example, or reasoning chain can arrive at labeler-specific signal the model was never exposed to.

The ceiling is set by calibration, not capability, which is why performance across models clusters within the same window. And the initial structural mismatch prevents post-training from being meaningful because of the intermediate noise.

The Solution: Decision Language Models

At Emissary, we're committed to bringing the best of ML to AI. LLMs are quick, ML is reliable.

So, we replace the LLM's language modeling head with a discriminative head. The LM backbone supplies zero-shot generalization; the head supplies closed outputs mapping to the judgement task. Inference is a single forward pass - no decoding, no format parsing, no prompt sensitivity, and direct loss signal is easy to learn against. Easily calibrated, fast, cheap.

But this gives rise to a transition problem - an uncanny valley between 0-100 samples, where the frontier LLM is better than the Decision-LM. So we created Semantic Initialization - a way to seed Decision-LMs with the inherent knowledge underlying LLMs, making them as smart at zero-shot while providing latency and cost gains.

Semantic Initialization enables the custom heads of language models to extract informational state from base LLMs through logit distribution analysis. By examining the logit behavior of models across a distribution of prompts, we can set the head weights in a manner that enables the models to match the performance of their base counterparts, with no labelled data.

Unlike ML, AI workloads are incremental, so we created three more graduated learning modes - few-sample (5-50) head-only training, warm-start LoRA (~100 labels, regression-based initialization), and full LoRA (1,000+ labels) - that AI engineers can seamlessly transition across as they generate more labelled feedback. Going from decent to good to great, all in one place.

Why These Four Modes

The four modes are not arbitrary. They map to the four data regimes any team building an LLM-as-Judge will be in:

Data RegimeModeWhy
No labels yetZero-shotDrop-in, matches frontier accuracy
A handful of labelsLow-sample (head-only)Stable, fast, no backbone risk
~100 labelsWarm-start LoRAClosed-form init unlocks small-N
>=1000 labelsFull LoRA fine-tuneBreaks the prompting ceiling

The progression is smooth. A team can start at zero-shot Emissary on day one, collect 100 labels over a week, correct mistakes and move to warm-start, then graduate to full fine-tune as their label set grows — without changing inference infrastructure or output interface.

Experimental Setup

Datasets

We evaluate on two standard LLM-as-Judge benchmarks, each cast as a balanced binary classification task.

  • MT-Bench. We use the score:1/10 split, mapping the original 10-point scale to a binary helpful / not-helpful label. We sample 1,000 examples for training (500 positive / 500 negative), 100 for the low-sample regime (50/50), and 1,000 for testing (500/500). All splits are class-balanced.
  • UltraFeedback 64K. We use the score:1/5 split, mapping the original 5-point scale to binary. Sampling and class balance match MT-Bench: 1,000 train / 100 low-sample / 1,000 test, all 50/50.

This paper covers binary judgment: helpful / not-helpful, pass / fail, safe / unsafe, prefer-A / prefer-B. Binary judgment covers a large share of LLM-as-Judge use cases in practice, and it is the cleanest setting in which to demonstrate that classification beats generation for closed-set decisions.

We are concurrently developing an Emissary variant for scoring judges - scalar and ordinal scores rather than binary labels. Results from that work will be published separately.

Baselines

We use Qwen3-8B as our base model. We compare Emissary's DLM against two baselines:

  1. Base-model prompting. Zero-shot and k-shot prompting using the same base model that we apply our techniques to. This isolates the effect of the classification-first approach from the effect of model choice.
  2. Frontier LLM APIs. GPT-5.5, Sonnet 4.6, Opus 4.7, and Gemini 3 Flash. Each is evaluated at 0-shot, 1-shot, and 2-shot, with thinking modes both enabled and disabled where applicable. The prompts used are reproduced in Appendix B.

Metrics

  • Accuracy. Agreement with held-out human labels. Because the task is calibrated to a specific labeler population, accuracy here is the most direct measure of human alignment for the judge.
  • Latency. End-to-end wall-clock time per judgment, reported as P50 and P99 in milliseconds. For both API baselines and Emissary, this includes network round-trip.
  • Cost. US dollars per 1,000 judgments at published API pricing as of the experiment date. Emissary's per-call inference cost is dominated by amortized infrastructure rather than per-call billing.

Results

Zero-Shot Parity: Frontier Capability Is Not the Bottleneck

Emissary's zero-shot configuration (Qwen3-8B as a classifier) sits inside the same accuracy cluster as frontier APIs - at a fraction of the latency and cost. The plateau is the prompting ceiling, set by the interface, not the model. More parameters, reasoning compute, or a different provider doesn't move it.

MethodMT-BenchUltraFeedbackP50 (ms)Cost / 1K
Emissary zero-shot0.9060.83465$0.06
Sonnet 4.6 0-shot0.9100.8351,148$2.43
Opus 4.7 0-shot0.9070.8381,433$5.82
Gemini 3 Flash 0-shot0.9010.8121,127$0.38
GPT-5.5 0-shot0.8640.789902$3.67

Post-Training Breaks the Ceiling

MethodMT-BenchUltraFeedback
Emissary post-train cold-10000.9790.955
Emissary post-train warm-10000.9720.949
Best frontier (Opus 4.7 2-shot, think ON)0.9450.838

When labeled data is available, Emissary DLMs move well past the prompting cluster. The UltraFeedback gap (+11.7 points over the best frontier) is the more telling number - that rubric's labeler-specific signal is only recoverable through training.

Warm-Start Makes 100 Labels Sufficient & decreases in value with increase in labels.

MethodMT-BenchUltraFeedback
Emissary post-train cold-1000.8570.833
Emissary post-train warm-1000.9300.889

DLM warm-start at 100 examples already exceeds every frontier API on UltraFeedback, collapsing the labeled-data requirement from "thousands" to "low hundreds" - achievable in a single labeling session. At 1000 samples (see section above), the value of warm starting starts to become less predictable.

Reasoning Compute Does Not Reliably Help

Thinking mode produced inconsistent, often negligible accuracy changes while multiplying latency 2-10x and cost 1.3-4x. GPT-5.5 thinking-ON scored lower than thinking-OFF across all shot counts. The pattern holds for both frontier APIs and Qwen3-8B generation.

MT-Bench - thinking ON vs. OFF (representative rows)

ModelThinkingAccuracyP50 (ms)P99 (ms)Cost / 1K
GPT-5.5 1-shotOFF0.8729061,850$7.81
GPT-5.5 1-shotON0.8622,35512,301$12.01
Sonnet 4.6 1-shotOFF0.9361,2314,520$5.20
Sonnet 4.6 1-shotON0.9281,21918,235$5.87
Opus 4.7 2-shotOFF0.9331,5515,336$16.65
Opus 4.7 2-shotON0.9451,7406,951$16.87
Qwen3-8B gen 1-shotOFF0.943284380$0.24
Qwen3-8B gen 1-shotON0.9483,89125,555$4.46

The Pareto Picture

Emissary configurations Pareto-dominate every frontier row: equal or better accuracy, ~15-25x lower latency, and effectively 100x lower cost. The Pareto frontier on this problem is occupied by purpose-built classifiers, not frontier LLMs.

ConfigurationAccuracyP50 (ms)Cost / 1K
Emissary post-train cold-10000.97965$0.06
Emissary post-train warm-1000.93065$0.06
Emissary zero-shot0.90665$0.06
Opus 4.7 2-shot, think ON0.9451,740$16.87
Sonnet 4.6 1-shot, think OFF0.9361,231$5.20
Sonnet 4.6 0-shot, think OFF0.9101,148$2.43
GPT-5.5 0-shot, think OFF0.864902$3.67

Conclusion

The diagnosis predicted three things, all confirmed: a flat accuracy cluster across frontier models, a ceiling break when trained on actual labels, and thinking modes failing to help. All three appear cleanly in the data.

Pareto dominance follows from using the right tool for the job. Accuracy improves by training on the exact decision boundary. Latency drops because a single forward pass replaces autoregressive decoding. Cost falls because inference runs on owned hardware rather than per-token billing. None of these gains require novel research - they require taking the structure of the problem seriously.

The standard advice - "use the strongest frontier model you can afford" - is wrong for closed-set judgment tasks. The binding constraint is calibration to labelers, not model capability, and calibration requires effective training, not scaling. Every additional dollar spent on a larger frontier model, a longer reasoning chain, or a more elaborate prompt is a dollar spent moving along a ceiling rather than through it. The teams that recognize this early will ship faster, evaluate more, and trust their pipelines more than the teams that don't.

The implications go beyond cost savings. A judge that runs in 65ms instead of 1,700ms is no longer a batch-time artifact - it becomes a real-time component you can put in the inference path, in safety guardrails, in router logic, in online RLHF loops. A judge that costs $0.06 per thousand calls instead of $16.87 makes 100x more evaluation economically viable, which changes what teams can measure and how often. Quality stops being something you sample and starts being something you observe continuously.

Get Started with Emissary

If you're running LLM judges in production today - for evals, monitoring, guardrails, or RLHF - you're almost certainly on the wrong side of the Pareto frontier. We'd love to help you move.

  • Try Emissary on your own data. Bring a metric + criteria and optionally, a few samples; we'll show you the zero-shot, warm-start, and full fine-tune numbers side-by-side against whatever frontier baseline you're using now. Or you can try it out here yourself: withemissary.com/demo
  • Talk to us about scoring judges. If your use case needs scalar or ordinal scores rather than binary labels, our forthcoming variant is in active development and we're onboarding design partners.

Reach out to us to book a technical deep-dive. The ceiling is real, but it isn't yours to live under.

Appendix A: Full Results

A.1 MT-Bench - Frontier APIs

ModelShotsThinkingAccuracyP50 (ms)P99 (ms)Cost / 1K
GPT-5.50OFF0.8649022,765$3.67
GPT-5.50ON0.8562,34412,604$7.80
GPT-5.51OFF0.8729061,850$7.81
GPT-5.51ON0.8622,35512,301$12.01
GPT-5.52OFF0.8709002,533$10.60
GPT-5.52ON0.8612,42115,116$15.04
Sonnet 4.60OFF0.9101,1484,244$2.43
Sonnet 4.60ON0.9061,22316,426$3.20
Sonnet 4.61OFF0.9361,2314,520$5.20
Sonnet 4.61ON0.9281,21918,235$5.87
Sonnet 4.62OFF0.9271,2813,745$7.32
Sonnet 4.62ON0.9261,26318,128$8.03
Opus 4.70OFF0.9071,4335,358$5.82
Opus 4.70ON0.9121,6118,928$6.18
Opus 4.71OFF0.9381,5195,253$11.68
Opus 4.71ON0.9361,6687,032$11.90
Opus 4.72OFF0.9331,5515,336$16.65
Opus 4.72ON0.9451,7406,951$16.87
Gemini 3 Flash0OFF0.9011,1273,894$0.38
Gemini 3 Flash0ON0.9022,62310,141$1.54
Gemini 3 Flash1OFF0.8871,1264,306$0.82
Gemini 3 Flash1ON0.9082,67212,636$1.99
Gemini 3 Flash2OFF0.8981,1794,868$1.03
Gemini 3 Flash2ON0.9142,40313,157$2.09

A.2 MT-Bench - Qwen3-8B Generation Baseline

ShotsThinkingAccuracyP50 (ms)P99 (ms)Cost / 1K
0OFF0.937284.27384.34$0.24
0ON0.93783,577.0320,394.03$4.07
1OFF0.943283.95380.34$0.24
1ON0.94783,891.2025,555.16$4.46
2OFF0.940285.89385.07$0.24
2ON0.94473,984.7929,678.37$4.66

A.3 MT-Bench - Emissary Decision-LM (Qwen3-8B backbone)

ModeTrain labelsAccuracyP50 (ms)P99 (ms)Cost / 1K
Zero-shot (Semantic Init)00.90665.41124.44$0.06
5-shot00.767---
Cold LoRA (5 ep)1000.85753.67136.87$0.05
Warm-start LoRA (5 ep)1000.930---
Cold LoRA (4 ep)1,0000.979---
Warm-start LoRA (5 ep)1,0000.972---

A.4 UltraFeedback - Frontier APIs

ModelAccuracy
GPT-5.50.789
Sonnet 4.60.835
Opus 4.70.838
Gemini 3 Flash0.8118

A.5 UltraFeedback - Qwen3-8B Generation Baseline

MethodAccuracy
Qwen3-8B generation (no-think)0.884
Qwen3-8B generation (think)0.8855

A.6 UltraFeedback - Emissary Decision-LM (Qwen3-8B backbone)

ModeTrain labelsAccuracy
Zero-shot (Semantic Init)00.834
5-shot00.8607
Cold LoRA (5 ep)1000.833
Warm-start LoRA (5 ep)1000.889
Cold LoRA (4 ep)1,0000.955
Warm-start LoRA (5 ep)1,0000.949

Appendix B: Prompts Used for Frontier API Baselines

B.1 MT-Bench Prompt

System:

1You are an impartial MT-Bench-style judge. Decide whether the candidate answer is helpful for the latest user request.
2
3Use MT-Bench-style response-quality factors: helpfulness, relevance, accuracy, depth, appropriate creativity, and level of detail.
4
5Return HELPFUL if the candidate answer substantially satisfies the latest user request, stays on task, is factually and logically correct, and provides enough useful detail for the task. Minor wording or style issues are acceptable only when they do not affect usefulness.
6
7Return NOT_HELPFUL if the candidate answer fails a key requirement, is incorrect, irrelevant, evasive, refuses without good reason, contradicts the provided context, or is too incomplete or shallow to be useful.
8
9For math, coding, reasoning, and factual tasks, prioritize correctness and whether the final answer solves the requested problem. If previous conversation context is provided, use it only to understand the latest user request; judge the candidate answer itself.
10
11Output exactly one token: HELPFUL or NOT_HELPFUL. No punctuation, no explanation.

User:

1Task / conversation context:
2Previous conversation:
3User: ...
4Assistant: ...
5
6Latest user request:
7...
8
9Candidate answer:
10<candidate answer>
11
12Final verdict (HELPFUL or NOT_HELPFUL):

Few-shot examples are appended in the standard Example N: ... Final verdict: ... format before the test item.

B.2 UltraFeedback Prompt

System:

1You are evaluating whether an AI assistant's response is helpful.
2
3A response is HELPFUL only when it meets ALL of these criteria:
41. Clarity and Relevance: Does the response directly address the task and remain on-topic?
52. Useful and Comprehensive Information: Does the response provide relevant background, reasoning, or detailed explanation that improves understanding?
63. Not Lengthy, No Repetition: Is the response concise without unnecessary repetition while still being comprehensive?
7
8If the response clearly satisfies all three, answer HELPFUL.
9If it clearly fails on any one - off-topic, vague/incomplete, or bloated/repetitive - answer NOT_HELPFUL.
10
11Output exactly one word: HELPFUL or NOT_HELPFUL. No punctuation, no explanation.

User:

1Evaluate the following instruction-response pair against the helpfulness criteria.
2
3Instruction:
4{question}
5
6Response:
7{answer}
8
9Verdict (HELPFUL or NOT_HELPFUL):