惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Schneier on Security
Schneier on Security
N
Netflix TechBlog - Medium
IT之家
IT之家
MongoDB | Blog
MongoDB | Blog
博客园_首页
S
SegmentFault 最新的问题
H
Help Net Security
P
Proofpoint News Feed
云风的 BLOG
云风的 BLOG
T
The Blog of Author Tim Ferriss
量子位
GbyAI
GbyAI
M
MIT News - Artificial intelligence
Recorded Future
Recorded Future
P
Privacy & Cybersecurity Law Blog
B
Blog
月光博客
月光博客
博客园 - 聂微东
Vercel News
Vercel News
罗磊的独立博客
腾讯CDC
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
A
Arctic Wolf
D
Darknet – Hacking Tools, Hacker News & Cyber Security
Stack Overflow Blog
Stack Overflow Blog
T
Threat Research - Cisco Blogs
Blog — PlanetScale
Blog — PlanetScale
L
Lohrmann on Cybersecurity
I
Intezer
小众软件
小众软件
T
The Exploit Database - CXSecurity.com
Jina AI
Jina AI
C
Check Point Blog
AWS News Blog
AWS News Blog
C
Cisco Blogs
Martin Fowler
Martin Fowler
The Last Watchdog
The Last Watchdog
www.infosecurity-magazine.com
www.infosecurity-magazine.com
宝玉的分享
宝玉的分享
S
Security Affairs
大猫的无限游戏
大猫的无限游戏
N
News and Events Feed by Topic
雷峰网
雷峰网
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
H
Hacker News: Front Page
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
F
Full Disclosure
P
Proofpoint News Feed
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
Microsoft Security Blog
Microsoft Security Blog

OpenAI Developers

API deployment checklist | OpenAI API Sora 2 Prompting Guide Codex Prompting Guide Docs MCP | OpenAI Developers Gpt-image-1.5 Prompting Guide GPT-5.2 Prompting Guide Transcribing User Audio with a Separate Realtime Request Modernizing your Codebase with Codex GitHub - openai/openai-sora-sample-app: Sample app to get started using the Video API with Sora GitHub - openai/openai-apps-sdk-examples: Example apps for the Apps SDK GitHub - openai/openai-chatkit-advanced-samples: Starter app to build with OpenAI ChatKit SDK GitHub - openai/openai-chatkit-starter-app: Starter app to build with OpenAI ChatKit + Agent Builder Rate limits | OpenAI API Web search | OpenAI API Getting started with datasets | OpenAI API Prompt optimizer | OpenAI API Verifying gpt-oss implementations How to run gpt-oss locally with LM Studio Fine-tuning with gpt-oss and Hugging Face Transformers How to run gpt-oss locally with Ollama Function calling | OpenAI API Models | OpenAI API Reasoning best practices | OpenAI API Reasoning models | OpenAI API Background mode | OpenAI API Batch API | OpenAI API Conversation state | OpenAI API File search | OpenAI API Flex processing | OpenAI API MCP and Connectors | OpenAI API Code Interpreter | OpenAI API Quickstart - OpenAI Agents SDK Build Hour: Agentic Tool Calling Build Hour: Built-In Tools Reasoning best practices | OpenAI API Graders | OpenAI API Working with evals | OpenAI API Guardrails - OpenAI Agents SDK Latency optimization | OpenAI API Optimizing LLM Accuracy | OpenAI API Agent orchestration - OpenAI Agents SDK Production best practices | OpenAI API Realtime transcription | OpenAI API Optimizing LLM Accuracy | OpenAI API Realtime and audio | OpenAI API Realtime conversations | OpenAI API Responses guide Migrate to the Responses API | OpenAI API Speech to text | OpenAI API Supervised fine-tuning | OpenAI API Tracing - OpenAI Agents SDK Vision fine-tuning | OpenAI API Audio and speech | OpenAI API GitHub - openai/openai-cs-agents-demo: Demo of a customer service use case implemented with the OpenAI Agents SDK Voice agents | OpenAI API Fine-tuning best practices | OpenAI API GitHub - openai/openai-agents-python: A lightweight, powerful framework for multi-agent workflows GitHub - openai/openai-agents-js: A lightweight, powerful framework for multi-agent workflows and voice agents Agents SDK | OpenAI API Using tools | OpenAI API Computer use | OpenAI API GitHub - openai/openai-cua-sample-app: Learn how to use CUA (our Computer Using Agent) via the API on multiple computer environments. GitHub - openai/openai-testing-agent-demo: Demo of a UI testing agent using the OpenAI CUA model and the Responses API. Model optimization | OpenAI API GitHub - openai/openai-fm: Code for openai.fm, a demo for the OpenAI Speech API Predicted Outputs | OpenAI API GitHub - openai/openai-realtime-console: React app for inspecting, building and debugging with the Realtime API Building Voice Agents GitHub - openai/openai-realtime-solar-system: Demo showing how to use the OpenAI Realtime API to navigate a 3D scene via tool calling GitHub - openai/openai-realtime-twilio-demo Reinforcement fine-tuning | OpenAI API GitHub - openai/openai-responses-starter-app: Starter app to build with the OpenAI Responses API Structured model outputs | OpenAI API GitHub - openai/openai-structured-outputs-samples: Sample apps to help developers get started with Structured Outputs Voice agents | OpenAI API Model optimization | OpenAI API GitHub - openai/openai-realtime-agents: This is a simple demonstration of more advanced, agentic patterns built on top of the Realtime API. GitHub - openai/openai-support-agent-demo: Demo of a customer support agent interface using NextJS and the OpenAI Responses API with File Search Building Voice Agents Generate images with high input fidelity AI app development: Concept to production Model optimization Building agents Eval Driven System Design - From Prototype to Production Multi-Agent Portfolio Collaboration with OpenAI Agents SDK o3/o4-mini Function Calling Guide Exploring Model Graders for Reinforcement Fine-Tuning Guide to Using the Responses API Reinforcement Fine-Tuning for Conversational Reasoning with the OpenAI API Evals API Use-case - Responses Evaluation Comparing Speech-to-Text Methods with the OpenAI API Generate images with GPT Image Multi-Tool Orchestration with RAG approach using OpenAI Multi-Language One-Way Translation with the Realtime API Doing RAG on PDFs using File Search in the Responses API How to use the Usage API and Cost API to monitor your OpenAI usage Leveraging model distillation to fine-tune a model Orchestrating Agents: Routines and Handoffs Prompt Caching 101 Developing Hallucination Guardrails
Evaluation best practices | OpenAI API
2025-07-21 · via OpenAI Developers

Generative AI is variable. Models sometimes produce different output from the same input, which makes traditional software testing methods insufficient for AI architectures. Evaluations (evals) are a way to test your AI system despite this variability.

This guide provides high-level guidance on designing evals. To get started with the Evals API, see evaluating model performance.

OpenAI is deprecating the Evals platform. Existing evals content remains available during the transition window. Evals will become read-only for existing users on October 31, 2026, and the platform is scheduled to shut down on November 30, 2026. See the deprecations page for the current timeline.

Evals are structured tests for measuring a model’s performance. They help ensure accuracy, performance, and reliability, despite the nondeterministic nature of AI systems. They’re also one of the only ways to improve performance of an LLM-based application (through fine-tuning).

Types of evals

When you see the word “evals,” it could refer to a few things:

  • Industry benchmarks for comparing models in isolation, like MMLU and those listed on HuggingFace’s leaderboard
  • Standard numerical scores—like ROUGE, BERTScore—that you can use as you design evals for your use case
  • Specific tests you implement to measure your LLM application’s performance

This guide is about the third type: designing your own evals.

How to read evals

You’ll often see numerical eval scores between 0 and 1. There’s more to evals than just scores. Combine metrics with human judgment to ensure you’re answering the right questions.

Evals tips


  • Adopt eval-driven development: Evaluate early and often. Write scoped tests at every stage.
  • Design task-specific evals: Make tests reflect model capability in real-world distributions.
  • Log everything: Log as you develop so you can mine your logs for good eval cases.
  • Automate when possible: Structure evaluations to allow for automated scoring.
  • It’s a journey, not a destination: Evaluation is a continuous process.
  • Maintain agreement: Use human feedback to calibrate automated scoring.

Anti-patterns


  • Overly generic metrics: Relying solely on academic metrics like perplexity or BLEU score.
  • Biased design: Creating eval datasets that don’t faithfully reproduce production traffic patterns.
  • Vibe-based evals: Using “it seems like it’s working” as an evaluation strategy, or waiting until you ship before implementing any evals.
  • Ignoring human feedback: Not calibrating your automated metrics against human evals.

There are a few important components of an eval workflow:

  1. Define eval objective. What’s the success criteria for the eval?
  2. Collect dataset. Which data will help you evaluate against your objective? Consider synthetic eval data, domain-specific eval data, purchased eval data, human-curated eval data, production data, and historical data.
  3. Define eval metrics. How will you check that the success criteria are met?
  4. Run and compare evals. Iterate and improve model performance for your task or system.
  5. Continuously evaluate. Set up continuous evaluation (CE) to run evals on every change, monitor your app to identify new cases of nondeterminism, and grow the eval set over time.

Let’s run through a few examples.

Example: Summarizing transcripts

To test your LLM-based application’s ability to summarize transcripts, your eval design might be:

  1. Define eval objective
    The model should be able to compete with reference summaries for relevance and accuracy.
  2. Collect dataset
    Use a mix of production data (collected from user feedback on generated summaries) and datasets created by domain experts (writers) to determine a “good” summary.
  3. Define eval metrics
    On a held-out set of 1000 reference transcripts → summaries, the implementation should achieve a ROUGE-L score of at least 0.40 and coherence score of at least 80% using G-Eval.
  4. Run and compare evals
    Use the Evals API to create and run evals in the OpenAI dashboard.
  5. Continuously evaluate
    Set up continuous evaluation (CE) to run evals on every change, monitor your app to identify new cases of nondeterminism, and grow the eval set over time.

LLMs are better at discriminating between options. Therefore, evaluations should focus on tasks like pairwise comparisons, classification, or scoring against specific criteria instead of open-ended generation. Aligning evaluation methods with LLMs’ strengths in comparison leads to more reliable assessments of LLM outputs or model comparisons.

Example: Q&A over docs

To test your LLM-based application’s ability to do Q&A over docs, your eval design might be:

  1. Define eval objective
    The model should be able to provide precise answers, recall context as needed to reason through user prompts, and provide an answer that satisfies the user’s need.
  2. Collect dataset
    Use a mix of production data (collected from users’ satisfaction with answers provided to their questions), hard-coded correct answers to questions created by domain experts, and historical data from logs.
  3. Define eval metrics
    Context recall of at least 0.85, context precision of over 0.7, and 70+% positively rated answers.
  4. Run and compare evals
    Use the Evals API to create and run evals in the OpenAI dashboard.
  5. Continuously evaluate
    Set up continuous evaluation (CE) to run evals on every change, monitor your app to identify new cases of nondeterminism, and grow the eval set over time.

When creating an eval dataset,

gpt-5.5

is useful for collecting eval examples and edge cases. Consider using it to help you generate a diverse set of test data across various scenarios. Ensure your test data includes typical cases, edge cases, and adversarial cases. Use human expert labellers.

Complexity increases as you move from simple to more complex architectures. Here are four common architecture patterns:

  • Single-turn model interactions
  • Workflows
  • Single-agent
  • Multi-agent

Read about each architecture below to identify where nondeterminism enters your system. That’s where you’ll want to implement evals.

Single-turn model interactions

In this kind of architecture, the user provides input to the model, and the model processes these inputs (along with any developer prompts provided) to generate a corresponding output.

Example

As an example, consider an online retail scenario. Your system prompt instructs the model to categorize the customer’s question into one of the following:

  • order_status
  • return_policy
  • technical_issue
  • cancel_order
  • other

To ensure a consistent, efficient user experience, the model should only return the label that matches user intent. Let’s say the customer asks, “What’s the status of my order?”

Nondeterminism introducedCorresponding area to evaluateExample eval questions
Inputs provided by the developer and user

Instruction following: Does the model accurately understand and act according to the provided instructions?

Instruction following: Does the model prioritize the system prompt over a conflicting user prompt?

Does the model stay focused on the triage task or get swayed by the user’s question?

Outputs generated by the model

Functional correctness: Are the model’s outputs accurate, relevant, and thorough enough to fulfill the intended task or objective?

Does the model’s determination of intent correctly match the expected intent?

Workflow architectures

As you look to solve more complex problems, you’ll likely transition from a single-turn model interaction to a multistep workflow that chains together several model calls. Workflows don’t introduce any new elements of nondeterminism, but they involve multiple underlying model interactions, which you can evaluate in isolation.

Example

Take the same example as before, where the customer asks about their order status. A workflow architecture triages the customer request and routes it through a step-by-step process:

  1. Extracting an Order ID
  2. Looking up the order details
  3. Providing the order details to a model for a final response

Each step in this workflow has its own system prompt that the model must follow, putting all fetched data into a friendly output.

Nondeterminism introducedCorresponding area to evaluateExample eval questions
Inputs provided by the developer and user

Instruction following: Does the model accurately understand and act according to the provided instructions?

Instruction following: Does the model prioritize the system prompt over a conflicting user prompt?

Does the model stay focused on the triage task or get swayed by the user’s question?


Does the model follow instructions to attempt to extract an Order ID?

Does the final response include the order status, estimated arrival date, and tracking number?

Outputs generated by the model

Functional correctness: Are the model’s outputs are accurate, relevant, and thorough enough to fulfill the intended task or objective?

Does the model’s determination of intent correctly match the expected intent?

Does the final response have the correct order status, estimated arrival date, and tracking number?

Single-agent architectures

Unlike workflows, agents solve unstructured problems that require flexible decision making. An agent has instructions and a set of tools and dynamically selects which tool to use. This introduces a new opportunity for nondeterminism.

Tools are developer defined chunks of code that the model can execute. This can range from small helper functions to API calls for existing services. For example, check_order_status(order_id) could be a tool, where it takes the argument order_id and calls an API to check the order status.

Example

Let’s adapt our customer service example to use a single agent. The agent has access to three distinct tools:

  • Order lookup tool
  • Password reset tool
  • Product FAQ tool

When the customer asks about their order status, the agent dynamically decides to either invoke a tool or respond to the customer. For example, if the customer asks, “What is my order status?” the agent can now follow up by requesting the order ID from the customer. This helps create a more natural user experience.

NondeterminismCorresponding area to evaluateExample eval questions
Inputs provided by the developer and user

Instruction following: Does the model accurately understand and act according to the provided instructions?

Instruction following: Does the model prioritize the system prompt over a conflicting user prompt?

Does the model stay focused on the triage task or get swayed by the user’s question?

Does the model follow instructions to attempt to extract an Order ID?

Outputs generated by the model

Functional correctness: Are the model’s outputs are accurate, relevant, and thorough enough to fulfill the intended task or objective?

Does the model’s determination of intent correctly match the expected intent?

Tools chosen by the model

Tool selection: Evaluations that test whether the agent is able to select the correct tool to use.

Data precision: Evaluations that verify the agent calls the tool with the correct arguments. Typically these arguments are extracted from the conversation history, so the goal is to validate this extraction was correct.

When the user asks about their order status, does the model correctly recommend invoking the order lookup tool?

Does the model correctly extract the user-provided order ID to the lookup tool?

Multi-agent architectures

As you add tools and tasks to your single-agent architecture, the model may struggle to follow instructions or select the correct tool to call. Multi-agent architectures help by creating several distinct agents who specialize in different areas. This triaging and handoff among multiple agents introduces a new opportunity for nondeterminism.

The decision to use a multi-agent architecture should be driven by your evals. Starting with a multi-agent architecture adds unnecessary complexity that can slow down your time to production.

Example

Splitting the single-agent example into a multi-agent architecture, we’ll have four distinct agents:

  1. Triage agent
  2. Order agent
  3. Account management agent
  4. Sales agent

When the customer asks about their order status, the triage agent may hand off the conversation to the order agent to look up the order. If the customer changes the topic to ask about a product, the order agent should hand the request back to the triage agent, who then hands off to the sales agent to fetch product information.

NondeterminismCorresponding area to evaluateExample eval questions
Inputs provided by the developer and userInstruction following: Does the model accurately understand and act according to the provided instructions?

Instruction following: Does the model prioritize the system prompt over a conflicting user prompt?

Does the model stay focused on the triage task or get swayed by the user’s question?

Assuming the lookup_order call returned, does the order agent return a tracking number and delivery date (doesn’t have to be the correct one)?

Outputs generated by the modelFunctional correctness: Are the model’s outputs are accurate, relevant, and thorough enough to fulfill the intended task or objective?Does the model’s determination of intent correctly match the expected intent?

Assuming the lookup_order call returned, does the order agent provide the correct tracking number and delivery date in its response?

Does the order agent follow system instructions to ask the customer their reason for requesting a return before processing the return?

Tools chosen by the modelTool selection: Evaluations that test whether the agent is able to select the correct tool to use.

Data precision: Evaluations that verify the agent calls the tool with the correct arguments. Typically these arguments are extracted from the conversation history, so the goal is to validate this extraction was correct.

Does the order agent correctly call the lookup order tool?

Does the order agent correctly call the refund_order tool?

Does the order agent call the lookup order tool with the correct order ID?

Does the account agent correctly call the reset_password tool with the correct account ID?

Agent handoffAgent handoff accuracy: Evaluations that test whether each agent can appropriately recognize the decision boundary for triaging to another agentWhen a user asks about order status, does the triage agent correctly pass to the order agent?

When the user changes the subject to talk about the latest product, does the order agent hand back control to the triage agent?

As you design your own evals, there are several specific evaluator types to choose from. Another way to think about this is what role you want the evaluator to play.

Metric-based evals

Quantitative evals provide a numerical score you can use to filter and rank results. They provide useful benchmarks for automated regression testing.

  • Examples: Exact match, string match, ROUGE/BLEU scoring, function call accuracy, executable evals (executed to assess functionality or behavior—e.g., text2sql)
  • Challenges: May not be tailored to specific use cases, may miss nuance

Human evals

Human judgment evals provide the highest quality but are slow and expensive.

  • Examples: Skim over system outputs to get a sense of whether they look better or worse; create a randomized, blinded test in which employees, contractors, or outsourced labeling agencies judge the quality of system outputs (e.g., ranking a small set of possible outputs, or giving each a grade of 1-5)
  • Challenges: Disagreement among human experts, expensive, slow
  • Recommendations:
    • Conduct multiple rounds of detailed human review to refine the scorecard
    • Implement a “show rather than tell” policy by providing examples of different score levels (e.g., 1, 3, and 8 out of 10)
    • Include a pass/fail threshold in addition to the numerical score
    • A simple way to aggregate multiple reviewers is to take consensus votes

LLM-as-a-judge and model graders

Using models to judge output is cheaper to run and more scalable than human evaluation. Start with gpt-5.5 when you need a strong LLM judge, then validate agreement against your human labels before optimizing for cost or latency.

  • Examples:
    • Pairwise comparison: Present the judge model with two responses and ask it to determine which one is better based on specific criteria
    • Single answer grading: The judge model evaluates a single response in isolation, assigning a score or rating based on predefined quality metrics
    • Reference-guided grading: Provide the judge model with a reference or “gold standard” answer, which it uses as a benchmark to evaluate the given response
  • Challenges: Position bias (response order), verbosity bias (preferring longer responses)
  • Recommendations:
    • Use pairwise comparison or pass/fail for more reliability
    • Use the most capable model to grade if you can. Start with gpt-5.5, then validate whether a specialized reasoning model performs better for your rubric or reference-answer set
    • Control for response lengths as LLMs bias towards longer responses in general
    • Add reasoning and chain-of-thought as reasoning before scoring improves eval performance
    • Once the LLM judge reaches a point where it’s faster, cheaper, and consistently agrees with human annotations, scale up
    • Structure questions to allow for automated grading while maintaining the integrity of the task—a common approach is to reformat questions into multiple choice formats
    • Ensure eval rubrics are clear and detailed

No strategy is perfect. The quality of LLM-as-Judge varies depending on problem context while using expert human annotators to provide ground-truth labels is expensive and time-consuming.

While your evaluations should cover primary, happy-path scenarios for each architecture, real-world AI systems frequently encounter edge cases that challenge system performance. Evaluating these edge cases is important for ensuring reliability and a good user experience.

We see these edge cases fall into a few buckets:

Input variability

Because users provide input to the model, our system must be flexible to handle the different ways our users may interact, like:

  • Non-English or multilingual inputs
  • Formats other than input text (e.g., XML, JSON, Markdown, CSV)
  • Input modalities (e.g., images)

Your evals for instruction following and functional correctness need to accommodate inputs that users might try.

Contextual complexity

Many LLM-based applications fail due to poor understanding of the context of the request. This context could be from the user or noise in the past conversation history.

Examples include:

  • Multiple questions or intents in a single request
  • Typos and misspellings
  • Short requests with minimal context (e.g., if a user just says: “returns”)
  • Long context or long-running conversations
  • Tool calls that return data with ambiguous property names (e.g., "on: 123", where “on” is the order number)
  • Multiple tool calls, sometimes leading to incorrect arguments
  • Multiple agent handoffs, sometimes leading to circular handoffs

Personalization and customization

While AI improves UX by adapting to user-specific requests, this flexibility introduces many edge cases. Clearly define evals for use cases you want to specifically support and block:

  • Jailbreak attempts to get the model to do something different
  • Formatting requests (e.g., format as JSON, or use bullet points)
  • Cases where user prompts conflict with your system prompts

When your evals reach a level of maturity that consistently measures performance, shift to using your evals data to improve your application’s performance.

Learn more about reinforcement fine-tuning to create a data flywheel.

For more inspiration, visit the OpenAI Cookbook, which contains example code and links to third-party resources, or learn more about our tools for evals: