惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

S
Schneier on Security
Security Archives - TechRepublic
Security Archives - TechRepublic
T
Threat Research - Cisco Blogs
G
GRAHAM CLULEY
P
Privacy & Cybersecurity Law Blog
C
CXSECURITY Database RSS Feed - CXSecurity.com
Cisco Talos Blog
Cisco Talos Blog
The Hacker News
The Hacker News
L
Lohrmann on Cybersecurity
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
C
Cyber Attacks, Cyber Crime and Cyber Security
Security Latest
Security Latest
Know Your Adversary
Know Your Adversary
P
Palo Alto Networks Blog
C
Cisco Blogs
AWS News Blog
AWS News Blog
T
Threatpost
L
LINUX DO - 热门话题
Simon Willison's Weblog
Simon Willison's Weblog
Scott Helme
Scott Helme
C
Cybersecurity and Infrastructure Security Agency CISA
T
Tor Project blog
Cyberwarzone
Cyberwarzone
P
Proofpoint News Feed
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
T
The Exploit Database - CXSecurity.com
The Register - Security
The Register - Security
D
Darknet – Hacking Tools, Hacker News & Cyber Security
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
罗磊的独立博客
云风的 BLOG
云风的 BLOG
V
Vulnerabilities – Threatpost
N
News | PayPal Newsroom
Project Zero
Project Zero
NISL@THU
NISL@THU
博客园_首页
MyScale Blog
MyScale Blog
V2EX - 技术
V2EX - 技术
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
F
Full Disclosure
T
Troy Hunt's Blog
Recorded Future
Recorded Future
N
Netflix TechBlog - Medium
P
Privacy International News Feed
H
Hackread – Cybersecurity News, Data Breaches, AI and More
A
Arctic Wolf
C
Check Point Blog
W
WeLiveSecurity
Apple Machine Learning Research
Apple Machine Learning Research
C
CERT Recently Published Vulnerability Notes

OpenAI Developers

API deployment checklist | OpenAI API Sora 2 Prompting Guide Codex Prompting Guide Docs MCP | OpenAI Developers Gpt-image-1.5 Prompting Guide GPT-5.2 Prompting Guide Transcribing User Audio with a Separate Realtime Request Modernizing your Codebase with Codex GitHub - openai/openai-sora-sample-app: Sample app to get started using the Video API with Sora GitHub - openai/openai-apps-sdk-examples: Example apps for the Apps SDK GitHub - openai/openai-chatkit-advanced-samples: Starter app to build with OpenAI ChatKit SDK GitHub - openai/openai-chatkit-starter-app: Starter app to build with OpenAI ChatKit + Agent Builder Rate limits | OpenAI API Web search | OpenAI API Prompt optimizer | OpenAI API Verifying gpt-oss implementations How to run gpt-oss locally with LM Studio Fine-tuning with gpt-oss and Hugging Face Transformers How to run gpt-oss locally with Ollama Function calling | OpenAI API Models | OpenAI API Reasoning best practices | OpenAI API Reasoning models | OpenAI API Background mode | OpenAI API Batch API | OpenAI API Conversation state | OpenAI API File search | OpenAI API Flex processing | OpenAI API MCP and Connectors | OpenAI API Code Interpreter | OpenAI API Quickstart - OpenAI Agents SDK Build Hour: Agentic Tool Calling Build Hour: Built-In Tools Reasoning best practices | OpenAI API Graders | OpenAI API Evaluation best practices | OpenAI API Working with evals | OpenAI API Guardrails - OpenAI Agents SDK Latency optimization | OpenAI API Optimizing LLM Accuracy | OpenAI API Agent orchestration - OpenAI Agents SDK Production best practices | OpenAI API Realtime transcription | OpenAI API Optimizing LLM Accuracy | OpenAI API Realtime and audio | OpenAI API Realtime conversations | OpenAI API Responses guide Migrate to the Responses API | OpenAI API Speech to text | OpenAI API Supervised fine-tuning | OpenAI API Tracing - OpenAI Agents SDK Vision fine-tuning | OpenAI API Audio and speech | OpenAI API GitHub - openai/openai-cs-agents-demo: Demo of a customer service use case implemented with the OpenAI Agents SDK Voice agents | OpenAI API Fine-tuning best practices | OpenAI API GitHub - openai/openai-agents-python: A lightweight, powerful framework for multi-agent workflows GitHub - openai/openai-agents-js: A lightweight, powerful framework for multi-agent workflows and voice agents Agents SDK | OpenAI API Using tools | OpenAI API Computer use | OpenAI API GitHub - openai/openai-cua-sample-app: Learn how to use CUA (our Computer Using Agent) via the API on multiple computer environments. GitHub - openai/openai-testing-agent-demo: Demo of a UI testing agent using the OpenAI CUA model and the Responses API. Model optimization | OpenAI API GitHub - openai/openai-fm: Code for openai.fm, a demo for the OpenAI Speech API Predicted Outputs | OpenAI API GitHub - openai/openai-realtime-console: React app for inspecting, building and debugging with the Realtime API Building Voice Agents GitHub - openai/openai-realtime-solar-system: Demo showing how to use the OpenAI Realtime API to navigate a 3D scene via tool calling GitHub - openai/openai-realtime-twilio-demo Reinforcement fine-tuning | OpenAI API GitHub - openai/openai-responses-starter-app: Starter app to build with the OpenAI Responses API Structured model outputs | OpenAI API GitHub - openai/openai-structured-outputs-samples: Sample apps to help developers get started with Structured Outputs Voice agents | OpenAI API Model optimization | OpenAI API GitHub - openai/openai-realtime-agents: This is a simple demonstration of more advanced, agentic patterns built on top of the Realtime API. GitHub - openai/openai-support-agent-demo: Demo of a customer support agent interface using NextJS and the OpenAI Responses API with File Search Building Voice Agents Generate images with high input fidelity AI app development: Concept to production Model optimization Building agents Eval Driven System Design - From Prototype to Production Multi-Agent Portfolio Collaboration with OpenAI Agents SDK o3/o4-mini Function Calling Guide Exploring Model Graders for Reinforcement Fine-Tuning Guide to Using the Responses API Reinforcement Fine-Tuning for Conversational Reasoning with the OpenAI API Evals API Use-case - Responses Evaluation Comparing Speech-to-Text Methods with the OpenAI API Generate images with GPT Image Multi-Tool Orchestration with RAG approach using OpenAI Multi-Language One-Way Translation with the Realtime API Doing RAG on PDFs using File Search in the Responses API How to use the Usage API and Cost API to monitor your OpenAI usage Leveraging model distillation to fine-tune a model Orchestrating Agents: Routines and Handoffs Prompt Caching 101 Developing Hallucination Guardrails
Getting started with datasets | OpenAI API
2025-08-13 · via OpenAI Developers

Evaluations (often called evals) test model outputs to ensure they meet your specified style and content criteria. Writing evals is an essential part of building reliable applications. Datasets, a feature of the OpenAI platform, provide a quick way to get started with evals and test prompts.

OpenAI is deprecating the Evals platform. Existing evals content remains available during the transition window. Evals will become read-only for existing users on October 31, 2026, and the platform is scheduled to shut down on November 30, 2026. See the deprecations page for the current timeline.

If you need advanced features such as evaluation against external models, want to interact with your eval runs via API, or want to run evaluations on a larger scale, consider using Evals instead.

First, create a dataset in the dashboard.

  1. On the evaluation page, navigate to the Datasets tab.
  2. Click the Create button in the top right to get started.
  3. Add a name for your dataset in the input field. In this guide, we’ll name our dataset “Investment memo generation.”
  4. Add data. To build your dataset from scratch, click Create and start adding data through our visual interface. If you already have a saved prompt or a CSV with data, upload it.

We recommend using your dataset as a dynamic space, expanding your set of evaluation data over time. As you identify edge cases or blind spots that need monitoring, add them using the dashboard interface.

Uploading a CSV

We have a simple CSV containing company names and actual values for their revenue from past quarters.

The columns in your CSV are accessible to both your prompt and graders. For example, our CSV contains input columns (company) and ground truth columns (correct_revenue, correct_income) for our graders to use as reference.

Using the visual data interface

After opening your dataset, you can manipulate your data in the Data tab. Click a cell to edit its contents. Add a row to add more data. You can also delete or duplicate rows in the overflow menu at the right edge of each row.

To save your changes, click Save button in the top right.

The tabs in the datasets dashboard let multiple prompts interact with the same data.

  1. To add a new prompt, click Add prompt.

    Datasets are designed to be used with your OpenAI prompts. If you’ve saved a prompt on the OpenAI platform, you’ll be able to select it from the dropdown and make changes in this interface. To save your prompt changes, click Save.

    Our prompts use a versioning system so you can safely make updates. Clicking Save creates a new version of your prompt, which you can refer to or use anywhere in the OpenAI platform.

  2. In the prompt panel, use the provided fields and settings to control the inference call:

  • Click the slider icon in the top right to control model temperature and top_p.
  • Add tools to grant your inference call the ability to access the web, use an MCP, or complete other tool-call actions.
  • Add variables. The prompt and your graders can both refer to these variables.
  • Type your system message directly, or click the pencil icon to have a model help generate a prompt for you, based on basic instructions you provide.

In our example, we’ll add the web search tool so our model call can pull financial data from the internet. In our variables list, we’ll add company so our prompt can reference the company column in our dataset. And for the prompt, we’ll generate one by telling the model to “generate a financial report.”

With your data and prompt set up, you’re ready to generate outputs. The model’s output gives you a sense of how the model performs your task with the prompt and tools you provided. You’ll then annotate the outputs so the model can improve its performance over time.

  1. In the top right, click Generate output.

    You’ll see a new special output column in the dataset begin to populate with results. This column contains the results from running your prompt on each row in your dataset.

  2. Once your generated outputs are ready, annotate them. Open the annotation view by clicking the output, rating, or output_feedback column.

    Annotate as little or as much as you want. Datasets are designed to work with any degree and type of annotation, but the higher quality of information you can provide, the better your results will be.

What annotation does

Annotations are a key part of evaluating and improving model output. A good annotation:

  • Serves as ground truth for desired model behavior, even for highly specific cases—including subjective elements, like style and tone
  • Provides information-dense context enabling automatic prompt improvement (via our prompt optimizer)
  • Enables diagnosing prompt shortcomings, particularly in subtle or infrequent cases
  • Helps ensure that graders are aligned with your intent

You can choose to annotate as little or as much as you want. Datasets are designed to work with any degree and type of annotation, but the higher quality of information you can provide, the better your results will be. Additionally, if you’re not an expert on the contents of your dataset, we recommend that a subject matter expert performs the annotation — this is the most valuable way for their expertise to be incorporated into your optimization process. Explore our cookbook to learn more about what we have found to be most effective in using evals to improve our prompt resilience.

Annotation starting points

Here are a few types of annotations you can use to get started:

  • A Good/Bad rating, indicating your judgment of the output
  • A text critique in the output_feedback section
  • Custom annotation categories that you added in the Columns dropdown in the top right

Incorporate expert annotations

If you’re not an expert on the contents of your dataset, have a subject matter expert perform the annotation. This is the best way to incorporate expertise into the optimization process. Explore our cookbook to learn more.

While annotations are the most effective way to incorporate human feedback into your evaluation process, graders let you run evaluations at scale. Graders are automated assessments that can produce a variety of inputs depending on their type.

TypeDetailsUse case
String checkCompares model output to the reference using exact string matchingCheck whether your response exactly matches a ground truth column
Text similarityUses embeddings to compute semantic similarity between model output and referenceCheck how close your response is to your ground truth reference, when exact matching is not needed
Score model graderUses an LLM to assign a numeric scoreMeasure subjective properties such as friendliness on a numeric scale
Label model graderUses an LLM to select a categorical labelCategorize your response based on fix labels, such as “concise” or “verbose”
Python code executionRuns custom Python code to compute a result programmaticallyCheck whether the output contains fewer than 50 words
  1. In the top right, navigate to Grade > New grader.
  2. From the dropdown, choose your grader type, and fill out the form to compose your grader.
  3. Reference the columns from your dataset to check against ground truth values.
  4. Create the grader.
  5. Once you’ve added at least one grader, use the Grade dropdown menu to run specific graders or all graders on your dataset. When a run is complete, you’ll see pass/fail ratings in your dataset in a dedicated column for each grader.

After saving your dataset, graders persist as you make changes to your dataset and prompt, making them a great way to quickly assess whether a prompt or model parameter change leads to improvements, or whether adding edge cases reveals shortcomings in your prompt. The datasets dashboard supports multiple tabs for simultaneously tracking results from automated graders across multiple variants of a prompt.

Datasets are great for rapid iteration. When you’re ready to track performance over time or run at scale, export your dataset to an Eval. Evals run asynchronously, support larger data volumes, and let you monitor performance across versions.

For more inspiration, visit the OpenAI Cookbook, which contains example code and links to third-party resources, or learn more about our evaluation tools:

Cookbook: Building resilient prompts with evals

Operate a flywheel of continuous improvement using evaluations.

Evaluate against external models, interact with evals via API, and more.

Use your dataset to automatically improve your prompts.

Build sophisticated graders to improve the effectiveness of your evals.