惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

P
Proofpoint News Feed
AI
AI
T
Threatpost
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
C
Cybersecurity and Infrastructure Security Agency CISA
Scott Helme
Scott Helme
AWS News Blog
AWS News Blog
P
Privacy & Cybersecurity Law Blog
G
GRAHAM CLULEY
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
C
CERT Recently Published Vulnerability Notes
Cyberwarzone
Cyberwarzone
NISL@THU
NISL@THU
P
Privacy International News Feed
Schneier on Security
Schneier on Security
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
SecWiki News
SecWiki News
T
Tor Project blog
W
WeLiveSecurity
Security Archives - TechRepublic
Security Archives - TechRepublic
Spread Privacy
Spread Privacy
H
Hacker News: Front Page
Latest news
Latest news
C
Cyber Attacks, Cyber Crime and Cyber Security
D
Darknet – Hacking Tools, Hacker News & Cyber Security
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
T
Troy Hunt's Blog
Cisco Talos Blog
Cisco Talos Blog
人人都是产品经理
人人都是产品经理
腾讯CDC
博客园 - 【当耐特】
Engineering at Meta
Engineering at Meta
The Hacker News
The Hacker News
Application and Cybersecurity Blog
Application and Cybersecurity Blog
PCI Perspectives
PCI Perspectives
罗磊的独立博客
阮一峰的网络日志
阮一峰的网络日志
N
News and Events Feed by Topic
The Cloudflare Blog
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
TaoSecurity Blog
TaoSecurity Blog
博客园 - 叶小钗
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
T
Threat Research - Cisco Blogs
量子位
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
美团技术团队
D
Docker
C
CXSECURITY Database RSS Feed - CXSecurity.com
T
Tenable Blog

The JetBrains Blog

Kotlin Turns 15: Celebrate the Kotlin Effect - The JetBrains Blog PhpStorm 2026.2 is Now Out - The JetBrains Blog Key Takeaways From PHPverse 2026 - The JetBrains Blog What's New in IntelliJ IDEA 2026.2 - The JetBrains Blog What’s fixed in IntelliJ IDEA 2026.2 - The JetBrains Blog CLion 2026.2 Is Here - The JetBrains Blog DataGrip 2026.2: AI Agent Skills, MCP Tools and CLI Commands for Data Source Management, Bundled JDBC Drivers, and Improved Session Control - The JetBrains Blog Download WebStorm 2026.2: TypeScript 7 Support, AI, and more GoLand 2026.2 Is Now Available! - The JetBrains Blog Code in Space: Redefining Tech Creation with AI and XR - The JetBrains Blog Rider 2026.2 Release Candidate Is Out! - The JetBrains Blog ReSharper 2026.2 Release Candidate Released! - The JetBrains Blog JetBrains GameDev Days 2026 – Call for Speakers - The JetBrains Blog MPS 2026.1 Has Been Released! - The JetBrains Blog IntelliJ Scala Plugin 2026.2 Is Out! - The JetBrains Blog What's New in ReSharper 2026.2 for VS Code-compatible editors  - The JetBrains Blog Debugging for .NET in VS Code and Cursor: The #1 Requested Feature Is Here - The JetBrains Blog dotInsights | July 2026 - The JetBrains Blog The History of Kodee, Kotlin’s Mascot - The JetBrains Blog JetBrains Academy – June Digest - The JetBrains Blog Introducing the Kotlin Benchmark for AI Coding Agents - The JetBrains Blog Best Object Detection Models for Machine Learning in 2026 - The JetBrains Blog What's Next for TeamCity – CI/CD by JetBrains - The JetBrains Blog The Benchmark Meaning Gap - The JetBrains Blog JetBrains AI for Teams and Organizations: From Fragmented AI Usage to Coordinated Software Development - The JetBrains Blog Java Annotated Monthly – July 2026  - The JetBrains Blog Shift-Left with JetBrains Qodana Natvis Comes to Linux and macOS: Visualize Your C++ Types Without Writing a Single Data Formatter - The JetBrains Blog Speaking to AI Agents like Cavemen Saves 65% of Tokens. We Test. In Conversation With the Golden Kodee Winners - The JetBrains Blog Toolbox App 3.6: Smarter Storage Cleanup, Windows installation diagnostics, and More - The JetBrains Blog IntelliJ IDEA 2026.1.4 Is Out! - The JetBrains Blog TeamCity 2026.1.2 and 2025.11.6 Are Now Available - The JetBrains Blog JetBrains Engineering Hiring Process Guide Kotlin Comes to BlueJ - The JetBrains Blog Improving Embedded Software Quality With Parasoft C/C++test, CLion, and AI - The JetBrains Blog Kodee’s Kotlin Roundup: Kotlin Turns 15, Kotlin 2.4.0, and the Kotlin Toolchain - The JetBrains Blog GitHub Copilot now an Integrated Agent in JetBrains IDEs - The JetBrains Blog JetBrains Air lands on Windows - The JetBrains Blog The Role of Static Code Analysis in Fintech Compliance Kotlin Notebook Sunset - The JetBrains Blog Open-Sourcing the LSP Client API in IntelliJ IDEA 2026.2 - The JetBrains Blog The Dev Containers Story: Introducing EelApi for Plugin Authors - The JetBrains Blog Cursor's $60B Acquisition - Qodana Codex is now the recommended agent in JetBrains IDEs - The JetBrains Blog SSH Connections Are Moving to JetBrains Daemon in the Toolbox App 3.6 EAP - The JetBrains Blog Your AI Agent Keeps Missing The Real Bottleneck. JetBrains Rider Can Fix It Now. - The JetBrains Blog Rust Web Development 2026: The Problems Nobody Talks About Our Research on Membership Inference Attacks and Preventing Privacy Leaks - The JetBrains Blog Explicit Lazy Imports Are Coming to Python 3.15 - The JetBrains Blog Kotlin Toolchain 0.11: The Next Step for Amper - The JetBrains Blog YouTrack Helpdesk Now Includes Customer Groups - The JetBrains Blog How to Win a Hackathon: Notes From the Judging Table - The JetBrains Blog How We Measure the ROI of JetBrains IDEs - The JetBrains Blog AWS Image Builder Plugin for TeamCity - The JetBrains Blog PHP Version Migration | Jetbrains Qodana Bamboo End of Life: How to Prepare and Choose the Right CI/CD Replacement - The JetBrains Blog Structuring IntelliJ Plugins with Optional Content Modules - The JetBrains Blog YouTrack Security Update: Upgrade Required for YouTrack Server - The JetBrains Blog Qodana Is a Finalist in the 2026 CODiE Awards for Best DevOps Tool - The JetBrains Blog JetBrains Marketplace Ecosystem Security Update: Addressing Malicious Third-Party AI Plugins - The JetBrains Blog Your JetBrains IDE Expertise, Now on LinkedIn - The JetBrains Blog The JetBrains AI Coding Agent moves to general availability The Anthropic Debate - The Qodana Blog dotInsights | June 2026 | The .NET Tools Blog Inside JetPride: How JetBrains Employees Built an LGBTQIA+ Community | The Life at JetBrains Blog MPS 2026.1 Release Candidate Arrives | The MPS Blog Best Python AI Frameworks in 2026 | The PyCharm Blog Contribute to the State of PHP Survey | The PhpStorm Blog The Rules of Zero, Three and Five - The Qodana Blog Modern C++ Support in CLion: What’s New | The CLion Blog Agentic AI Governance: Designing for Accountability and Control | The JetBrains AI Blog JetBrains Plugin Developer Conf 2026 – Call for Speakers | The JetBrains Platform Blog Fewer False Positives in RustRover 2026.2|The RustRover Blog Rider 2026.2 EAP 5: Code Quality Checks for Your AI Agents, and More. | The .NET Tools Blog Why Zig Isn’t 1.0 (Yet) | The JetBrains Blog Java Annotated Monthly – June 2026  | The IntelliJ IDEA Blog IntelliJ IDEA 2026.1.3 Is Out! | The IntelliJ IDEA Blog RustRover at RustWeek 2026 | The RustRover Blog WPF Hot Reload Is Here: Edit Your XAML and Watch It Update Live in Rider | The .NET Tools Blog Kotlin 2.4.0 Released | The Kotlin Blog IntelliJ IDEA 2025.3.6 Is Out! | The IntelliJ IDEA Blog Async VFS Content Writes - What Plugin Authors Need to Know | The JetBrains Platform Blog Top Agentic Frameworks for Building Applications 2026 | The PyCharm Blog Toolbox App 3.5: Better Remote Development Observability, More Reliable Enterprise Configuration, and Smoother Everyday Interactions | The Toolbox App Blog Stop Pasting Tokens: OAuth2 Login for JetBrains IDE Plugins | The JetBrains Platform Blog Fix Common TypeScript Issues | The Qodana Blog Mellum2 Goes Open Source: A Fast Model for AI Workflows | The JetBrains AI Blog What Does It Actually Take for an IDE to Understand Rust? Hibernate 7.4 New Features | The IntelliJ IDEA Blog How We Use AlphaEvolve to Make Complex IDE Algorithms Faster | The JetBrains AI Blog JetBrains Academy – May Digest | The JetBrains Academy Blog TeamCity 2026.1.1 Is Now Available | The TeamCity Blog The Upcoming Sunset of DataSpell | The DataSpell Blog Deprecating dotMemory Unit | The .NET Tools Blog Koog 1.0 Is Out: Stable Core, Better Interop, and Multiplatform Observability | The JetBrains AI Blog Introducing the Cloud9 JetStream Theme for JetBrains IDEs | The JetBrains Blog Build a Live Object Detection App for the Reachy Mini With TensorFlow and PyCharm | The PyCharm Blog IntelliJ IDEA 2026.2 EAP Is Open | The IntelliJ IDEA Blog How AI Agents Can Work with TeamCity | The TeamCity Blog
Step Rejection Fine-Tuning: Squeezing More Signal from Noisy Agent Trajectories - The JetBrains Blog
Igor Slinko · 2026-06-17 · via The JetBrains Blog

JetBrains Research

Research is crucial for progress and innovation, which is why at JetBrains we are passionate about both scientific and market research

Research

If you want to dive straight into the technical details, you can read our full paper here.

Imagine you are mentoring a junior developer. If they make a single logical error on line 42 of a 100-line script, do you throw away the entire file and tell them they learned nothing? Of course not. You point out the specific mistake and acknowledge what they got right.

Yet, when training large language model (LLM) agents, the standard practice is exactly that: We opt to discard the entire attempt if the final outcome isn’t perfect. In complex tasks, agents fail a lot, meaning we are constantly throwing away a massive amount of potentially valuable data.

Why is this data so valuable? Even when an agent fails to solve a task, many of its steps – such as exploring the directory structure, reading relevant files, and writing initial test scripts – are completely correct. By discarding the entire run, we throw away all of those high-quality examples of correct behavior.

To bring order to this inefficiency, our team at JetBrains Research developed Step Rejection Fine-Tuning (SRFT). It is a simple, practical technique to help models learn from their failed attempts without picking up bad habits. Our paper introducing this work has been accepted to the Deep Learning 4 Code (DL4C) workshop, co-located with ICML in South Korea this July.

In this blog post, we will:

  • Unpack traditional LLM agent training and see why standard methods waste data..
  • Uncover the hidden value inside unsuccessful trajectories.
  • Introduce SRFT and explain how it uses a “critic” to mask harmful steps.
  • Share our experimental results showing how SRFT boosts performance.

The problem with perfect trajectories

There are two main approaches to training LLM-based agents. The first is reinforcement learning, most commonly implemented using algorithms like Group Relative Policy Optimization (GRPO). In this approach, the model learns through trial and error. It receives a reward if the entire trajectory leads to a successful resolution, and is penalized if it fails.

The second approach involves knowledge distillation from a stronger teacher model. Here, a powerful (and usually expensive) model generates solutions, and a smaller student model learns to imitate its behavior. When using the distillation approach, the standard practice is Rejection-sampling Fine-Tuning (RFT). You generate a bunch of trajectories from the teacher to solve a task, throw away the ones that failed, and then train your student model only on the successful ones.

To give you an idea, a single trajectory is essentially the full conversation history of an agent trying to solve a problem. It consists of a sequence of steps where the agent reasons, takes an action (like running a command or editing a file), and receives an observation from the environment. On average, a trajectory in complex coding tasks contains dozens of such steps.

Crucially, we can usually only determine whether a trajectory was successful at its very end, as a typical trajectory concludes with the generation of a code patch. In standard benchmarks, pre-written test suites are run to verify whether this final patch resolves the original issue. Consequently, while we obtain complete, binary feedback on the success of the trajectory as a whole, we lack any test-level information regarding which specific steps taken by the agent were actually helpful, and which ones led to the incorrect patch.

Below is an example of what a stepwise-labeled trajectory looks like in practice. In the third column, SP represents system prompt, UP represents user prompt (which contains the issue description), the rows labeled with the letter A and a number represent the AI assistant step, and those labeled with the letter O and a number represent the corresponding output.

step rejection fine-tuning: example trajectory

In this example, assistant Step #3 (A3) was marked as unnecessary because the agent viewed a file that wasn’t related to the bug introduced in the issue description. Step #4 (A4) was marked as a mistake because the agent started fixing code before reproducing the bug, which directly contradicts the instruction given in the system prompt (SP). Additionally, Step #7 (A7) was labeled as “recover” because it corrects an error made in Step #5 (A5) during the agent’s attempt to reproduce the bug. We chose not to label Step #5 as a mistake because the replication script created in that step was otherwise completely correct, with only a single line containing an error.

It is worth noting that this specific trajectory was successful because it ultimately resolved the bug correctly, despite doing so in a suboptimal manner. While even successful trajectories are not always completely free of errors, unsuccessful trajectories always contain harmful steps that we may identify and label as mistakes.

Because standard RFT uses only successful trajectories, it discards tons of data. For instance, the recent SWE-smith project generated a large-scale dataset of agent trajectories for software engineering tasks. This dataset was then used to train an agentic model. Because they used standard RFT, they had to discard approximately 61% of all collected runs for training. That is a huge amount of potentially informative data lost just because the final outcome wasn’t perfect.

The hidden value of unsuccessful trajectories

Our core hypothesis is that these unsuccessful trajectories are not entirely erroneous – rather, they often consist of correct and useful steps interspersed with errors.

To test this, we conducted a manual analysis of 20 failed trajectories from the SWE-smith dataset.

We discovered that even in completely failed runs, only up to 24% of the steps could actually be classified as going in the wrong direction. The remaining 76% of the steps consisted of productive exploration, codebase navigation, or harmless tool actions.

To understand how these unsuccessful trajectories can be valuable, we first need to understand why distillation, within which RFT is standard practice, works at all. When we train a student model on a teacher’s trajectories, the performance boost comes from two distinct sources:

  1. Learning “smart” tokens: The student learns from a much smarter, more capable model. It absorbs better ways to reason, to understand tasks, and to use the provided tools.
  2. Learning the path to success: By filtering only successful trajectories (as in standard RFT), we bias the model to choose actions that actually lead to a resolved task.

As mentioned above, standard RFT throws away unsuccessful trajectories because they lack that second source of improvement. In other words, they would teach the model to imitate the mistakes that led to failure.

But what if we train only on unsuccessful trajectories generated by a strong teacher model? Will it boost the model’s performance?

Before we answer this question, let’s set the stage with our experimental setup. To make things easier, we’ll present the complete table with all our results right after the experiment’s preliminaries, and then we will walk you through each experiment, starting with the answer to this very question.

We tested our approach on SWE-bench Verified, a challenging benchmark that tasks AI agents with solving real-world GitHub issues in large Python repositories. It thoroughly tests an agent’s ability to navigate codebases, edit files, and run tests.

For the training data, we used trajectories from the SWE-smith dataset to fine-tune the Qwen2.5-Coder-32B-Instruct model, running all the experiments on the SWE-agent scaffold. To filter out the random noise of individual runs and ensure our conclusions are reliable, we repeated each experiment seven times. For more details on the methodology, see our paper.

The table below shows the results of our experiments. The Training data column indicates which part of the SWE-smith dataset was used to fine-tune the model; each subset is built from pools of 5,000 resolved, unresolved, or unresolved (masked) trajectories, used either individually or combined. The Resolved column shows the resolved rate across 500 SWE-bench Verified tasks averaged over the seven consecutive runs, along with the standard deviation. The experiments are sorted in ascending order of this main metric. Consequently, the Δ vs. Prev. column represents the improvement in the main score over the previous row.

Now, let’s look at the results to answer our question about whether training on unsuccessful trajectories actually helps.

As it turns out, yes! Because of the first source of improvement (learning “smart” tokens), even unsuccessful trajectories significantly boost the model’s performance! As you can see in the table above, when training only on unsuccessful trajectories (Experiment #2) the resolution rate jumps from the base model’s 7.0% up to 27.7%. The student is still learning how to use tools like the smart teacher, even if the final patch of a failed run didn’t resolve the issue.

Okay, so we got 27.7% using only unresolved trajectories. But we also have resolved trajectories at our disposal. What happens if we add them to the mix?

As you can see in Experiment #3 (Naïve Distillation), simply combining 5,000 unresolved and 5,000 resolved trajectories increases the resolution rate to 28.5% (a modest 0.8% boost). While there is a slight improvement, the added benefit is quite small.

Now, look at Experiment #5 (RFT, or Rejection-sampling Fine-Tuning), where we train the model only on the 5,000 resolved trajectories. It achieves 30.9%, which is better than mixing them with unresolved ones. This is the core philosophy behind standard RFT: You should only train on successful, high-quality trajectories and discard the unsuccessful ones, because adding failed attempts back into the mix actually degrades the model’s performance.

Yet, we can clearly see that unsuccessful trajectories still hold massive potential. They genuinely teach the model useful skills, as demonstrated by the huge boost in Experiment #2. Is there a way to extract this valuable information from failed runs while avoiding steering the model toward making mistakes?

As it turns out, there is! This is exactly what Experiment #6 (SRFT) is all about. As you can see in the table, SRFT outperforms standard RFT (32.2% vs. 30.9%), yet it relies on a remarkably simple trick.

Step Rejection Fine-Tuning

So, how do we extract the good parts of a failed attempt without having the model learn the bad parts?

Our solution, Step Rejection Fine-Tuning (SRFT), works as follows. We use another LLM as a “critic” to analyze unsuccessful trajectories step by step. The critic’s job is to diagnose each action and flag which steps were actually harmful (like introducing a bug or going down a completely wrong path) and which steps were productive.

Labeling these steps with a critic model is incredibly cheap compared to the massive compute and API costs required to generate the agent trajectories in the first place. This is because the critic analyzes the entire trajectory in a single pass (requiring just one model call), and its output is extremely concise – simply a list of step numbers with their corresponding labels: good, unnecessary, mistake, or recover.

Now that we have labels for each step, how do we actually use them?

In theory, there are several ways to handle this. We could take a prefix of the trajectory – training the model only on the initial good steps and cutting it off at the first mistake. Alternatively, we could modify and transform the trajectory by completely removing the mistake steps to create synthetic, “clean” trajectories.

But there is a much simpler and more elegant approach: we can just skip calculating the training loss on the mistake steps.

Why is this method superior? First, we don’t generate any synthetic data, which means we avoid inventing artificial scenarios that never actually happened. Second, the model still sees the entire trajectory and learns from the full context, but we simply don’t train it to predict the tokens inside the mistake steps.

This means the model sees the mistake happen in the context, but it isn’t trained to reproduce it. Furthermore, if the agent managed to recover from that mistake later in the trajectory, the model will actually learn how to perform that recovery!

From a technical perspective, during training, we “mask” the tokens inside these mistake steps so they don’t contribute to the training loss. If you are familiar with standard next-token prediction training, masking is a very common technique. For example, user messages (prompts) are usually masked so the model doesn’t learn to predict them, while the assistant’s responses are not masked. We aren’t doing anything overly complex here; we are simply applying this standard masking technique to specific, harmful steps of the assistant. During training, we mask the loss for this specific mistake step, while keeping the rest of the steps intact.

Referring back to our example trajectory above, this means that while Step #4 (A4) will not have its loss calculated, it will still remain in the context when the model calculates the training loss for the subsequent Steps #5, #6, #7, #8, and #9.

It is also worth noting that this masking approach works incredibly well even if you apply it only to unsuccessful trajectories. This brings us to the only remaining experiment in our table that we haven’t discussed yet: Experiment #4 (unresolved masked).

As you can see, training only on 5,000 unresolved trajectories with masked mistake steps yields a 29.7% resolution rate. This is actually better than naïvely mixing 5,000 unresolved and 5,000 resolved trajectories together (28.5%)! This means you can take purely unsuccessful trajectories, filter them with a critic, and still get a massive performance boost. Of course, if you already have successful trajectories, you should definitely include them in the training mix. But it was highly encouraging to see that our step-masking approach delivers such strong results even in a failed-runs-only scenario.

Conclusion

Step Rejection Fine-Tuning (SRFT) offers a practical way to squeeze more value out of your training data. Instead of throwing away hard-earned trajectories just because they didn’t perfectly solve the task, we can use a critic to filter out the noise and learn from the signal.

Of course, the exact benefit depends on your specific task, the ratio of successful to unsuccessful trajectories you have, and how well your critic model can identify the harmful steps. The strictness of your critic is a crucial balance to strike:

  • If the critic is too lenient, you might leave harmful steps in the training data, which will degrade the model’s quality (similar to the naïve mixing approach).
  • If the critic is too strict, you throw away too many potentially useful steps, losing the benefit of including unresolved trajectories in the first place.

This strictness is usually determined by the critic’s prompt. Alternatively, you can ask the critic to output a confidence score for its judgment and filter steps based on a specific threshold. Either way, this balance needs to be tuned for each specific dataset. But overall, once tuned, it’s a straightforward technique that can noticeably improve your agent’s performance.

Subscribe to JetBrains Research blog updates

Discover more