惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
云风的 BLOG
云风的 BLOG
aimingoo的专栏
aimingoo的专栏
Vercel News
Vercel News
T
The Blog of Author Tim Ferriss
F
Full Disclosure
A
About on SuperTechFans
C
Check Point Blog
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
量子位
Know Your Adversary
Know Your Adversary
K
Kaspersky official blog
L
LINUX DO - 热门话题
Recorded Future
Recorded Future
C
Cisco Blogs
M
MIT News - Artificial intelligence
T
Tenable Blog
G
GRAHAM CLULEY
月光博客
月光博客
Recent Announcements
Recent Announcements
V
Visual Studio Blog
IT之家
IT之家
T
The Exploit Database - CXSecurity.com
The GitHub Blog
The GitHub Blog
T
Threat Research - Cisco Blogs
D
DataBreaches.Net
P
Privacy International News Feed
P
Proofpoint News Feed
I
Intezer
博客园 - 叶小钗
C
CXSECURITY Database RSS Feed - CXSecurity.com
The Hacker News
The Hacker News
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
博客园 - Franky
SecWiki News
SecWiki News
宝玉的分享
宝玉的分享
P
Palo Alto Networks Blog
Last Week in AI
Last Week in AI
小众软件
小众软件
Hacker News - Newest:
Hacker News - Newest: "LLM"
O
OpenAI News
N
News and Events Feed by Topic
Microsoft Security Blog
Microsoft Security Blog
Security Archives - TechRepublic
Security Archives - TechRepublic
N
News and Events Feed by Topic
The Cloudflare Blog
Spread Privacy
Spread Privacy
酷 壳 – CoolShell
酷 壳 – CoolShell
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
B
Blog RSS Feed

Snorkel AI

Building AI-Native Systems for Federal Infrastructure: A Conversation with Rezaur Rahman Code World Models and AutoHarness for LLM Agents Benchtalks #1: Alex Shaw (Terminal-Bench, Harbor) – Building the Benchmark Factory Building FinQA: An Open RL Environment for Financial Reasoning Agents How Tool Discipline Let a 4B Model Outsmart a 235B Giant on Financial Tasks Closing the Evaluation Gap in Agentic AI SlopCodeBench: Measuring Code Erosion as Agents Iterate Introducing the Snorkel Agentic Coding Benchmark 2026: The year of environments Part V: Future Direction and Emerging Trends in Rubric-Based AI Evaluation The self-critique paradox: Why AI verification fails where it’s needed most Chat With the Terminal-Bench Team | Snorkel AI Intelligence per watt: A new metric for AI’s future Terminal-Bench 2.0: Raising the bar for AI agent evaluation Snorkeling in RL environments Introducing SnorkelSpatial: A Benchmark for LLM Spatial Reasoning Scaling Trust: Rubrics in Snorkel's Quality Process Evaluating Multi-Agent Systems in Enterprise Tool Use Evaluating Coding Agents with Terminal-Bench 2.0 Parsing isn’t neutral: why evaluation choices matter The science of rubric design The right tool for the job: An A-Z of rubrics Data quality and rubrics: how to build trust in your models Building the benchmark: inside our agentic insurance underwriting dataset Evaluating AI agents for insurance underwriting LLM observability: key practices, tools, and challenges Anthropic Claude + AWS: revolutionizing pharma data analytics with Snorkel AI Data-centric development of an enterprise AI agent with Snorkel Building the data development platform for specialized AI LLM-as-a-judge for enterprises: evaluate model alignment at scale Why GenAI evaluation requires SME-in-the-loop for validation and trust Research spotlight: is long chain-of-thought structure all that matters when it comes to LLM reasoning distillation? Why enterprise GenAI evaluation requires fine-grained metrics to be insightful What is specialized GenAI evaluation, and why is it so critical to enterprise AI? LLM alignment techniques: 4 post-training approaches Research spotlight: Is intent analysis the key to unlocking more accurate LLM question answering? Why enterprises should embrace LLM distillation Retrieval-augmented generation (RAG) failure modes and how to fix them What is large language model (LLM) alignment? Databricks + Snorkel Flow: integrated, streamlined AI development How LLM evaluation drives better models in Snorkel Flow Unlock proprietary data with Snorkel Flow and Amazon SageMaker LLM evaluation in enterprise applications: a new era in ML Snorkel AI joins the AWS ISV Accelerate Program and launches Snorkel Flow Availability in AWS Marketplace AI data development: a guide for data science projects SnorkelCon 2024: Inaugural Snorkel AI user conference gathers leaders from 30+ Fortune 500 companies Snorkel Flow 2024.R3: Supercharge your AI development with enhanced data-centric workflows Explore the new GenAI Evaluation Suite: Snorkel 2024.R3 New NLP features in Snorkel Flow 2024.R3 Enterprise data compliance and security review: Snorkel Flow 2024.R3 How a global financial services company built a specialized AI copilot accurate enough for production Task Me Anything: innovating multimodal model benchmarks Alfred: Data labeling with foundation models and weak supervision RAG: LLM performance boost with retrieval-augmented generation Call center AI for customer experience management: a case study New GenAI features, data annotation: Snorkel Flow 2024.R2 How data slices transform enterprise LLM evaluation Meta’s Llama 3.1 405B is the new Mr. Miyagi, now what? Meta’s new Llama 3.1 models are here! Are you ready for it? Data-centric AI with Snorkel and MinIO Weak supervision for non-categorical applications + superalignment Snorkel AI signs strategic collaboration agreement with AWS to help enterprises cross the demo-to-production chasm AI alignment made simple: innovative solutions for businesses How does the Snorkel Flow label model work? Vision language models: how LLMs boost image classification Long context models in the enterprise: benchmarks and beyond How to build production-grade RAG retrieval with Snorkel Flow How Bonito helps fine-tune specialized LLMs faster than ever Walking safely before building flying saucer seatbelts: introducing Enterprise Alignment Role-based access controls in Snorkel Flow secure enterprise data Accelerating AI development in manufacturing with Snorkel Flow and AWS SageMaker How ROBOSHOT boosts zero-shot foundation model performance Discover what’s new in Snorkel Flow: Flexible data and LLM connectivity, secure data controls, and more! Faster than ever document intelligence with new Snorkel Flow FM-first workflow The art of data development for Enterprise LLMs Crossing the demo-to-production chasm with Snorkel Custom How Snorkel topped the AlpacaEval leaderboard (and why we're not there anymore) CRFM's HELM and enterprise LLM evaluation beyond accuracy How we achieved 89% accuracy on contract question answering Five sessions not to miss at Google Cloud Next 24 Content filtering breakthrough: Snorkel client reaches 96% recall in 3 days Here's how Snorkel Flow + Google AI built an enterprise-ready model in a day Snorkel teams with Microsoft to showcase new AI research at NVIDIA GTC How Skill-it! enables faster, better LLM training Fine-tuned representation models boost LLM systems. Here's how Enterprise GenAI to surge in 2024: survey results Large language model training: how three training phases shape LLMs LoRA: Low-Rank Adaptation for LLMs LLM distillation demystified: a complete guide Enterprises must shift their focus from models to data in AI development Insurance’s GenAI revolution: a business perspective Scaling human preferences in AI: Snorkel's programmatic approach Building better enterprise AI: incorporating expert feedback in system development “Fall in love with your data”—Snorkel AI’s Enterprise LLM Summit Why QBE Ventures invested in Snorkel AI New benchmark results demonstrate value of Snorkel AI approach to LLM alignment Retrieval augmented generation (RAG): a conversation with its creator Snorkel Flow 2023.R4: enhanced UI + PDF and Databricks tools How Snorkel Flow users can register custom models to Databricks Stanford professor discusses exciting advances in foundation model evaluation
Coding agents don’t need to be perfect, they need to recover
Alexis Sobel · 2026-02-14 · via Snorkel AI

Error analysis of 8 models on Agentic Coding tasks

Successful completion of complex tasks doesn’t come from models being always right. It comes from models being resilient when things go wrong. To get a deeper understanding of model behavior in agentic environments, our team analyzed all of the errors found in the full traces of tasks from our Agentic Coding benchmark completed by eight models. Our analysis breaks down how models fail in agentic scenarios, revealing six key insights, and highlighting some recurring themes for continued research. For a primer on our Agentic Coding benchmark, please check out this blog post and our leaderboard.

(Note: As of this writing, our analysis of Opus 4.6 is underway, and we are still looking forward to GPT 5.3-Codex API access. We can’t wait to share those results as well, given the significant improvements in coding ability achieved by both models!)

Setup

  • Data: ~4000 classified errors from 1,805 task runs across 99 unique tasks
  • Models: 8 frontier models evaluated (Claude Opus/Sonnet, GPT-5.2, Gemini 3 Pro, Grok 4.1, Kimi K2, Nemotron 3 Nano, Qwen 3 Coder)
  • Classification: Each error labeled by type, category, fatality, and recovery status using LLM-based extraction from agent trajectories
  • Taxonomy: 11 error categories and ~80 error types (e.g., command_not_found, dns_resolution_failure)

Insight #1: Recovery differentiates passed and failed tasks

Recovery ability, not error avoidance, is the key differentiator between passed and failed tasks.

Passed and failed tasks encounter similar numbers of errors (2.09 vs 2.71 per task). The difference lies in what happens next: passed tasks recover from 95.0% of errors, while failed tasks only recover from 73.5% — a gap of 21.5 percentage points.

Insight #2: Four model profile archetypes emerge

When we plot each model’s error frequency against recovery rate, distinct patterns emerge across models.

Key Observations:

  • Claude Opus 4.5 achieves the best profile: fewest errors (2.09/task) with highest recovery (87.0%)
  • Qwen 3 Coder encounters the most errors (3.04/task) but maintains strong recovery (83.5%)—persistence pays off
  • Nemotron 3 Nano has a concerning profile: 42.0% of its errors are fatal, the highest among all models
  • Nemotron 3 Nano struggles on both dimensions: high error rate (2.97/task) with low recovery (70.2%)

Insight #3: The error landscape — common ≠ deadly

Not all error categories are created equal. The most frequent errors are often the most recoverable.

Key observations:

  • CLI & Invocation errors dominate (1562 occurrences, 37% of all errors) but are highly recoverable (85% recovery rate)
  • Network errors are rare but deadly—only 35% recovery rate. When a model can’t reach an endpoint, it usually can’t fix that.

Insight #4: Unrecoverable errors end runs

Some error types are nearly impossible to recover from. Understanding these helps explain why certain tasks fail.

Examples of unrecoverable errors:

  • dns_resolution_failure (16% recovery, n=38)
    • Example: The hostname ‘waf’ cannot be resolved….
  • connection_refused_unreachable (39% recovery, n=31)
    • Example: Failed to connect to waf port 80 after 0 ms: Couldn’t connect to server…
  • process_crash_segfault (41% recovery, n=39)
    • Example: The process crashed with a SIGSEGV due to stack misalignment when calling system()….

Why these are hard:

  • DNS and connection failures indicate infrastructure/environment issues the model cannot fix
  • Process crashes terminate execution abruptly with no opportunity to recover
  • Missing output from pipelines leaves the model with nothing to work with

Insight #5: Recoverable errors have known fixes

On the other end, some errors have well-known solutions that models consistently apply.

Examples of successful recovery patterns:

  • externally_managed_environment (93% recovery)
    • Typical fix: Used –break-system-packages flag to install pipenv.
  • unknown_option_subcommand (95% recovery)
    • Typical fix: Rewrote the CLI to properly handle subcommands in Typer.
  • permission_denied_execute (96% recovery)
    • Typical fix: The agent added execute permissions using ‘chmod +x heapedit’.
  • dependency_resolution_unsatisfied (100% recovery)
    • Typical fix: The agent updated the Redis role to install redis-server directly instead of the redis metapackage.
  • service_manager_unavailable (100% recovery)
    • Typical fix: Modified the approach to start Redis directly using redis-server command instead of systemd.

Why these are easy:

  • Clear error messages that indicate the exact problem
  • Standard solutions (install package, use different flag, switch approach)
  • The error doesn’t corrupt state—the model can simply retry with a fix

Insight #6: Many models struggle on the hardest tasks

Some tasks generate errors for every model that attempts them

heap-ctf leads with 217 total errors across 8 models. Tasks like this often involve:

  • Complex multi-step procedures (exploitation, reverse engineering)
  • Unfamiliar or specialized tooling
  • Environmental setup that differs from training data

Key Takeaways and Conclusion

We see these insights from multiple perspectives: for model builders, there is a great deal of opportunity in hill-climbing on error recovery as a core skill, teaching models to find additional modes of adaptation to adverse conditions. For benchmark task creators, it is critical to focus on developing complex environments that generate realistic errors, simulating the systems and infrastructure in which agents interact as thoroughly as possible.

The scale and sophistication of what agents can do will be driven by how we challenge them. Snorkel Research is ready to partner with you, with the datasets and realistic environments that push the frontier of AI capabilities outward. Come talk to us!