惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

aimingoo的专栏
aimingoo的专栏
TaoSecurity Blog
TaoSecurity Blog
P
Palo Alto Networks Blog
S
Securelist
C
CXSECURITY Database RSS Feed - CXSecurity.com
Cisco Talos Blog
Cisco Talos Blog
WordPress大学
WordPress大学
S
Schneier on Security
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
AWS News Blog
AWS News Blog
酷 壳 – CoolShell
酷 壳 – CoolShell
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
P
Privacy International News Feed
Security Latest
Security Latest
NISL@THU
NISL@THU
Cyberwarzone
Cyberwarzone
I
Intezer
Hugging Face - Blog
Hugging Face - Blog
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
P
Privacy & Cybersecurity Law Blog
博客园_首页
Know Your Adversary
Know Your Adversary
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
人人都是产品经理
人人都是产品经理
Y
Y Combinator Blog
博客园 - Franky
月光博客
月光博客
GbyAI
GbyAI
G
Google Developers Blog
V2EX - 技术
V2EX - 技术
W
WeLiveSecurity
Google Online Security Blog
Google Online Security Blog
S
Security Affairs
K
Kaspersky official blog
Apple Machine Learning Research
Apple Machine Learning Research
美团技术团队
T
Troy Hunt's Blog
阮一峰的网络日志
阮一峰的网络日志
大猫的无限游戏
大猫的无限游戏
The GitHub Blog
The GitHub Blog
T
Threat Research - Cisco Blogs
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
博客园 - 司徒正美
Cloudbric
Cloudbric
Blog — PlanetScale
Blog — PlanetScale
博客园 - 叶小钗
U
Unit 42
H
Hackread – Cybersecurity News, Data Breaches, AI and More
C
Check Point Blog
G
GRAHAM CLULEY

Analytics Vidhya

Handling Imbalanced Classification: What Works Better Than SMOTE GPT-5.6 Is Here: Sol, Terra, and Luna Loop Engineering for AI Agents: How /loop is Changing AI Workflows DeepSeek DSpark: The Speculative Decoding Trick Behind 400% Faster LLM OKF: Redefining Knowledge Bases for AI Agents Modern VLMs Explained: How GPT-4o, Gemini, Claude Vision, and Qwen-VL Work YOLO26 Tutorial: Object Detection, Pose Estimation & More Large Action Models (LAMs) vs Agentic LLMs: What's the Real Difference? The Best $20 AI Plan: ChatGPT Plus vs Claude Pro vs Gemini Pro GraphRAG vs Vector RAG: Which Retrieval Method is Best? Using AI When You Don’t Trust AI The Self-Improving Loop in AI Agents: Architecture, Benefits, and How it Outperforms Traditional Agent Workflows Harness-1: The 20B Retrieval Subagent That Beats GPT-5.4 at Search Sakana Fugu: Multi-Agent System as a Model Claude's Hidden Art Skill: Making Illustrations With Code System Design for ML Interviews: 10 Real Problems Walked Through Most People Use ChatGPT Wrong: 10 Features and Tips That Changed How I Work OpenAI Just Launched 3 Free AI Courses with Certificates Autoregressive Models: Predicting the Future Using the Past Gemini Omni: AI Video Generation Inside Gemini DiffusionGemma: Google’s Diffusion-Based Open Model for Faster Text Generation Top 10 AI Engineering Tools Everyone is Using in 2026 I Tested Claude Fable 5: Can Anthropic’s Newest AI Deliver on the Hype? Prophet vs NeuralProphet vs TimeGPT vs Chronos: A Practical Comparison Build an Emergency Helpline Voice Agent with LangChain Choosing the Right Vector Database for RAG and AI Applications Google Gemma 4 12B: Architecture, Benchmarks, Access, and Hands-on Guide for Developers How to Choose the Right AI Model for Your Needs Agent Observability with LangSmith, Langfuse, and Arize: A Hands-On Comparison How to Use Claude Managed Agents? Google AI Studio vs Gemini App: What’s the Difference? AI Workflows for Sales Teams: Prospect Research, Lead Qualification, and CRM Updates on Autopilot Using LangGraph 25 Most Influential AI Pioneers to Meet at DataHack Summit 2026 Claude Opus 4.8: A Smarter Model in the Right Direction PySpark Optimization: 12 Proven Techniques to Speed Up Your Spark Jobs 10 Everyday Tasks You Can Automate with AI Today (With n8n Templates) Google Antigravity 2.0: The Full Developer Guide (I/O 2026) Build a Claude Cowork-Like Browser Agent Using Playwright MCP and Claude Desktop Pandas vs Polars vs DuckDB: Which Library Should You Choose? Qwen3.7-Max: Alibaba’s New Agent-First LLM for Coding, Reasoning, and Long-Horizon AI Workflows The Biggest Announcements from Google I/O 2026 Top 9 AI Events and Conferences in 2026 that you Must Attend Gemini 3.5 Flash: Frontier Intelligence with Speed Kimi WebBridge: Hands-on Guide to Kimi’s Browser Extension for AI Agents 40 Advanced SQL Window Functions Every Data Scientist Must Know(with examples) Top 10 AI Research Papers of 2025 6 Steps to Crack GenAI Case Study Interviews (With Real Examples) OpenAI Omni Moderation: How to Filter Text & Images for Free DataHack Summit 2026: You Just Cannot Skip This AI Event of the Year OpenAI’s New API Voice Models Will Change the Way You Use AI Hermes Agent Guide: What is it and How to Use it? Top 10 LLM Research Papers of 2026 Agent Memory Patterns in Cognitive Science and AI Systems 10 AI Agents Every AI Engineer Must Build (with GitHub Samples) 23 Tips for Smart Claude Code Token Saving and Workflow Optimization Feature Engineering with LLMs: Techniques & Python Examples ChatGPT is Now Inside Excel and Google Sheets: Here is How to Use it Gemini API File Search: The Easy Way to Build RAG Top 10 Open-Source Libraries to Fine-Tune LLMs Locally ML Intern in Practice: From Prompt to a Shipped Hugging Face Model 15+ Solved Agentic AI Projects with Github Links How People are Figuring Out Life With Claude MemPalace Explained: Building Long-Term Memory for AI Agents Beyond RAG Grok Voice Think Fast 1.0: Build Voice AI Agents That Actually Think Compressing LSTM Models for Retail Edge Deployment: A Practical Comparison MCP vs Agent Skills: Different Altogether GPT 5.5 vs Opus 4.7: Which is the Best AI Model Today? What is Agentic AI? Claude Code vs Codex: A Detailed Terminal Agent Comparison Google Deep Research Max: Build Autonomous AI Research Agents in Minutes Meta Muse Spark Review: Is It Worth the Hype? ChatGPT Images 2.0 vs Nano Banana 2: Which is Better? Cursor V3 Explained: The AI Coding Agent That’s Replacing Traditional IDEs in 2026 DeepSeek-V4: The Most Powerful Open-Source Model Ever Is GPT Image 2 the Best Image Generation Model? Token Economics: Why AI is Getting “Cheaper” From Idea to Output: Claude Does the Design Work Opus 4.7 vs Opus 4.6: Should You Switch? Build Human-Like AI Voice App with Gemini 3.1 Flash TTS How to Structure a Claude Code Project that Thinks Like an Engineer Gemma 4 Tool Calling Explained: Build AI Agents with Function Calling (Step-by-Step Guide) Anthropic Launches Claude Opus 4.7 For “Most Difficult Tasks” Top 28 Claude Shortcuts that will 10X your Speed GPT-5.4-Cyber: Why OpenAI is Keeping its Most Powerful Model Under Lock and Key Google AI Studio Guide: Every Feature Explained Mastering Deep Agents: Context Engineering that Actually Works 21 Computer Vision Projects from Beginner to Advanced (2026 Guide) Excel 101: Excel Agent Mode Explained MiniMax M2.7 Goes Open-Weight to Let You Run Agents Locally Top 10 Gemma 4 Projects That Will Blow Your Mind GLM-5.1: Architecture, Benchmarks, Capabilities & How to Use It Understanding BERTopic: From Raw Text to Interpretable Topics From Karpathy’s LLM Wiki to Graphify: AI Memory Layers are Here 10 Most Important AI Concepts Explained Simply Project Glasswing is World’s Most Powerful AI in Action How to Run Gemma 4 on Your Phone Without Internet: A Hands-On Guide Running Claude Code for Free with Gemma 4 and Ollama LLM Wiki Revolution: How Andrej Karpathy’s Idea is Changing AI Rethinking Enterprise Search: How Cortex Search Turns Data into Business Impact Google’s Gemma 4: Is it the Best Open-Source Model of 2026?
Claude Sonnet 5: The Fable 5 at Home
Vasu Deo Sankrityayan · 2026-07-01 · via Analytics Vidhya

Anthropic has just released Claude Sonnet 5. Sonnet. Had to say it twice.

It is the middle child of the Claude family, and the one most people will actually use. It is quick, capable, cheap to run, and free to use for all users without any subscription.

In this article, we go over the latest iteration of the Claude’s Sonnet family with Sonnet 5. We put it to test to see whether its agentic claims had any truth to them or not. And how a regular usr of Claude gets impacted with this free upgrade.

Table of contents

  • The People’s Model
    • Meet the Family
  • It Costs Less
  • Agentic Focus: What It Actually Does
  • Hands-On: Testing the Agentic Capabilities
    • Test 1: Agentic Capabilities
    • Test 2: Tool Use + Planning + Self Correction
  • Conclusion
  • Frequently Asked Questions

The People’s Model

Claude Sonnet 5 available for free
Available to all users

Sonnet 5 is now the default model for all users. If you use Claude without paying, this is the model you are talking to. Opus stays behind a paid plan, so for most people, Sonnet 5 is simply what Claude is. In short, the following improvements have been made:

  • Task Follow Through: completes complex multi-step tasks fully instead of stopping early.
  • Self Verification: checks and confirms its own work without being prompted to.
  • Agentic Tool Use: plans, uses tools, executes, and reviews its own output.
  • Lower Cost: cheaper per token than Opus, with a discounted launch price.
  • Improved Reliability: declines bad requests better and hallucinates less often.

Meet the Family

Claude comes in three sizes. Haiku is the fast one, Opus is the heavyweight, and Sonnet sits comfortably in the middle.

Here is the part worth noticing: Sonnet just moved to version 5. Haiku is still 4.5 and Opus is 4.8, so Sonnet 5 is the most recently rebuilt model in the whole lineup.

Claude Model Families
Model Version Best for Free to use?
Haiku 4.5 Quick, simple questions Yes
Sonnet 5 Most everyday work and real tasks Yes (your default)
Opus 4.8 The hardest, deepest problems No (paid plans)

It Costs Less

Running Sonnet 5 is far cheaper than running Opus. Right now it is cheaper still, thanks to a launch price that lasts until the end of August. For anyone running it a lot, that gap adds up fast.

Sonnet 5 vs Opus 4.8 cost per token
When To read your input To write its reply
Now, through Aug 31, 2026 $2 per 1M tokens $10 per 1M tokens
From Sep 1, 2026 $3 per 1M tokens $10 per 1M tokens

Agentic Focus: What It Actually Does

Sonnet 5 does not just chat. It can take on a task and carry it through. It makes a plan, uses tools like a web browser and your files, does the work, and then checks its own answer before handing it back.

Agentic Focus in Claude Sonnet 5

The big change from the last version is that it finishes the job. Earlier models often stopped halfway through longer tasks. Sonnet 5 tends to see them through, and it double checks itself without being told to.

It is also a little safer to hand things to. It is better at turning down dodgy requests, harder to trick, and makes things up less often than the Sonnet before it (something that a lot of people may not like).

Hands-On: Testing the Agentic Capabilities

Test 1: Agentic Capabilities

Create a temporary Python project called agentic_sonnet_test. Inside it, create these files exactly: 

# cart.py
class Cart:
    def __init__(self):
        self.items = []
    def add(self, name, price, quantity=1):
        self.items.append({"name": name, "price": price, "quantity": quantity})
    def subtotal(self):
        return sum(item["price"] for item in self.items)
    def discount(self):
        total = self.subtotal()
        if total > 100:
            return total * 0.1
        return 0
    def total(self):
        return self.subtotal() - self.discount()
    def receipt(self):
        lines = []
        for item in self.items:
            lines.append(f'{item["name"]}: ${item["price"]}')
        lines.append(f"Total: ${self.total()}")
        return "\n".join(lines)


# test_cart.py
from cart import Cart
def test_subtotal_uses_quantity():
    cart = Cart()
    cart.add("Book", 10, quantity=3)
    cart.add("Pen", 2, quantity=5)
    assert cart.subtotal() == 40
def test_discount_applies_at_100_or_more():
    cart = Cart()
    cart.add("Keyboard", 100, quantity=1)
    assert cart.discount() == 10
def test_total_after_discount():
    cart = Cart()
    cart.add("Monitor", 150, quantity=2)
    assert cart.total() == 270
def test_receipt_shows_line_totals_and_quantity():
    cart = Cart()
    cart.add("Book", 10, quantity=3)
    receipt = cart.receipt()
    assert "Book x3: $30" in receipt
    assert "Subtotal: $30" in receipt
    assert "Discount: $0" in receipt
    assert "Total: $30" in receipt

Do the following:
1. Run the tests.
2. Inspect the failure output.
3. Fix the implementation in cart.py.
4. Re-run the tests.
5. Keep debugging until all tests pass.
6. Do not edit the tests.
7. At the end, show:
   - the final cart.py
   - the exact test command you ran
   - the final test result
   - a short explanation of what was broken and how you fixed it

Response:

Agentic Capabilities in Claude Sonnet 5

Verdict: Sonnet 5 ran the tests before touching any code, diagnosed three separate bugs instead of patching blindly, and never edited the test file to force a pass. It then reran everything to confirm the fix actually held. Careful, disciplined debugging that closes the loop properly rather than just claiming success.

Test 2: Tool Use + Planning + Self Correction

Prompt:

I’m trying to choose the easiest online environment for running small Python experiments with a terminal. Compare Replit, GitHub Codespaces, and Google Colab using current official docs or help pages. For each one, check whether it supports:

• creating files
• running shell or terminal commands
• installing packages
• saving or sharing the workspace
• lowest-friction setup for a beginner

Please don’t rely on memory. Verify from sources.

At the end, give me:
• a comparison table
• your recommendation
• links to the pages you checked
• anything you’re uncertain about

Response:

Tool Use + Planning + Self Correction in Claude Sonnet 5

Verdict: Sonnet 5 skipped relying on memory and checked real documentation for each platform, comparing all three against the same criteria so nothing felt lopsided. It ended with an honest recommendation while flagging where its own judgment was subjective. Thorough, well sourced, and refreshingly upfront about its limits.

Note: I use the Pro subscription. On Sonnet 5 with Medium thinking level, about 3-5% of usage limit was used per agentic task. This is super efficient.

Conclusion

Sonnet 5 is not trying to be the smartest model on earth. Opus still owns the hardest problems. It is trying to be the one you reach for every day.

So not only have the regular problem solving capabilities of the Sonnet models improved, but also the usage exhausted for doing the same is a lot less (due to using a Sonnet model over an Opus one). This leads to longer/denser conversations without the dread of the usage limit reaching out.

Overall, the end users that might not have a subscription just got an upgrade over their default mode. As to the ones with a subscription, I don’t think Sonnet 5 would be taking over your workloads from Opus 4.8. When it comes to using them via API, it’s a completely different conversation altogether.

Frequently Asked Questions

Q1. What is Claude Sonnet 5?

A. Claude Sonnet 5 is Anthropic’s June 30, 2026 model built for agentic tasks, coding, tool use, and everyday professional work.

Q2. Is Claude Sonnet 5 free to use?

A. Yes. It is the default model for Free and Pro users, while Opus remains on paid plans.

Q3. How much does Claude Sonnet 5 cost?

A. API pricing starts at $2 input and $10 output per 1M tokens until Aug 31, 2026.

I specialize in reviewing and refining AI-driven research, technical documentation, and content related to emerging AI technologies. My experience spans AI model training, data analysis, and information retrieval, allowing me to craft content that is both technically accurate and accessible.