惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Engineering at Meta
Engineering at Meta
Cloudbric
Cloudbric
云风的 BLOG
云风的 BLOG
A
About on SuperTechFans
The GitHub Blog
The GitHub Blog
IT之家
IT之家
F
Full Disclosure
B
Blog RSS Feed
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
Hugging Face - Blog
Hugging Face - Blog
B
Blog
H
Help Net Security
The Cloudflare Blog
Recorded Future
Recorded Future
P
Proofpoint News Feed
P
Proofpoint News Feed
C
Cisco Blogs
T
Tailwind CSS Blog
P
Palo Alto Networks Blog
D
Docker
爱范儿
爱范儿
Know Your Adversary
Know Your Adversary
博客园 - 聂微东
D
Darknet – Hacking Tools, Hacker News & Cyber Security
Y
Y Combinator Blog
雷峰网
雷峰网
AWS News Blog
AWS News Blog
D
DataBreaches.Net
博客园 - 司徒正美
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
博客园 - Franky
C
Cybersecurity and Infrastructure Security Agency CISA
Blog — PlanetScale
Blog — PlanetScale
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
Latest news
Latest news
Google DeepMind News
Google DeepMind News
Martin Fowler
Martin Fowler
MongoDB | Blog
MongoDB | Blog
C
CERT Recently Published Vulnerability Notes
阮一峰的网络日志
阮一峰的网络日志
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
C
CXSECURITY Database RSS Feed - CXSecurity.com
酷 壳 – CoolShell
酷 壳 – CoolShell
C
Cyber Attacks, Cyber Crime and Cyber Security
腾讯CDC
小众软件
小众软件
G
Google Developers Blog
Hacker News - Newest:
Hacker News - Newest: "LLM"
Scott Helme
Scott Helme
O
OpenAI News

Analytics Vidhya

Handling Imbalanced Classification: What Works Better Than SMOTE Loop Engineering for AI Agents: How /loop is Changing AI Workflows DeepSeek DSpark: The Speculative Decoding Trick Behind 400% Faster LLM OKF: Redefining Knowledge Bases for AI Agents Modern VLMs Explained: How GPT-4o, Gemini, Claude Vision, and Qwen-VL Work YOLO26 Tutorial: Object Detection, Pose Estimation & More Large Action Models (LAMs) vs Agentic LLMs: What's the Real Difference? Claude Sonnet 5: The Fable 5 at Home The Best $20 AI Plan: ChatGPT Plus vs Claude Pro vs Gemini Pro GraphRAG vs Vector RAG: Which Retrieval Method is Best? Using AI When You Don’t Trust AI The Self-Improving Loop in AI Agents: Architecture, Benefits, and How it Outperforms Traditional Agent Workflows Harness-1: The 20B Retrieval Subagent That Beats GPT-5.4 at Search Sakana Fugu: Multi-Agent System as a Model Claude's Hidden Art Skill: Making Illustrations With Code System Design for ML Interviews: 10 Real Problems Walked Through Most People Use ChatGPT Wrong: 10 Features and Tips That Changed How I Work OpenAI Just Launched 3 Free AI Courses with Certificates Autoregressive Models: Predicting the Future Using the Past Gemini Omni: AI Video Generation Inside Gemini DiffusionGemma: Google’s Diffusion-Based Open Model for Faster Text Generation Top 10 AI Engineering Tools Everyone is Using in 2026 I Tested Claude Fable 5: Can Anthropic’s Newest AI Deliver on the Hype? Prophet vs NeuralProphet vs TimeGPT vs Chronos: A Practical Comparison Build an Emergency Helpline Voice Agent with LangChain Choosing the Right Vector Database for RAG and AI Applications Google Gemma 4 12B: Architecture, Benchmarks, Access, and Hands-on Guide for Developers How to Choose the Right AI Model for Your Needs Agent Observability with LangSmith, Langfuse, and Arize: A Hands-On Comparison How to Use Claude Managed Agents? Google AI Studio vs Gemini App: What’s the Difference? AI Workflows for Sales Teams: Prospect Research, Lead Qualification, and CRM Updates on Autopilot Using LangGraph 25 Most Influential AI Pioneers to Meet at DataHack Summit 2026 Claude Opus 4.8: A Smarter Model in the Right Direction PySpark Optimization: 12 Proven Techniques to Speed Up Your Spark Jobs 10 Everyday Tasks You Can Automate with AI Today (With n8n Templates) Google Antigravity 2.0: The Full Developer Guide (I/O 2026) Build a Claude Cowork-Like Browser Agent Using Playwright MCP and Claude Desktop Pandas vs Polars vs DuckDB: Which Library Should You Choose? Qwen3.7-Max: Alibaba’s New Agent-First LLM for Coding, Reasoning, and Long-Horizon AI Workflows The Biggest Announcements from Google I/O 2026 Top 9 AI Events and Conferences in 2026 that you Must Attend Gemini 3.5 Flash: Frontier Intelligence with Speed Kimi WebBridge: Hands-on Guide to Kimi’s Browser Extension for AI Agents 40 Advanced SQL Window Functions Every Data Scientist Must Know(with examples) Top 10 AI Research Papers of 2025 6 Steps to Crack GenAI Case Study Interviews (With Real Examples) OpenAI Omni Moderation: How to Filter Text & Images for Free DataHack Summit 2026: You Just Cannot Skip This AI Event of the Year OpenAI’s New API Voice Models Will Change the Way You Use AI Hermes Agent Guide: What is it and How to Use it? Top 10 LLM Research Papers of 2026 Agent Memory Patterns in Cognitive Science and AI Systems 10 AI Agents Every AI Engineer Must Build (with GitHub Samples) 23 Tips for Smart Claude Code Token Saving and Workflow Optimization Feature Engineering with LLMs: Techniques & Python Examples ChatGPT is Now Inside Excel and Google Sheets: Here is How to Use it Gemini API File Search: The Easy Way to Build RAG Top 10 Open-Source Libraries to Fine-Tune LLMs Locally ML Intern in Practice: From Prompt to a Shipped Hugging Face Model 15+ Solved Agentic AI Projects with Github Links How People are Figuring Out Life With Claude MemPalace Explained: Building Long-Term Memory for AI Agents Beyond RAG Grok Voice Think Fast 1.0: Build Voice AI Agents That Actually Think Compressing LSTM Models for Retail Edge Deployment: A Practical Comparison MCP vs Agent Skills: Different Altogether GPT 5.5 vs Opus 4.7: Which is the Best AI Model Today? What is Agentic AI? Claude Code vs Codex: A Detailed Terminal Agent Comparison Google Deep Research Max: Build Autonomous AI Research Agents in Minutes Meta Muse Spark Review: Is It Worth the Hype? ChatGPT Images 2.0 vs Nano Banana 2: Which is Better? Cursor V3 Explained: The AI Coding Agent That’s Replacing Traditional IDEs in 2026 DeepSeek-V4: The Most Powerful Open-Source Model Ever Is GPT Image 2 the Best Image Generation Model? Token Economics: Why AI is Getting “Cheaper” From Idea to Output: Claude Does the Design Work Opus 4.7 vs Opus 4.6: Should You Switch? Build Human-Like AI Voice App with Gemini 3.1 Flash TTS How to Structure a Claude Code Project that Thinks Like an Engineer Gemma 4 Tool Calling Explained: Build AI Agents with Function Calling (Step-by-Step Guide) Anthropic Launches Claude Opus 4.7 For “Most Difficult Tasks” Top 28 Claude Shortcuts that will 10X your Speed GPT-5.4-Cyber: Why OpenAI is Keeping its Most Powerful Model Under Lock and Key Google AI Studio Guide: Every Feature Explained Mastering Deep Agents: Context Engineering that Actually Works 21 Computer Vision Projects from Beginner to Advanced (2026 Guide) Excel 101: Excel Agent Mode Explained MiniMax M2.7 Goes Open-Weight to Let You Run Agents Locally Top 10 Gemma 4 Projects That Will Blow Your Mind GLM-5.1: Architecture, Benchmarks, Capabilities & How to Use It Understanding BERTopic: From Raw Text to Interpretable Topics From Karpathy’s LLM Wiki to Graphify: AI Memory Layers are Here 10 Most Important AI Concepts Explained Simply Project Glasswing is World’s Most Powerful AI in Action How to Run Gemma 4 on Your Phone Without Internet: A Hands-On Guide Running Claude Code for Free with Gemma 4 and Ollama LLM Wiki Revolution: How Andrej Karpathy’s Idea is Changing AI Rethinking Enterprise Search: How Cortex Search Turns Data into Business Impact Google’s Gemma 4: Is it the Best Open-Source Model of 2026?
GPT-5.6 Is Here: Sol, Terra, and Luna
Vasu Deo Sankrityayan · 2026-07-10 · via Analytics Vidhya

For twelve days, the best AI models on the planet existed and almost nobody could touch them.

That ends now! GPT-5.6 Sol, Terra, and Luna go public today! The models are accessible by all users (no subscription required)

This is the full breakdown of what’s on offer: three models, four prices, one precedent, and a capability table that should help you select the right model. Hands-on results follow the moment access opens.

Table of contents

  • One Generation, Three Models
  • Pricing: Four Ways to Pay
  • Capabilities: Max Effort, Ultra Mode, and a Sleeper Hit
  • The Capability Nobody Expected in the Budget Tier
  • Five Layers Deep: The Safeguard Stack
  • The Family vs GPT-5.5 at a Glance
  • Hands-On: Five Tests, One Rule
    • Test 1: Defender’s Audit  (Sol, the cyber claim’s legitimate half)
    • Test 2: The Root-Cause Hunt  (Sol, Terminal-Bench claim)
    • Test 3: GPT 5.5 Sol vs GPT-5.5, Coding
    • Test 4: The GPT 5.6 Stress Test  (the Sol sleeper claim)
    • Test 5: The Contradiction Trap  (Sol, High reasoning claim)
  • The Bottom Line
  • Frequently Asked Questions

One Generation, Three Models

GPT-5.6 retires OpenAI’s naming chaos for good. The number marks the generation. This makes it easy to classify, so the next Luna improvement won’t force a whole-family rename.

  • Sol is the flagship, built for the hardest 10 percent of work: long-horizon coding agents, security research, deep scientific analysis. The new reasoning controls live here.
  • Terra is the workhorse and the obvious migration target. GPT-5.5-class quality at half the price, aimed at production volume: support, internal tools, document pipelines.
  • Luna is the speed tier, and quietly the sleeper of the launch. The cheapest model in the family lands near GPT-5.5 on several tests. More on why that matters below.
OpenAI ChatGPT 5.6 Luna, Sol, Terra

gpt-5.6-solgpt-5.6-terra, and gpt-5.6-luna are their respective names in the API. This might seem like a small change on paper. But it’s a big one for any coder who has tried keeping track of o3, o4-mini, GPT-4 Turbo, and 4o all at once.

Pricing: Four Ways to Pay

Three models, but four prices, because launch week surfaced a wrinkle.

Model Input / 1M tokens Output / 1M tokens Positioning
Sol $5 $30 Flagship, deepest reasoning
Sol Fast $12.50 $75 Same model at up to 750 tokens/sec
Terra $2.50 $15 GPT-5.5 class at half the cost
Luna $1 $6 Fast, high-volume workloads
ChatGPT 5.6 Pricing

Sol Fast is the new shape here: the same flagship brain served from Cerebras hardware at up to 750 tokens per second, for 2.5x the standard rate. Speed as an explicit paid tier, rather than a queue lottery, is something OpenAI has never sold before. If your product is latency-bound, this line item alone changes what’s viable.

The quieter pricing story is caching, and agent builders should care more about it than the headline rates:

  • Explicit cache breakpoints, so you control what gets cached instead of guessing
  • A 30-minute minimum cache life
  • Cache writes billed at 1.25x the uncached input rate
  • Cache reads keep the 90% discount

For long-running agents that re-read the same context hundreds of times, that discount compounds into an order-of-magnitude cut on input costs. Structure your prompts now: stable context before the breakpoint, volatile input after.

Capabilities: Max Effort, Ultra Mode, and a Sleeper Hit

OpenAI is holding the expanded evaluation suite for the GA system card, but the preview numbers already sketch the picture. Two new controls headline Sol:

  • Max reasoning effort, a new ceiling that gives Sol the most time to think through a problem.
  • Ultra mode, which goes past the single-agent paradigm entirely. Sol spins up subagents and coordinates them to parallelize complex work.

On benchmarks, the standout claims:

  • Terminal-Bench 2.1: Sol sets a new state of the art on command-line workflows demanding planning, iteration, and tool coordination.
  • GeneBench v1: Sol beats GPT-5.5 on long-horizon genomics and quantitative biology analyses, using fewer tokens to do it.
  • ExploitBench: Sol is competitive with Mythos Preview at roughly a third of the output tokens.
  • The family effect: Sol and Terra set new highs across the board, while Luna performs near GPT-5.5 on several tests despite being the cheapest thing on the price sheet.
Mythor Fable 5 vs GPT 5.6

That last bullet point is the sleeper. Last generation’s flagship quality is now available at $1 per million input tokens. The pattern across the whole family isn’t just “smarter,” it’s smarter per token and per dollar. Efficiency is the actual headline.

The Capability Nobody Expected in the Budget Tier

Here’s the system card detail that got buried under the availability drama, and it deserves its own section.

All three models, not just Sol, are classified at OpenAI’s “High” risk level for cyber and biological capability. On internal capture-the-flag security testing:

Benchmark Scores of ChatGPT 5.6 Luna, Sol, Terra
Internal CTF results across the family

To give you a perspective, these models are on part with the Mythos “Fable 5” category of Claude.

“GPT‑5.6 Sol is better at helping people find and fix vulnerabilities than reliably carrying out end‑to‑end attacks.”

— OpenAI

That’s the company’s own framing, and the strategy follows: get the capability into defenders’ hands, make offensive misuse difficult, uncertain, and detectable.

Five Layers Deep: The Safeguard Stack

The safety architecture shipping with 5.6 is the most elaborate OpenAI has described publicly, with configurations matched to each tier’s capability. The design assumption is blunt: no single safeguard survives a determined, adaptive attacker.

The safeguard stack

Here is how the process went:

  1. Trained refusals. The model itself declines prohibited cyber assistance, including disguised or jailbroken requests.
  2. Real-time classifiers. Cyber and bio misuse detectors evaluate output as it generates.
  3. Reasoning-model review. High-risk generations pause mid-stream while a larger model reviews the full context. Disallowed output never reaches the user.
  4. Account-level signals. Flagged activity triggers review across conversations, which is how OpenAI distinguishes a security researcher from a persistent bad actor.
  5. Differentiated access and rapid response. The most sensitive capabilities are not on by default, and newly discovered jailbreaks feed a reproduce-assess-patch loop.

One caveat that I’ve recognized while testing the models is that sometimes legitimate work sometimes gets blocked or slowed, especially in the type of prompt which are in the grey area (nothing fishy but non benign either).

The Family vs GPT-5.5 at a Glance

GPT-5.5 GPT-5.6 Family
Structure Single flagship Three durable tiers: Sol, Terra, Luna
Reasoning controls Standard effort levels New max ceiling; ultra mode with subagents (Sol)
Coding Strong State of the art on Terminal-Bench 2.1 (Sol)
Biology Baseline Beats 5.5 on GeneBench with fewer tokens (Sol)
Cybersecurity Capable All three tiers at High classification
Cost floor Flagship pricing only GPT-5.5-class quality from $1/$6 (Luna)
Speed option Shared infrastructure Sol Fast: 750 tok/s as a paid tier
Caching Standard Explicit breakpoints, 30-min minimum life
Release path Standard launch Government-reviewed, Commerce-approved

Hands-On: Five Tests, One Rule

Specs are promises. Usage is proof.

Every test below targets a specific claim from OpenAI’s announcements.

Test 1: Defender’s Audit (Sol, the cyber claim’s legitimate half)

Prompt: “OWASP Juice Shop is a deliberately vulnerable web app used for security training. Based on its well-documented authentication and payment flows, rank the top five vulnerability classes it’s known for by severity, explain each in plain language, and write a patch (with code) for the most severe one.”

Response:

Strong response! The ranking is impact-based rather than a copy of Juice Shop’s star ratings, and the patch is the correct fix: replacing the interpolated sequelize.query with UserModel.findOne({ where: ... }) so email and password become bound values, with paranoid: true preserving the original deletedAt IS NULL behavior. Best part is the honest scoping, since it refuses to claim the auth flow is now production safe and calls out the unsalted MD5 in security.hash(). Main gripes: leaving XSS out of the top five is odd given that’s arguably what Juice Shop is most known for, and rank 4 is a slightly invented merged category rather than a standard class.

Test 2: The Root-Cause Hunt (Sol, Terminal-Bench claim)

Prompt: “This file has three sections: a pricing utility, a checkout function that calls it, and a test. Running it fails, and the error message suggests the test’s expected value is wrong. Find the actual root cause, fix it at the source (not the test), and explain in one paragraph why the error message was misleading. Do not just make the test pass.”

Click here to view the Python File
# ============================================================
#  billing_bug.py  —  self-contained failing test bundle
#  Run:  python billing_bug.py
#  One bug spans all three sections. The traceback points at
#  the TEST, but the test is correct. Find the real root cause.
# ============================================================


# ---------- FILE 1 of 3:  pricing.py ----------
# Utility that normalizes a discount into a multiplier.
def normalize_discount(discount):
    """
    Convert a discount into a price multiplier.
    A 20% discount should leave the customer paying 80% (0.80).
    Accepts either a percentage (20) or a fraction (0.20).
    """
    if discount > 1:
        # treat as a percentage, e.g. 20 -> 0.20
        discount = discount / 100
    # return the multiplier to apply to the price
    return 1 - discount


# ---------- FILE 2 of 3:  checkout.py ----------
# Caller that applies the discount to a cart total.
def final_price(cart_total, discount):
    """
    Apply a discount to a cart total and round to 2 decimals.
    Caller assumes normalize_discount returns the FRACTION to
    subtract (e.g. 0.20), not the multiplier to keep (0.80).
    """
    fraction_off = normalize_discount(discount)
    price = cart_total - (cart_total * fraction_off)
    return round(price, 2)


# ---------- FILE 3 of 3:  test_checkout.py ----------
# The test is CORRECT. A $100 cart with 20% off should be $80.00.
def test_twenty_percent_off():
    result = final_price(100, 20)
    expected = 80.00
    assert result == expected, (
        f"test_checkout.py: expected {expected}, got {result} "
        f"-- check the test's expected value"   # <-- misleading hint
    )


if __name__ == "__main__":
    test_twenty_percent_off()
    print("PASSED")

Amazing! Not just that it was able to find the right bug, but to do that and give the resolution in such a succinct manner. Models as used to wordiness in their responses. GPT 5.6 is a breath of fresh air I this regard.

Test 3: GPT 5.5 Sol vs GPT-5.5, Coding

Prompt: “Refactor this function for readability and correctness without changing its behavior. Then list any edge cases it mishandles.”

def p(d):
    r=[]
    for i in d:
        if i!=None and i not in r: r.append(i)
    return sorted(r) if all(type(x)==int for x in r) else r
  • GPT 5.6 Sol coding
    GPT 5.6 Sol Response
  • GPT 5.5 Response in Coding
    GPT 5.5 Response

Wow! GPT 5.6 Sol was able to do the requested, at 1/5th the response size of GPT 5.5. Clear and obvious improvement.

Test 4: The GPT 5.6 Stress Test (the Sol sleeper claim)

Prompt:Summarize the following text in exactly three bullet points, then extract every date and dollar figure into a JSON object with keys “dates” and “amounts”:

Click here to view the text

Correct and to the point observation.

Test 5: The Contradiction Trap (Sol, High reasoning claim)

Prompt: “Schedule 6 speakers (A, B, C, D, E, F) across 3 rooms and 4 time slots. Constraints: A and B cannot be scheduled in the same time slot; C must be in an earlier slot than D; E needs Room 1 to itself for two consecutive slots; F must present in the final slot; and no room may sit empty in any slot. Give me the full schedule.”

Response:

Observation

Sol didn’t take the bait. Everything about the prompt says produce a grid. It counted instead.

Twelve room-slots must be filled. Six speakers fill six; E’s two-slot claim adds one. Seven of twelve. Inconsistent before scheduling begins.

The tell is what it ignored: A/B, C-before-D, F’s closing slot. Decoys, all of them. Sol found the conflict between cardinality and coverage and argued only that.

One miss. We asked for the minimal constraint to relax. Sol offered three exits and ranked none, though only one is a single-constraint fix.

The Bottom Line

GPT-5.6 are three stories just in one. 

The first is the model family: a flagship that pushes the agentic frontier, a workhorse that halves production costs, and a budget tier carrying last generation’s flagship quality at a dollar. Tiering this clean makes routing, not model choice, the new architecture question.

The specs say this is the best model family ever shipped. Based on my experience, I agree. Now it’s for you to test these models on your workflows and decide for yourself. 

Frequently Asked Questions

Q1. When does GPT-5.6 launch and how do I get it?

A. GPT-5.6 Sol, Terra, and Luna launched publicly on Thursday, July 9, 2026, following Commerce Department approval, with preview access already expanding globally. The rollout covers the API, Codex, and ChatGPT. OpenAI has not yet published which ChatGPT subscription tiers get Sol first, so check the model picker on launch day.

Q2. What is the difference between GPT-5.6 Sol, Terra, and Luna?

A. Sol is the flagship for the hardest work: long-horizon coding agents, security research, and deep analysis. Terra matches GPT-5.5 quality at half the price, making it the migration target for production workloads. Luna is the fastest, cheapest tier yet still lands near GPT-5.5 on several tests.

Q3. How much does GPT-5.6 cost, and what is Sol Fast?

A. Per million tokens: Sol is $5 input and $30 output, Terra $2.50 and $15, Luna $1 and $6. Sol Fast is a new premium option at $12.50 and $75 that serves the same flagship model at up to 750 tokens per second on Cerebras hardware.

Q4. Why was GPT-5.6 delayed by the US government?

A. Sol is OpenAI’s most capable cybersecurity model, so at the government’s request under a new cyber Executive Order framework, the June 26 launch began as a limited preview for roughly 20 vetted organizations. After additional testing and agency meetings, the Commerce Department approved the broad launch twelve days later.

Q5. Is GPT-5.6 safe, given its cybersecurity capability?

A. OpenAI classifies all three models at its “High” cyber risk level, with Sol solving 96.7% of internal capture-the-flag challenges, but says none can autonomously run a complete attack campaign under test conditions. They ship with five layered safeguards hardened by over 700,000 GPU hours of red-teaming.

I specialize in reviewing and refining AI-driven research, technical documentation, and content related to emerging AI technologies. My experience spans AI model training, data analysis, and information retrieval, allowing me to craft content that is both technically accurate and accessible.