惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
Recorded Future
Recorded Future
Apple Machine Learning Research
Apple Machine Learning Research
博客园_首页
S
SegmentFault 最新的问题
博客园 - 司徒正美
Last Week in AI
Last Week in AI
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
云风的 BLOG
云风的 BLOG
雷峰网
雷峰网
博客园 - 叶小钗
The GitHub Blog
The GitHub Blog
MyScale Blog
MyScale Blog
腾讯CDC
博客园 - 聂微东
D
DataBreaches.Net
博客园 - Franky
人人都是产品经理
人人都是产品经理
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
博客园 - 【当耐特】
量子位
宝玉的分享
宝玉的分享
D
Docker
T
Tailwind CSS Blog
IT之家
IT之家
Engineering at Meta
Engineering at Meta
P
Proofpoint News Feed
C
CERT Recently Published Vulnerability Notes
Scott Helme
Scott Helme
Project Zero
Project Zero
Microsoft Azure Blog
Microsoft Azure Blog
AWS News Blog
AWS News Blog
Google DeepMind News
Google DeepMind News
H
Heimdal Security Blog
W
WeLiveSecurity
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
有赞技术团队
有赞技术团队
Simon Willison's Weblog
Simon Willison's Weblog
NISL@THU
NISL@THU
C
Cybersecurity and Infrastructure Security Agency CISA
Google DeepMind News
Google DeepMind News
T
Threatpost
TaoSecurity Blog
TaoSecurity Blog
N
News and Events Feed by Topic
aimingoo的专栏
aimingoo的专栏
Recent Commits to openclaw:main
Recent Commits to openclaw:main
www.infosecurity-magazine.com
www.infosecurity-magazine.com
SecWiki News
SecWiki News
S
Securelist

OfficeChai

These Are The 10 Cheapest AI Models In The World [June 2026] 18 Best AI Tools For English Speaking (With Examples) [2026] AI Impact? Vacancy Rates For US Office Properties Are Now Highest Since The 2008 Crisis KPMG Pulls Report Praising AI After It Was Found To Have Fake AI-Generated Citations India's Sarvam Raises $234 Million At $1.5 Billion Valuation After SpaceX Stock Pops 20%, Musk Has Made More Money In The Last 24 Hours Than Warren Buffett Made In His Entire Career OfficeChai Nobody Is Using AI Better Than Meta: NVIDIA CEO Jensen Huang 21 Best AI Tools For Animation (With Examples) [2026] 22 Best AI Tools For Architecture (With Examples) [2026] Datacenter Construction Spending Has Eclipsed Public Transportation Spending In The US China Scraps 12,000 Degree Courses, Mainly In Arts And Humanities, To Prepare For AI Age OfficeChai There Is No Job Loss With AI: David Friedberg Loop Between Human Capital And "Token Capital" Will Be The New IP For Firms, Says Satya Nadella How to Reduce Dependency on Key Employees 8 Google Index Checker Use Cases Beyond New Blog Posts Memory Squeeze? Smartphone Purchases Are Down Globally 21 Best AI Tools For Accounting (With Examples) [2026] AI For Voice Generation: 22 Best Options (With Examples) [2026] These Are The Most Popular Image Generation Models On OpenRouter [June 2026] Search Traffic For Websites Is Down 25% Over The Last Year Because Of AI: a16z Data Agentic Coding Has Led To A 50% Increase In Number Of Apps, But Most Are Finding Very Few Users: SimilarWeb Data OpenRouter Launches Fusion API, Which Uses A Combination Of Models To Achieve Fable-Like Performance At Half The Price Dario Amodei Refused To De-Deploy Or Fix Vulnerabilities In Fable Before US Export Controls, Says David Sacks 23 Best AI Tools For Notes Making (With Examples) [2026] 16 Best AI Tools For Astrology (With Examples) [2026] How Jensen Huang Once Had To Ask SEGA's CEO To Pay NVIDIA For A Technology That Didn't Work ChatGPT Already Has 11% Of The Search Market: OpenAI CFO Sarah Friar SpaceX Has Now Launched More Satellites Than Rest Of Humanity Combined Across History Globalization Is Dead, Time For India To Wake Up Says Sridhar Vembu After US Bans Anthropic Mythos And Fable Models For Foreign Users Elon Musk Becomes World's First Trillionaire After Record SpaceX IPO Anthropic Suspends Access To Mythos And Fable Models Following US Govt Directive Against Foreign Users 27 Best AI Tools For Market Research (With Examples) [2026] Why Jeff Bezos Makes Important Decisions Early in The Morning Education And Healthcare IT Have Been The Hardest Areas To Invest In: Peter Thiel Giving AI Long-Term Goals Could Lead To The Emergence Of Self-Preservation: Geoffrey Hinton Your Startup Doesn't Have a Hardware Problem. It Has an Accountability Problem Cyber Incidents Rarely Start With a Hacker: The Weak Links Businesses Overlook What Makes an App Worth Returning to Every Day? 21 Best AI Tools For Lead Generation (With Examples) [2026] How NBA Player Shaquille O'Neal Became An Early Investor In Ring AI For Kids Learning: 22 Best Options (With Examples) [2026] These Are The Most Popular AI Model Companies On OpenRouter [June 2026] Advanced Fintech and NeoBank Software Development Solutions: Building the Digital Banks of Tomorrow TRON Payments: Integrating AML Checks Into Business Workflows 18 Best AI Tools For Resume (With Examples) [2026] 16 Best AI Tools For UI Design (With Examples) [2026] These Are Top 10 Countries Generating The Most Internet Traffic How to Choose the Best Magento Agency for Your Store These Are The Best AI Models For Creative Writing [June 2026] AI For Managers: 28 Best Tools (With Examples) [2026] 17 AI Tools For Trading (With Examples) [2026] AI Has Led To An Explosion Of New Apps, But Nearly None Have Managed To Garner Significant Usage Cloudflare CEO Matthew Prince Says Vinod Khosla Asked Him To Fire His Co-founders For Him To Invest In His Company Australia’s AirTrunk To Invest $30 Billion To Develop Datacenters In India Anthropic Says That Their Employees Are Using AI To Write 8x More Code Compared To 18 Months Ago Anthropic Is Extremely Expensive, Many Are Urgently Looking For Alternatives: Microsoft AI CEO Mustafa Suleyman Sergey Tokarev on creating DIY “Beehives” and a free guidebook AI Crypto Price Prediction: How Accurate Are Machine Learning Models? Why Anthropic Could Find It Hard To Maintain Its $965 Billion Valuation Startup CEO Says They're Saving "Millions Of Dollars" By Replacing Anthropic Models With DeepSeek Ola Cabs' Valuation Falls 99% From Peak, Now Valued At Just $70 Million By Vanguard After TCS Case, Former Wipro Employee Alleges Attempt At Religious Conversion By Coworkers Bot Traffic Has Surpassed Human Traffic On The Internet For The First Time In History, Clouflare Says ChatGPT's Free Users Do 7 Queries Per Day, Those On $20 Plan Do 3x More: CFO Sarah Friar How Keith Rabois Had Been "Highly Skeptical" In 2023 That Anthropic Would Be Worth More Than $5 Billion In 10 Years How to Install AdGuard Home with Docker Step by Step We're Running Out Of Training Data, But Not Too Worried Because There Are Alternate Approaches: Google's Jeff Dean JioHotstar Is Hiring For 75 AI Roles Amid AI Content Push NVIDIA's Nemotron 3 Becomes Most Intelligent Open Weights Model From The US Hackers Allegedly Fooled Meta's AI To Take Over Accounts By Simply Asking It To Change User Emails Manchester Super Giants' AI Promotional Video Gets Panned As "Slop" For Glaring Cricketing Errors AI Reducing Jobs Is "Complete Nonsense": NVIDIA CEO Jensen Huang MiniMax Releases MiniMax M3, Is Competitive With Frontier Models On Many Benchmarks IIT Delhi-Incubated BotLab Dynamics Lights Up Skies With Lord Shiva Themed Drone Show During IPL Final NVIDIA Introduces RTX Spark, A New Chip Optimized For AI Agents For Windows Laptops And PCs NVIDIA Introduces Vera, A New CPU Chip For AI Agents That Is 80% Faster Than x86 CPUs OpenAI's Codex Reaches 5 Million Users, Resets Rate Limits For Users Key Factors That Influence Personal Loan Approval in India AI Is Allowing Me To Experiment And Try Crazier Things: Mathematician Terrance Tao Efficiency Of Human Learning Is Still A Thousand Times Better Than LLM Learning, Need Algorithmic Advances To Improve It: Jeff Dean San Francisco Home's Zillow Listing Says It'll Accept OpenAI Or Anthropic Stock As Payment Open-Source Models Currently Lag Proprietary Models By Just 4 Months: Epoch AI Self-Improvement Possible In AI Models Within A Year, Say Google's Top AI Leaders Digital Minds: Preparing for a Moral Challenge Before It Arrives Nearly 30% Of US-Based Y-Combinator Founders Are Of Indian Origin: SF Chronicle Data "A New Era Of PC": NVIDIA, Microsoft Windows Tease New Collaboration At Least 146,000 AI Hallucinated Citations In Papers Published In 2025, Finds Paper AI Doesn't Undergo Experiences, Has No Moral Conscience: Pope Leo XIV Claude Opus 4.8 Tops Artificial Analysis Intelligence Index, Edges Out GPT 5.5 With Score Of 61.4 Anthropic Says Its Annual Revenue Run-rate Has Now Touched $47 Billion Anthropic Raises $65 Billion At $965 Billion Valuation, Is Now Worth More Than OpenAI Claude Opus 4.8 Is Better Than Opus 4.7 But Not As Good As Mythos Preview, Says Anthropic Claude Opus 4.8 Beats GPT 5.5 On GDPval-AA Benchmark For Real World Tasks Anthropic Releases Claude Opus 4.8, Beats Opus 4.7, GPT-5.5 On Many Benchmarks GTM for Tech Startups Explained How to Use an AI Picture Generator to Create Professional Images Anthropic Is Now Generating 35% More Revenue Than OpenAI: The Information SK Hynix, Micron Join $1 Trillion Club Following AI-Led Memory Shortages
OpenAI Launches GPT-5.6 Sol, Beats Mythos On TerminalBench
OfficeChai Team · 2026-06-27 · via OfficeChai

OpenAI has announced the GPT-5.6 series — Sol, Terra, and Luna — in a limited preview beginning today. Sol is the flagship, Terra is pitched as a capable mid-tier option at half the cost of Sol, and Luna is the economy option for high-volume, cost-sensitive workloads. Broad availability across ChatGPT, Codex, and the API is promised “in the coming weeks,” though OpenAI has chosen a phased rollout coordinated with the U.S. government rather than an open launch.

The benchmark that will draw the most attention is TerminalBench 2.1, which tests command-line workflows requiring multi-step planning, tool coordination, and iteration. GPT-5.6 Sol scores 88.8% on that benchmark — behind only GPT-5.6 Sol Ultra, a new compute-intensive mode that hits 91.9%. More significant for the competitive picture: Claude Mythos 5, Anthropic’s restricted frontier model, sits at 88.0% on the same benchmark, and Claude Fable 5 — Anthropic’s current publicly available flagship — scores 84.3%, tied with GPT-5.6 Terra. GPT-5.6 Sol clears Mythos by nearly a full point on TerminalBench. That’s a meaningful result given how much of Mythos’s reputation has rested on coding and agentic capability.

On biology, OpenAI says GPT-5.6 Sol outperforms GPT-5.5 on GeneBench v1, a genomics and quantitative biology benchmark, while using fewer tokens. GPT-5.5 had already set a strong mark when it launched in April, and GPT-5.6 Sol improving on it with greater efficiency points to a model that is running smarter rather than just longer.

Cybersecurity is where OpenAI is being most deliberate in its framing. On ExploitBench, GPT-5.6 Sol is described as competitive with Claude Mythos while using roughly a third of the output tokens. On ExploitGym, a benchmark developed with UC Berkeley researchers, all three GPT-5.6 models show “strong improvements” in cyber capabilities as reasoning effort increases. OpenAI does clarify that Sol does not cross what it calls the “Cyber Critical” threshold — in tests against Chromium and Firefox, it found bugs and exploitation primitives but did not autonomously produce a functional full-chain exploit. Whether that’s a limitation or a deliberate design choice is hard to tell from the outside, but OpenAI is clearly aware of the optics of putting a model with Mythos-level cyber capability into general distribution.

The safeguard stack OpenAI describes for GPT-5.6 is the most layered it has shipped. There are model-level refusals, real-time output classifiers for cyber and biology misuse, a “pause and review” mechanism where a larger reasoning model can evaluate flagged outputs before they reach the user, and account-level review that looks across conversations rather than just individual prompts. OpenAI says it dedicated over 700,000 A100-equivalent GPU hours to automated red-teaming specifically aimed at finding universal jailbreaks — attacks that generalize across many prompts rather than exploiting one narrow pattern. Human red-teaming is ongoing through the preview period, with a rapid-response process to turn discovered weaknesses into updated evaluations.

The new model also introduces two new operational modes. max reasoning effort gives Sol more time to reason deeply before responding — a toggle that maps to what competitors have called “extended thinking.” ultra mode goes further, coordinating multiple subagents in parallel to tackle complex work. GPT-5.6 Sol Ultra’s 91.9% on TerminalBench reflects that mode’s output.

On pricing, Sol comes in at $5 per million input tokens and $30 per million output tokens — matching GPT-5.5’s pricing exactly. Terra is $2.50/$15, and Luna is $1/$6. OpenAI has also redesigned how prompt caching works: cache writes now cost 1.25x the base input rate, cache reads remain at a 90% discount, and there’s a minimum 30-minute cache lifetime with support for explicit cache breakpoints — which should make costs more predictable for developers running long agentic sessions. In July, OpenAI is also launching Sol on Cerebras at up to 750 tokens per second for select customers, which would be a significant speed advantage for latency-sensitive deployments.

The naming system is new. GPT-5.6 is the generation identifier, while Sol, Terra, and Luna are intended as durable capability tiers that can each advance on their own release cadence — meaning future updates to any tier may not require bumping the generation number.

The release arrives at a competitive moment. Chinese models have been gaining ground on US labs across several key benchmarks, and Anthropic has held the coding model leadership position for much of the past several months. TerminalBench 2.1 gives OpenAI a specific, credible result to point to — one that puts Sol ahead of Mythos on a test that matters to the agentic developer market. Whether that translates into benchmark leadership on the broader Artificial Analysis Intelligence Index will become clearer once those results are published alongside the full model release.

For now, GPT-5.6 Sol is available to a limited group of partners and organizations through the API and Codex. General availability is expected in the coming weeks.