惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

博客园_首页
博客园 - 【当耐特】
博客园 - 叶小钗
阮一峰的网络日志
阮一峰的网络日志
WordPress大学
WordPress大学
D
Docker
T
The Blog of Author Tim Ferriss
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
Microsoft Azure Blog
Microsoft Azure Blog
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
月光博客
月光博客
M
MIT News - Artificial intelligence
H
Hackread – Cybersecurity News, Data Breaches, AI and More
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
云风的 BLOG
云风的 BLOG
F
Fortinet All Blogs
罗磊的独立博客
小众软件
小众软件
A
About on SuperTechFans
MyScale Blog
MyScale Blog
D
DataBreaches.Net
The GitHub Blog
The GitHub Blog
C
Check Point Blog
L
LangChain Blog

Interesting Engineering

US firm to scale laser-based nuclear fusion ‘breakthrough’ with new partnership Military Archives - Interesting Engineering World’s first non-nuclear lead-cooled reactor to generate electricity begins installation US scientists devise new process to turn sewage sludge into 99% pure natural gas US firm unveils submarine-hunting drone with 9,200-mile-range, 35 mph top speed Military Archives - Interesting Engineering Supercomputer finds lithium-titanium tweak to boost sodium-ion batteries for grids Lockheed Martin demonstrates vertical launch missile system for mobile drone defense China’s 1116 MWe Taipingling Unit 1 reactor goes online, set to generate 9bn kWh yearly ChatGPT Images 2.0 update combines reasoning, research, and design with 2K output US Navy tests plug-and-play laser system on USS Bush carrier, downs drones at sea China’s CATL reveals 621-mile EV battery, under-7-minute charging to challenge BYD US uses world’s first exascale supercomputer to model supernovae, fusion reactors AI and Robotics Archives - Interesting Engineering First-in-human study confirms safety of graphene-based brain interface Tesla’s Optimus humanoid robot greets runners, poses for photos at Boston Marathon Interlocking materials offer high strength and flexibility for robotics, infrastructure US redeploys 100,000-ton nuclear-powered aircraft carrier in Red Sea after repairs US scientists unveil concept for ‘world’s first neutrino laser’ to unlock breakthroughs New military tech can maintain communication in contested electronic warfare environments Got a dark personality? Psychologists can help you choose your career wisely Humidity boosts performance of 3D-printed nanogenerator instead of degrading it China demonstrates microwave beam that recharges drones in flight, continues power delivery Scientists run compact free-electron laser for eight hours, cracks FEL stability problem China’s PLA considers to use minelaying underwater drones to enforce Taiwan blockade: Report 1-ton sharks may struggle for survival in waters exceeding 62.6°F, study suggests US firm’s thorium nuclear fuel bundles move to manufacturing for commercial reactors Tesla hits 0% charge in remote Chilean desert as YouTuber uses hood-mounted solar Humanoid robot surpasses human world record in Beijing half-marathon, clocking 50:26 mins New method extracts maximum work from unknown quantum states using symmetry tricks
GPT-5.5 crushes Claude Opus 4.7 in agentic coding with 82...
Aamir Kholla · 2026-04-24 · via Interesting Engineering

OpenAI has introduced GPT-5.5, positioning it as its most capable and intuitive model yet, with a focus on helping users complete complex, multi-step tasks more independently.

The release marks a continued push toward “agentic” AI systems that can plan, execute, and refine work with minimal human intervention.

The company said the model improves how users interact with AI across coding, research, and general knowledge work.

Instead of guiding every step, users can now assign broader tasks and rely on the model to navigate ambiguity and complete workflows.

“GPT-5.5 understands what you’re trying to do faster and can carry more of the work itself,” the company stated.

Stronger agentic coding

GPT-5.5 shows major gains in coding, especially in complex workflows that require planning and tool coordination.

On Terminal-Bench 2.0, it achieved 82.7% accuracy, a state-of-the-art score.

On SWE-Bench Pro, it reached 58.6%, solving more real-world GitHub issues in a single pass than earlier versions.

The model also outperformed its predecessor in long-horizon engineering tasks measured by internal benchmarks.

These tasks often take human developers up to 20 hours to complete.

Introducing GPT-5.5

A new class of intelligence for real work and powering agents, built to understand complex goals, use tools, check its work, and carry more tasks through to completion. It marks a new way of getting computer work done.

Now available in ChatGPT and Codex. pic.twitter.com/rPLTk99ZH5

— OpenAI (@OpenAI) April 23, 2026

OpenAI said the improvements go beyond benchmarks. Early testers reported that GPT-5.5 better understands system architecture and failure points.

It can identify where fixes belong and predict downstream impacts across a codebase.

The company emphasized efficiency alongside capability. GPT-5.5 matches GPT-5.4’s per-token latency despite higher intelligence.

It also uses fewer tokens to complete the same tasks, lowering computational cost.

“GPT-5.5 delivers this step up in intelligence without compromising on speed,” OpenAI noted. It added that the model performs at a higher level while maintaining real-world responsiveness.

Expanding real-world use

Beyond coding, GPT-5.5 expands its role in everyday knowledge work.

The model can move across tasks such as gathering information, analyzing data, and generating structured outputs like documents and spreadsheets.

1. We believe in iterative deployment; although GPT-5.5 is already a smart model, we expect rapid improvements. Iterative deployment is a big part of our safety strategy; we believe the world will be best equipped to win at the team sport of AI resilience this way.

2. We believe…

— Sam Altman (@sama) April 23, 2026

OpenAI said this reflects a broader shift toward AI systems that can actively operate software and tools.

The model can interpret interfaces, take actions, and transition between workflows with minimal friction.

Internal adoption highlights these capabilities.

More than 85% of OpenAI employees now use Codex weekly across departments, including engineering, finance, and marketing.

In one example, the communications team used GPT-5.5 to process six months of speaking request data.

The system built a scoring and risk framework and helped automate low-risk approvals.

In finance, the model reviewed 24,771 K-1 tax forms totaling over 71,000 pages.

The workflow excluded personal data and reduced processing time by two weeks.

Another team automated weekly business reporting, saving between five and ten hours each week.

OpenAI also stressed safety in the rollout.

The company said it deployed its strongest safeguards so far, including red-teaming, advanced testing, and feedback from nearly 200 early-access partners.

“Today, GPT-5.5 is rolling out to Plus, Pro, Business, and Enterprise users in ChatGPT and Codex,” the company said.

API access will follow after additional safety and scaling requirements are met.

The launch signals OpenAI’s continued focus on building infrastructure for agentic AI.

GPT-5.5 delivers this step up in intelligence without compromising on speed.

GPT-5.5 matches GPT-5.4 per-token latency in real-world serving, while performing better across nearly every evaluation we measured.

It also uses significantly fewer tokens to complete the same Codex… pic.twitter.com/5mR46SM7mW

— OpenAI (@OpenAI) April 23, 2026

The company aims to expand how people and businesses use AI to complete complex work across domains.