惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

S
SegmentFault 最新的问题
S
Secure Thoughts
Google DeepMind News
Google DeepMind News
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
S
Security Affairs
TaoSecurity Blog
TaoSecurity Blog
Cloudbric
Cloudbric
Cisco Talos Blog
Cisco Talos Blog
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
H
Heimdal Security Blog
The Last Watchdog
The Last Watchdog
T
Threatpost
Hacker News: Ask HN
Hacker News: Ask HN
Security Latest
Security Latest
Know Your Adversary
Know Your Adversary
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
S
Securelist
Microsoft Azure Blog
Microsoft Azure Blog
The GitHub Blog
The GitHub Blog
阮一峰的网络日志
阮一峰的网络日志
D
Docker
V
Vulnerabilities – Threatpost
Attack and Defense Labs
Attack and Defense Labs
Hugging Face - Blog
Hugging Face - Blog
W
WeLiveSecurity
Engineering at Meta
Engineering at Meta
aimingoo的专栏
aimingoo的专栏
Last Week in AI
Last Week in AI
L
LINUX DO - 热门话题
NISL@THU
NISL@THU
D
Darknet – Hacking Tools, Hacker News & Cyber Security
The Cloudflare Blog
博客园_首页
P
Privacy International News Feed
Scott Helme
Scott Helme
N
News and Events Feed by Topic
WordPress大学
WordPress大学
宝玉的分享
宝玉的分享
T
Tenable Blog
H
Hacker News: Front Page
N
News and Events Feed by Topic
罗磊的独立博客
Google Online Security Blog
Google Online Security Blog
S
Security @ Cisco Blogs
Hacker News - Newest:
Hacker News - Newest: "LLM"
A
About on SuperTechFans
有赞技术团队
有赞技术团队
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
G
GRAHAM CLULEY
Application and Cybersecurity Blog
Application and Cybersecurity Blog

TestingCatalog

SpaceXAI gearing up for upcoming Grok 4.5 release ByteDance set to launch Seedance 2.5 with 3-minute output Meta prepares Scheduled tasks for Meta AI users on web Google tests new Gemini Inbox section for Workspace triage OpenAI might be preparing GPT-5.6 for next week's release Mistral releases Leanstral 1.5 model for proof engineering xAI debuts Grok Voice Agent Builder for Enterprises Vellum adds agent-to-agent AI collaboration for Slack Early look at Anthropic's Claude Science app for researchers Google launches Nano Banana 2 Lite and Gemini Omni Flash Google might be testing Gemini Flash upgrade on LM Arena Anthropic may impose KYC restrictions for Fable 5 access Anthropic launches Claude Sonnet 5 model on Claude and APIs NoimosAI launches Creative Agent for brand assets Apify lets AI Agents pay via Coinbase x402 for web tools Bloome launches chat platform for AI agent teams Meituan launches LongCat-2.0 1.6T parameter model Cursor releases its iOS app for vibe coding on the go OpenAI prepares upgraded Office controls for Codex OpenAI tests gifting Codex credits as new growth strategy Microsoft launches MAI-Code-1-Flash on GitHub Copilot Google adds Computer Use to Gemini 3.5 Flash Google tests notebook collections for NotebookLM OpenAI launches GPT-5.6 Sol preview for select partners Microsoft adds Copilot finance tools to Excel for M365 users DeepReinforce releases Ornith-1.0 open-source coding models Gemini to get voice dictation and Magic Pointer on desktop Meta launches AI glasses with three new styles from $299 Anthropic launches Claude Tag on Team and Enterprise plans Mistral launches OCR 4 for multilingual document extraction ClickUp rolls out Brain² AI with deep workspace context Latitude launches open-source platform to monitor AI agents OpenAI prepares bidirectional voice mode for rollout Anthropic prepares Cowork support for mobile apps Google tests literature review matrix tool for NotebookLM OpenAI launches new security tools and updates GPT-5.5-Cyber Sakana AI releases Fugu Ultra system to rival top AI labs Perplexity releases Brain Memory for Perplexity Computer Anthropic launches live Artifacts for Claude Code Anthropic launches managed connector access with Okta OpenAI prepares real-time voice mode for Pets in Codex OpenAI prepares GPT-5.6 models for the upcoming release Microsoft evaluates different open models for Cowork Zeta Labs brings AI employee Viktor to Microsoft Teams Mistral AI to get Code and Apps features on Vibe Z AI launches GLM-5.2 open-weight model with 1M context Microsoft launches Copilot Cowork globally for Microsoft 365 OpenAI readies ChatGPT for Science subscription plan Nitrosend launched AI-native email platform for agents Google readies Personalization and AI Editing for NotebookLM OpenAI prepares major ChatGPT voice upgrade with GPT-Bidi-1 Mistral embraces cat mascot after Le Chaton Fat goes viral ICYMI: OpenAI released CDP support for browser use on Codex Google develops Personalization controls for Gemini xAI is working on Automations feature for Grok Cutback launches AI tool to automate long-form video editing Telegram launches watch apps and rich formatting for bots Anthropic suspends Fable 5 and Mythos 5 after export order Google is working on Skills Marketplace for Gemini Business MiniMax M3 launches on NVIDIA platform with Free Endpoint Meta AI to get new modes for Deep Research, Social, and more Maket launches floor plan upload for faster project kick-off NoimosAI launches autonomous AI marketing team Claude Code Managed Agents and model selector for Voice Mode Mora launches AI analytics platform with SQL transparency Maket debuts Auto-Complete for generating floor plans Microsoft rolls out Scout AI agent to Frontier users Anthropic started red teaming new Mythos models OpenSquilla lets AI agents organize their own skills Microsoft Build 2026 recap OpenAI makes its next hardware move with Opal Electronics Google tests planning mode for NotebookLM Video Overviews TinyFish Bigset turns text prompts into live datasets What releases to expect from Anthropic in coming weeks Exclusive: New screenshots of upcoming Copilot Super App Microsoft released MAI voice and image models for Build 2026 3 upcoming NotebookLM features we all should be waiting for Perplexity tests daily Digest for its Computer agent Explee launches AutoGTM AI Agent for outbound sales Anthropic launches dynamic workflows for Claude Code Nano Banana 2 and Pro are now in General Availability Anthropic launches Claude Opus 4.8 and new effort selector Sesame debuts iOS app in preview with personal voice agents Google expands Gemini for Business with shareable Projects Anthropic to expand Claude Voice Mode to more languages Alook launches open-source AI team orchestration platform Helio launches invite-only AI-powered team workspace Anthropic to introduce AI Fluency scorecard in Claude Capafy launches AI Skills Marketplace for creators Anthropic plans Claude memory update with new Memory Files Anthropic prepares Mythos 1 for Claude Code and Security Perplexity open-sources Bumblebee security scanner Google unveils 24/7 Gemini Spark AI Agent for advanced tasks Google launches Gemini 3.5 Flash AI model to all users Google rolls out Gemini Omni AI for video generation Anthropic launches secure sandboxes and private MCPs How to watch Google I/O 2026 and what to expect Cursor released Composer 2.5 with up to 10x cost efficiency Manus released Scheduled Tasks 2.0 upgrade for all users Exclusive: Early look at the next Gemini desktop upgrade
Condense launches proxy to cut AI coding agent bills by 66%
https://www.facebook.com/nero.soares.9/ · 2026-07-03 · via TestingCatalog

Condense.chat has opened public access to a context-compression proxy for coding agents, a system that sits between an agent and the upstream model, shrinking each request before billing. It targets a line item most teams never inspect. In a working agent loop, the model is re-sent the system prompt plus the whole conversation on every turn, so by the middle of a long session, it has re-read the same early context hundreds of times. Across twelve real coding sessions and 18,333 assistant turns, Condense puts cache reads at 67.7 percent of a typical bill, the cost the proxy is built to remove.

We're giving you 100M free tokens to prove your agent is wasting context.

Most of what your agent sends upstream is dead weight. We built state-of-the-art compaction to strip it, and today we published the receipts.

One real session, run to full depth, benchmarked against… pic.twitter.com/40ZdASrHoj

— condense.chat (@densechat) July 3, 2026

The proxy runs two compression models in sequence. Helene 1, an extractive model, scores every token and keeps the survivors verbatim as a strict subset of what the agent saw, stripping content before it reaches cache. Adeline 1, a diffusion-based rewriter, takes settled agent loops the session has moved past and packs each into a short summary that holds intents, file paths, identifiers, errors, and code, landing at roughly 9 percent of the original tokens. The skeleton, meaning the system prompt, user messages, final answers, and paired tool calls with their results, is preserved byte for byte and never rewritten. Only the aged interior of past loops is stripped or packed, so the working set the model is reasoning over stays raw.

Condense.chat

On one real session replayed to 938 turns, Condense reports the bill falling 72.3 percent, taking a Sonnet run from 154 dollars to 43 and an Opus run from 771 to 214, with dollar-weighted savings across every session size near 66 percent. The company states the effect compounds with depth, since a larger context means a larger frozen prefix and a smaller slice of history re-read each turn, passing 53 percent on sessions with chains above 400,000 tokens. Answer faithfulness against uncompressed transcripts is reported at 94.2 percent.

Condense.chat

Setup is a single command that drops the harness in front of an existing agent with no key swap, and the drop-in API speaks both the OpenAI and Anthropic SDKs, so a team can point its base URL at a provider route and keep its own key. The proxy currently drives Claude Code, Codex, and OpenCode across macOS, Linux, and Windows. It is aimed at developers running agent workflows at scale, where re-read costs dominate, and the deepest sessions incur the largest bills.

Condense is giving TestingCatalog readers 100M saved tokens.

Test it out!

Condense is built by engineers with backgrounds shipping large-model infrastructure at Nord Security, Kilo Health, and Nexos AI. The full measurement pipeline is published as an open harness on GitHub, including the cost-split study and the replay tool, so the figures can be recomputed without a key.