惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

D
DataBreaches.Net
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
Google DeepMind News
Google DeepMind News
博客园 - 聂微东
Microsoft Azure Blog
Microsoft Azure Blog
V
Visual Studio Blog
IT之家
IT之家
博客园 - 【当耐特】
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
B
Blog
爱范儿
爱范儿
阮一峰的网络日志
阮一峰的网络日志
云风的 BLOG
云风的 BLOG
Vercel News
Vercel News
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
H
Hackread – Cybersecurity News, Data Breaches, AI and More
H
Help Net Security
J
Java Code Geeks
aimingoo的专栏
aimingoo的专栏
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
B
Blog RSS Feed
Blog — PlanetScale
Blog — PlanetScale
S
SegmentFault 最新的问题
Apple Machine Learning Research
Apple Machine Learning Research

SiliconANGLE

Will agentic AI governance run amok? The lesson of Asimov’s Three Laws - SiliconANGLE AI + quantum, Amazon vs. Starlink and the wide-open US-China internet battle - SiliconANGLE Team Cymru launches Total Insights Feed to replace legacy threat intelligence lists - SiliconANGLE AI Mode in Chrome adds split-screen view to enhance the web search experience - SiliconANGLE Resolve AI raises $40M at $1.5B valuation to optimize production environments - SiliconANGLE How Zscaler and OpenAI turn zero-trust security into an AI accelerator - SiliconANGLE OpenAI ratchets up Codex's agentic capabilities to rival Claude Code - SiliconANGLE Anthropic launches Claude Opus 4.7 with coding, visual reasoning improvements - SiliconANGLE Slash raises $100M at a $1.4B valuation to expand AI-powered banking platform for online businesses - SiliconANGLE Canva unveils Canva AI 2.0, recasting its platform as an agentic system for work - SiliconANGLE Data center, consumer device chips boost TSMC’s revenue - SiliconANGLE Mission-critical security cannot be bolted on, says Oracle - SiliconANGLE Agentic infrastructure reshapes enterprise AI - SiliconANGLE Data quality, and data freedom, foundational for AI success - SiliconANGLE Data trust is a bedrock in successful, scalable AI outcomes - SiliconANGLE Google introduces new agentic AI-ready tools and resources for Android developers  - SiliconANGLE Agentic AI orchestration separates winners from laggards - SiliconANGLE Data-driven tools turning the tide against human trafficking - SiliconANGLE Achieving trusted AI development goes beyond 'vibes' - SiliconANGLE Impinj boosts edge computing power in updated R700 RAIN RFID reader - SiliconANGLE Certinia powers professional services with AI - SiliconANGLE Antioch prepares to accelerate simulated testing for autonomous robots after raising $8.5M - SiliconANGLE Developer tooling startup Expo nabs $45M investment - SiliconANGLE Solidroad lands $25M to bring AI to customer support interactions - SiliconANGLE DuploCloud lands compliance and AI governance certifications as enterprise buyers tighten scrutiny - SiliconANGLE Lua lands $5.8M to help businesses build and manage AI agent workforces - SiliconANGLE Best of frenemies: Oracle's and AWS' clouds unite with dedicated, private connectivity - SiliconANGLE NIST shifts National Vulnerability Database to risk-based triage as CVE submissions hit record levels - SiliconANGLE Cisco goes to the races with new Churchill Downs multiyear partnership - SiliconANGLE Susecon 2026 will tackle the future of open-source platforms - SiliconANGLE
Thinking Machines drops a new, highly responsive model de...
Mike Wheatle · 2026-05-12 · via SiliconANGLE

Thinking Machines drops a new, highly responsive model designed for humanlike interactions in real-time

Thinking Machines Lab Inc., the artificial intelligence research startup founded by former OpenAI Group PBC Chief Technology Officer Mira Murati, wants to move beyond the era of “turn-based” AI interactions. The company has just announced a research preview of its first “interaction models,” which are a new class of multimodal AI systems designed to avoid the inevitable pauses that characterize human interactions with AI systems.

As anyone who uses AI regularly knows, the basic interaction is a spotty one, at best: The user provides an input, such as text or an image upload, then waits anywhere from a few milliseconds to several minutes, depending on the model used, before finally receiving the output.

This occurs because existing models need to wait for their users to finish asking a question or complete the sentence they’re saying before they can start processing a response. To get around this, Thinking Machines has created an entirely new model architecture that enables “full-duplex” communication, which means AI that can listen, see and talk simultaneously.

Thinking Machines argues that the back-and-forth interactions with current models forces human users to “contort themselves” to the interface. Over multiple months of use, humans have learned to phrase their questions like emails and batch their thoughts, because they know the AI they’re using cannot handle interruptions or deal with the subtle “backchanneling,” or the “mhmms” and “I sees” that exist in truly natural human interactions. But if AI is to become a true humanlike collaborator in high-stakes applications like medical surgery, it has to find a way to ditch that lag.

The company’s answer is a new model architecture that drops the standard alternating token sequence in favor of a larger, multistream micro-turn-based design. The way it works is the system processes inputs and outputs in tiny 200-millisecond chunks, enabling it to react in real-time to any visual or auditory cues it picks up on, even when it’s already speaking. The startup says this “dual-model” architecture is designed to balance speed with deep reasoning.

The first component of this new architecture is TML-Interaction-Small, a 276-billion parameter Mixture-of-Experts model that’s designed to manage dialogue, presence and immediate follow-ups with rapid speed. It’s paired with an asynchronous agent that’s meant to work behind the scenes, so while the Interaction Model keeps the conversation flowing, the Background Model takes care of all of the heavy lifting – the complex reasoning, web searches and tool calls required to get things done or work things out. It can then send its findings to the Interaction Model when it’s ready, and these will be woven into the live chat.

In a blog post, the company explained that instead of using heavy external encoders to translate audio or video into signals the model can understand, it utilizes “encoder-free early fusion” that takes in raw signals directly through a lightweight embedding layer. Everything is processed rapidly within the transformer, which is what gives it such an advantage in terms of latency.

Thinking Machines claims that this dual-model architecture delivers some impressive results. On FD-bench, a benchmark designed to measure AI interaction quality, TML-Interaction-Small achieved a turn-taking latency of less than 0.4 seconds, well ahead of Google LLC’s Gemini-3.1-flash-live, which clocked in at 0.57 seconds, and GPT-realtime-2.0, which achieved a score of just 1.18 seconds.

While speedier chatbots will be appreciated by most people, the most significant implications could be found in enterprise applications. Models that can see and react in real-time pave the way for possibilities that simply don’t exist when dealing with the latency prevalent with today’s models. For instance, a native interaction model could be set up to monitor a video feed in a laboratory or a manufacturing facility and alert humans the moment a safety violation occurs, rather than waiting for a human supervisor to stroll past and see it with their own eyes. In customer service, the lower latency can help to make calls feel more like real conversations.

What’s especially useful is that Thinking Machine’s models have an internal sense of time, which allows them to manage time-sensitive requests. A user in a lab could tell a model to “alert me if this chemical reaction takes longer than the last one,” without needing to provide any timestamps in the prompt.

Thinking Machines says TML-Interaction-Small and its partnering background model are only being made available to a select number of partners during the research preview phase, with a public release slated for later in the year.

Main image: SiliconANGLE/Gemini

A message from John Furrier, co-founder of SiliconANGLE:

Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities.

  • 15M+ viewers of theCUBE videos, powering conversations across AI, cloud, cybersecurity and more
  • 11.4k+ theCUBE alumni — Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network.

About SiliconANGLE Media

SiliconANGLE Media is a recognized leader in digital media innovation, uniting breakthrough technology, strategic insights and real-time audience engagement. As the parent company of SiliconANGLE, theCUBE Network, theCUBE Research, CUBE365, theCUBE AI and theCUBE SuperStudios — with flagship locations in Silicon Valley and the New York Stock Exchange — SiliconANGLE Media operates at the intersection of media, technology and AI.

Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.