Briefing

Thinking Machines Unveils 276‑Billion‑Parameter Real‑Time Voice Model

ai-dev
by smhx ·

Build an interaction model that processes 200 ms micro‑turns across audio, video, and text, delegating long‑horizon tasks to an asynchronous background model.

What to do now

Prototype the micro‑turn interaction model and integrate it with a background reasoning model for real‑time multimodal collaboration.

Summary

Thinking Machines, a startup known for pushing the boundaries of conversational AI, announced the launch of its latest model, TML‑Interaction‑Small, during a keynote delivered by CEO Neil Zeghidour. The new system is a 276‑billion‑parameter Mixture‑of‑Experts architecture that activates only 12 billion parameters for any given interaction, allowing it to respond within 200 milliseconds. The announcement coincided with Zeghidour’s discussion of real‑time voice capabilities, underscoring the company’s ambition to bring conversational agents closer to production‑ready performance.

On the technical front, TML‑Interaction‑Small outperforms leading competitors such as GPT‑Realtime‑2 and Gemini 3.1‑Flash across a suite of benchmarks, including BigBench Audio, IFEval, and FD‑bench. Its architecture employs encoder‑free early fusion, enabling simultaneous processing of images and audio in under 200 ms—a design choice that mirrors Meta’s Chameleon model. The system’s speed and multimodal proficiency position it as a strong contender for applications that demand instant, context‑aware responses.

To validate its time‑sensitive capabilities, Thinking Machines introduced five proprietary internal benchmarks: TimeSpeak, CueSpeak, RepCount‑A, ProactiveVideoQA, and Charades. These tests assess the model’s ability to initiate speech at user‑specified times, respond to code‑switching cues, maintain continuous visual tracking, answer video‑based questions proactively, and localize temporal actions. Demonstrations revealed that the model can begin speaking at precise moments and adapt its output when users switch languages mid‑conversation, showcasing a level of flexibility that is rare in current real‑time voice systems.

Looking ahead, the company hinted at a roadmap that pairs background agents with interactive models, suggesting future collaborations that blend multimodal data streams. If realized, such integrations could unlock new use cases in customer service, accessibility, and beyond, bringing real‑time voice AI closer to widespread commercial deployment.

Key changes

  • Interaction model processes 200 ms micro‑turns of input and output streams
  • Model can interrupt and respond immediately while the user speaks
  • Supports simultaneous speech (full‑duplex) across audio, video, and text
  • Model has a direct sense of elapsed time and time‑aware context
  • Background model handles sustained reasoning, tool use, and longer‑horizon work asynchronously
  • Interaction model can perform concurrent tool calls, search, and UI generation
  • No separate dialog‑management component; the model tracks conversation implicitly
  • Architecture splits into an interaction model and a background model that share context

Affects

internal

Source angles · 2 perspectives

Hacker News (front page)
Independent angle

Interaction Models

Open
Latent Space
Independent angle

Thinking Machines Unveils TML‑Interaction‑Small, 276B‑Parameter MoE for Real‑Time Voice

Open

Customer impact

Analyzing matches…

Ask about this story

Impact on an agency? Which customers? Compare historically Risks of waiting