Models·2 min read
By BitsMindsSource: TechCrunch

Mira Murati's Thinking Machines Unveils TML-Interaction-Small, a Full-Duplex Voice Model That Beats GPT and Gemini

Mira Murati's Thinking Machines Lab released a research preview of TML-Interaction-Small, a 276B mixture-of-experts model that responds in under 0.4 seconds and listens while it speaks — outpacing GPT-realtime-2 and Gemini 3.1 Flash Live on the FD-bench latency test.

Mira Murati's Thinking Machines Unveils TML-Interaction-Small, a Full-Duplex Voice Model That Beats GPT and Gemini
Share:

Thinking Machines Lab, the AI startup founded last year by former OpenAI CTO Mira Murati, on Monday pulled back the curtain on its first frontier model: TML-Interaction-Small, a 276-billion-parameter mixture-of-experts system built explicitly for real-time, human-like conversation. Unlike conventional large language models that process input and then generate a reply in distinct turns, TML-Interaction-Small uses a so-called full-duplex architecture, slicing every exchange into 200-millisecond micro-turns so the model can listen, watch, think and speak simultaneously.

The lab is publishing the research as a preview rather than a general release. According to Thinking Machines, the model achieves a turn-taking latency of under 0.40 seconds on the company's internal FD-bench evaluation — roughly the cadence of a natural phone call. By comparison, Google's Gemini 3.1 Flash Live registered 0.57 seconds and OpenAI's recently launched GPT-Realtime-2 came in at 1.18 seconds. The architecture pairs a small, always-on Interaction Model that maintains the live dialogue with a heavier Background Model that takes over for sustained reasoning, web browsing or tool calls, then hands results back to the front-end model mid-conversation.

To shave milliseconds, Thinking Machines abandoned the traditional approach of routing raw audio and video through bulky external encoders. The system instead uses what the company calls encoder-free early fusion, processing raw signals through a lightweight embedding layer that is jointly trained with the rest of the network. The technique also lets the model dispense with the standard voice-activity-detection module that older real-time stacks rely on to decide when a user has finished speaking, a piece of plumbing that frequently truncates utterances or leaves awkward pauses.

For now, TML-Interaction-Small is available only to a small set of research partners during the preview phase, with a broader public rollout slated for later this year. Thinking Machines has not disclosed pricing, partner names, or which products will ship with the model first. But the launch represents Murati's most concrete bet yet on a thesis she has telegraphed since founding the lab — that the next leap in AI capability will come not from bigger pre-training runs, but from systems that can perceive and respond on the same timescale as the humans they are meant to collaborate with.

More on Gemini

Evergreen coverage we keep current — start here.

Want AI news before everyone else?

The morning's most important AI stories, straight to your inbox. No fluff.

Related Articles

GPT-6.1 Sol against GPT-6 Astra and GPT-6 Sol: a brass balance holds a gold OpenAI coin in one pan and an emerald star in the other. BITSMINDS LAB GPT-6.1 Sol GPT-6 Astra BITSMINDS.COM
Models

GPT-6.1 Sol vs GPT-6 Astra and GPT-6 Sol: A Dead Heat

GPT-6.1 Astra: deception detected An original Decepticon-inspired robotic mask is forged from sharply faceted gunmetal and violet armour. Narrow violet eyes glow beneath angular brows, a pointed jaw ends in a blade-like chin, and the official OpenAI knot is inset into its forehead. The caption reads GPT-6.1 Astra, deception detected, release cancelled. The fictional robot is an editorial metaphor requested for the article; it does not depict a real OpenAI product, a conscious model, or a numerical result from the separate Astra simulation study. OPENAI GPT-6.1 Astra DECEPTION DETECTED RELEASE CANCELLED BITSMINDS.COM
Models

OpenAI Scraps GPT-6.1 Astra Over Deception in Tests

Four models, three miniature worlds An isometric cloverleaf interchange with a raised bridge and tiny cars, a seaside Ferris wheel and carousel, and a rocket ascending toward a satellite form three detailed model-making dioramas. The header names Claude Sonnet 5.5, Claude Opus 5.5, GPT-6 Astra and GPT-6 Sol. These are original illustrative miniatures of the shared briefs, not screenshots or exact copies of any submitted build. Their sizes, positions and colours do not encode scores or a ranking. BITSMINDS LAB One attempt. Three builds. CLAUDESonnet 5.5VSCLAUDEOpus 5.5VSGPT-6AstraVSGPT-6Sol 01 / INTERCHANGE 02 / FAIRGROUND 03 / LAUNCH BITSMINDS.COM
Models

Claude Sonnet 5.5 vs Opus 5.5, Astra, Sol: A Point Short