COMPANY
Thinking Machines Lab
Overview
Thinking Machines Lab (TML) is an AI research lab founded by Mira Murati, former CTO of OpenAI. The lab focuses on interactive AI systems, particularly voice and video interaction with sub-half-second latency. Its architectural thesis is that true interactivity requires end-to-end system design rather than scaffolding TTS/VAD components onto a text model.
Timeline
- 2026-05-13-AI-Digest — TML released its first model, TML-Interaction-Small, on May 12 as a limited research preview (partner access only, no public GA, no pricing disclosed). The model is a 276B-parameter mixture-of-experts with 12B active parameters, processing audio and video in 200ms parallel chunks and self-deciding when to interject; it hits a 0.40s response latency floor versus GPT-Realtime-2’s 1.18s minimum. TML’s framing — “interactivity is what OpenAI gets wrong about voice” — positions the lab as an architectural critic of scaffolded voice approaches from the first ship.
- 2026-07-16-AI-Digest — TML ships Inkling, a 975B-parameter mixture-of-experts with ~41B active trained on 45T multimodal tokens across text, image, audio, and video — first open-weights release from the lab — paired with the Tinker fine-tuning platform and a dial-able “thinking effort” that trades quality for latency. Notable inclusion: Mira Murati‘s lab explicitly concedes Inkling isn’t the strongest general model and is betting enterprises want customizability, on-prem inference, and calibrated uncertainty over leaderboard wins. Existing ~$2B seed at ~$10–12B valuation (closed pre-Inkling with a16z and NVIDIA on cap table) frames this as a distribution move, not a fresh raise. Narrow read: Inkling is a real US frontier-lab open-weights entrant but the disclaim-the-frontier framing matters — this is not a bet that open-source wins the Aider leaderboard where GPT-5 variants still hold four of five slots. Structural read: the corpus two-leaderboards frame (2026-07-15-AI-Digest‘s Chinese open vs US closed distribution split) now needs sharpening to a three-way split — Chinese open frontier / US open below-frontier / US closed frontier — with the interior question being whether US enterprise fine-tunes push customized Inkling past Chinese open-weight peers on domain evals. 60-day watch: whether the first credible US-enterprise Inkling fine-tune lands and posts a comparable domain-eval score.
Key Developments
-
TML-Interaction-Small (May 2026): 276B-parameter MoE (12B active), 200ms audio/video chunks, 0.40s response latency floor — the first model release from the lab and a direct architectural counter-thesis to scaffolded voice systems. Limited research preview; partner access only.
-
Inkling — 975B Open-Weights MoE Disclaiming the Frontier (July 16, 2026): TML’s first open-weights release, 975B MoE with ~41B active on 45T multimodal tokens across text/image/audio/video, paired with Tinker fine-tuning platform and dial-able “thinking effort.” The lab explicitly concedes Inkling isn’t the strongest general model — the bet is that enterprises want customizability, on-prem inference, and calibrated uncertainty over leaderboard wins. Sharpens the corpus two-leaderboards frame to a three-way split: Chinese open frontier / US open below-frontier / US closed frontier. Interior question: whether US enterprise fine-tunes push customized Inkling past Chinese open-weight peers on domain evals.
- 2026-07-17-AI-Digest — TML pushed Inkling‘s Tinker fine-tuning platform to a scheduled price increase today — roughly ~50% on prefill and sample inference, ~10% on training. First meaningful cost-adjustment signal from a frontier fine-tuning platform, landing the same news slot Anthropic bookrunners began pre-roadshow investor meetings on the $965B S-1 — the compute-market tightening the Tinker hike signals is the backdrop Anthropic’s first-profitable-quarter posture is being priced against. Narrow read: single-platform price adjustment on a scheduled cadence, not a broad frontier-fine-tuning re-pricing yet. Structural read the corpus carries: the Tinker hike is the first datapoint on compute-market tightening from the fine-tuning-platform side — the enterprise Inkling-fine-tune adoption question from 2026-07-16-AI-Digest now has a paired cost variable, and the fine-tune-vs-domain-eval math shifts in exactly the direction that raises the bar for a US-open-below-frontier customization thesis.
- Tinker Scheduled Price Increase — First Frontier Fine-Tuning Cost-Adjustment Signal (July 17, 2026): ~50% hike on prefill and sample inference, ~10% on training on Inkling’s Tinker fine-tuning platform. Reads as compute-market-tightening evidence from the fine-tuning-platform side rather than a TML-specific move, and lands the same news slot Anthropic bookrunners began pre-roadshow investor meetings on the $965B S-1. Raises the bar on the 2026-07-16-AI-Digest US-enterprise-Inkling-fine-tune thesis by shifting the fine-tune-vs-domain-eval math directionally against customization.
- 2026-07-21-AI-Digest — TML’s Inkling resurfaces in today’s TechCrunch-cited framing that consolidates the launch shape: released 2026-07-15 as a 975B MoE (41B active) open-weight model under Apache 2.0, monetised through the Tinker fine-tuning platform rather than per-token API charges — an explicit bet that enterprises want to modify and self-host rather than rent tokens. TML says explicitly that Inkling “is not the strongest overall model available today” — unusually calibrated launch language for a first-model announcement. Structural read the digest carries: a first-model release that ships open-weight, foregrounds Tinker as the revenue lane, and openly concedes it isn’t the frontier is doing pricing power differently than OpenAI and Anthropic do — TML is building the customisation-surface business rather than the token-margin business. Extends the 2026-07-16-AI-Digest three-way-split reframe (Chinese open frontier / US open below-frontier / US closed frontier) with the sharpest read yet on the why of the US-open-below-frontier leg’s monetisation shape.