There is a charming perversity in raising $2 billion to make something less mysterious. Artificial intelligence companies usually sell awe: the model knows more, reasons longer, speaks more smoothly. Thinking Machines Lab sells a wrench. Its first commercial product, Tinker, gives researchers four low-level functions and says, in effect, you may open the hood. The company keeps the GPU cluster, scheduling and failure recovery out of sight. The customer keeps the data, the training recipe and the consequential decisions.
- What it does: managed model training through Tinker, plus open-weight Inkling models and real-time interaction research.
- Who buys: AI teams, researchers and universities with their own examples, evaluations and domain judgment.
- What it costs: usage-priced training and sampling; storage is $0.10 per GB-month. The company itself has raised $2 billion in confirmed seed funding.
- The catch: customization rewards good data and measurable outcomes. It cannot manufacture expertise a customer never possessed.
Mira Murati created the company in February 2025 after leaving OpenAI to make room for what she called her “own exploration.” She recruited five fellow OpenAI veterans: John Schulman, Barret Zoph, Lilian Weng, Andrew Tulloch and Luke Metz. The biographies were the first product. Andreessen Horowitz led a $2 billion seed round five months later, valuing a company with no released product at $12 billion.
A cluster reduced to four verbs
Tinker arrived in October 2025. It did not answer questions, draw pictures or pretend to be a friend. It exposed forward_backward, optim_step, sample and save_state. Those verbs let a team fine-tune an open model with supervised learning or reinforcement learning while Thinking Machines operates the distributed infrastructure underneath. Changing from a small model to a much larger mixture-of-experts model can be as simple as changing one string.
The bill is equally unromantic. Tinker charges by the million tokens for input processing, sampling and training, with rates depending on the model. Cached prompts receive a discount; checkpoint storage is $0.10 per gigabyte each month. Enterprise terms exist. The open models and Apache-licensed cookbook attract experimenters, while the meter runs on the managed machinery.
The customer brings the secret ingredient
Early users make the theory concrete. Glean trained Waldo, a search model built from anonymized production traces. Inside Glean's system it cut search latency by 50 percent and used 25 percent fewer tokens. Chroma trained a small search subagent that matched or beat several frontier models on web, email, finance and legal retrieval while running up to ten times as many tokens per second. Mantic used roughly 10,000 forecasting questions to turn a model that trailed the frontier into one that edged past it.
These are not miracles of scale. They are victories of specificity. The general model has read the world; the customer knows which documents mattered, which forecast resolved and which tool call wasted time. Thinking Machines turns that private judgment into training signal.
“Inkling is not the strongest overall model available today, open or closed.”Thinking Machines Lab, introducing Inkling
The honest model
Inkling, released with full weights in July 2026, contains 975 billion parameters, with 41 billion active for each token. It can reason across text, images and audio, accepts a context window up to one million tokens and lets a user vary how much computation it spends thinking. Inkling-Small followed two weeks later with 276 billion total parameters and 12 billion active.
The unusual sentence in the launch was the disclaimer above. A weaker model sounds like a commercial defect only if the model is the finished product. Thinking Machines treats it as material. Inkling is broad enough to shape, efficient enough to run and open enough to inspect. Customers can download the weights, fine-tune them on Tinker and deploy the result on infrastructure they control. The moat, if one forms, lies in the loop between a customer's proprietary evidence and the lab's training machinery.
What failed first
Failure appears in the product stories, which is where it belongs. Trajectory tried a naive importance-sampling fix for off-policy reinforcement learning. Training collapsed. Its replacement clipped importance ratios and capped advantages; on the APEX-Agents benchmark, the resulting recipe raised a model's pass rate from 5 percent to 25 percent and halved wall-clock time. Mantic's researchers found open-source reinforcement-learning frameworks too slow to iterate with before moving their experiments to Tinker.
The company has suffered a more public failure of cohesion. Tulloch left for Meta in 2025. Zoph and Metz returned to OpenAI in January 2026; Soumith Chintala became CTO. Weng stepped down in July and also returned to OpenAI. Four of six announced cofounders were gone within roughly eighteen months. No amount of compute makes that ordinary. Yet Tinker reached general availability, two models shipped, research continued and the NVIDIA partnership expanded. The live question is whether the institution has become more durable than its founding cast.
The industrial and the intimate
In March 2026, NVIDIA agreed to help deploy at least one gigawatt of Vera Rubin systems, targeted to begin in 2027, and made another investment. A reported Google Cloud agreement, worth several billion dollars and not exclusive, added GB300 systems. This is the industrial side of making AI personal: someone must first purchase an astonishing quantity of electricity and silicon.
The intimate side is the lab's interaction research. Most agents are designed to disappear with a task and return later. Thinking Machines argues that useful work is rarely specified perfectly at the start. Its preview model processes audio, video and text in simultaneous micro-turns so a person can interrupt, point, correct and continue. The ambition is less like sending a memo to a clever intern and more like sharing a desk with one.
Choose one narrow job with abundant examples and a score you trust. Start with the smallest capable open model. Train on your own traces, reward both correctness and efficiency, and compare against the untouched base model. Do not begin with “we need fine-tuning.” Begin with a repeated error you can measure.
This approach will not work everywhere. A vague creative task has no clean reward. Sparse or biased examples merely automate sparse or biased judgment. A company that wants the best general assistant tomorrow may be happier renting a closed frontier API. And teams without the discipline to evaluate their model will acquire a custom model without acquiring knowledge.
That is what changed between Murati's old job and this one. At OpenAI, the landmark achievement was a general system delivered to everyone. At Thinking Machines, the wager is that “everyone” is not a use case. A chemist, a search engineer and a financial forecaster do not need identical intelligence. They need a capable base, their own evidence and the means to alter the machine. The company has spent lavishly to make that alteration feel almost mundane. In this business, mundanity may be the more radical product.