Thinking Machines Lab, a startup founded by former OpenAI chief technology officer (CTO) Mira Murati, has launched its first internal artificial intelligence model called Inkling.
Inkling is an open-weight AI model that is expected to provide accurate and well-calibrated answers, follow user instructions and offer configurable reasoning levels. The users can make the model’s “thinking efforts” go up or down based on their needs for more accuracy or speed of response.
The model is versatile enough to be used for coding, reasoning and AI agent workflows. Currently, the output of Inkling includes text such as codes, structured data, and other information presented in format. Its versatility makes it a tool that can be useful in different types of workflows instead of being focused only on benchmarks.
Contrary to OpenAI, Anthropic, and Google, Thinking Machines Lab has released Inkling under an open source license. Thus, the developers and companies can download model weights from Hugging Face and customize them.Organizations are expected to get better value from AI models that can be customized to fit their unique use cases than universal AI solutions. To back up this assertion, Inkling can be customized using the startup’s platform, called Tinker, while the firm expects to make money from hosting and customization rather than accessing the AI model.
Inkling is a large language model that contains 975 billion parameters with a context window of one million tokens. This AI model uses Mixture-of-Experts (MoE) architecture, where it activates about 41 billion parameters for all its tasks.
The model was pretrained from scratch on 45 trillion tokens of multimodal content like text, images, audio, and video, offering natural multimodal reasoning. In addition, during the development of this AI, Thinking Machines also relied on open-weight models like Kimi 2.5 by Moonshot AI to train early data through reinforcement learning.
Inkling was trained using Nvidia’s GB300 NVL72 systems due to a recent collaboration agreement between the two firms.In performance testing, Thinking Machines stated that Inkling was comparable to Nvidia’s Nemotron 3 Ultra in code completion tasks but used just about a third of the tokens. It also enables controllable reasoning capabilities in code and AI agents.
Thinking Machines admitted that Inkling is not the strongest AI model today. Rather, it markets the model as a very good open-weighted base for enterprise modifications that incorporate multimodality, reasoning, and tuning via Tinker.
In addition to Inkling, Thinking Machines also launched a smaller version called Inkling-Small with 12 billion active parameters and low latency.



