Thinking Machines Lab Releases Inkling-Small: A 276B Total, 12B Active Open Weights Multimodal MoE Model
What it does
Thinking Machines Lab has released Inkling-Small, a new multimodal Mixture of Experts model with a total capacity of 276 billion parameters. Unlike monolithic models, Inkling-Small activates only 12 billion parameters at a time, drastically reducing compute requirements during inference. The model supports multiple input types and is provided with open weights, including an efficient NVFP4 checkpoint that can run on a single NVIDIA B300 GPU. This makes it significantly more accessible for smaller-scale deployments while matching the performance of its larger Inkling predecessor at about one-quarter of the size.
Why it matters
Inkling-Small’s design forces a rethink of scaling laws in large models. By activating a fraction of its total parameters on demand, it lowers the hardware bar for running advanced multimodal models. This reduces costs and complexity for businesses and researchers needing powerful AI without access to large GPU clusters. As a result, operational budgets and infrastructure demands can shrink, enabling more institutions and startups to experiment with cutting-edge MoE architectures. The open weights also facilitate transparency and customization, encouraging broader adoption and competition in this space.
Who it is for
This model serves AI developers and teams focused on multimodal applications—combining text, images, and possibly other inputs. Organizations with limited GPU resources but requiring next-gen AI performance can integrate Inkling-Small more feasibly than full-scale models. Investors and businesses tracking AI infrastructure economics should note the pressure this puts on hardware vendors and cloud providers to adapt pricing models. Additionally, labs experimenting with MoE architectures gain a practical demonstration that large parameter counts don’t have to translate directly to computational cost spikes.
The catch
While Inkling-Small runs on a single NVIDIA B300 GPU in NVFP4 precision, performance specifics beyond matching Inkling at smaller size remain to be independently validated. Mixed precision and MoE routing introduce complexity in deployment and reliability that can challenge builders not versed in these architectures. The engineering effort to optimize for uncommon hardware setups might limit immediate off-the-shelf use. Also, open weights require responsible use given the model’s large capacity and multimodal capabilities, which raise standard concerns about misuse in image and text generation.
What to watch next
Monitor the uptake of Inkling-Small in real-world use cases that demand efficient multimodal AI. Watch whether more labs follow with similar sparse MoE designs focused on lowering GPU requirements. Pay attention to benchmarks comparing Inkling-Small with other open models on efficiency and inference speed, especially on modest hardware. Also track ecosystem responses such as cloud providers adjusting offerings for sparse MoE models, and whether this shifts AI deployment economics in the coming quarters.
AI Quick Briefs Editorial Desk