Open Source

Cursor Open-Sources Mixture-of-Kittens (MoK): A Deterministic MoE Training Megakernel for GB300 NVL72 Racks

· August 4, 2026
Cursor Open-Sources Mixture-of-Kittens (MoK): A Deterministic MoE Training Megakernel for GB300 NVL72 Racks

What changed

Cursor Research open-sourced Mixture-of-Kittens (MoK), a deterministic training megakernel designed for mixture-of-experts (MoE) models. MoK integrates all MoE communication and computation into a single GPU kernel, streamlining the training pipeline. According to Cursor, MoK achieves up to 2.37 times the speed of the best publicly available baseline when run on GB300 NVL72 racks equipped with Blackwell SM100 or SM103 GPUs.

Why builders should care

MoE models promise massive scale without a linear rise in compute needs, but their training complexity often bottlenecks performance. Cursor’s MoK tackles this by collapsing multiple MoE operations—routing, expert computation, and communication—into one deterministic kernel, cutting overhead and synchronization delays. The result is significant acceleration on advanced hardware, directly improving training efficiency for large-scale MoE architectures.

However, MoK requires very specialized hardware: the new Blackwell SM100 or SM103 GPUs on GB300 NVL72 racks. This limits adoption to organizations with access to the latest Nvidia infrastructure, sidelining smaller players or those on older GPUs.

The practical takeaway

For teams running massive MoE models on cutting-edge Nvidia systems, MoK offers a valuable tool to squeeze better throughput and reduce training time for complex models like Cursor’s Composer. The open-source release invites tune-ups and wider experimentation, potentially pushing MoE performance forward. But for builders without GB300 NVL72 racks and Blackwell GPUs, the advantage remains out of reach, reinforcing the hardware arms race in large-scale AI training.

What to watch next

Watch whether Cursor or others will adapt MoK’s megakernel approach to more common hardware or support other GPU types. Also, track if performance improvements translate into real-world cost savings or faster training cycles for production MoE systems outside experimental setups. MoK could push Nvidia and cloud providers to prioritize GB300 NVL72 capacity as MoE training demand grows.

AI Quick Briefs Editorial Desk

Stay ahead of AI Get the most important AI news delivered to your inbox — free.