Aleph Alpha Releases Kolibri: A 78.1B Open-Weight English-German MoE Model With Only 3.46B Active Parameters
What changed
Aleph Alpha unveiled Kolibri, a new English-German Mixture-of-Experts (MoE) model with 78.1 billion parameters. Unlike standard large models that activate all weights for each input, Kolibri activates only 3.46 billion parameters per token. This selective activation lowers computational demand while preserving the advantages of a large parameter space. Kolibri also supports a one-million-token context window, allowing it to handle very long documents and contexts efficiently. The model weights use Apache 2.0 licensed FP8 format, making them runnable on a single B200 or H200 AI accelerator.
Why builders should care
Large MoE models promise big performance without the full compute cost of standard models, but have been tricky to deploy at scale because of their complexity and infrastructure requirements. Kolibri changes that by offering an open-weight model that runs on a single, modern AI chip with much lower active parameter counts during inference. Builders working on English-German language tasks or multilingual pipelines now have a scalable, efficient MoE option that balances capacity and speed. The long context window also suits document-centric applications in legal, academic, or technical fields that need consistent reasoning across large inputs.
The practical takeaway
Kolibri’s design forces a rethink of model scaling beyond parameter count alone. This approach lets operators save on expensive compute while still getting the benefits of expansive model capacity. Businesses building translation, summarization, or knowledge management tools can embed Kolibri to process more data faster without paying the usual price in hardware or cloud costs. The Apache 2.0 licensing further removes traditional barriers, enabling more experimentation and deployment flexibility. However, the effectiveness will depend on real-world benchmarks outside Aleph Alpha’s internal tests and how well developers can integrate the mixture-of-experts routing into their existing stacks.
What to watch next
Keep an eye on independent evaluations of Kolibri’s performance and cost-efficiency on real workloads. Also watch Aleph Alpha’s broader AI ecosystem developments, including any APIs or frameworks that leverage Kolibri and simplify MoE model adoption. The ability to run this scale model on single-chip solutions pressures competitors relying on full-parameter dense models to optimize for cost-effective scaling. Finally, observe if Kolibri’s approach triggers more open-weight MoE releases, as well as the uptake in industries with multilingual and long-context demands.
AI Quick Briefs Editorial Desk