Liquid AI Releases LFM2.5-2.6B: An On-Device Agentic Model With 128K Context, Tool Calling, And Open Weights
What it does
Liquid AI has released LFM2.5-2.6B, a large language model designed to run entirely on-device while supporting agentic capabilities. The model includes 2.69 billion parameters, arranged with 22 double-gated short convolution blocks and 8 gated quadratic attention blocks spread across 30 layers. It can process an extreme context window of 131,072 tokens, handling very long inputs without trimming. The model runs efficiently on Apple’s M5 Max chip, decoding at 220 tokens per second while using less than 2.5 GB of RAM. Open weights are distributed in multiple formats including GGUF, MLX, and ONNX, facilitating integration and experimentation.
Why it matters
LFM2.5-2.6B pushes device-resident AI beyond typical limits by combining a massive context window with agentic features like planning and tool calling. This shifts more autonomy from cloud servers to end-user devices, cutting latency, reducing dependency on external infrastructure, and tightening data privacy since sensitive data never needs to leave the device. Builders get a model capable of complex, multi-step task workflows that could power smarter assistants, apps, or automation on laptops, desktops, or mobile platforms without cloud costs or connectivity concerns. The open weights also mean developers can tailor or optimize the model for bespoke on-device use cases.
Who it is for
This model targets AI builders and developers focused on edge inference and advanced agentic functionality in resource-constrained environments. Businesses wanting on-device automation or enhanced local AI with large context handling can leverage LFM2.5-2.6B to decrease reliance on central servers and reduce operational costs. It also offers researchers and startups a high-quality open-weight alternative optimized for tools and planning workflows, lowering barriers to experimenting with complex, lengthy-context on-device AI.
The catch
Despite the advanced feature set, LFM2.5-2.6B still requires a capable device like an Apple M5 Max chip to run smoothly, limiting adoption on less powerful hardware. Its slightly unusual architecture and multi-format releases may require some pipeline adjustment for integration. Running large context windows at speed always stresses memory and compute, restricting real-world throughput in some applications. Users should verify feasibility against their target devices and workflows before full deployment.
What to watch next
Liquid AI’s next steps will be telling—whether they extend support to more chipsets and OS platforms or develop a broader ecosystem for real-world on-device agentic AI tools. Observing community uptake around the open-weight release will shed light on developer interest in extending on-device autonomy. Progress in model optimization and tooling integration will define how quickly builders can put these capabilities into production and what new workflows emerge from accessible large-context agentic local AI.
AI Quick Briefs Editorial Desk