Models & Research

Liquid AI Releases LFM2.5-Encoder-230M and LFM2.5-Encoder-350M: Bidirectional Encoders That Stay Fast at 8K…

· July 29, 2026
Liquid AI Releases LFM2.5-Encoder-230M and LFM2.5-Encoder-350M: Bidirectional Encoders That Stay Fast at 8K…

What it does

Liquid AI has launched two bidirectional encoder models, LFM2.5-Encoder-230M and LFM2.5-Encoder-350M, designed to handle very long contexts up to 8,192 tokens. These models are based on the LFM2 hybrid backbone architecture. The 230M model executes an 8,000-token forward pass on a CPU in roughly 28 seconds, highlighting efficient performance in large-context settings without relying on GPUs. The 350M model ranks fourth among 14 evaluated models on a comprehensive set of 17 tasks spanning GLUE, SuperGLUE, and multilingual benchmarks, trailing only larger and potentially slower models.

Why it matters

Handling 8,192 tokens with a bidirectional encoder at practical speeds on CPU opens new opportunities for organizations and developers who need to process long documents without expensive hardware. Most bidirectional transformer models often struggle to scale beyond typical 512 or 1,024-token limits or require GPU acceleration. Liquid AI’s approach offers a cost-effective way to embed large context efficiently, improving applications like document search, summarization, and multi-turn dialogue systems, especially where infrastructure budgets are constrained. The strong performance of the 350M indicates that the model is well-balanced in accuracy and scalability, making it a credible alternative to heavier encoders.

Who it is for

These models suit developers and teams working with CPU-bound environments or edge computing devices where GPU availability is limited or power is constrained. Enterprises dealing with large textual datasets but seeking to avoid high cloud or hardware costs will find these encoders useful. Also, companies building multilingual or multitask applications can leverage the 350M’s high ranking on several language understanding benchmarks for improved model quality without scaling to massive parameter counts.

The catch

Although the 230M model can run 8K-token passes in about 28 seconds on CPU, this speed is still slow for real-time applications. Users prioritizing latency may still need GPU acceleration or smaller context windows. Additionally, while the 350M ranks well, it does not outperform the largest models on all tasks, so some accuracy trade-offs persist for the sake of speed and efficiency. As open-weight models, integration and fine-tuning will also require technical expertise.

What to watch next

Monitor how Liquid AI’s long-context encoders perform in real-world deployments, especially in retrieval-augmented generation and conversational AI tasks. Watch for community adoption and fine-tuning utilities that make these models easier to customize. Competitors may respond with similar long-context CPU-optimized encoders, tightening cost-performance trade-offs in this space. Finally, updates that reduce inference time further or expand the parameter scale without sacrificing speed could shift operational standards for scalable bidirectional encoders.

AI Quick Briefs Editorial Desk

Stay ahead of AI Get the most important AI news delivered to your inbox — free.