Models & Research

Liquid AI Releases LFM2.5-DSpark Draft Models That Deliver Up to 3.18x Faster Decoding Without Changing Mod…

· August 20, 2026
Liquid AI Releases LFM2.5-DSpark Draft Models That Deliver Up to 3.18x Faster Decoding Without Changing Mod…

What changed

Liquid AI has introduced three draft models based on their LFM2.5 large language framework, each with roughly 300 million parameters. These new draft models implement speculative decoding techniques, which accelerate the decoding process by up to 3.18 times compared to standard greedy decoding. Crucially, this speed-up does not alter the final output, preserving the identical greedy generation quality users expect.

Why builders should care

Speculative decoding splits the generation workload by first producing “draft” output tokens more quickly, then verifying them with a slower but accurate model. By combining smaller, faster drafters with the established LFM2.5 framework, Liquid AI offers a way to accelerate inference without degrading response fidelity. For developers and AI operators managing cost and latency, this means faster results without the risk of quality loss or retraining larger models.

The practical takeaway

Operators running LFM2.5 models can integrate these draft models to cut latency for real-time applications or batch generation tasks. The draft models act as a low-overhead front line, reducing compute time and lowering infrastructure costs without the common tradeoff in output accuracy. This enhancement is particularly relevant for startups and businesses seeking higher throughput on mid-sized models without jumping to larger, more expensive architectures.

What to watch next

It will be important to track Liquid AI’s rollout of these speculative decoding models across different LFM versions and whether the approach scales to larger parameter counts. Adoption hinges on how easily these draft models integrate into existing pipelines and their performance in diverse use cases. Competitors might respond by pushing similar draft or multi-model decoding strategies to optimize latency and cost.

AI Quick Briefs Editorial Desk

Stay ahead of AI Get the most important AI news delivered to your inbox — free.