Models & Research

Liquid AI Releases LFM2.5-VL-3B-DSpark: Speculative Decoding for Vision-Language Models With Up to 3.13x Fa…

· September 25, 2026
Liquid AI Releases LFM2.5-VL-3B-DSpark: Speculative Decoding for Vision-Language Models With Up to 3.13x Fa…

What it does

Liquid AI has introduced LFM2.5-VL-3B-DSpark, a draft vision-language model featuring speculative decoding. This model contains 279.5 million parameters and builds on the existing LFM2.5-VL-3B by accelerating the decoding step in generating text outputs from combined visual and language inputs. Speculative decoding, which anticipates likely next steps in the output sequence, speeds up the process significantly without changing the final output when using greedy decoding. Reported performance gains reach 3.13 times faster decoding on Apple M5 Max chips and 2.66 times on Nvidia H100 GPUs.

Why it matters

Decoding speed is a critical bottleneck in deploying vision-language models, especially in applications requiring real-time or near-real-time responses like image captioning, visual question answering, or multimodal chatbots. Slower decoding can inflate infrastructure costs and introduce latency that harms user experience. By cutting decoding time by more than threefold on common hardware, LFM2.5-VL-3B-DSpark makes running these models more efficient and cost-effective, making it easier for businesses and developers to embed advanced multimodal AI into products and services without needing outsized computing resources.

Who it is for

This release matters most to AI developers and product teams working on vision-language applications where speed and resource efficiency are priorities. Builders running models on edge devices like Apple’s silicon or in cloud environments with H100 GPUs will benefit from faster throughput and reduced compute budgets. The model’s availability in widely used frameworks like llama.cpp, MLX-VLM, and SGLang supports easy integration and experimentation, enabling quicker adoption by operators who need to accelerate vision-language workloads.

The catch

While LFM2.5-VL-3B-DSpark accelerates decoding significantly, it remains a draft model. Speculative decoding optimizations often come with trade-offs such as increased complexity in implementation or edge case output behavior, though this release claims identical outputs under the specified decoding method. Users should evaluate stability and compatibility in their specific deployment contexts before full production use.

What to watch next

Expect further refinements to speculative decoding techniques and expanded support across more hardware and software stacks. Watch for more operational metrics from real-world deployments, including cost savings versus traditional decoding. Also, see if Liquid AI extends this approach to larger or more complex vision-language models, pushing speed gains further while maintaining output quality and ease of use.

AI Quick Briefs Editorial Desk

Stay ahead of AI Get the most important AI news delivered to your inbox — free.