Models & Research

Liquid AI Releases LFM2.5-VL-3B: A 3B Vision-Language Model That Reads Screens, Grounds Objects, and Calls …

· August 13, 2026
Liquid AI Releases LFM2.5-VL-3B: A 3B Vision-Language Model That Reads Screens, Grounds Objects, and Calls …

What it does

Liquid AI launched the LFM2.5-VL-3B, a compact 3.1 billion-parameter vision-language model designed for running entirely on-device. It can analyze screen content like apps or webpages, accurately identify and ground objects visually, and now supports function calling to interact with tools locally. The model’s size is about 3 GB, and it processes around 228 tokens per second on an Apple M5 Max chip.

Why it matters

This model significantly raises the bar for on-device vision-language AI by combining visual understanding and tool interaction without relying on cloud computation. That cuts latency, reduces data privacy risks, and lowers cloud infrastructure costs for developers and businesses building AI-powered apps. The improvements in ScreenSpot-v2 accuracy (80.7 average) and RefCOCO object grounding (jumping from 57.1 to 87.9) show it can interpret screen and scene details with far greater precision. Function calling boosting ToolSandbox performance more than doubles utility in practical tasks like tool manipulation.

Who it is for

Developers building mobile or edge applications that require visual input combined with natural language commands will find LFM2.5-VL-3B useful. It should appeal to teams prioritizing privacy, speed, and offline capabilities for use cases ranging from AR interfaces to smart assistants on personal devices. Investors and tech buyers focused on scalable AI deployments may see this as a sign that powerful vision-language tech is moving away from cloud dependency.

The catch

Despite strong performance, the model’s 3+ GB size still demands a high-end device with ample memory and processing power, limiting usability on entry-level hardware. Function calling is new to this series and likely needs further tool ecosystem development to realize its full potential. The metrics reported are specialized benchmarks, so real-world effectiveness depends on integration and application context.

What to watch next

Keep an eye on how Liquid AI’s approach influences other vision-language model makers toward on-device tool interaction. Watch for new developer tools and APIs to unlock function calling capabilities in actual workflows. Also, monitor adoption by mobile AI applications that require tight integration of vision and language under privacy and latency constraints.

AI Quick Briefs Editorial Desk

Stay ahead of AI Get the most important AI news delivered to your inbox — free.