Liquid AI Releases LFM2.5-VL-3B: A 3B Vision-Language Model That Reads Screens, Grounds Objects, and Calls …
What it does
Liquid AI launched the LFM2.5-VL-3B, a compact 3.1 billion-parameter vision-language model designed for running entirely on-device. It can analyze screen content like apps or webpages, accurately identify and ground objects visually, and now supports function calling to interact with tools locally. The model’s size is about 3 GB, and it processes around 228 tokens per second on an Apple M5 Max chip.
Why it matters
This model significantly raises the bar for on-device vision-language AI by combining visual understanding and tool interaction without relying on cloud computation. That cuts latency, reduces data privacy risks, and lowers cloud infrastructure costs for developers and businesses building AI-powered apps. The improvements in ScreenSpot-v2 accuracy (80.7 average) and RefCOCO object grounding (jumping from 57.1 to 87.9) show it can interpret screen and scene details with far greater precision. Function calling boosting ToolSandbox performance more than doubles utility in practical tasks like tool manipulation.
Who it is for
Developers building mobile or edge applications that require visual input combined with natural language commands will find LFM2.5-VL-3B useful. It should appeal to teams prioritizing privacy, speed, and offline capabilities for use cases ranging from AR interfaces to smart assistants on personal devices. Investors and tech buyers focused on scalable AI deployments may see this as a sign that powerful vision-language tech is moving away from cloud dependency.
The catch
Despite strong performance, the model’s 3+ GB size still demands a high-end device with ample memory and processing power, limiting usability on entry-level hardware. Function calling is new to this series and likely needs further tool ecosystem development to realize its full potential. The metrics reported are specialized benchmarks, so real-world effectiveness depends on integration and application context.
What to watch next
Keep an eye on how Liquid AI’s approach influences other vision-language model makers toward on-device tool interaction. Watch for new developer tools and APIs to unlock function calling capabilities in actual workflows. Also, monitor adoption by mobile AI applications that require tight integration of vision and language under privacy and latency constraints.
AI Quick Briefs Editorial Desk