Liquid AI ships LFM2.5-VL-3B, a 3B vision model built for the edge

Liquid AI has released LFM2.5-VL-3B, a roughly three billion parameter vision-language model designed to run on phones, laptops, and other edge hardware instead of in a data center. The model pairs the LFM2.5-2.6B text backbone with a SigLIP2 NaFlex image encoder. Unlike many recent models, it does not reason step by step. It answers directly, which keeps latency low for real-time and on-device use, and it handles a 32,000 token context. What it is built to do Liquid says the release adds fo

2 min
Liquid AI ships LFM2.5-VL-3B, a 3B vision model built for the edge

Liquid AI has released LFM2.5-VL-3B, a roughly three billion parameter vision-language model designed to run on phones, laptops, and other edge hardware instead of in a data center.

The model pairs the LFM2.5-2.6B text backbone with a SigLIP2 NaFlex image encoder. Unlike many recent models, it does not reason step by step. It answers directly, which keeps latency low for real-time and on-device use, and it handles a 32,000 token context.

What it is built to do

Illustration of a phone screen with an AI marking objects on it

Liquid says the release adds four capabilities over its earlier LFM2-VL-3B: reading digital screens across mobile, web, and desktop; grounding objects to coordinates from a text query; taking multiple images as input; and triggering actions from text or image prompts. Those are the skills that let a model act as an interface layer between what a device sees and what software should do next.

How it scores

On ScreenSpot-v2, which measures how well a model understands UI screens, it averages 80.7, ahead of Google's Gemma-4-E4B at 51.2 and Qwen 3.5 4B at 78.5, and close to the larger InternVL-3.5-4B at 84.1, according to Liquid's own benchmarks. It reaches 87.9 on RefCOCO-avg for grounding, up from 57.1 in the prior release, and its function-calling scores more than double on ToolSandbox and BFCL v4.

Liquid reports the model runs in about three gigabytes of memory, decoding 228 tokens per second on an M5 Max and 20 tokens per second on a Galaxy S26 Ultra, so it can run fully offline on consumer hardware.

Availability

LFM2.5-VL-3B is available on Hugging Face and Liquid's Playground under the LFM Open License v1.0, which allows free commercial use only for companies under 10 million dollars in annual revenue. The release follows a wave of small, efficient vision models aimed at on-device assistants and private inference.

Sources

Liquid AI blog, LFM2.5-VL-3B: https://www.liquid.ai/blog/lfm2-5-vl-3b

Hugging Face blog: https://huggingface.co/blog/LiquidAI/lfm2-5-vl-3b

Liquid Docs: https://docs.liquid.ai/lfm/models/lfm25-vl-3b

MarktechPost: https://www.marktechpost.com/2026/08/13/liquid-ai-lfm2-5-vl-3b-on-device-vision-language-model

Written by

More to read

  • Anthropic Demonstrates Autonomous De Novo Protein Design and Chemical Analysis with Claude

    Anthropic Demonstrates Autonomous De Novo Protein Design and Chemical Analysis with Claude Anthropic has published experimental results demonstrating Claude's ability to autonomously design de novo protein binders with physical wet-lab validation and automate complex analytical chemistry workflows. The findings show frontier LLMs acting as autonomous agents across computational biology and molecular characterization pipelines. In the primary experiment, Anthropic evaluated Claude Mythos Previe

    1 min
  • Cerebras Unveils CS-4 Rack-Scale System Powered by Three WSE-3 Turbo Chips and Nexus Architecture

    Cerebras Unveils CS-4 Rack-Scale System Powered by Three WSE-3 Turbo Chips and Nexus Architecture Cerebras Systems has announced the CS-4, a rack-scale AI accelerator system designed around three of its next-generation Wafer Scale Engine 3 Turbo (WSE-3 Turbo) chips and a modular hardware architecture dubbed Nexus. Cerebras confirmed that initial customer shipments for the CS-4 are scheduled to begin in the current quarter. The new system marks a structural shift from Cerebras's single-wafer CS

    1 min
  • AI FinOps: Cutting LLM Inference Costs by 30-60% Through Model Tiering, Caching, and GPU Optimization

    AI FinOps: Cutting LLM Inference Costs by 30-60% Through Model Tiering, Caching, and GPU Optimization Inference costs have become the second-largest line item in enterprise AI budgets, trailing only talent spend according to RapidData's State of Enterprise AI 2026. This shift represents a fundamental inversion from the 2021-2023 era when training dominated AI expenditure. The compounding nature of serving costs—accumulating every hour as long as users hit the API—means that even modest producti

    1 min