Liquid AI’s newest vision-language model, LFM2.5-VL-3B, targets jobs that run on a device: reading screens, driving GUI automation, pulling text off invoices and translating street signs. The company calls it its most capable vision model yet and says it beats models twice its size on vision benchmarks while running faster across CPUs and GPUs. Weights are live on Hugging Face and in Liquid’s playground.
Tool use is a first for the vision line. ToolSandbox more than doubles to 59.5 from 26.4, and BFCL v4 climbs from 20.5 to 32.5. Screen understanding averages 80.7 on ScreenSpot-v2, comfortably ahead of the much larger Gemma-4-E4B at 51.2 and Qwen3.5-4B at 78.5, while grounding precision on RefCOCO jumps about 30 points to 87.9.
Multi-image reasoning improves as well, with BLINK rising from 50.2 to 61.5 and MuirBench from 34.9 to 58.3. The non-reasoning model answers directly to keep latency low, fits in roughly 3GB of memory, and ships in native, GGUF, ONNX and MLX formats with day-one support for llama.cpp, MLX, vLLM, SGLang and ONNX runtimes. A SigLIP2 vision tower handles the image side.
Liquid AI keeps the model open under its LFM Open License: commercial use is free until a company passes $10M in annual revenue, at which point a commercial agreement kicks in, while research and nonprofit use carry no revenue limits. The company lists on-device screen agents, GUI test automation and OCR as primary targets.