METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

Liquid AI unveils screen-reading model that runs in 3GB

LFM2.5-VL-3B hits 20 tokens/sec on Galaxy S26 Ultra, 228 tokens/sec on M5 Max

Liquid AI unveils screen-reading model that runs in 3GB

Summary

  • Liquid AI released LFM2.5-VL-3B, a lightweight vision-language model that reads screens, documents, and real-world images, on Hugging Face
  • At about 3GB of memory under Q4_K_M quantization, it delivered 228 tokens/sec on an Apple M5 Max and 20 tokens/sec on a Galaxy S26 Ultra
  • On a single H100 running vLLM 0.26, the company said it achieved a 34ms time-to-first-token for multi-image input and roughly 11,000 tokens/sec output throughput under high concurrency

Eyes that fit in 3GB

A model running on a single Galaxy S26 Ultra reads an image and writes an answer at 20 tokens per second. No cloud, no account, no upload. That's the measured result Liquid AI published on the 12th for its vision-language model LFM2.5-VL-3B. It has 3.1 billion parameters and occupies about 3GB of memory at 4-bit quantization. Given that mid-range smartphones today typically carry 8-12GB of RAM, a model that reads photos has now shrunk to a size small enough to live inside a phone.

What the model reads

The company positions the model for three use cases. First, screens — reading mobile, web, and desktop UI. Second, text and tables in documents and charts. Third, real objects captured through a camera. Two additional capabilities come attached. One is grounding — the ability to point to a specific object in an image by coordinates. Because it answers with "this coordinate on the screen" rather than "this button," it can serve as the hands of an agent that operates a screen on a user's behalf. The other is tool calling, which the company said works not only with text but also with image inputs.

AMD 라이젠, 애플 M5 맥, 퀄컴 스냅드래곤별 여러 AI 모델의 첫 토큰 생성 시간, 디코드 속도, 메모리 사용량 비교 그래프
이미지: @liquidai (X)

Device-by-device benchmarks — alongside rival models

The attached benchmark table used a 512×512 image plus 1,024 tokens of text as input, generating 64 tokens. Notably, the model doesn't always top the speed rankings. Looking purely at decode speed on the M5 Max, InternVL3.5-2B (260) and Qwen3.5-2B (237) are faster. Instead, LFM2.5-VL-3B strikes a balance between time-to-first-token and memory footprint.

Apple M5 MaxTime to first token (ms) ↓Decode (tok/s) ↑Memory (MB) ↓
LFM2.5-VL-3B2482283,257
gemma-4-E2B2711724,938
gemma-4-E4B4411117,204
InternVL3.5-2B4382602,759
Qwen3.5-4B4061265,172

Things look different on phones. On a Snapdragon-based Galaxy device (SM-S948U1), time to first token was 13,092ms — 13 seconds. Still, InternVL3.5-4B, one of the comparison models, took 65 seconds. The numbers show that on phones, image understanding is still a "wait for it" task rather than an instant answer.

A different picture on servers

Using a single H100 SXM5 with vLLM 0.26, the company presented three figures. Time to first token for multi-image input was 34ms — compared with roughly 200ms for the Gemma family, used as a benchmark. Under high-concurrency conditions, output throughput reached about 11,000 tokens/sec, roughly twice that of the 4B-class models tested. And on a single GPU, the setup could handle about 1 billion output tokens per day. The claim is that a small model's value isn't limited to "running on a phone." Handling more requests on the same GPU also lowers the unit cost of document- and screen-processing pipelines.

vLLM adds same-day support for Nemotron 3.5 Lightning, with three inference acceleration features

GPU H100에서 동시 처리량에 따른 여러 AI 모델의 출력 처리량 변화를 나타낸 선 그래프
이미지: @liquidai (X)

How it was built

The model builds on the company's own LFM2.5, with expanded visual capacity this time. Vision pretraining tokens were quadrupled, and the training data mixed image-caption pairs, OCR, grounding, and instruction-following data, drawn from both curated and synthetic sources. Rather than building a new tokenizer, the company expanded the existing one, doubling the vocabulary to 128,000 tokens. This is seen as contributing to its strong multilingual scores.

Internal evaluationGeneral averageMultilingual
LFM2.5-VL-3B70.181.2
LFM2-VL-3B (previous version)69.179.1
gemma-4-E4B-it (8B)63.275.8
InternVL3.5 4B7069.7–
Qwen3.5-4B7069.7–
gemma-4-E2B-it (5.1B)56.268.0

These figures come from the company's own internal testing, so some caution is warranted. Still, the reported gains over the previous version — 1.0 point on General and 2.1 points on Multilingual — indicate the scale of this update.

What actually changes

Until now, "AI that sees a screen and clicks" has generally required uploading screenshots to a server. Many of those images — internal documents, banking app screens — aren't things people want to send elsewhere. If a 3GB-class model can run on laptops and phones, it opens room to process such images without sending them out at all. This is the mirror image of what Naver did on August 9, when it published its 32-billion-parameter vision-language model HyperCLOVAX-SEED-Think-32B on Hugging Face. One aims for deep reasoning on servers; this one aims for fast answers on-device.

The company itself drew the line clearly: it's suited for tasks needing quick responses, such as screen and UI agents, document and chart reading, grounding, and multi-image tasks, while problems requiring step-by-step reasoning should be handled through standard fine-tuning pipelines. The model is available for download now on Hugging Face.

Naver unveils 32B-class reasoning vision-language model SEED-Think

Comments