METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

AI GlossaryㅋWords you meet while using AI

Qwen3.8-Flash-Next

An open-weight 125B-parameter model released by Alibaba's Qwen, compressible to run locally with just 75GB of memory.

In plain words

Qwen3.8-Flash-Next is an AI model from Alibaba's Qwen that's originally huge, but has been released in a compressed form that can run on personal computers. Think of it like shrinking a thick 355GB encyclopedia down to a 75GB travel-sized summary, yet reportedly keeping accuracy at around 79% of the original.

Because its size has been cut down so drastically, it can run on Macs or certain PCs with enough memory, without needing expensive server-grade GPUs. It can process text and images together, and its context window is large enough to take in very long documents or entire code repositories at once.

In comparison charts released by Alibaba's Qwen and Unsloth, which adapted the model for local use, this model scored higher than Anthropic's Claude Opus 4.6 Max on several task-based benchmarks. However, these are figures released directly by the developer and related companies, and have not been independently reproduced or verified by outside organizations.

How it shows up in the news

Articles use it in sentences like "Qwen3.8-Flash-Next was run locally with just 75GB of memory." Going by the number in its name alone, it might look like a small model, but its actual parameter count is a sizable 125B — meaning it also involved compression (quantization) techniques to make local use possible.

Try it yourself

  1. Search Hugging Face for a repository under this model's name.
  2. Look for a distribution like Unsloth's that provides files and guides for local use.
  3. Choose a compression (quantization) tier that matches your computer's memory. Pick a lighter tier if you have less memory, or a more accurate tier if you have plenty.
  4. Load the file into a local inference tool and test it with prompts like summarizing a document or explaining code.

See also

Stories using this term

Browse every entry