AI GlossaryMTechnical words in the news
MXFP4
A low-precision format that represents an AI model's numbers (weights) in as little as 4 bits, cutting down on size and compute
In plain words
MXFP4 is a way of squeezing the huge number of values inside an AI model into a very small number of digits. Normally, when a model is built, each number is stored with around 16 digits of precision, but MXFP4 shrinks that down to just 4 digits. It's similar to how a high-resolution photo takes up a lot of space, but converting it to a compressed file shrinks the size while losing fine detail.
However, MXFP4 doesn't just uniformly cut down the digits. Instead, it groups similar numbers together, sets a single shared scale (a common reference point) for that group, and then represents each individual number only within that scale. This loses less information than simply shrinking every digit evenly across the board. Doing this greatly reduces the memory and compute needed to run the model, making it possible to run bigger models on the same hardware or operate services more cheaply.
The catch is that squeezing numbers down this way usually also reduces the model's judgment — its ability to reason or calculate. So after converting to MXFP4, a separate recovery step to restore performance is often needed.
How it shows up in the news
Articles have described things like "the student model is half the size and runs on MXFP4" or "verified by quantizing further down to MXFP4 with QAH." A common misunderstanding here is assuming that shrinking to MXFP4 automatically improves performance. In reality, cutting precision down to 4 bits usually degrades performance, and the results reported in such articles are exceptional outcomes achieved only because a new recovery technique was applied alongside the quantization.
See also
Stories using this term
- DeepSeek v4 Flash Gets GGUF Build for DwarfStar, Lowering the Bar for Local DeploymentAI · 2026.08.01
- LG AI Research unveils 750B-parameter K-EXAONE 2.0 FP8 modelAI · 2026.08.10
- DeepSeek v4 Flash 0731 run locally with RTX 4090 and Tesla P40 comboAI · 2026.08.10
- NVIDIA unveils 'Ising Calibration 1.5' VLM for automated quantum computer calibrationAI · 2026.08.09
- Liquid AI's 300M Draft Model Speeds Up Decoding by Up to 3.18xAI · 2026.08.21
- Liquid AI unveils screen-reading model that runs in 3GBAI · 2026.08.13
