METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

AI GlossaryNTechnical words in the news

NVFP4 Quantization

A compression method that shrinks the numbers an AI model uses down to a simple 4-bit format, drastically cutting size and compute needs.

In plain words

NVFP4 quantization is a technique that shrinks the huge number of values stored inside an AI model into a simpler, shorter form, reducing the model's overall size. The original model represents each number as precisely as a scale with very fine markings, but switching to a scale with coarser markings saves storage space and speeds up computation. This can introduce tiny errors, but if designed well, the size can be cut dramatically without any noticeable drop in performance.

By shrinking the model this way, something that once required several large servers can now run on a single computer, or even a small dedicated device sitting on a desk. This is also what makes it possible to bring AI directly into places like labs or field sites where large equipment isn't practical.

Companies like NVIDIA are increasingly releasing this compressed version alongside the original full-size version of their models, giving people who want a lighter way to use the same model an extra option.

How it shows up in the news

The article notes that "an NVFP4-quantized version is also provided, making it possible to deploy directly at lab sites using a single GPU or an NVIDIA DGX Spark setup." One easy point of confusion here: quantization doesn't turn a model into a completely different model. It only simplifies how the same model's numbers are represented, and while performance isn't exactly identical to the original, it's designed to stay close.

See also

Stories using this term

Browse every entry