AI GlossaryITechnical words in the news
INT8 quantization
A compression technique that shrinks a model's internal numbers down to 8-bit integers, cutting both size and computation.
In plain words
INT8 quantization is a way to shrink an AI model's numbers into coarser units, reducing storage size and computation. It's similar to saving a photo as a compressed file instead of the raw original — the file size drops a lot, but the image looks almost the same to the eye. Models are usually made up of billions of numbers with fine decimal precision, and rounding those down to just 256 integer levels (8 bits) can shrink the model to as little as a quarter of its original size while also speeding up computation significantly.
A lighter model like this can run smoothly even on devices with limited storage and processing power, like smartphones or laptops. But because it's a rounding-off process, some precision is inevitably lost, so care must be taken in choosing what and how much to simplify in order to keep quality close to the original model.
Try it yourself
Try searching a model repository like Hugging Face for a version of the same model with 'int8' or '8bit' added to its name. It's often listed right alongside the original model, and you can see the file size drop to less than half.
See also
Stories using this term
- Mid-Size Model Tier Gets Crowded Again, with 70-80B as the Next BattlegroundAI · 2026.08.01
- Multiverse Computing shrinks a model to 4-bit and gets one smarter than the originalAI · 2026.08.25
- NVIDIA unveils 'Ising Calibration 1.5' VLM for automated quantum computer calibrationAI · 2026.08.09
- DeepSeek v4 Flash Gets GGUF Build for DwarfStar, Lowering the Bar for Local DeploymentAI · 2026.08.01
- Local LLM Community Flags a Gap in New 8B-12B ModelsAI · 2026.08.11
- AI agent memory needs different prescriptions by model size to boost performanceAI · 2026.08.19
