LLMs GuruLLMsGuru
Find your AI
AI ModelsAI ToolsLearnGlossaryAI NewsFree ToolsFind your AISpeakRightAstrology
Explore All Tools
← The AI Dictionary
Plain-English definition

Quantization

In one breath

Shrinking a model by storing its numbers less precisely — like compressing a photo. It gets smaller and faster, with a small hit to quality.

In more detail

Quantization shrinks a model by storing its billions of parameters less precisely — rounding long, exact numbers down to shorter, coarser ones. It’s very much like compressing a photo: the file becomes dramatically smaller and faster to work with, at a small cost in fidelity.

It matters because it’s the magic that moves AI out of data centers. A model too large for anything but server hardware can, once quantized, run on a gaming laptop or even a phone — slightly blunter, but private, cheap and entirely yours. Much of on-device AI depends on it.

📌 See it in action

An enthusiast downloads an open-weights model that normally demands a rack of server hardware. Quantized, the very same model squeezes into an ordinary laptop’s memory. Side by side, its answers are barely distinguishable from the original’s — like a well-compressed photo: smaller, and still sharp where it counts.

Goes with
DistillationOn-Device AIParameters

Now put the vocabulary to work.

Meet the models →Which AI is mine?