I used to think quantization was scary math stuff. It is not. Not really. The weights are numbers. Numbers can be stored in fewer bits. Fewer bits, less memory, smaller model. The catch is the rounding loses a little precision, and precision is, well, the model.

In practice the tradeoff is kinder than it sounds. At 4-bit, good models keep the vast majority of their quality while shrinking a lot. That is why 4-bit is the default advice for local runs. Huge memory savings, small quality cost, most people cannot tell the difference. Most local setups live at 4-bit and never look back. Mine does.

Go lower and the deal changes. 3-bit, 2-bit, the footprint keeps shrinking but the quality degradation stops being subtle. Casual chat? You might not care. Code or precise reasoning? You will. Match the quantization to the task, not to your ambition. Your ambition is not a benchmark.

One thing people miss: quantization is not the only memory consumer. The context window needs memory too, and it scales with length. A quantized model with a giant context can still blow past your VRAM if you feed it a novel. Headroom is cheap insurance against the weirdest errors you will ever debug.

Keeping it visible is the modern move. A private AI console SaaS showing running models, quantization levels, real memory footprint, one view. Plus a local LLM management app for iOS and Android and you can check the setup from the couch. Hobbyist install becomes infrastructure you trust.

Some people just want a quantized LLM VRAM requirements chart. Fine, ours maps footprint by parameter count and precision across 200 models. Yes, a chart exists, even though this article promised you would not need one. How much VRAM for a 70B model, a 13B, a 7B, it is all there. Rather not think about any of this? I do private AI setup for businesses at privateaiagent.fyi. Your data never leaves.