Why “Quantization” is being discussed
The most recent Hacker News stories and comments contributing to this topic's mentions.
>As for why 4 bit is the limit, the ParetoQ paper has some interesting theories: https://arxiv.org/abs/2502.02631 Its not so clear to me what their real expl…
by cpldcpu · Oct 8, 2026
The capacity of your model scales with total number of bits in your weights. Shannon entropy. Yes, there are many way to distribute the error. And LLM are usually not trained to…
by cpldcpu · Oct 8, 2026
Which quantization are you using there?
by loehnsberg · Oct 8, 2026
It needs 3-4 Sparks to run well (at an acceptable quantization and sufficient KV cache): https://github.com/christopherowen/spark-ds41f
by jonsoft · Oct 8, 2026
It's dangerous to go alone. Take this: [1] I reimplemented most of the features of the Deepseek v4.1 flash paper (apart from quantization aware training which doesn't make sense…
by cookiengineer · Oct 8, 2026
Compression doesn't really work for model weights. Model quantization and model distillation are two techniques to reduce model size.
by tintor · Oct 8, 2026
Interest
Proportion of Hacker News items mentioning "quantization" over time.
Mentions
Total number of Hacker News items mentioning "quantization" over time.
Quantization is the process of constraining an input from a continuous or otherwise large set of values to a discrete set. The term quantization may refer to: Read more on Wikipedia