Applied AI·Local and self-hosted inference
rounding the model's numbers to save space; round moderately and quality barely notices.
Quantisation
Storing weights at lower numeric precision, such as 4-bit integers instead of 16-bit floats, shrinking memory and speeding inference for a small quality cost. The reason a 27B model runs on your 16GB laptop at all.