jargon

Applied AI·Local and self-hosted inference

rounding the model's numbers to save space; round moderately and quality barely notices.

Quantisation

Storing weights at lower numeric precision, such as 4-bit integers instead of 16-bit floats, shrinking memory and speeding inference for a small quality cost. The reason a 27B model runs on your 16GB laptop at all.

Commonly confused with