jargon

Applied AI·Local and self-hosted inference

you halve the number of bits per weight and the memory you need roughly halves with it.

Precision (FP32, FP16, BF16, INT8, INT4)

Draft summary, pending review

The numeric formats weights and computation use, from 32-bit floats down to 4-bit integers. Training runs at FP16 or BF16; local inference commonly at 4 to 8 bit. Halving precision roughly halves memory.