Applied AI·Local and self-hosted inference
you halve the number of bits per weight and the memory you need roughly halves with it.
Precision (FP32, FP16, BF16, INT8, INT4)
Draft summary, pending review
The numeric formats weights and computation use, from 32-bit floats down to 4-bit integers. Training runs at FP16 or BF16; local inference commonly at 4 to 8 bit. Halving precision roughly halves memory.