Applied AI·Local and self-hosted inference
you compare two quantisations of the same model with a single number, and remember it is not a task-level eval.
Perplexity
Draft summary, pending review
A statistical measure of how well a model predicts text, lower being better. In local circles it is mainly used to compare quantisation levels of the same model; treat it as a rough quality proxy, not a task-level eval.