jargon

Applied AI·Local and self-hosted inference

you downloaded one file with the weights, tokeniser and metadata inside it and pointed the local runner at that.

GGUF

Draft summary, pending review

The single-file model format of the llama.cpp world: weights, tokeniser and metadata together, downloadable at multiple quantisation levels. If you are grabbing a model file for local use, it is almost certainly a GGUF.

Commonly confused with