jargon

Applied AI·Local and self-hosted inference

you measured tokens per second on your own hardware at your own context length, and it was nothing like the published number.

Throughput benchmarking

Draft summary, pending review

Measuring your actual stack: tokens per second at your context lengths, TTFT, memory headroom, behaviour under concurrent load. Published numbers rarely match your hardware and quant; script-8 style measurement does. This is precisely portfolio project four.