Applied AI·Local and self-hosted inference
you measured tokens per second on your own hardware at your own context length, and it was nothing like the published number.
Throughput benchmarking
Draft summary, pending review
Measuring your actual stack: tokens per second at your context lengths, TTFT, memory headroom, behaviour under concurrent load. Published numbers rarely match your hardware and quant; script-8 style measurement does. This is precisely portfolio project four.