Applied AI·Fine-tuning and customisation
you fine-tuned a mid-sized model on one consumer GPU, because the base underneath the adapter was quantised to four bits.
QLoRA
Draft summary, pending review
LoRA applied over a 4-bit quantised base model, shrinking memory enough to fine-tune mid-sized models on a single consumer GPU. The technique that democratised fine-tuning.