Tag
← all experiments
A from-scratch 5M-parameter classifier plateaued at 30% recall on real speech. A LoRA fine-tune of a 230M pretrained model, on the same data, didn't.
How small can a model be and still learn a structured capability?