Tag
A from-scratch 5M-parameter classifier plateaued at 30% recall on real speech. A LoRA fine-tune of a 230M pretrained model, on the same data, didn't.
Read →Division of labour — a 14M model maps language to structure; ordinary code does the arithmetic
Read →