Tag
Prism shipped a 27B that fits a laptop. We put it through a battery of agentic traps to see if it reasons like its size.
Read →A 27B model whose weights are all −1, 0 or +1 — what it takes to run, and how fast it goes.
Read →A from-scratch 5M-parameter classifier plateaued at 30% recall on real speech. A LoRA fine-tune of a 230M pretrained model, on the same data, didn't.
Read →Haku's code search reads each language's syntax tree, indexes real symbols, and embeds them on the Neural Engine — running inside the terminal, nothing leaving the Mac.
Read →Division of labour — a 14M model maps language to structure; ordinary code does the arithmetic
Read →One layer of the filter between raw accessibility data and the agent — element routing, on the Neural Engine.
Read →A 60M-parameter T5, distilled from a larger model, that cleans up raw speech on-device — split across the Neural Engine and the GPU.
Read →A small adapter turns a frozen left-to-right model into a one-pass, both-directions corrector.
Read →Making a text-to-speech model decode in one pass instead of sixteen, so it runs fast on-device. With audio you can play.
Read & listen →Decoding 16 residual codebooks in one pass instead of sixteen. On-device.
Read →How small can a model be and still learn a structured capability?
Read →Where ternary {−1, 0, +1} weights are the right tool for small, on-device models — and where they aren't.
Read →