Tiny Tales
A language model small enough to run on your phone writes a children's story. It runs on this device: the model is a few MB, and no server or GPU is involved.
How this works
The model is a 4-layer transformer trained for about 10 minutes on a laptop on 40 MB of TinyStories, short stories written in simple English. It uses rotary positions, RMSNorm, QK-norm, ReLU² and the Muon optimizer. The weights are stored as 8-bit integers. Your browser runs it with about 150 lines of plain JavaScript: no WebGPU, no libraries.
It doesn't know facts and it loses the plot after a paragraph or so. It is fluent, though, and that's the surprising part at this size. The experiments behind it are in teg/tinylm-lab.