A small transformer trained from scratch on one novel, Cormac McCarthy's Blood Meridian. It writes one character at a time, and it runs here in your browser.
This is a GPT-style transformer trained on the text of Blood Meridian and nothing else. It starts from random weights, so everything it knows about English it learned from this one book. It reads and writes single characters: its whole vocabulary is the 94 distinct characters in the novel.
The architecture follows Karpathy's nanoGPT. The text was split 90/10 into training and validation sets (571,476 and 63,498 characters), and the model was trained on a single T4 GPU on Modal with a cosine learning-rate schedule for 5,000 steps. On the held-out tenth it scores 2.10 nats per character, against 0.17 on its own training text, so it has overfit the book. Its samples still rarely copy it: about 1% of eight-word runs in a generated passage appear in the novel.
This page loads the trained weights as an ONNX file and runs them with ONNX Runtime Web. Each new character is drawn from the model's 40 most likely choices, scaled by the temperature you set. Nothing is sent to a server.
Before training, the novel was measured so that generated text could be scored against it. Across 117,484 words:
The model gets the word level close. Its output is around four-fifths monosyllables and uses no quotation marks. It misses the sentence level: it leans on “and” too often and its sentences run long, without the book's mix of short and long. A 256-character window holds about 50 words, which is probably too little to learn where a sentence should end.
| v0 | v1 | |
|---|---|---|
| Status | Runs on this page | Next architecture, defined in the repo |
| Layers, width | 6 layers, 256 dims | 8 layers, 8 heads, 512 dims |
| Context | 256 characters | 768 characters |
| Position | Learned embeddings | ALiBi distance bias |
| Feed-forward | GELU | SwiGLU |
| Training | Cosine learning-rate decay | Curriculum over length 256, 512, 768; stochastic depth |
v1 is aimed at the sentence-level gap. Three times the context lets the model see where sentences end, and ALiBi biases attention by distance instead of learning a fixed position table.
The next piece is an agent that speaks as Judge Holden. A model that has read the book can recite his famous lines; the goal is a Judge who makes up parables of his own, setting an idea against whatever scene he is placed in. Because the novel has no quotation marks, even finding his speech takes work: it is recovered from tags like “said the judge”. The scenes holding his best-known lines are held out, so if one of those lines comes back he is reciting. Each answer is scored on whether it lands near the book in subject and feeling for that scene, and whether he stays in character.