Blood Meridian GPT

A small transformer trained from scratch on one novel, Cormac McCarthy's Blood Meridian. It writes one character at a time, and it runs here in your browser.

Loading model (21 MB)...
Try
Lower is steadier, higher is stranger.
Characters to generate.
Output
Generated text will appear here once the model has loaded.

The model

This is a GPT-style transformer trained on the text of Blood Meridian and nothing else. It starts from random weights, so everything it knows about English it learned from this one book. It reads and writes single characters: its whole vocabulary is the 94 distinct characters in the novel.

The architecture follows Karpathy's nanoGPT. The text was split 90/10 into training and validation sets (571,476 and 63,498 characters), and the model was trained on a single T4 GPU on Modal with a cosine learning-rate schedule for 5,000 steps. On the held-out tenth it scores 2.10 nats per character, against 0.17 on its own training text, so it has overfit the book. Its samples still rarely copy it: about 1% of eight-word runs in a generated passage appear in the novel.

This page loads the trained weights as an ONNX file and runs them with ONNX Runtime Web. Each new character is drawn from the model's 40 most likely choices, scaled by the temperature you set. Nothing is sent to a server.

4.81M
parameters
6
transformer layers
256
characters of context
94
characters of vocabulary

McCarthy's fingerprint

Before training, the novel was measured so that generated text could be scored against it. Across 117,484 words:

  1. 81.7%of words are one syllable. The average is 1.22 syllables per word.
  2. 5.8%of all words are “and”. Clauses are chained with it again and again.
  3. 2double quotation marks in the whole book. Speech is marked by “said”, never by quotes.
  4. 48%of sentences are ten words or fewer. The mean is 15.7 words, and the longest runs to 245.

The model gets the word level close. Its output is around four-fifths monosyllables and uses no quotation marks. It misses the sentence level: it leans on “and” too often and its sentences run long, without the book's mix of short and long. A 256-character window holds about 50 words, which is probably too little to learn where a sentence should end.

Two models

v0v1
StatusRuns on this pageNext architecture, defined in the repo
Layers, width6 layers, 256 dims8 layers, 8 heads, 512 dims
Context256 characters768 characters
PositionLearned embeddingsALiBi distance bias
Feed-forwardGELUSwiGLU
TrainingCosine learning-rate decayCurriculum over length 256, 512, 768; stochastic depth

v1 is aimed at the sentence-level gap. Three times the context lets the model see where sentences end, and ALiBi biases attention by distance instead of learning a fixed position table.

Next: the Judge

The next piece is an agent that speaks as Judge Holden. A model that has read the book can recite his famous lines; the goal is a Judge who makes up parables of his own, setting an idea against whatever scene he is placed in. Because the novel has no quotation marks, even finding his speech takes work: it is recovered from tags like “said the judge”. The scenes holding his best-known lines are held out, so if one of those lines comes back he is reciting. Each answer is scored on whether it lands near the book in subject and feeling for that scene, and whether he stays in character.