It turns out he was right
Decades later, a research team at the University of Vermont actually tried it.
They fed more than a thousand English-language novels through a computer, plotted each one’s
arc, and the shapes were there: six of them, three basic moves and their mirror images.
Vonnegut himself drew a seventh: Hamlet, whose line stays flat because nobody, not the prince
and not the audience, can tell the good news from the bad. He thought that was the most
truthful shape of all. And our own data demanded an eighth: stories that reverse so many times
the line becomes a braid. Below, each shape is drawn bold, with the stories from this page that follow
it traced faintly behind it.
What about the rest of the world?
Vonnegut drew Western stories, and the Vermont team tested Western novels. So we
asked the obvious next question: do the oldest stories from the rest of the world take the same
shapes? We drew twenty-two classics from eight traditions: Gilgamesh from Mesopotamia, the
Ramayana and the Mahabharata from India, the Persian romances, China’s four great novels, the
court tales of Japan, Korea’s Chunhyang, Vietnam’s Tale of Kieu. They do: the falls, the
recoveries, and the rises are all here, and most of these stories would sit comfortably among
the shapes above. The differences live in the endings. Several of these epics keep going past
the moment a Western telling would stop; the Ramayana reaches its triumphant coronation and
then continues into loss, which you can see in its chart below. And Gilgamesh, the oldest story
we have, ends where almost no Western story chooses to end: not in triumph, not in ruin, but in
acceptance.
How these lines were made
So how do you turn a book into a line? The University of Vermont researchers, the
team that first tested Vonnegut’s claim, did it with a trick both crude and clever. A computer
cannot follow a plot, but it can look up words: thousands of common English words have been
scored for how pleasant they feel, by panels of human readers. Slide a window through a book,
average the scores of the words inside it, and the book traces a curve out of its own
vocabulary. Watch the window slide:
That trick has a blind spot: it hears the mood of the
language, not the luck of the hero. So the lines on this page were made differently: from the
plot, not the prose. For each book we took the published plot synopsis (or the full text, where
it is freely available), broke the story into its events, and scored each event for how things
stand for the hero: safety, freedom, love, prospects. A beautifully written murder scores as a
murder.
The difference is not subtle, and you can see it. Below, each book carries two
lines: dashed for what the words sound like, solid for what actually happens. A Christmas
Carol is the clearest case. Its merriest vocabulary crowds the middle of the book, at
Fezziwig’s party and the Cratchits’ dinner, so the word line peaks in the middle and sags at
the end, while the story itself bottoms out at Scrooge’s own grave and ends at its absolute
peak. Emma’s word line sinks through the gossip and dread of her final chapters and misses the
weddings entirely. The Hound of the Baskervilles reads to a word-counter like a decline; the
events rise to the solution any reader remembers. The Secret Garden sounds like a tragedy and
blooms anyway.
what the words sound like
what actually happens (colored by shape)
And to check ourselves, we scored nine books twice, once
from the full text and once from the synopsis alone: seven of the nine traced nearly the same
line, the two sprawling epics being the exception. The full fine print follows.
How every line was made
Every fortune line in the atlas was scored by an AI agent judging the story’s events against
one shared rubric (1 = catastrophic, 5 = ordinary life, 9 = triumphant; prose mood never counts).
Lines are drawn as a gently smoothed envelope of those scores (beat scoring exaggerates each event
into a full swing), and every raw beat score is available on hover.
The difference between books is the evidence the agent worked from:
Full text (9 works): the agent read the entire book in equal slices (twenty for novels and
epics, sixteen for the two shorter works), scoring each slice as it went.
Published synopsis (105 works): the agent scored beat-by-beat from the book’s public plot
synopsis (linked on each book’s page), so every score is checkable against the same source. Event
positions along the book are approximate; synopses compress unevenly. Known wrinkles, kept visible:
one book (A Man Called Ove) was scored from its faithful film adaptation’s synopsis because the
novel’s page lacks one; Into Thin Air’s thin plot section was supplemented with the 1996 Everest
disaster article; The Hunger Games was re-scored after a first pass accidentally covered the whole
trilogy; the validation gates exist because these things happen.
Word-mood lines (the 2016 method: average happiness of a moving chunk of vocabulary,
Reagan et al., EPJ Data Science 2016, labMT lexicon) appear only as toggleable overlays on
37 pre-1923 books, for comparison: they are the book’s soundtrack, not its plot. A scale note:
event scores use absolute anchors (1 = catastrophic anywhere in literature, 9 = fairy-tale peak;
across all 1,771 scored beats only 2% are 9s), while the 2016 lines have no absolute units (each
book is normalized to its own range), so overlays are rescaled to the book’s event range and
compare shape only.
Validation of the synopsis method
Before trusting synopsis-scored lines, we tested them against ground truth: nine works that had
been scored from their full texts were independently re-scored from synopses alone.
Seven of nine reproduce the full-text shape (r = +0.61 to +0.87, median +0.68). The two
failures are the two epics, whose synopses summarize a different selection of episodes than the
condensed translations the full-text reader used, so epic-scale works are the method’s known
weak spot, and the two epics in this atlas display their full-text lines, not synopsis lines.
Separately, adversarial audits re-checked 199 beats across 12 random books against their cited
synopses: one fabricated event (corrected), no order errors beyond one, and a tail of
embellished details beyond the synopsis text (corrected). Treat the lines as careful readings
with known error bars, not ground truth.
What we verified ourselves
- Re-ran the 2016 decomposition on the corpus: three basic curves ± their mirrors explain ~84%
of all arc variance. The six shapes are real.
- Word-mood vs reader-scored fortune on 5 books: near-zero agreement on 4 (r = −0.14…+0.16);
Dorian Gray the exception (r = +0.90). Failures concentrate at endings.
- No new shapes in 9 reader-scored stories, Western or Indian.
- The “Ramayana” in one common anthology is Book One only; always check a text covers the
full arc before trusting its line.
Sources
- Kurt Vonnegut, “The Shapes of Stories” lecture (1985); rejected master’s thesis, U. Chicago
- Reagan et al. 2016 · EPJ Data Science 5:31 · github.com/andyreagan/core-stories
- Dodds et al. 2011 (labMT lexicon)
- Boyd, Blackburn & Pennebaker 2020 · Science Advances (narrative structure beyond sentiment)
- R. C. Dutt’s condensed Ramayana & Mahabharata (1898–99); Ryder’s Shakuntala; Arnold’s Nala and Damayanti
- Campbell, The Hero with a Thousand Faces: his monomyth, traced as fortune, is the valley shape