The Shape of Our Words
Conversation Fractal and the Structure of Writing
It is odd to think that our words, which we consider to be ours uniquely, have a defined global structure that can be mapped as a predictable form.
However, research suggests that our words have a fractal-like internal structure that can be mapped. Even more interesting, this same structure has shown up across the languages and text types studied so far.
For decades, researchers studying the statistics of written language have found that text carries a kind of memory that can be described mathematically. One measure of it is the Hurst exponent. A value above 0.5 points to a long-range memory in the sequence, meaning what comes next is shaped in part by what came before rather than drawn at random. This tendency has turned up in many of the languages and text types studied so far, though its strength varies from writer to writer and genre to genre. The conversation fractal is one way to map that shape and compare it across texts.
How This Tool Works
This conversation fractal tool maps text to analyze its internal structure shape. Each sentence is broken down into twelve data points. These include word count, average word length, lexical diversity, punctuation density, question density, exclamation density, first person density, second person density, negation density, number density, mid-sentence capitalization, and syllabic complexity. Those twelve points are assembled as an aggregate and analyzed using Principal Component Analysis to establish three axes where the largest number of variations occur. The sentences are plotted on a 3D axis and connect as they occur in the text.
To the side, you can see two calculations. The Zipf slope describes how word frequency falls off as words get rarer: a handful of words are very common and a long tail are rare, and that fall-off follows a power law rather than a straight line. The Hurst exponent describes the sentence sequence. Above 0.5 it leans persistent, meaning sentences tend to carry the pattern of the ones before them; near 0.5 it reads as independent, with no memory; below 0.5 it leans alternating, tending to reverse rather than continue.
What is displayed is a map of your entire text. Some writing traces a path that folds back on itself and holds a recognizable shape; other writing reads much like its own shuffled version, with no memory the test can tell apart from chance. The tool reports which case yours falls into rather than assuming every text carries structure. Longer passages give the more reliable reading, since the memory test needs many sentences to separate signal from noise.
This tool is designed to show the beauty of the words we write.
Directions for Use
To use the Conversation Fractal tool, enter your text in the box or upload a file, and then click analyze.
The tool will return a 3D image, with Zipf and Hurst value to the side, as well as an interpretation of the results.
Frequently Asked Questions
What is the Zipf Law applied to language?
Zipf’s Law applied to language says that when words are ranked from most to least frequent, a word’s frequency falls off roughly in proportion to one over its rank: the most common word appears about twice as often as the second, three times as often as the third, and on down. Mandelbrot’s refinement, the Zipf-Mandelbrot Law, adds two parameters so the curve can bend to fit the very top and the long tail, where plain Zipf matches real text poorly. Both describe a power-law relationship between rank and frequency.
What is the Hurst variable in language analysis?
The Hurst exponent is a value related to the fractal dimension of data. Values between .5 and 1 indicate that the data follows a memory of the past (like a fractal holding to its pattern). A value of .5 is the random case, with no memory. Values below .5 mean the data tends to reverse direction rather than carry its pattern forward.
What are twelve text points that the tool analyzes to graph variance?
Word count — the length of the sentence
Average word length — the mean number of characters per word
Lexical diversity — the ratio of unique words to total words
Punctuation density — frequency of commas, semicolons, and colons
Question density — how often question marks appear
Exclamation density — how often exclamation marks appear
First-person density — frequency of I, me, my, we, our, and related words
Second-person density — frequency of you, your, yourself, and related words
Negation density — frequency of not, never, no, can’t, won’t, and related words
Number density — how often numerical figures appear
Mid-sentence capitalisation — unexpected capital letters indicating proper nouns or emphasis
Syllabic complexity — the average number of syllables per word
What do the three color modes show?
Sequence mode colors each sentence point by the order it appears in the text; early sentences are blue, moving through to gold as the text progresses. This shows you the path of the text over time.
Energy mode colors each point by lexical complexity; sentences with longer, more unusual, and more varied words glow brighter.
Personal mode colors each point by first and second person density, how much the words I, me, my, we, and you appear in each sentence.