Voynich Manuscript, Part 2: Inside the Language No One Can Crack

Close-up study of the unknown Voynich manuscript script with magnifying glass.

In Part 1 of this series, we traced the Voynich Manuscript from its rediscovery by Wilfrid Voynich in 1912 back through the hands of Holy Roman Emperor Rudolf II, and confirmed, via radiocarbon dating, that the book is genuinely a product of the early 15th century. What we didn't do is look closely at the one thing that makes this manuscript famous in the first place: the writing itself. Roughly 170,000 characters of it, spread across 240 pages, in an alphabet nobody has ever matched to anything else on Earth. This is the part of the story where linguists, wartime codebreakers, and now artificial intelligence have all taken their turn and all come away with essentially nothing.

What Does the Voynich Script Actually Look Like?

Researchers refer to the manuscript's writing system as "Voynichese." It's built from an alphabet of roughly 20 to 30 distinct glyphs, most of them fast, fluid, one or two-stroke shapes that flow across the page with none of the hesitation you'd expect from someone copying an unfamiliar script letter by letter. Certain characters show up almost exclusively at the beginning of words, others cluster in the middle, and others almost always end a word, a level of internal structure that looks nothing like random doodling.

A handful of words appear constantly throughout the text, strings transcribed by researchers as daiin, chol, chor, and qokeedy recur so often they function almost like scaffolding for the rest of the script. In the 1970s, researcher Prescott Currier went further and identified two distinct statistical "dialects" within the manuscript, now called Currier A and Currier B, different vocabularies and glyph frequencies that dominate different sections of the book. Language A shows up mostly in the herbal section; Language B takes over in the bathing and later portions. Paired with the handwriting analysis covered in Part 1, this suggests the manuscript wasn't just copied by multiple scribes, it may represent two related but distinguishable versions of whatever this script actually is.

Does Voynichese Actually Behave Like a Real Language?

This is the question that has kept the "it's just meaningless gibberish" theory alive for a century, and also the question that keeps knocking that theory back down. Statisticians have thrown some of the same mathematical tools used to analyze real languages at the Voynich text, and the results are strange in a very specific way: the manuscript behaves too much like a language to easily dismiss.

The clearest example is Zipf's Law, a rule that holds true across virtually every natural language on Earth, the most frequent word in a text tends to appear roughly twice as often as the second most frequent word, three times as often as the third, and so on down a predictable curve. Random gibberish typically doesn't follow this pattern. Voynichese does. Multiple independent studies have confirmed the manuscript's word frequencies fall into a Zipfian distribution consistent with genuine written language, which is a difficult thing to fake by accident.

At the same time, the text's entropy, a measure of how predictable each character is based on the ones before it, comes out unusually low compared to real languages, more in line with an outlier like Hawaiian, which has an especially small set of possible syllable shapes. That combination is exactly why the manuscript is so hard to categorize: it clears the bar for "structured enough to be a real language," while also being strange enough that no matching real language has ever been found.

Who Has Tried to Crack the Voynich Code?

The list of people who've attempted this is long, and it includes some of the best cryptographic minds of the 20th century.

The first major claimed breakthrough came in 1921, when University of Pennsylvania philosopher William Newbold announced he had translated the manuscript and identified its author as Roger Bacon. Newbold's method involved reading tiny secondary symbols he claimed were hidden within each letter's pen strokes, readable only under magnification, a theory that would have required Bacon to have invented the microscope roughly three centuries early. After Newbold's death, fellow scholar John Manly demonstrated that what Newbold had been "reading" were simply natural cracks and irregularities in the dried ink, and the theory collapsed entirely.

Two decades later, William and Elizebeth Friedman, among the most accomplished codebreakers of the 20th century, credited with cracking Japan's "Purple" cipher during World War II, spent years working the Voynich Manuscript in their spare time. They never solved it either, though their working notes and transcription systems are still referenced by researchers today.

More recent claims have followed a familiar cycle: a confident announcement, a wave of headlines, and a quiet unraveling shortly after. In 2014, applied linguist Stephen Bax announced a partial decipherment identifying ten words, including several plant names, by cross-referencing illustrations with medieval herbal manuscripts, a provisional claim covering a tiny fraction of the text that was never widely accepted. In 2019, University of Bristol researcher Gerard Cheshire made international headlines claiming the manuscript was written in an extinct "proto-Romance" language, describing its content as everything from herbal remedies to relationship advice. The University of Bristol itself pulled its press release within days after other linguists publicly dismantled the methodology, pointing out that Cheshire had assigned meanings to symbols with no consistent underlying rule, arbitrary enough that his critics called the results essentially uncorrectable.

Not every serious researcher believes there's a language to find at all. Psychologist Gordon Rugg has spent years arguing the manuscript could be an elaborate hoax, demonstrating that a simple table-and-grille method, mechanically combining syllables using a physical template, could have let a 15th-century scribe generate Voynichese-looking gibberish by hand, without needing to understand a word of it.

Can AI Finally Decode the Voynich Manuscript?

The manuscript has become something of a proving ground for computational linguists and, more recently, machine learning researchers, with results that are intriguing without being conclusive. In 2018, a University of Alberta team used algorithmic decipherment techniques to suggest the underlying language might be a heavily altered form of Hebrew, with vowels stripped and letters reordered, a theory that generated headlines but no verified translation. Later machine-learning approaches have applied topic-modeling techniques, normally used to sort documents by subject, directly to the manuscript's pages, successfully clustering sections by vocabulary in ways that roughly track the book's existing herbal, astronomical, and pharmaceutical divisions, useful for mapping the text's internal structure, but still no closer to reading it.

A genuinely new kind of evidence emerged in 2024, when researchers using multispectral imaging on the manuscript's opening page uncovered faded, previously invisible columns of writing, ordinary alphabetic letters alongside a column of Voynichese script, believed to be the work of a 17th-century owner attempting to build a personal decoding key. It's the first physical evidence that someone who actually handled the manuscript centuries ago was trying to crack it using the same side-by-side comparison method researchers still rely on today and it apparently didn't work for them either.

In 2025, a study published in the journal Cryptologia proposed something more modest and arguably more useful than a decipherment: a hand-executable substitution cipher, nicknamed "Naibbe," that could encrypt ordinary Latin or Italian text into ciphertext statistically similar to the real manuscript. It isn't a claim that the mystery is solved, the researchers themselves have described it as a benchmark, not an answer, but it demonstrates that a scribe working entirely by hand could plausibly have produced Voynichese-like text without any modern technology at all.

As for large language models trained specifically on the manuscript's text, several teams have quietly experimented with the approach, and at least one has publicly claimed a breakthrough, though as of this post, no such claim has been independently verified or published through peer review. Given the track record covered in this article, that's reason enough for healthy skepticism until real evidence surfaces.

So Is It a Language, a Cipher, or a Hoax?

That's the exact question this series turns to next. In Part 3, we'll lay out every major theory about what the Voynich Manuscript actually is, genuine lost language, elaborate cipher, medieval hoax, or something else entirely and weigh each one against the evidence covered so far.


Frequently Asked Questions

What alphabet is the Voynich Manuscript written in? The manuscript uses an unidentified script called "Voynichese," made up of roughly 20 to 30 distinct glyphs. It doesn't match any known historical alphabet, and no other document written in the same script has ever been found.

Has anyone successfully deciphered the Voynich Manuscript? No. Every claimed decipherment, including well-publicized ones from 2014 and 2019, has been challenged and rejected by other researchers after closer scrutiny. As of 2026, no translation of the manuscript is considered credible or accepted by scholars.

Is Voynichese a real language or a hoax? Statistical analysis shows the text follows Zipf's Law and has structural patterns consistent with real languages, arguing against a simple hoax. However, researchers have also shown a hand-executable hoax method could replicate many of these same patterns.

Can artificial intelligence decode the Voynich Manuscript? AI and computational linguistics have helped map the manuscript's internal structure and test cipher theories, but no AI-based decipherment has been independently verified or peer-reviewed as of 2026. Claims of AI breakthroughs should be treated cautiously until confirmed.

Who were the Friedmans, and why do they matter to the Voynich Manuscript? William and Elizebeth Friedman were leading 20th-century codebreakers, famous for cracking Japan's "Purple" cipher in World War II. They spent years attempting to decode the Voynich Manuscript and, despite their expertise, never succeeded.


Read the Full Voynich Manuscript Investigation Series

Comments