The Voynich Manuscript (cryptography)

The content on this page was written by AI under human supervision.

The Voynich manuscript is an illustrated fifteenth-century book in a script no one has been able to read. In the paper summarized here, every proposed reading that can be made mechanical, and every proposed way of producing such text without meaning, is expressed as a probability model. Each model is fitted on half of a scribe's pages and judged on the other half, which is withheld from the fitting. The readings all fail on the withheld pages or pass unselectively; such a failure excludes a reading only where the test is shown to detect disguised real text of the same kind, so far only scripture-like prose. A keyless writing routine, reusing words, now and then copying a recent one, and coining others by letter habit, reproduces nearly all the summary statistics of the two principal scribal hands. Neither it nor any published generator reproduces one feature of the text, a statistical coupling between the last glyph of each word and the first glyph of the next. This routine is the paper's account of how the text was made, and it sharpens a picture that has been emerging since the work of Gordon Rugg and of Torsten Timm and Andreas Schinner.

A photograph of folio 26v of the Voynich manuscript: nine lines of looping brown script above a painted plant with green fern-like leaves and blue-and-white flower heads. Three boxed callouts on the right magnify details: the word transliterated qokedy, written twice in one line and again two lines below; a run of four near-identical words transliterated qoteedy, qokody, qotedy, qotedy, the last two identical; and one flower head, captioned as a plant matching no known species. A footer notes that transliterations follow the EVA alphabet and that the photograph is from Yale's Beinecke Library, public domain.

Folio 26v of the manuscript, a page from the herbal section, with three details magnified: a word written twice in one line and again two lines below, a run of near-identical words differing by a stroke or two, and one of the book’s plants, most of which match no known species. The same page, in transliteration, appears in the last figure below beside a page written by the procedure the paper recovers. (Site illustration; image Yale University, Beinecke Library, public domain.)

A book no one can read

The manuscript, MS 408 of Yale's Beinecke Library, is a small book of about 240 pages of vellumParchment prepared from animal skin, the usual writing surface of a fine medieval book., radiocarbon-dated to between 1404 and 1438, with script and drawings that flow around each other, evidently made as one project. Drawings of plants, most matching no known species, fill 129 of the 227 text pages. The rest show astronomical and zodiac diagrams, bathing figures in green pools, labeled jars and plant parts, star-marked paragraphs usually read as recipes, and a few pages of text alone. The script occurs in no other surviving document. Its basic alphabet has roughly two dozen glyph shapes, though inventories of over forty have been defended, and the paper takes its counts under two different divisions into characters. The writing is fluent, with very few corrections, so the scribes were doing something familiar.

The book's documented history begins only in the early 1600s at the Prague court of Rudolf II. A secondhand report, which the paper treats as unverified, has Rudolf paying 600 ducats for it as a work of Roger Bacon. The book reached the Jesuit scholar Athanasius Kircher in Rome in the 1660s, stayed with the Jesuits until Wilfrid Voynich bought it in 1912, and has been at Yale since 1969. The ink and pigments indicate that the writing is as old as the vellum, which excludes Bacon, who lived in the thirteenth century, and a forgery by John Dee or Edward Kelley in the sixteenth.

Four centuries of readers

In 1921 William Romaine Newbold announced that the manuscript was Bacon's notebook, written in a microscopic shorthand hidden in the ink strokes. John Manly dismantled the reading in 1931: the shorthand was cracks in the drying ink, and the readings came from anagramming so free that Manly produced counter-readings of his own. William Friedman and John Tiltman, among the best cryptanalysts of the twentieth century, both worked on it. Friedman concluded, in a 1959 footnote written as an anagram, that it was an early constructed language and not a cipher, and Tiltman endorsed no solution at all.

While attempts at decipherment failed, statistical facts about the text accumulated: the characters are more predictable than in any European language tried, and words have a rigid structure of prefix, core and suffix. Prescott Currier showed in the 1970s that the pages divide into two statistical "dialects," A and B, which change with the handwriting. Lisa Fagin Davis has since distinguished five scribal hands, an attribution the paper adopts without testing it and which Torsten Timm disputes in favor of a single scribe whose hand drifted. The unusual statistics notwithstanding, a 2021 review by Claire Bowern and Luke Lindemann judged the text most likely a natural language in some non-standard encoding.

A second camp holds that there is nothing to read. In 2004 Gordon Rugg showed that a Cardan grille, a card with cut windows slid over a table of word fragments, mass-produces Voynich-like text, though the grille is a sixteenth-century device. Timm's 2014 study How the Voynich Manuscript was created, whose title the present paper echoes, proposed a copy-and-modify process. He and Andreas Schinner formalized it in 2020 as a "self-citation" generator, which reproduces several of the text's statistics. Cipher proposals kept arriving regardless, most recently the Naibbe cipher of 2025, a scheme executable by hand that turns Latin into Voynich-like text. A cipher still has defenders and Bowern and Lindemann favor a language, but much of the recent quantitative work has come to favor a writing procedure carrying no message, and the paper sets out to make that picture precise.

Weighing one explanation against another

Announced readings have generally come from searching many keys, languages and rearrangements for language-like output, which Manly showed can be found in almost any text given enough freedom. The self-citation generator, in turn, had not, so far as the paper's author knows, been tested on text withheld from its fitting. The paper puts both to the same test: each reading and each generator becomes a probability model for the glyphs on a page, fitted on half of a scribe's pages and judged on the half withheld from the fit. Most tests run on the pages of the two principal hands, Scribe 1, who writes dialect A, and Scribe 2, who writes dialect B. Models are compared by their evidence, the probability each assigns to the text once its adjustable parameters are averaged out:

$$Z_M \;=\; P(\text{text} \mid M) \;=\; \int P(\text{text}\mid\theta, M)\, P(\theta \mid M)\, d\theta .$$

Here $M$ is a model, say "Latin under a substitution cipher with an unknown key," $\theta$ collects its unknown parameters, $P(\text{text}\mid\theta, M)$ is the probability of the text at one setting of them, and $P(\theta \mid M)$ is their prior spread. The average runs over every setting instead of the best one. A model flexible enough to produce almost any text therefore earns little credit for fitting this one; a free search pays no such penalty. For the paper's count-based models the integral is exact. The paper keeps three kinds of number apart. The evidence itself is used to classify the dialect of each page and to make the whole-corpus comparisons. The reading tests use the predictive score of a decode of the withheld pages under a language's word and letter statistics. The writing procedures are compared by the bits per word a fitted generator needs to predict the withheld pages. Evidence ratios are quoted in bits per word, the base-2 logarithm of the ratio divided by the number of words, and one bit per word doubles the odds with every word read.

Every solver, the search program that fits a reading's key or table, first had to recover a known answer from real Latin, Hebrew or Arabic disguised by the mechanism under test. Three such checks fell short, and the paper marks the results they affect. The reading tests are calibrated on scripture in six languages, Latin, Hebrew, Arabic, Italian, German and English, cut to the manuscript's page layout. The production comparisons use a Latin herbal, biblical Hebrew, medical Arabic and Dante, with 23 further texts, recipes, glossaries and a botanical catalog among them, added for word order. The manuscript itself enters the analysis as a standard transliterationThe Zandbergen–Landini transcription of the script into the Extensible Voynich Alphabet (EVA), whose Latin letters name glyph shapes without any claim about their sound; maintained by volunteer scholars. The paper counts everything under two different divisions of the script into characters. of its 227 text pages, about 38,000 words.

The readings fail, and what that excludes

A century of proposed readings falls into five families: substitution, one glyph per letter; verbose substitution, a group of glyphs per letter; consonantal writing with the vowels left out; Latin abbreviation; and codebooks, each word a token in a private table. The paper adds composites and the Naibbe cipher built from its published tables. Counted across languages, hands and character divisions, that makes 130 hypotheses, each judged by whether its decode of the withheld pages has the letter and word statistics of the language claimed. Ten alternative reading orders, and character divisions chosen freely by the solver, are searched together with the mapping in further runs judged against control texts. Free anagramming, on which Newbold's reading relied, is left out on purpose: at this text length it admits almost any answer, so no test could contradict it. Anagramming by a fixed rule within each word, the last stage of a 2016 Hebrew proposal by Bradley Hauer and Grzegorz Kondrak, is testable, and the paper tests it.

None of the 130 hypotheses passes on the withheld pages. The nearest simple substitution, into Hebrew, falls 2.8 bits per word short of the passing mark, a word-level score calibrated on scripture in each language. The nearest verbose reading falls 1.1 short, the nearest consonantal one 2.1, the Naibbe mapping 1.5, Latin abbreviation 5.7. The ordinary reading order fits better than any of the ten alternatives, and re-dividing the script into characters brings the text no closer to any language.

What a failure proves depends on whether the test would pass a true reading of the same kind. The paper measures this on 22 real texts in the six languages (scripture, herbal, medical, technical and narrative prose, recipes), disguised by up to sixteen encodings at four levels of transcription noise, with five trials of each combination. The test reliably accepts a correct reading only of scripture-like prose at low noise, and there only for substitution, fixed-rule anagramming, abbreviation and narrow homophonic ciphers (each letter written by any of a few alternative signs), in some languages. Because its passing mark was set on scripture, it accepts a correct reading of the other genres in about one trial in a hundred with no transcription noise and in none at 2 percent noise or more. A matched control text containing no reading passed in none of 6,520 trials. The exclusions therefore hold for scripture-like text and carry no weight for herbal, medical, recipe or narrative text in any language. Two further limits apply: full-width verbose tables were searched for Latin alone, and on the Naibbe class of cipher the solver showed no power at all. The solver recovered nothing in 1,370 trials on texts enciphered with freshly randomized tables of the same design, so nothing is claimed about that class either way. Fixed-rule anagramming into Latin, Italian, German or English is rejected on both hands by 2.2 to 5.0 bits per word. Hebrew is not rejected and Arabic is accepted, but the test also accepts an Arabic reading of disguised German and English texts, so for both Semitic languages it is uninformative.

A further experiment in the paper suggests why readings keep being announced. Modern decipherment attempts search many mappings for the one whose output looks most like language, and such optimizers, run unconstrained on the manuscript, "succeed" every time. But the decoded "Hebrew" and "Arabic" use eight to ten of some twenty letters, and the same pages decode equally well into both languages, which no genuine reading permits. The paper then requires the output to be language-shaped, with a real language's letter and vocabulary statistics, and compares the manuscript with a meaningless self-copying text searched with the same freedom. The manuscript's advantage over that control collapses, from 4.1 to 0.26 bits per word in the Hebrew test and from 6.3 to 2.6 in the Arabic. The Arabic residue rests on a decode that fails the shape requirement on the withheld half. Three to four bits per word of apparent signal appear to come from the search's reward for language-like output alone, enough, the paper suggests, to account for much of the history of announced readings.

Two bar-chart panels titled scribe A core (left) and scribe B core (right). Vertical axis: advantage of the solver's best decode on withheld text over shuffled text, in bits per word, from 0 to 8. Blue bars are the manuscript text, gray bars a structureless control (a matched self-copying stream). In each panel the left pair of bars is the unconstrained solver and the right pair the decode forced to language shape. Left panel: unconstrained, manuscript 4.99 and control 0.92, bracket labeled gap 4.06; constrained, manuscript 3.28 and control 3.02, gap 0.26; an annotation notes that forcing language shape lifts even the control from 0.92 to 3.02. Right panel: unconstrained, manuscript 6.29 and control minus 0.01, gap 6.31; constrained, manuscript 2.69 and control 0.12, gap 2.57.

The decipherment illusion. Each pair of bars compares an optimizing solver’s best “reading” of the manuscript (blue) with its best reading of a meaningless self-copying control text searched with the same freedom (gray), in bits per word of advantage over shuffled text; left, the Hebrew test on Scribe 1’s dialect-A pages, right, the Arabic test on Scribe 2’s dialect-B pages. Unconstrained, the solver opens a gap of 4.1 and 6.3 bits per word over the control; once its output must have the letter and vocabulary statistics of a real language the gap falls to 0.26 and 2.6, and the Arabic residue rests on a decode that fails those requirements on the test half. Between three and four bits per word of the apparent signal appear to come from the freedom of the search alone. (Figure 4 of the paper.)

How the pages were written

With no reading passing its test, the paper asks how the text was made, working on those pages of the two principal hands that are unmixed in dialect. The candidate processes, among them a habit chain in which each word depends on the one before, Rugg's grille, and Timm and Schinner's copying from nearby lines, each became a working generator. They are compared against a reference with no procedure at all, a memorized vocabulary drawn on regardless of order (a "bag of words"). Two statistics, both built from cross-entropies (the average bits a fitted model needs per word of the withheld half), then separate the manuscript from each of the four comparison languages on which these models were run:

$$G_{\mathrm{order}} \;=\; H_{\mathrm{vocab}} - H_{\mathrm{chain}}, \qquad H_M \;=\; -\frac{1}{N}\,\log_2 P_M(w_1 w_2 \cdots w_N \mid \text{layout}) .$$

Here $w_1 \cdots w_N$ are the $N$ withheld words with their line layout given, $P_M$ is the probability that model $M$ assigns them, and "vocab" and "chain" are the bag of words and the habit chain. So $G_{\mathrm{order}}$ is the predictive gain from knowing the previous word. For Latin, Hebrew, Arabic and Italian laid into the manuscript's page templates it is 0.15 to 1.06 bits per word; on the two hands it is 0.010 to 0.022. The companion gain of a memorized word list over spelling habits is 1.3 to 3.0 bits in the four languages and at most 0.18 here. To cover lists as well as prose, the paper laid 23 further texts into the same templates, among them technical, herbal and medical Latin, recipe collections, glossaries, an index and a botanical catalog. No model class it fitted gives either hand more than 0.022 bits, below every one of these texts, and the botanical catalog shows the most word order of all. The nearest text, a list-like stretch of Pliny on trees, reaches 0.039. Adjacent words do share something at their edges: the last glyph unit of one word predicts the first unit of the next by about 0.12 and 0.21 bits on the two hands, against at most 0.11 in the four comparison languages. Only some recipe collections and the catalog show as much of this edge coupling or more.

Two dot-plot panels sharing a vertical list of corpora: Latin (A pages), Arabic (B pages), Arabic (A pages), Hebrew (A pages), Hebrew (B pages) as orange squares above a dashed divider, then Voynich A, Voynich A (atomic parsing), Voynich B, Voynich B (atomic parsing), Voynich A+B as blue circles. A title reads: language leaves signatures the manuscript does not have (held-out cross-entropy differences, same models, same layouts). Panel (a), does word order carry information: horizontal axis the word-order gain, the vocabulary model’s cross-entropy minus the chain model’s on withheld pages, in bits per token, 0 to 3; languages at +0.15, +0.27, +0.59, +0.93, +1.06 with the note real languages in the same page templates +0.15 to +1.06; manuscript points at +0.01, +0.01, +0.02, +0.02, +0.04 with the note manuscript +0.01 to +0.04. Panel (b), is the vocabulary more than its letter habits: horizontal axis the vocabulary surplus, the letter model’s cross-entropy minus the vocabulary model’s on withheld pages, in bits per token; languages at +2.29, +2.11, +2.95, +1.71, +1.78 (real languages +1.7 to +3.0); manuscript at +0.18, +0.02, +0.15, +0.02, +0.10 (parity, 0.00 to +0.18).

Two signatures of language that the manuscript lacks, measured as differences in cross-entropy (the bits a model needs to predict each word) on withheld pages with the same models on the same page layouts; these are not evidences. (a) The word-order gain of the displayed equation, in bits per word: orange squares are Latin, Arabic and Hebrew cut to the manuscript’s page templates (0.15 to 1.06), blue circles the two principal hands under both divisions of the script into characters, and both hands pooled (0.01 to 0.04). (b) The gain of a memorized word list over a letter-habit model: 1.7 to 3.0 bits per word for the three languages plotted (1.3 with Italian, which is not shown), 0.00 to 0.18 for the manuscript. In the manuscript, knowing the previous word tells you almost nothing about the next, and the vocabulary is little more than its spelling habits. (Figure 5 of the paper.)

By exact evidence, whether the bag of words or the habit chain ranks first depends on the prior placed on the chain, and where the chain wins its margin is a few hundredths of a bit per word. Independent of the prior, the grille, as parametrized, fits worse than the bag on both hands, and copying from nearby lines is a real component, decisively better than its shuffled control. For genuine output of Timm and Schinner's generator, scored the same way, copying is preferred to the bag and the copy rate is above either hand's, so self-citation in its published, local form does not appear to be the manuscript's process. Copying from much further back, though, is statistically the same as reusing a growing personal stock, so Timm's broader account, one writer citing his own drifting text, is not excluded by any statistic of this kind.

The process that fits the text is a loop of two or three small decisions per word. Most words are reused, drawn by frequency from the scribe's own stock: 72.9 percent of Scribe 1's withheld text and 76.7 percent of Scribe 2's. A few are copied from one of the last twenty or so words, about one in fifty (2.0 percent) for Scribe 1 and one in twenty (4.9 percent) for Scribe 2, whose copies are nearly all verbatim. Scribe 1's fitted copying cannot be separated from what page-to-page variation in word frequencies alone would produce, whereas Scribe 2's can. The rest, a quarter of Scribe 1's words and under a fifth of Scribe 2's, are coined afresh in the scribe's letter habits, so closely that a letter model predicts them as well as the vocabulary does.

A flow diagram titled the recovered procedure, one self-contained loop per word, with fitted settings and withheld-page branch shares shown per scribe. Boxes from top to bottom joined by arrows: next word; check position on the line (line-start boosts p, f/s, t, A times 4.2 and B times 6.7 for p; line-start words run short, ramp A 0.59 to 1.35, B 0.17 to 2.36); lean toward the previous word's length (length momentum, lag-1 correlation 0.16 for A and 0.15 for B, worth +0.008 and +0.017 bits per token); choose how to say it, which branches three ways: REUSE of an existing type (72.9 percent A, 76.7 percent B of withheld tokens, a frequency-weighted draw from the scribe's own vocabulary), LOCAL COPY within about 20 words (2.0 percent A, 4.9 percent B of withheld tokens, a copy of one of the preceding 20 or so words; fitted copy weight 2.4 percent A, 5.8 percent B), and COINAGE of a new type (25.1 percent A, 18.4 percent B of withheld tokens; same letter habits, same prefix-core-suffix template). The three branches rejoin at a final box: write it, then repeat (word order carries almost nothing more). A footer states that the branch shares are posterior probabilities of each branch averaged over each hand's withheld tokens, with A denoting Scribe 1 and B Scribe 2.

The recovered writing procedure, drawn with the fitted settings of the two principal hands (A is Scribe 1, B is Scribe 2). Each word is one pass through the loop: a check of the position on the line, a lean toward the previous word’s length, then one of three branches, reuse of a word from the scribe’s own vocabulary (about 73 and 77 percent of withheld words), a copy of a recent word (2.0 and 4.9 percent), or a fresh coinage in the same letter habits (25 and 18 percent); the smaller numbers in the boxes are secondary fitted settings. No step involves a key or a source text, so text written this way leaves no message to decipher. (Figure 6 of the paper.)

Two smaller habits complete the procedure: a preference for particular words at the start and end of a line, and a tendency for each word's length to follow the previous word's. With these added, the complete model beats reuse, copying and coinage alone by 0.14 to 0.19 bits per word on fresh page splits. Pages it writes fall outside the manuscript's own page-to-page range on two and three of 22 summary statistics for the two hands, counts that pages simulated from the fitted model itself reach with probability 0.16 and 0.06. Simple controls miss six to eleven of the 22, Timm and Schinner's output seven and five, and Naibbe ciphertext nine and thirteen. The model still misses a few statistics outside those 22, above all the edge coupling between adjacent words, which no generator yet proposed reproduces either.

Three blocks of monospaced transliterated text, four lines each. Top block, headed REAL, manuscript page f26v (withheld from the fit): pchedar qodary daiiin pcheety s air shedy ypchedy ypchdy qopy shdy / saraiir chekedy qokedy otedy sar y etedy qokedy or ai'he alys chedy / pchdar opar dar cheeol ofchdy otedy odar chedy ytedy okchdy g / yckheody qokedy deey saldy okedor or cheos oraiin okeo chekaiin. Middle block, headed GENERATED, the fitted procedure on the same page layout: pdshey chedal qeedy r dy qokain sheotaiin r kal y dal / pchedy lokar shes sheckhy alol qokain dychedy olsheedy qokykal qokeey qokaiin sain / sain chedy dainy oiir okaiin aiin y ar qokar qokam qedaiokaiin / olkchdy qokain odaiin aral dy shedy otedy cheykain kor sheol. Bottom block, headed CONTROL, same procedure with the word table replaced by a random draw, tagged visibly wrong: syqy shay drshch qol aehee qiat qtd shera oesch ttdr qiat / yis eetoor tiins qokor dcheek shay sodey qotar el rr iinqkh cheeas / sceor qda shedy al othy shay oqlsheor kddar oldokr shtey ol / othy pq okhy qoty al qopol cheetr ychytyl chchapqee pt. In every block words beginning qo- are set in blue and words beginning ch- or sh- in green.

Real lines, generated lines and a broken control. Top: the first four lines of folio 26v in transliteration, a page in Scribe 2’s hand that the fitting never used. Middle: the first four lines of a page written in the same layout by the complete fitted procedure. Bottom: four lines from the same procedure with its vocabulary table replaced by a random draw. Blue marks words opening qo- and green ch-/sh-, Scribe 2’s two commonest word beginnings. The real page’s mix of word skeletons and its runs of near-identical words recur in the generated lines and vanish in the control; the paper presents the comparison as an illustration, the quantitative test being the count of missed statistics. (Figure 10 of the paper, Appendix C.)

The two hands' vocabularies, of 3,155 and 2,389 distinct words, are built alike. The paper splits every word by one fixed rule into prefix, core and suffix: the first two glyph units, whatever lies between, and the last two. Each hand's prefixes and suffixes form small inventories, and the two hands' inventories largely coincide (an overlap index of 0.65 for prefixes and 0.48 for suffixes) while their cores barely do (0.23). But two halves of either scribe's own pages split the same way (cores 0.23 to 0.28), as do two of the minor hands against the principal ones and every comparison language cut in two. Any productive spelling habit therefore yields this pattern, and the pattern does not show that tables passed between writers or that anything was written down. The cores look like coinages of each scribe's letter habit. Families of one-glyph variants such as kcheo, kheo, pcheo, scheo, ssheo, lsheo, theo make up a little under 60 percent of each core vocabulary (a part of the model checked only after fitting). An ordering statistic shows no sign that the cores were copied out from a written grid. A family's members, a slight same-page excess for Scribe 1 aside, first appear no closer together in the text than unrelated cores, so they do not seem to have been coined in one run.

Two faint traces of the pictures

Meaning could still lie in the labels or in word choice, and the paper tests for it with scribe and page order controlled; two weak traces survive. The 864 labels beside the drawings form a register of their own. On the pharmacy page f88r the six top-band labels read otorchety, osal, orald, oldar, otoky, otaly, all opening with o-. Across the book three labels in five are o-initial, against one word in three in matched running text. Names, though, would recur with their objects, and the labels repeat across pages less often than chance: 75 types on more than one page against 85 to 111 expected. Within every well-populated class of pictured object, stars included, reuse is at or below a matched chance level. A naming system sparser than about a fifth of repeat mentions could have escaped the test, and a name used once for one star cannot be tested at all.

The second trace is a faint tint of vocabulary by subject. The paper examined seven channels: which words a page uses, how long they are, and which glyph units fill each of five positions in the word. Only the first, the mix of word types on a page, tracks the illustrated section beyond scribe and page order, and it survives fake contiguous sections and fake sections aligned to the quiresThe folded gatherings of sheets from which a codex is sewn together. In this book the sections, the hands and the quires are physically entangled, so a signal that follows the binding can masquerade as one that follows the subject.. The effect is about 0.2 to 0.3 bits per page, marginal under the harshest controls, and it comes from the astronomical and bathing sections; the herbal section, the largest, shows none. One survivor out of seven channels tried, the paper adds, is weaker support than an independent confirmation. Beyond these two traces the tests find no dependence between words past the coupling of adjacent word edges, no stable names, and no mapping to any language or mechanism examined, each within the sensitivity of its test.

What kind of book, and what would change the answer

The reconstruction is offered as a proposal, the process most consistent with the measurements, fitted on two hands and extended to the others largely by inference. The vocabulary statistics cannot distinguish several writers trained in one spelling habit from one writer whose habit drifted. The routine needs practice, as the fluent penmanship suggests, but no literacy in any particular language, no knowledge of ciphers, and no apparatus such as Rugg's table and grille. Why the book was made can only be inferred, and the paper marks that part as less secure. Assuming ordinary scribal rates, the text is 40 to 80 hours of writing, little beside preparing a hundred leaves of vellum and the drawings. The low cost of the writing and the secret-knowledge genres its sections imitate suggest, in the paper's phrase, a book made to be looked at and believed, as a practitioner's credential or for sale to a collector of marvels.

On what the text says, the results are negative within a stated space. Where the test is shown to pass true readings of the same kind, the scripture-like cases described above, the failures are the manuscript's own. Elsewhere the test rejects even correct readings or, as for wide homophonic tables and the Naibbe class, has no power. Such readings are constrained only by the requirement that any encoding yield characters as predictable as the manuscript's, and by the production evidence. Free anagramming, untestable at this length, lies outside the space altogether.

On how the text was produced, the measurements sharpen, in the paper's words, a picture that has been emerging since the work of Rugg and of Timm and Schinner. Rugg's grille as parametrized fits worse than reuse alone; Timm and Schinner's local copying is a minor component, while their broader drifting self-citation is not excluded; and the edge coupling is reproduced by no generator yet proposed, the paper's included. The two traces of content are read as signs that the writers tracked the pictures; neither shows that the text describes them. Meaning carried by anything but word identity and order would be invisible to every test used. A confirmed reading of any passage would overturn the reconstruction, if it passed where past claims failed: on withheld text, against matched controls, specific to one language. So would a passage that predicted the imagery in detail, or a natural language that matched the manuscript on the word-order and vocabulary comparisons, though no text compared comes close at the manuscript’s vocabulary sparsity. So, too, would evidence from the physical book, beyond the published study of its quires and binding, that the text was not written onto finished drawings.

The paper

Supplementary material

The transliteration of the manuscript is public and is not redistributed here; the script below verifies a downloaded copy against a recorded fingerprint before it computes anything.

References

W. R. Newbold, The Cipher of Roger Bacon, University of Pennsylvania Press, Philadelphia (1928)announced the book as Bacon’s notebook in a microscopic shorthand hidden in the ink
J. M. Manly, Roger Bacon and the Voynich MS, Speculum 6 (1931) 345dismantled that reading: the shorthand was cracks in the ink, the anagramming unconstrained
W. F. Friedman, Anagrammed conclusion on the Voynich manuscript, footnote in Philological Quarterly (1959)his verdict, deposited as an anagram: an early constructed language, not a cipher
J. H. Tiltman, The Voynich manuscript: “The most mysterious manuscript in the world”, address to the Baltimore Bibliophiles (1967)cataloged the rigid structure of the words and endorsed no solution
P. H. Currier, Papers on the Voynich manuscript, in New Research on the Voynich Manuscript: Proceedings of a Seminar, ed. M. E. D’Imperio (1976)the two statistical dialects, A and B, which change with the handwriting
G. Rugg, An elegant hoax? A possible solution to the Voynich manuscript, Cryptologia 28 (2004) 31a Cardan grille slid over a table of word fragments mass-produces Voynich-like text
T. Timm, How the Voynich manuscript was created, arXiv:1407.6639 (2014)the copy-and-modify proposal whose title the paper echoes
B. Hauer and G. Kondrak, Decoding anagrammed texts written in an unknown language and script, Trans. Assoc. Comput. Linguist. 4 (2016) 75the Hebrew proposal with fixed-rule anagramming within each word, which the paper tests
L. Fagin Davis, How many glyphs and how many scribes? Digital paleography and the Voynich manuscript, Manuscript Studies 5 (2020) 164the five scribal hands, an attribution the paper adopts without testing
T. Timm, One hand, five labels: A critical examination of the five-scribe hypothesis for the Voynich manuscript, Zenodo preprint, doi:10.5281/zenodo.19169450 (2026)the objection that a single scribe wrote the book in a hand that drifted
T. Timm and A. Schinner, A possible generating algorithm of the Voynich manuscript, Cryptologia 44 (2020) 1the “self-citation” generator; on withheld pages its local copying is a minor component
C. L. Bowern and L. Lindemann, The linguistics of the Voynich manuscript, Annu. Rev. Linguist. 7 (2021) 285the review judging the text most likely a natural language in a non-standard encoding
M. A. Greshko, The Naibbe cipher: A substitution cipher that encrypts Latin and Italian as Voynich manuscript-like ciphertext, Cryptologia (2025) 1the Naibbe cipher, executable by hand, tested here from its published tables
R. Zandbergen, The Voynich manuscript, www.voynich.nu (accessed 28 July 2026)the volunteer-maintained site of the Zandbergen–Landini transliteration on which all counts are taken, also cited for the book’s early history

← back to the web summaries