Reference check: Kinnison (2026), "The Voynich Manuscript's positional entropy collapse"
The new Cryptologia paper by Rod Kinnison (published 7 September, doi:You are not allowed to view links.
Register or
Login to view.) argues that a verbose substitution cipher reproduces the manuscript's statistical profile where natural languages do not. My first step with any new paper about the Voynich manuscript is to check the references. I did the same for Layfield and Davis (2026) and put the results on Zenodo (doi:You are not allowed to view links.
Register or
Login to view.). The results here are of the same kind, so I am posting them in the same descriptive spirit: what the entry says, what the source says, no argument about the paper's conclusions.
The bibliography has nine entries.
1. Bennett 1971. Given as "Scientific and engineering problems of the Voynich Manuscript, Proceedings of the American Philosophical Society 115(5):395–412", and it carries the paper's opening empirical claim about low conditional entropy. You are not allowed to view links.
Register or
Login to view., 15 October 1971, runs pp. 335–421 and contains six items: Baumgartner 335–340, Newman 341–349, Schultz 350–365, Kaminsky 366–397, True and Núñez 398–421, plus front and back matter. There is no Bennett in the volume and nothing on the Voynich Manuscript. The page range 395–412 does not correspond to any article; it straddles Kaminsky and True. The intended source is almost certainly You are not allowed to view links.
Register or
Login to view., W. R., Scientific and Engineering Problem-Solving with the Computer, Prentice-Hall 1976.
2. You are not allowed to view links.
Register or
Login to view.. Given as "Bowern, C., and S. Lindemann. 2021. The linguistic structure of the Voynich Manuscript. Annual Review of Linguistics 7:315–37." Four things are wrong: the co-author's initial (Luke, not S.), the title ("The Linguistics of the Voynich Manuscript"), the pages (285–308), and the missing DOI.
The content claim also does not hold. The citing sentence says the positional rigidity "remains stable across manuscript sections and transcription conventions and is not readily explained by known linguistic typologies, including agglutinative or highly regular constructed languages."
Of the four assertions there:
- Stable across sections: they report h2 = 2.17 for Currier A and 2.01 for Currier B, attribute the difference to the qo- and -dy affixes, and note the two values converge (2.23, 2.24) only after those affixes are removed.
- Across transcription conventions: they state that conditional entropy is to some extent affected by how characters are divided, and give the -iin case (three characters in EVA, one in Currier). Their claim is that character division cannot lower entropy to Voynich levels, which is a different and narrower statement.
- Agglutinative and constructed languages: the word "agglutinative" does not occur in the paper. Their conlang discussion covers Balaibalan, Lingua Ignota and Enochian as hypotheses about the manuscript, not as controls. The Figure 4 sample is natural languages in alphabets and abjads. The clause names precisely Kinnison's own two controls, Finnish and Esperanto.
- Not readily explained by known typologies: this one is supported. They do write that the entropy of Voynichese is unlike that of any other language or script.
3. You are not allowed to view links.
Register or
Login to view. 2013. Cited twice, and in my view this is the substantive problem. First, for the claim that quantitative studies converged on properties that "sharply diverge from those of known natural languages"; second, that "later information-theoretic studies placed Voynichese as an extreme outlier in comparative entropy analyses of written language."
They performed no such analysis. Their measure is an information quantity in bits per word describing how the distribution of a word across text partitions identifies those partitions. There is no conditional character entropy in the paper, no bigram or trigram analysis, no comparative entropy ranking of scripts. And their comparative result is the reverse of what is attributed: the Voynich maximum falls slightly above English and well below Chinese, peaking at a scale close to the natural-language examples, while the yeast DNA sequence and the Fortran source code fall outside the language cluster. Their abstract states that the manuscript's word organization is compatible with real language sequences and that this supports a genuine message in the book.
Their Discussion additionally states that any fabrication model must explain how the linguistic-like structures arose, and that the statistical structure requires an explanation going beyond local features such as word forms and local word sequences. Kinnison's evidence is entirely word-internal character adjacency and edit distance between neighbouring word forms.
4. You are not allowed to view links.
Register or
Login to view.. The entry is correct (NSA Technical Journal 12(3):41–85; the declassified scan is on media.defense.gov). The use is not. It is cited as work "subsequent" to Bennett 1971 that "confirmed that this effect persists across different transcription systems and analytical methods."
Tiltman's paper is an expanded version of a talk to the Baltimore Bibliophiles on 4 March 1967, which he describes as an introduction for readers coming to the manuscript for the first time. Contents: provenance and the Marci letter, the Newbold, Feely and Strong solutions, the text of his 1951 report to Friedman, digressions on Cave Beck and on the history of herbals, thirty-one plates, and a three-item bibliographical note. There is no entropy calculation and no frequency table. The 1951 report is qualitative — positional behaviour of the commonest symbols, roots and suffixes, the A-groups — and he opens it by saying there is no opinion in it to which he cannot himself find plenty of contradiction.
On transcriptions: the paper contains one, his own, in which he reduced the script to 17 symbols with arbitrary letters and figures, similar to a scheme Friedman had used. That is not a comparison across systems. This is hardly surprising since multiple parallel transcriptions of the whole text did not exist in 1967. Currier's dates from the mid-1970s; Takahashi's, Stolfi's and EVA from the 1990s. Three of the four transcriber layers in Kinnison's own robustness table (C, F, H, U in the Landini–Stolfi Interlinear) postdate the source credited with the finding. Conditional-entropy analysis of Voynichese begins with Bennett in 1976.
5. You are not allowed to view links.
Register or
Login to view.. Cited for the claim that "character distributions at word boundaries differ markedly from those in medial positions, with medial transitions exhibiting especially high predictability." I have not been able to consult the article and make no claim about everything it contains. What can be established is how the field cites it. Bowern and Lindemann invoke Landini (2001) for higher-level document structure and, jointly with Reddy and Knight, for the finding that Voynich word frequencies follow Zipf's law, illustrated by a rank-frequency chart. Montemurro and Zanette write that Landini showed the collection of tokens satisfies Zipf's law, with a smooth, approximately inverse relation between occurrences and rank, and group this with the token-level rather than the character-based studies. Both of these papers are in Kinnison's own bibliography. The standard citation context for Landini 2001 is word-frequency distribution; I have not found it cited anywhere for word-internal positional character statistics, and its method — spectral analysis — targets periodicity and long-range correlation rather than position within words. Anyone with the Cryptologia article to hand is welcome to correct this.
6. Zandbergen 2025. Cited as the source of the transcription: "All Voynich analyses in this study use the LSI-format transcription of the Voynich Manuscript (Zandbergen, 2025)." The transcription used is Takahashi's, carried as transcriber layer H in the Landini–Stolfi Interlinear file, which is distributed in IVTFF format. Takahashi is not cited anywhere in the paper. The entry for Zandbergen gives no URL, no file version and no access date, although the interlinear file is periodically revised and its contents determine which transcriber layers are available for the robustness check in Section 4.
7–9. You are not allowed to view links.
Register or
Login to view. 1978 is cited for the folio ranges of the sections, which needs no source but is not wrong. You are not allowed to view links.
Register or
Login to view. is correct. "Hart, n.d." is a citation to You are not allowed to view links.
Register or
Login to view. as an institution rather than to any text: the four control corpora appear only as bare identifiers in table headers (pg57184, pg18837, pg11940, 345-0), in two different formats, with no titles, authors, lengths or genres, and they are not among the files listed in the data availability statement.
On what is missing. Six references. Absent are You are not allowed to view links.
Register or
Login to view., You are not allowed to view links.
Register or
Login to view. and You are not allowed to view links.
Register or
Login to view. — the generative-mechanism literature, most of it in the same journal — in a paper whose stated aim is to narrow the space of plausible generative explanations. Also absent are You are not allowed to view links.
Register or
Login to view., the entropy paper whose results are attributed to the review paper of Bowern and Lindemann 2021, You are not allowed to view links.
Register or
Login to view., and You are not allowed to view links.
Register or
Login to view., although the discussion contains a paragraph on what Currier A and B represent.
I am aware I am an interested party regarding one of the omissions and I raise it as a matter of coverage, not priority.
Summary. One entry does not exist as printed. One is wrong in four particulars, with a page range that again matches no article. Three are used for claims their sources do not support, one of them in the opposite sense. One is not a citation to a text. Two are sound.
All of this is checkable in a few minutes by anyone with the volumes to hand, and all of it is correctable in an erratum. I make no claim here about the paper's methodology or its conclusions; that is a separate discussion, and I may come back to it.