The Voynich Ninja

Full Version: The Voynich Manuscript: Understanding the Untranslatable with Lisa Fagin Davis
You're currently viewing a stripped down version of our content. View the full version with proper formatting.
Pages: 1 2 3 4 5
(13-09-2026, 02:19 PM)Torsten Wrote: You are not allowed to view links. Register or Login to view.
(13-09-2026, 01:46 PM)asteckley Wrote: You are not allowed to view links. Register or Login to view....

You write that I "know full well the many reasons why Davis holds a fairly common opinion that it seems to be Indo-European." I would be interested to hear what those many reasons are and where they are published.

I would like to know those reasons, too.
EDIT: ok, summation arrived right before my question, thx.
(13-09-2026, 03:02 PM)Stefan Wirtz_2 Wrote: You are not allowed to view links. Register or Login to view.I would like to know those reasons, too.
See previous post  -- it slipped in before yours.
(12-09-2026, 10:22 PM)Torsten Wrote: You are not allowed to view links. Register or Login to view.In this podcast Lisa Davis makes several claims about the Voynich text. ...

I agree with several of your criticisms. (And I do not think they are "AI generated").  But I disagree in others.

1. The Voynichese alphabet is "too small for a natural language" only if one assumes that each distinct phoneme is encoded as a distinct single glyph.  But this need not be true.  Even in English several phonemes are encoded as digraphs, and a single letter or digraph can represent several very distinct phonemes. 

The structure of the Voynichese words suggests that e is a modifier for the preceding glyph, so that Ch and Che are different "letters", and so probably are k and ke, ee and eeein and iin and iiin.  If that possibility is taken into account, the size of the alphabet is easily 20 or more.

2 and 3. If anything, the VMS word structure and statistics exclude all "European" languages, all Semitic languages, and probably all the languages of Africa, West and Central Asa, Oceania, and the Americas.  But they are quite compatible with the monosyllabic languages of East Asia, such as Sino-Tibetan, Vietnamese, etc.

4. The methods used to identify consonants and vowels (like Sukhotin's algorithm, that Guy unsuccessfully tried) are based on the principle that Cs and Vs tend to alternate in polysyllabic words.  But the structure of Voynichese words is definitely not compatible with multiple syllables.  

5. Indeed the fact that the word frequency distribution of a text follows Zipf's Law is just a necessary but not sufficient condition for it being natural language, either unencrypted or encrypted one-to-one on words -- that is, in such a way that each word type is mapped to a single word type.  

Thus the fact that Voynichese follows the law does let us rule out is certain ciphers like Vigenère, that are not one-to-one on words.  That fact also lets us rule out some methods to generate random gibberish, like Rugg's "grille" method".  But I suspect it would also exclude the "copy and mutate" method, because obtaining a Zipfian distribution with it seems to require a rather nontrivial "mutate" procedure.

6. If the text is a natural language, unencrypted or encrypted one-to-one on words, we expect that each section will have a significantly different word frequency distribution, with rather abrupt changes across section boundaries.  Whereas gibberish generated by a random method (including "copy and mutate") would probably generate either the same word distribution through the whole book, or a gradually drifting one, oblivious of section boundaries. 

This has always been a very popular topic of statistical research, but the results seem hard to interpret, because the historical scrambling of the bifolios can both homogenize the distribution within each section and create abrupt transitions between sections. 

7. Full-word repeats are indeed rare in the "usual suspect" languages.  But not in East Asian languages -- especially in imperfect transcriptions that may omit tones or other phonetic details. 

The fact that "the most experienced NSA researchers in the history of Voynich studies" have concluded that the word repeats rule out natural language only reminds us that "all the experts" on anything can often be quite wrong.  

(Ironically, someone recently criticized the Chinese theory by pointing out that the Starred Parags section has too few repeated sequences compared to the Shennong Bencao...)

All the best, --stolfi
(13-09-2026, 03:01 PM)asteckley Wrote: You are not allowed to view links. Register or Login to view.
  • Morphological profile. Luke Lindemann’s broad comparison of Voynichese with 160 languages found that the closest family-level profiles were mostly Indo-European. Iranian ranked first for the full manuscript and for both Currier varieties; Germanic ranked second for the full text and Currier B; Indic and Romance also ranked highly. This is the most direct quantitative reason for preferring an Indo-European candidate, though it rests on only two surface measures and cannot identify a particular language.

When you apply any metric to an unknown text, you will always find a closest match. But a closest match is not an identification. For instance You are not allowed to view links. Register or Login to view. (2016) identified Hebrew as closest match. You are not allowed to view links. Register or Login to view. (2022) did found Iranian as closest match. But he also hedges his result by writing: "While character-level measures mark Voynichese as a complete outlier among historical European manuscripts, the word-level measures for Voynich place it comfortably between the Old Slavonic and Hebrew historical texts on one side and the Italian and Welsh texts on the other."

Anyway, it is interesting that you see the linguistic results of Tiltman, Bennett, Friedmann, Currier, D'Imperio, Stallings, Reddy & Knight, and Bowern & Lindemann as debatable opinions about "open questions" while pointing to non-language-related arguments like 'European manuscript environment' as counter-arguments.
(12-09-2026, 10:22 PM)Torsten Wrote: You are not allowed to view links. Register or Login to view.4. Davis further argues that "there's a linguist who's doing work trying to figure out which are the vowels and which are the consonants." However, Guy (1991) and Reddy & Knight (2011) both attempted vowel-consonant separation and found it does not work.

I don’t think that Guy found that vowel-consonant separation did not work. This is what he wrote in “Statistical properties of two folios of the Voynich manuscript” (1991). Of course, one can disagree with his conclusions.

Guy Wrote:Sukhotin's algorithm identified ( c), ( C) (i.e. (cc)), (o), (g), (a) and (i) [EVA:e ee o y a i - e ee o y a i] as representing vowels in both transcription systems used here.
First, (o) and (a) of the Voynich manuscript are similar to the corresponding vowels of our modern scripts.
Second, the Spanish script known as "Visigothic" and the Italian script known as "Beneventan" or "Beneventine"', both in use from the VIIth to the XIlIth century, are characterised by a form of 't' identical to the group (ct) [ch] in the Voynich manuscript, identified as a consonant by Sukhotin's algorithm.
Third, that same Visigothic script has a form of 'a' very similar to the Voynich (cc) [ee] sequence, identified as a vowel; while the 'a' of Beneventan resembles a ‘c' or an 'o' fused to a following 'c'. Irish and Anglo-Saxon manuscripts of the same period exhibit both forms of 'a' (open, as in Visigothic, and closed as in Beneventan) along with forms like our modern cursive small 'a'.
Unless these are mere coincidences, they suggest that the Voynich manuscript either is of medieval origin, probably from Italy or Spain, or is a later forgery by someone with a knowledge of medieval palaeography.
Quote:When you apply any metric to an unknown text, you will always find a closest match. But a closest match is not an identification.

I agree. That's some weakness of statistical tool and models. They will always give you something - closest language, consonant and vowels, groups of similar words that behave in the same way, sections and places where "topic" changes etc. Are they real or just statistical artifacts is another story.

My belief, after spending some time on this forum, is that currently existing statistical models are unable to tell real language with meaningful information from structured gibberish. 

I haven't listen to Lisa podcast but if actually she said some stuff then some criticism of Thorsten is justified.

On the other hand, sorry to say that but it feels a bit too intensive. Thorsten, I like your autocitation theory but please don't become a psychofan of Lisa  Wink
I would like to point out that I do make an effort in interviews such as this one to equivocate - use words like SUGGESTS and MAY BE and MIGHT INDICATE. In fact, I clearly did so in four of Torsten's seven complaints. There are indeed moments where I don't - and that is useful feedback for the future. I hope you all have noticed that I do not EVER claim to know what the manuscript IS. That's because I don't know. None of us do, at least not with certainty.
(14-09-2026, 01:54 PM)LisaFaginDavis Wrote: You are not allowed to view links. Register or Login to view.I would like to point out that I do make an effort in interviews such as this one to equivocate - use words like SUGGESTS and MAY BE and MIGHT INDICATE. In fact, I clearly did so in four of Torsten's seven complaints. There are indeed moments where I don't - and that is useful feedback for the future. I hope you all have noticed that I do not EVER claim to know what the manuscript IS. That's because I don't know. None of us do, at least not with certainty.

Taking a step back, it's actually kind of silly to expect you (or anyone else) to have to mention all of the references, reasoning and justifications that ground their opinions in a spoken conversation. It's a podcast, not a scientific paper. You should be able to say what you think and feel, which is something which people find valuable and interesting. 

Obviously, there is a line where things would be inappropriate (like saying it was definitely written by aliens or something), but speculation and unfiltered opinions are entirely appropriate in this setting imo. For them to be picked apart on technicalities as if they were part of an officially published paper is against the spirit of the content.
Thank you, I agree with you absolutely. This is a podcast. It isn't a peer-reviewed essay or a conference lecture. It's a public-facing informal podcast, so while we should all take pains to be cautious in how we speak about the VMS, we also need to consider the context and audience of any presentation or publication.
(14-09-2026, 02:21 PM)LisaFaginDavis Wrote: You are not allowed to view links. Register or Login to view.Thank you, I agree with you absolutely. This is a podcast. It isn't a peer-reviewed essay or a conference lecture. It's a public-facing informal podcast, so while we should all take pains to be cautious in how we speak about the VMS, we also need to consider the context and audience of any presentation or publication.

But it would be best to not spread misinformation, like "He (Voynich) was Jewish" (in this podcast and in the Atlas Obscura podcast, 16 may 2023) and "There's some letters from Kircher where he says, oh, looks like Glagolitic to me".
Pages: 1 2 3 4 5