New user, I'm sure this idea has been debunked already but on the tiny chance that it hasn't been:
What if what we currently consider to be "letters" aren't letters at all but akin to "strokes" in hanzi? So no connection to phonemes, but each word is still a word, or at least a concept akin to a single Chinese character. This could explain why the inventory of characters is so low (and strangely distributed) for a natural language (Hanzi has around 8 or so basic strokes, which combined in 2d space make a character, Voynichese could be similar but in 1d), and no attempt at pulling phonemic values from them has been successful.
Obviously there is absolutely nothing in the historic record akin to this, and I'm not even sure if this has been repeated in the world of conlanging, but it could theoretically work as a system?
This would however make it significantly harder to actually solve, as the "words" are then relatively arbitrary when it comes to associated meanings.
Any thoughts or ideas as to how this could be proved/disproved?
This is pure and idle speculation, but I'm trying to think of ways that aren't the extremely debunked substitutions, but still would (I believe) conform with the statistical analyses.
Hello you-all!
I first want to say I have no qualifications besides finding the VMS an incredibly interesting book (and being an occasional conlanger), and I apologize if this theory has already been stated, restated, and debunked ad nauseam, and for any unorthodox terminology used.
My idea is this: although exactly what counts as a letter as opposed to a variation, something in complementary distribution, etc., in the VMS's script is endlessly debated, all the theories that I have seen on the matter agree with 10-14 unique "letters," which is then argued as to whether it is an alphabet, an abjad, etc., but no system can explain the especially low entropy of the VMS text, in comparison to any Latin, German, Italian, or otherwise text in any language whatsoever, to the best of my knowledge.
So instead, assuming that the text has any meaning to it at all, perhaps the "letters" do not represent sound, but instead could be used, sort of as a logography, by combining certain letters and ascribing that sequence to a word, perhaps in combination with certain letters being a semantic category of sorts. For example, I could make "I" A, "sit" B, "on" AA, "the" AB, and "chair" BB, resulting in the sequence "A B AA AB BB" for "I sit on the chair," which, although just to my intuition, would result in a lower entropy than a representation of phonemes in an order, no matter the system devised to turn the words into sequences of characters, as human-assigned patterns would probably have a higher degree of predictability than phonemes. You could then augment this system with, say, "Q" showing that a word is inanimate, and "X" showing that a word is animate, resulting in "XA B AA AB QBB," which may explain the gallows?
I confess to not knowing much of anything when it comes to how text is statistically analyzed, but in assuming this is true, it is still probably impossible to ascertain which letters are "semantic," if such characters do exist, and which are the "building blocks," until and unless the MS is translated, so perhaps one could give a unique identifier to each "word" in the MS (though I know precisely discriminating them is a hard task) and comparing the "relations" of words to one another and then to that of a natural language, which in concrete terms could mean assessing the probability that given some word from the MS, how likely another certain word would follow? And comparing that to the results of some other natural language? That would be my best guess as to falsifying the theory on the merits of statistics.
On problems with the theory, I will say, this mean of statistical analysis seems deceptively simple to the point I would think somebody to have already done this sort of assessment on the MS. In addition this method of "cipher," or perhaps shorthand, to my knowledge is unattested as a method thereof, and in any case would work much better on a more analytical language like modern English rather than Latin, Arabic, or to a lesser extent, German, a sentiment I find strong enough, especially in consideration of a statistic I heard, which may be false, that there are 6000 unique sequences of characters in the MS, that I believe reconciling this fact would require the MS writer to forego articles, declension, and conjugation of whatever language they wrote in, or more likely simplify it considerably, in order to achieve that figure. So testing this hypothesis may require a considerable simplification of the target language one thinks likely to be that of the MS, probably to the point of implausibility. Now, if that figure were wrong, and there was, maybe, 8000, 10000+ "words" in the MS, I would find it more plausible, but this (im)plausibility does rely on an incomplete understanding of the languages it could have been written in, that is to say, if it were in German one might just have to remove the articles to get a "Voynich-like" look, but this borderlines on irrelevancy.
I would like to hear your thoughts on this, if for nothing else than to learn more about the VMS.
Thank you for reading
I would like to share with you the latest revision of my work on the statistical infrastructure of Beinecke MS 408.
I tried to structure the paper following a strict, step-by-step logical thread.
I also want to take this opportunity to sincerely thank Prof. @Jorge_Stolfi for his insights in previous threads, which I found incredibly useful for better calibrating the morphological parameters.
You can download the PDF + the Python codes for verification directly from Zenodo at this link:
[KG: link to hallucination-filled paper redacted]
For those who don't have the time to read everything right away, here is a quick summary of the 5 key points (all fully supported by their respective tests and charts inside the paper):
1. Reduplications are not random (Logistic Regression)
Running Monte Carlo tests (1,000 iterations, being careful to avoid data leakage across single dialects), the numbers speak clearly. Word reduplication is not a distraction error by the scribe, but responds to precise syntactic rules. Furthermore, cross-linguistic benchmarking shows that this behavior is quantitatively superimposable on that of isolating/Austronesian languages (like Indonesian).
2. 1-to-1 Compression and Greenberg Index
I applied Prof. Stolfi's Core Alphabet to clean the text from false entropy peaks generated by the EVA system. By feeding the data to the Zellig Harris algorithm, we get an Index of Synthesis of 2.86. Considering we are talking about very short words (3-4 glyphs on average), this certifies that we are dealing with 3 rigid morphological slots (Onset + Nucleus + Coda). It looks exactly like the profile of a monosyllabic isolating language.
3. The anomaly of the missing pages
By mapping the syntax, we can see that the density of the markers (the reduplicated words) is not uniform, but it spikes exactly near the 14 folios that were historically removed from the quire. This leads me to hypothesize that they acted as references/pointers to reading keys or legends that are now lost. (Note: this analysis was validated by cross-referencing the dialects across the entire volume)
4. Falsification Test I (Copy-Mutate & Skew Squares)
I rigorously tested the stochastic Copy-Mutate model (Null Hypothesis) and it fails the Contextual Asymmetry test: it cannot replicate the real cross-dependencies we see between prefixes (like o-, y-) and suffixes (-edy, -eedy)
5. Falsification Test II (Ghost Roots)
I tried to permute the letters within the dictionary. The text rejects these permutations with a ratio of 200 to 1 (12,450 attested real words vs. only 60 "ghost" ones). It is very hard to believe in a random generation with such strict phonotactic constraints.
I would really appreciate knowing what you think of these specific mathematical metrics, and I would also be absolutely thrilled to see these tests pushed to their limits
-Alfredo
Hello all,
Almost a year and a half ago I registered on this site for the first time and made a hello post where I said I was going to make a copy of the entire manuscript by hand. Never posted or commented since, though I've done a lot more lurking and reading.
And I've done it!
Almost.
I have about 2 pages left to ink, and 20 pages plus the Rosette fold out to paint. And of course I have to bind the entire thing. I'm hoping I'll be completely done in two or three weeks.
I'm in the process of scanning and formatting the pages I've completed while I finish the last physical pages, and I think at least one person was interested in seeing it? Here's a link to a goole drive folder with a sampling of completed pages if that's you: You are not allowed to view links. Register or Login to view.
Let me know if there are any particular pages anyone is interested in and I'll add them.
Thanks for looking!
-michelle
Hi! It's @JustAnotherTheory again. This time I submit to you the theory that Johannes von Gmünden was the writer of the VMS marginalia.
Johannes von Gmünden (1380/1384 - 1442) was an astronomer in Austria (You are not allowed to view links. Register or Login to view.). Of particular note is that he was the student of Heinrich von Langenstein, who worked personally with Nicolas Oresme in Paris (link to You are not allowed to view links. Register or Login to view.).
Not many writings of his own hand survive, unfortunately. One of them that does, is Augsburg, Universitätsbibl., Cod. III.1.4° 1 (You are not allowed to view links. Register or Login to view.).
In this manuscript, we can see some startling things. Here are a few words that remind one of the VMS folio 17r:
Others ressemble the VMS folio 116v:
We also have an MS from him where he draws some kind of star map with the sun and the moon having human faces:
(07-07-2026, 09:39 PM)Stefan Wirtz_2 Wrote: You are not allowed to view links. Register or Login to view.Those entropy calculations are still very depending upon the type of transliteration; some more aggressive "experts" here are favourizing an "h2" of 2.3 for Voynichese, while the change of transliteration file can produce a value of ~2.9.
As @nablator already pointed out, this is not correct. You wrote that this value was reported for the v101 transliteration. I just recomputed this with my own tools, stripping the transliteration of all annotations, but keeping the full character set.
The v101 file uses 158 different characters, if one counts the space also as a character.
Computing the entropy values, also counting space as a character yields:
Single char entropy H1 = 4.036
Character pair entropy H2 = 6.608
Conditional ch entropy h2 = 2.572
More realistic is not to count space as a character (it isn't). Then:
Single char entropy H1 = 4.131
Character pair entropy H2 = 6.501
Conditional ch entropy h2 = 2.370
In this case, character pairs spanning a space are not included in the statistics.
The most representative value is 2.37.
Comparison values may be found here: You are not allowed to view links. Register or Login to view.
and here: You are not allowed to view links. Register or Login to view.
Trying to create a cipher based on Roman numerals that would give a result similar to VMS, I came across the fact that even after all the manipulations with the text, it was not quite similar to Voynichese. But by adding various third-party rules, such as turning er into or in the right places, I gradually crept up on the most Voynich-like result.
I couldn't create an accurate algorithm that would immediately give me the desired result, no matter how much I wanted to (if anything, I made a cipher based on Roman numerals).
In the course of these attempts, I wondered how accurate the Voynichese itself was. How perfect and accurate is this algorithm from the point of view of cryptology? Is it an algorithm at all?
These questions call into question the usually unconditional "presumption of perfection" (my name for this cliche ), which is implied, in particular, by referring to the "scribe's mistake". For some reason, errors on the part of the author are often omitted.
One of the difficulties for analyzing Voynichese may be that it is not suitable for analysis as a cipher due to the lack of a clear encryption algorithm. By a clear algorithm, I mean a strictly defined and mathematically clear technique that gives and gives the same result (APPLE - GVVRK, the shift is 6; if I do the same with the word GVVRK, I get APPLE back). You can try to complicate this algorithm by adding non-mathematical rules. In the example with the Caesar cipher, I may want the first consonant to be I, the middle consonant is shifted by 2 again, and the final consonant is not encrypted. So I'll get GIXXRE from GVVRK. It became more difficult to determine the encryption algorithm and the specific substitution method that I used, the letters G and R, unaffected by my crazy rules, will not give me anything (I can shift them back to 6, there will be AIXXLE. Of course, although it's easy to guess the word APPLE, we may be confused by the two XX's in the middle and by the I. Or maybe it's not APPLE, but ACCEPT, but E is put at the end, C = X, and P and T are put in other places?). Briefly, we have the opportunity to blur the algorithm, making it difficult for a stranger to decipher it.
And now we ask ourselves the question - why should the Voynichese be a clear cipher? Why do we believe that the author was so brilliant and talented that he created a great undecipherable method unknown to science? In general, a person who is ignorant of cryptology is not expected to create the cleanest algorithm, but to use conditional substitution as a basis, having previously complicated it. There is nothing supernatural in the fact that Voynichese is unlike other ciphers of its time, because I can also mock an ordinary substitution to bring it to "perfection" without changing the algorithm itself. At this rate, I won't have any trouble, but...
By abusing these so-called "rules", I can't complicate the cipher, but distort it, making decryption difficult even for a person who knows the cipher, especially if these rules are unclear (for example, not "after every first consonant I", but "after some first consonants I"). In theory, it is possible to complicate the cipher so much that decryption will be a great challenge for the expert (that is, I will ruin it), which should not be. And there's no guarantee that Voynichese is not a corrupted cipher. There's no guarantee that even with an algorithm, we will immediately decipher the entire manuscript without difficulty. It is possible that it is impossible to decipher Voynichese a priori (more precisely, it is possible, but the decoding will be incomplete). Perhaps that it's impossible to decrypt Voynichese a priori (or rather, it is possible, but the decryption will turn out to be incomplete) due to the fact that the author knew very little about the logic of ciphers and, without realizing it himself, made his possibly rather simple cipher very difficult to decrypt.
In general, questioning the presumption of perfection applies not only to cipher theory. Some linguistic hypotheses also depend on it, for example, Chinese theory, but to a lesser extent (most likely, the author was not a professional linguist, and his transcription of Chinese could be completely unreadable. Historical example: You are not allowed to view links. Register or Login to view. for Armenian and, maybe, Hebrew, which requires knowledge of the language itself to read).
What are some historical examples of a "corrupted cipher"?
Codex Copiale (but here, rather, an example of a complicated substitution with the addition of zeros) and You are not allowed to view links. Register or Login to view. (still not deciphered, but it is believed that this is a complicated Polybius square).
How did the owners read this manuscript?
By owners, I mean a team of people who understood what VMS was about and may have participated in its creation. In this case, they have an advantage - they know in advance what should be written, and therefore at some points they can, for example, guess undeciphered words from the context, or even intuitively understand the meaning of what is written without deciphering it (for example, if the balneological section is anatomy, then the reader can take a look look at the illustration, remember what kind of organ it depicts, and from there remember what he knows about this organ. In this case, he probably won't need to decrypt anything anymore).
What should I do with VMS now?
It's worth thinking not about how to create the most warlike cipher possible, but about how we can achieve this result in more mundane and accessible ways. What we can learn from these attempts can be applied directly to the text of the manuscript to find any coincidences or something else...
In any case, it seems to me to be a satisfying food for thought...
This unique symbol appears in the text several times, but it seems to me that given this number of appearances, it can't be called "unique".
There were also his appearances in botany, but I couldn't find them. Most often, there is no additional loop (as in the first two screenshots).
Outwardly, it looks vaguely like an f with a second leg, but I think that's not the case, because we already have an f with a second leg: