09-07-2026, 08:12 PM
Hello you-all!
I first want to say I have no qualifications besides finding the VMS an incredibly interesting book (and being an occasional conlanger), and I apologize if this theory has already been stated, restated, and debunked ad nauseam, and for any unorthodox terminology used.
My idea is this: although exactly what counts as a letter as opposed to a variation, something in complementary distribution, etc., in the VMS's script is endlessly debated, all the theories that I have seen on the matter agree with 10-14 unique "letters," which is then argued as to whether it is an alphabet, an abjad, etc., but no system can explain the especially low entropy of the VMS text, in comparison to any Latin, German, Italian, or otherwise text in any language whatsoever, to the best of my knowledge.
So instead, assuming that the text has any meaning to it at all, perhaps the "letters" do not represent sound, but instead could be used, sort of as a logography, by combining certain letters and ascribing that sequence to a word, perhaps in combination with certain letters being a semantic category of sorts. For example, I could make "I" A, "sit" B, "on" AA, "the" AB, and "chair" BB, resulting in the sequence "A B AA AB BB" for "I sit on the chair," which, although just to my intuition, would result in a lower entropy than a representation of phonemes in an order, no matter the system devised to turn the words into sequences of characters, as human-assigned patterns would probably have a higher degree of predictability than phonemes. You could then augment this system with, say, "Q" showing that a word is inanimate, and "X" showing that a word is animate, resulting in "XA B AA AB QBB," which may explain the gallows?
I confess to not knowing much of anything when it comes to how text is statistically analyzed, but in assuming this is true, it is still probably impossible to ascertain which letters are "semantic," if such characters do exist, and which are the "building blocks," until and unless the MS is translated, so perhaps one could give a unique identifier to each "word" in the MS (though I know precisely discriminating them is a hard task) and comparing the "relations" of words to one another and then to that of a natural language, which in concrete terms could mean assessing the probability that given some word from the MS, how likely another certain word would follow? And comparing that to the results of some other natural language? That would be my best guess as to falsifying the theory on the merits of statistics.
On problems with the theory, I will say, this mean of statistical analysis seems deceptively simple to the point I would think somebody to have already done this sort of assessment on the MS. In addition this method of "cipher," or perhaps shorthand, to my knowledge is unattested as a method thereof, and in any case would work much better on a more analytical language like modern English rather than Latin, Arabic, or to a lesser extent, German, a sentiment I find strong enough, especially in consideration of a statistic I heard, which may be false, that there are 6000 unique sequences of characters in the MS, that I believe reconciling this fact would require the MS writer to forego articles, declension, and conjugation of whatever language they wrote in, or more likely simplify it considerably, in order to achieve that figure. So testing this hypothesis may require a considerable simplification of the target language one thinks likely to be that of the MS, probably to the point of implausibility. Now, if that figure were wrong, and there was, maybe, 8000, 10000+ "words" in the MS, I would find it more plausible, but this (im)plausibility does rely on an incomplete understanding of the languages it could have been written in, that is to say, if it were in German one might just have to remove the articles to get a "Voynich-like" look, but this borderlines on irrelevancy.
I would like to hear your thoughts on this, if for nothing else than to learn more about the VMS.
Thank you for reading
I first want to say I have no qualifications besides finding the VMS an incredibly interesting book (and being an occasional conlanger), and I apologize if this theory has already been stated, restated, and debunked ad nauseam, and for any unorthodox terminology used.
My idea is this: although exactly what counts as a letter as opposed to a variation, something in complementary distribution, etc., in the VMS's script is endlessly debated, all the theories that I have seen on the matter agree with 10-14 unique "letters," which is then argued as to whether it is an alphabet, an abjad, etc., but no system can explain the especially low entropy of the VMS text, in comparison to any Latin, German, Italian, or otherwise text in any language whatsoever, to the best of my knowledge.
So instead, assuming that the text has any meaning to it at all, perhaps the "letters" do not represent sound, but instead could be used, sort of as a logography, by combining certain letters and ascribing that sequence to a word, perhaps in combination with certain letters being a semantic category of sorts. For example, I could make "I" A, "sit" B, "on" AA, "the" AB, and "chair" BB, resulting in the sequence "A B AA AB BB" for "I sit on the chair," which, although just to my intuition, would result in a lower entropy than a representation of phonemes in an order, no matter the system devised to turn the words into sequences of characters, as human-assigned patterns would probably have a higher degree of predictability than phonemes. You could then augment this system with, say, "Q" showing that a word is inanimate, and "X" showing that a word is animate, resulting in "XA B AA AB QBB," which may explain the gallows?
I confess to not knowing much of anything when it comes to how text is statistically analyzed, but in assuming this is true, it is still probably impossible to ascertain which letters are "semantic," if such characters do exist, and which are the "building blocks," until and unless the MS is translated, so perhaps one could give a unique identifier to each "word" in the MS (though I know precisely discriminating them is a hard task) and comparing the "relations" of words to one another and then to that of a natural language, which in concrete terms could mean assessing the probability that given some word from the MS, how likely another certain word would follow? And comparing that to the results of some other natural language? That would be my best guess as to falsifying the theory on the merits of statistics.
On problems with the theory, I will say, this mean of statistical analysis seems deceptively simple to the point I would think somebody to have already done this sort of assessment on the MS. In addition this method of "cipher," or perhaps shorthand, to my knowledge is unattested as a method thereof, and in any case would work much better on a more analytical language like modern English rather than Latin, Arabic, or to a lesser extent, German, a sentiment I find strong enough, especially in consideration of a statistic I heard, which may be false, that there are 6000 unique sequences of characters in the MS, that I believe reconciling this fact would require the MS writer to forego articles, declension, and conjugation of whatever language they wrote in, or more likely simplify it considerably, in order to achieve that figure. So testing this hypothesis may require a considerable simplification of the target language one thinks likely to be that of the MS, probably to the point of implausibility. Now, if that figure were wrong, and there was, maybe, 8000, 10000+ "words" in the MS, I would find it more plausible, but this (im)plausibility does rely on an incomplete understanding of the languages it could have been written in, that is to say, if it were in German one might just have to remove the articles to get a "Voynich-like" look, but this borderlines on irrelevancy.
I would like to hear your thoughts on this, if for nothing else than to learn more about the VMS.
Thank you for reading


