The Voynich Ninja

Full Version: Is there any protocols to identify what sort of document MS-408 is?
You're currently viewing a stripped down version of our content. View the full version with proper formatting.
Pages: 1 2 3
(16-07-2026, 04:08 PM)Oscroft Wrote: You are not allowed to view links. Register or Login to view.Interesting... a colang should look like a real lang, but with less of the apparent randomness that builds up with natural langs over the centuries. So, more consistent structure,

I agree on this, well said. 

(16-07-2026, 04:08 PM)Oscroft Wrote: You are not allowed to view links. Register or Login to view. less entropy... yes, makes sense, thanks.

About entropies I'm not so sure. Lower-than-usual entropy is of course possible but I'd rather measure it on real conlangs before saying anything about it.
(16-07-2026, 04:47 PM)Mauro Wrote: You are not allowed to view links. Register or Login to view.About entropies I'm not so sure. Lower-than-usual entropy is of course possible but I'd rather measure it on real conlangs before saying anything about it.
This is just a matter of creating a conlang where the letters order such that the uncertainty is lower than most naturalistic languages, which could be done. In Voynichese, entropy is almost certainly downstream of the qualitative lettering-ordering effects and whatever constraints created those. You could probably target a pretty accurate entropy value just by defining your word generation process to have exactly as much entropy as you wanted the final text to have and hit pretty close. Again, this would be bizarre, like Voynichese, but there really isn't a conceptual hurdle to restricting word formation to certain letter orders and, has been amply discussed, some real-world languages like Hawai'ian have alphabetic representations where we do see these constraints giving lower entropy values. It also would not capture some of the other oddities, like linestart words and apparent positional variation, but neither does the entropy summary statistic fully, so that's all good.
(16-07-2026, 05:17 PM)rikforto Wrote: You are not allowed to view links. Register or Login to view.This is just a matter of creating a conlang where the letters order such that the uncertainty is lower than most naturalistic languages, which could be done.

Surely it can be done, but the reverse could be done too. I don't know what the entropies of actual conlangs are and I wouldn't dare say anything about Esperanto, Volapük or You are not allowed to view links. Register or Login to view. entropies without measuring them first.
Quote:I don't know what the entropies of actual conlangs are

Esperanto will be probably very similar to Latin or Italian. Any constructed language based on one or more real languages ( so called a posteriori conlang) will have have similar statistical properties to the real languages being the inspiration.

See: You are not allowed to view links. Register or Login to view.

On the other hand there are some really weird constructed languages which may behave differently, for example:

You are not allowed to view links. Register or Login to view.
You are not allowed to view links. Register or Login to view.

But yet on the another hand any "weird" constructed language, strongly different from natural languages, would be quite anachronistic for the 1400s.
(16-07-2026, 05:25 PM)Mauro Wrote: You are not allowed to view links. Register or Login to view.I don't know what the entropies of actual conlangs are and I wouldn't dare say anything about Esperanto, Volapük or You are not allowed to view links. Register or Login to view. entropies without measuring them first.
I don't know for sure, but it might be slightly lower than natural.
As I said, comparing it to Esperanto or Volapük is pointless, because Voynichese was not intended for communication.

(16-07-2026, 05:48 PM)Rafal Wrote: You are not allowed to view links. Register or Login to view.But yet on the another hand any "weird" constructed language, strongly different from natural languages, would be quite anachronistic for the 1400s.
I don't think so... Of course, it would be more of an anachronism to have a spoken artificial language in the 15th century. It could have been created for specific purposes (such as the VMS itself) and based on a rigid principle.
(16-07-2026, 03:09 AM)Nyeogmi Wrote: You are not allowed to view links. Register or Login to view.I don't think this information is currently known.


If it's a natural language, it either has a very rigid word and syllable structure (and a small number of sounds) or the letters do not encode that. I think the "monosyllabic tonal language" theory is the most plausible version of this, but it raises questions given that the book is full of western European cultural images.

If it is a conlang, it's probably not especially similar to German or Latin. In a "low entropy" conlang, we would expect that it takes four or five times as many words for Voynich Writer Guy to get their meaning across.

If it's a cipher, we're not sure what technology was used. It would probably be of higher complexity than common ciphers from the time period. Weirder things have happened. I would expect the construction to feel a lot less artificial than something like the Naibbe cipher, but if it is discovered to mean anything, this is kind of what I'm expecting.

When I see "qokedy qokeedy qokechdy lol qokedy qokeedy qokechdy lol" my immediate thought is "That person is inexactly copying and manipulating symbols from elsewhere on the page.

Is it possible the author used a cipher with the conlang of a known language yet still obscured the language by not using words from that language in it's range at least 90% of it?  This would make people if they stumbled across the cipher to not except it, because most of the language remained unknown.  The low entropy repeating would mean the document does not contain in depth information.
(16-07-2026, 05:25 PM)Mauro Wrote: You are not allowed to view links. Register or Login to view.and I wouldn't dare say anything about Esperanto, Volapük or You are not allowed to view links. Register or Login to view. entropies without measuring them first

There are several good points in this thread, and this is one.
People usually underestimate how significant a change you need to a text in order to appreciably change the character-level entropy values. 

(16-07-2026, 05:17 PM)rikforto Wrote: You are not allowed to view links. Register or Login to view.In Voynichese, entropy is almost certainly downstream of the qualitative lettering-ordering effects and whatever constraints created those.

Another good point. Entropy (contrary to age) is just a number. There is a whole frequency distribution of fractions adding up to one behind it. There are infinitely many very different distrubutions with the same entropy.
(17-07-2026, 12:46 AM)ReneZ Wrote: You are not allowed to view links. Register or Login to view.Another good point. Entropy (contrary to age) is just a number. There is a whole frequency distribution of fractions adding up to one behind it. There are infinitely many very different distrubutions with the same entropy.

This is true, and the VMS's low entropy also places fundamental information-theoretic constraints. If you are trying to reconcile Voynichese with meaningful European text via some kind of letter- or phoneme-based cipher, you can have a lot of plaintext rearranging (e.g., alphabetization), high verbosity within the cipher alphabet(s), or some intermediate mix of both. Abbreviations could shape things on the margins, but they alone cannot be driving the low entropy.
(17-07-2026, 12:46 AM)ReneZ Wrote: You are not allowed to view links. Register or Login to view.People usually underestimate how significant a change you need to a text in order to appreciably change the character-level entropy values.

Indeed: deleting, rearranging, and duplicating words will usually have very little effect on character entropy, whether simple (h0) or conditional (h1,h2,..).  To lower the entropy one would have to delete words with high character entropy and replicate words with low entropy many times.  The banana ananas text ananas banana banana would banana become daiin daiin raiiin rather daiin daiin boring daiin kaiin.

But changes in the spelling (or encryption method, or gibberish generator) can easily change any hk to any desired value, from near zero to the maximum of log2(N) for an alphabet of N characters (~4.7 for lowercase Latin letters plus space). If the h0 of the plaintext is (say) 2.0, one can raise it to ~(2.0+4.7)/2 = ~3.3 by inserting a random letter after each original letter or space; or reduce it to ~1.0 by inserting an "s" after each letter or space.  

If the plaintext is all in lower case, one can add 1 to every hk by randomly changing half the letters to upper case.   

One can even raise every hk to the maximum, without changing the length of the text or the alphabet, by encrypting it with a Vigenère cipher with the 26-letter key ABCDE...XYZ.

Conversely, if h0 is greater than h1, one can lower h0 without changing the length of the message by encoding each letter with a simple cipher that depends on the previous letter.  For example, suppose the plaintext strictly alternates between the "consonants" K,T and the "vowels" A, O, and KA,KO,TA,TO  have the same frequency (~25%).  Then h0 will be ~2 bits/letter but h1 will be ~1.  Then one can lower h0 to ~1 by replacing K☛A, T☛O except at the first letter of each word.

All the best, --stolfi
Pages: 1 2 3