What if it's a "Lossy Encoder" for 15th-Century Latin?
Pointless.. > 2 hours ago
I have been considering the possibility that the Voynich manuscript uses some form of lossy encoding.
By “lossy,” I mean that the source text may originally have been written in Latin, but that only part of its linguistic information was preserved in writing. Such a system would assume that the intended reader already possessed enough practical knowledge, contextual information, or interpretive tools to reconstruct the meaning. Where the result might otherwise be ambiguous, the system could provide additional cues to guide the reader toward the correct interpretation.
Shannon’s information theory suggests that natural languages contain a considerable amount of redundancy, sometimes estimated at around 50 percent. Latin, as a highly inflected language, carries much of its grammatical information in word endings. Nouns, adjectives, and verbs are marked according to their roles in the sentence, meaning that some structural information is communicated more than once.
In a practical text, however, much of this information may not be strictly necessary. If a practitioner already knows how to prepare a decoction, grammatical case endings and many function words may add relatively little to the basic operational meaning.
To see how this might work in practice, consider a short apothecary-style recipe in Latin:
"Ad dolorem capitis: Recipe radices hellebori, et tere in mortario. Misce cum aceto et oleo rosarum. Ungue frontem patientis ante somnum."
In ordinary English, the passage means roughly:
"For headache: Take hellebore roots and grind them in a mortar. Mix with vinegar and rose oil. Anoint the patient’s forehead before sleep."
If we reduce the Latin words to their basic dictionary forms, we get something like this:
“ad dolor caput recipere radix helleborus et terere in mortarium miscere cum acetum et oleum rosa unguere frons patiens ante somnus”
We can then remove prepositions such as ad, in, cum, and ante, conjunctions such as et, and predictable instructional verbs such as recipere and miscere. What remains is the raw referential payload:
"dolor caput radix helleborus terere mortarium acetum oleum rosa unguere frons patiens somnus"
The result is an awkward, almost “caveman-style” form of Latin. To a modern linguist looking for conventional morphology and syntax, it appears broken. To a trained fifteenth-century apothecary, however, much of the intended meaning would still be recoverable. The practical context constrains the possible relationships among hellebore roots, a mortar, vinegar, rose oil, headache, and the patient’s forehead.
This is important because practical knowledge does part of the interpretive work that grammar would normally perform. The reader does not have to consider every theoretically possible relationship among the words. Existing knowledge of the craft or some other external knowledge narrows the range of plausible interpretations. The system expects that the receiver has sufficient information to handle the context recovery, transmitted beforehand.
The difference between the information carried by ordinary Latin and the apparently smaller capacity of the encoded text might therefore be explained, at least in part, by removing inflectional endings, function words, predictable instructions, and other contextually recoverable material.
Once Latin is reduced in this way, however, the result no longer looks much like Latin. Ambiguities also inevitably arise. For example, simply listing an ingredient and an action does not always tell the reader what is being acted upon, in what order, or under which conditions.
To preserve meaning without ordinary inflections, the encoding system would therefore need an alternative way of organizing the remaining roots or semantic units. One possibility is a rigid positional framework in which an element’s location determines its function. A position might indicate whether a root refers to an ingredient, an action, a quantity, a body part, a preparation method, or some other category.
In effect, position would take over part of the work normally performed by Latin morphology and syntax. The resulting text would not simply be abbreviated Latin. It would be Latin material reorganized within a different information system, one that relies on shared professional knowledge, predictable context, and positional structure to recover what has been omitted.
Any views or comments on this?
Thank you.
-Pointless