![]() |
|
What if it's a "Lossy Encoder" for 15th-Century Latin? - Printable Version +- The Voynich Ninja (https://www.voynich.ninja) +-- Forum: Voynich Research (https://www.voynich.ninja/forum-27.html) +--- Forum: Voynich Talk (https://www.voynich.ninja/forum-6.html) +--- Thread: What if it's a "Lossy Encoder" for 15th-Century Latin? (/thread-6041.html) Pages:
1
2
|
What if it's a "Lossy Encoder" for 15th-Century Latin? - Pointless.. - 24-08-2026 I have been considering the possibility that the Voynich manuscript uses some form of lossy encoding. By “lossy,” I mean that the source text may originally have been written in Latin, but that only part of its linguistic information was preserved in writing. Such a system would assume that the intended reader already possessed enough practical knowledge, contextual information, or interpretive tools to reconstruct the meaning. Where the result might otherwise be ambiguous, the system could provide additional cues to guide the reader toward the correct interpretation. Shannon’s information theory suggests that natural languages contain a considerable amount of redundancy, sometimes estimated at around 50 percent. Latin, as a highly inflected language, carries much of its grammatical information in word endings. Nouns, adjectives, and verbs are marked according to their roles in the sentence, meaning that some structural information is communicated more than once. In a practical text, however, much of this information may not be strictly necessary. If a practitioner already knows how to prepare a decoction, grammatical case endings and many function words may add relatively little to the basic operational meaning. To see how this might work in practice, consider a short apothecary-style recipe in Latin: "Ad dolorem capitis: Recipe radices hellebori, et tere in mortario. Misce cum aceto et oleo rosarum. Ungue frontem patientis ante somnum." In ordinary English, the passage means roughly: "For headache: Take hellebore roots and grind them in a mortar. Mix with vinegar and rose oil. Anoint the patient’s forehead before sleep." If we reduce the Latin words to their basic dictionary forms, we get something like this: “ad dolor caput recipere radix helleborus et terere in mortarium miscere cum acetum et oleum rosa unguere frons patiens ante somnus” We can then remove prepositions such as ad, in, cum, and ante, conjunctions such as et, and predictable instructional verbs such as recipere and miscere. What remains is the raw referential payload: "dolor caput radix helleborus terere mortarium acetum oleum rosa unguere frons patiens somnus" The result is an awkward, almost “caveman-style” form of Latin. To a modern linguist looking for conventional morphology and syntax, it appears broken. To a trained fifteenth-century apothecary, however, much of the intended meaning would still be recoverable. The practical context constrains the possible relationships among hellebore roots, a mortar, vinegar, rose oil, headache, and the patient’s forehead. This is important because practical knowledge does part of the interpretive work that grammar would normally perform. The reader does not have to consider every theoretically possible relationship among the words. Existing knowledge of the craft or some other external knowledge narrows the range of plausible interpretations. The system expects that the receiver has sufficient information to handle the context recovery, transmitted beforehand. The difference between the information carried by ordinary Latin and the apparently smaller capacity of the encoded text might therefore be explained, at least in part, by removing inflectional endings, function words, predictable instructions, and other contextually recoverable material. Once Latin is reduced in this way, however, the result no longer looks much like Latin. Ambiguities also inevitably arise. For example, simply listing an ingredient and an action does not always tell the reader what is being acted upon, in what order, or under which conditions. To preserve meaning without ordinary inflections, the encoding system would therefore need an alternative way of organizing the remaining roots or semantic units. One possibility is a rigid positional framework in which an element’s location determines its function. A position might indicate whether a root refers to an ingredient, an action, a quantity, a body part, a preparation method, or some other category. In effect, position would take over part of the work normally performed by Latin morphology and syntax. The resulting text would not simply be abbreviated Latin. It would be Latin material reorganized within a different information system, one that relies on shared professional knowledge, predictable context, and positional structure to recover what has been omitted. Any views or comments on this? Thank you. -Pointless RE: What if it's a "Lossy Encoder" for 15th-Century Latin? - oshfdk - 24-08-2026 This falls of the spectrum of many theories starting with abbreviated Latin and ending with purely functional/algorithmic recipes. I can't say why this one won't work, because there is little information in your post. If you provide more details about how specifically this should work, I can try to explain why specifically this won't work. RE: What if it's a "Lossy Encoder" for 15th-Century Latin? - ololololo - 24-08-2026 There is no point in this procedure. With such text, a conventional decipherer won’t be able to understand it properly. But in reality, there is a possibility that the encryption algorithm may distort the text, and the decipherer will end up with something different from what they would have wanted to see. In the case of a book that has remained undeciphered for about 600 years, this is indeed possible... RE: What if it's a "Lossy Encoder" for 15th-Century Latin? - Pointless.. - 24-08-2026 Thank oshfdk you for short comment. I will post more details soon, step by step. It is important to distinguish this idea from ordinary medieval Latin abbreviation. Traditional abbreviation makes the writing shorter; a lossy system would make the message itself shorter. Medieval scribes left out letters, syllables, endings, and sometimes whole words, but the Latin sentence was generally still there underneath. A reader who knew the conventions could expand the abbreviations and recover something close to the original wording and grammar. The system I am proposing would go further. It would leave out grammatical information that an informed reader did not need, preserving only enough of the semantic content to reconstruct the intended meaning. The exact original Latin sentence might no longer be recoverable, but the practical instruction would be. I am not claiming that the source language was definitely Latin. Latin simply fits this model particularly well because it carries so much grammatical information in its endings. If the reader could recover that information from context, practical knowledge, or the position of words within the encoded sequence, much of it would not need to be written down at all. In that sense, this would not be heavily abbreviated Latin. It would be something more radical: Latin material reorganized into a different system, one designed to preserve meaning rather than the exact original wording. RE: What if it's a "Lossy Encoder" for 15th-Century Latin? - Pointless.. - 24-08-2026 (24-08-2026, 04:50 PM)ololololo Wrote: You are not allowed to view links. Register or Login to view.There is no point in this procedure. Think of it this way. I have a writing system that does not record every detail of the original text. Instead, it preserves the essential information and adds markers or hints where the meaning might otherwise be unclear. The intended reader has instructions explaining how the system works, how the markers should be interpreted, and how the missing information can be reconstructed from context. The result may not reproduce the original wording exactly, but it should allow the reader to recover the intended meaning. Anyone who does not know the rules would see only an unfamiliar and ambiguous text. They might recognize patterns, but they would not know what had been omitted, how the remaining elements were organized, or how the markers resolved ambiguities. The decoding procedure would therefore function rather like a key, making the text difficult for an outsider to interpret. This could also provide a degree of protection against unintended readers. The secrecy would not lie only in the symbols themselves, but in the combination of the writing system, the reconstruction rules, and the specialized knowledge needed to apply them. RE: What if it's a "Lossy Encoder" for 15th-Century Latin? - oshfdk - 24-08-2026 (24-08-2026, 05:05 PM)Pointless.. Wrote: You are not allowed to view links. Register or Login to view.In that sense, this would not be heavily abbreviated Latin. It would be something more radical: Latin material reorganized into a different system, one designed to preserve meaning rather than the exact original wording. There is a very simple test, you can take a small text of 200-300 words and convert it to Voynichese using your proposed method. If the result looks anything like Voynichese, this will be remarkable. Otherwise, to me your method doesn't look particularly different from many proposed solutions of the past. RE: What if it's a "Lossy Encoder" for 15th-Century Latin? - ololololo - 24-08-2026 (24-08-2026, 05:15 PM)Pointless.. Wrote: You are not allowed to view links. Register or Login to view.Well, then it’s possible... But even such a system would be much easier to implement. You can simply use substitution so as not to get bogged down. In any case, even if the system is transparent and understandable to a legitimate decrypter, it would be foolish to rely on their ingenuity...(24-08-2026, 04:50 PM)ololololo Wrote: You are not allowed to view links. Register or Login to view.There is no point in this procedure. Perhaps your idea will become a starting point in the field of research into Friedman’s artificial language. It fits perfectly into the logic of philosophical language. For example, for each such word, you can add prefixes such as d-, ch-, or qo-, which will make it easier to identify missing words. RE: What if it's a "Lossy Encoder" for 15th-Century Latin? - Pointless.. - 24-08-2026 Dear oshdfk, I am not proposing any solution nor am I seeking one. But I am trying to find a possible explanation for several Voynich anomalies, such as low character entropy, bimodal frequencies, repetition, and immediate reduplication. I am also interested in glyph dependencies, weak word-to-word relationships, positional patterns, and the differences between Currier A and B. I have some ideas about how these features might fit together, but I would welcome help developing and polishing them. I am comfortable with the statistical side, but I am not an expert in linguistics or cryptography. My basic premise is that the glyph strings we see in the manuscript are not a language as such. This would mean that there may be no directly recoverable grammar or identifiable language at the visible-text level. RE: What if it's a "Lossy Encoder" for 15th-Century Latin? - Rafal - 24-08-2026 Some observations: You don't say anything about encoding (the used cipher which hypothetically turns let's say "radix" into "chedy"), just about the structure of the source text. Do you have any ideas about the cipher? Do you think that if the text lacked basic words like "and", "so", "with" etc. but was encrypted in some simple way, then it would be harder to crack? Personally I don't think so. The structure of Voynichese isn't also compatible with what you suggest. It has a high percentage of hapax legomena (unique) words which is similar to Latin with declension and much higher than let's say English. It has very common words like "daiin" which would be good candidates for words like "and" (but it doesn't work). If there is no "and" then what "daiin" could be? RE: What if it's a "Lossy Encoder" for 15th-Century Latin? - Pointless.. - 24-08-2026 (24-08-2026, 07:02 PM)ololololo Wrote: You are not allowed to view links. Register or Login to view.Well, then it’s possible... But even such a system would be much easier to implement. You can simply use substitution so as not to get bogged down. In any case, even if the system is transparent and understandable to a legitimate decrypter, it would be foolish to rely on their ingenuity... Thank you for your comment. This is exactly what I am looking for here: an exchange of ideas, discussion of different possibilities, and perspectives I may not have considered. My current thinking is that the manuscript is not language as such. It may contain information derived from a language but encoded in a way that removes or obscures many of its linguistic features. Of course, I may be wrong about some of the premises or methods, but I hope the idea can at least lead to a fruitful discussion and perhaps open up new lines of research. |