![]() |
|
What if it's a "Lossy Encoder" for 15th-Century Latin? - Printable Version +- The Voynich Ninja (https://www.voynich.ninja) +-- Forum: Voynich Research (https://www.voynich.ninja/forum-27.html) +--- Forum: Voynich Talk (https://www.voynich.ninja/forum-6.html) +--- Thread: What if it's a "Lossy Encoder" for 15th-Century Latin? (/thread-6041.html) Pages:
1
2
|
RE: What if it's a "Lossy Encoder" for 15th-Century Latin? - JoJo_Jost - 24-08-2026 That, too, would in a sense – at least if it were a simple substitution – require a kind of natural language that would be relatively easy to crack and would not fit the statistical anomalies of the VMS. In fact, this cipher, if indeed it is one, is based on a highly organised structure. Look at the families (the ‘E’ family, ‘aiin’, ‘air’, ‘Familien’, ‘ol’ / ‘or’, bigrams ‘ar’, ‘al’); this means that the VMS is either a more complex cipher that wouldn’t need such an abbreviation, or it is a hoax – though statistically, there is much to suggest it is not a hoax. So the conclusion remains: it is a simple yet more complex cipher, or – though I can now more or less rule this out, though not entirely – Chinese...
RE: What if it's a "Lossy Encoder" for 15th-Century Latin? - oshfdk - 24-08-2026 (24-08-2026, 07:19 PM)Pointless.. Wrote: You are not allowed to view links. Register or Login to view.But I am trying to find a possible explanation for several Voynich anomalies, such as low character entropy, bimodal frequencies, repetition, and immediate reduplication. I am also interested in glyph dependencies, weak word-to-word relationships, positional patterns, and the differences between Currier A and B. I don't understand how it's possible to explain these without actual tests. The anomalies you list are quantitative. It's not just "low character entropy", but a specific range of character entropy, specific binomial type/token length distributions, etc. I think arguing that method X can lower entropy or unskew the length distribution is of little value unless it's possible to model the results and show whether the resulting entropies and distribution are similar to Voynichese. RE: What if it's a "Lossy Encoder" for 15th-Century Latin? - Pointless.. - 25-08-2026 (24-08-2026, 07:29 PM)Rafal Wrote: You are not allowed to view links. Register or Login to view.Some observations: Thank you for these comments, appreciated. You are right, the Voynich manuscript appears to have an unusually hapax-rich vocabulary. Depending on the transcription and tokenization method, I understand that around 70% of its distinct space-delimited forms occur only once. This is quite high compared with ordinary prose of similar length, although the result depends heavily on whether Voynich spaces actually mark word boundaries and whether its glyph strings correspond to linguistic words at all, as my current working hypothesis on this idea is that that the glyph strings are not linguistic. Under a lossy positional model, this anomalous hapax rate would sense. If Voynich spaces do not separate normal spoken words, but instead separate clusters of positional data or semantic parameters, the number of unique combinations will naturally explode. Think of it like combining our earlier stripped Latin roots with specific disambiguation cues or positional padding. If a scribe combines a core root with a highly specific modifier to create a single space-delimited string, that exact string might only ever be needed once in the entire book. The 70% hapax rate isn't a sign of an impossibly large vocabulary; it is the statistical footprint of a modular, combinatorial system where tiny tweaks to a string create thousands of single-use variants. And regarding the mysterious "daiin daiin", I think will test case for this whole writing system idea - it's the single token-pair where the padding hypothesis, the classifier typology, and the autocopy-null all make competing, checkable predictions. How does this yet a bit vague fidea sound? RE: What if it's a "Lossy Encoder" for 15th-Century Latin? - Pointless.. - 25-08-2026 (24-08-2026, 09:50 PM)oshfdk Wrote: You are not allowed to view links. Register or Login to view.(24-08-2026, 07:19 PM)Pointless.. Wrote: You are not allowed to view links. Register or Login to view.But I am trying to find a possible explanation for several Voynich anomalies, such as low character entropy, bimodal frequencies, repetition, and immediate reduplication. I am also interested in glyph dependencies, weak word-to-word relationships, positional patterns, and the differences between Currier A and B. I am still at the thought-experiment stage, exploring different ideas and possibilities. If the discussion leads to concrete and testable hypotheses, they can and should be tested. I hope that collaboration and the exchange of different perspectives will help move the idea forward. RE: What if it's a "Lossy Encoder" for 15th-Century Latin? - Rafal - 25-08-2026 Quote:my current working hypothesis on this idea is that that the glyph strings are not linguistic Not linguistic or not phonetic? If Voynichese "words" mean things like leaf, flower, star or water then it's still a language. You may not read them loud as folium, flos, stella and aqua. You may not read them loud at all. But they remain words from some language. Even if they mean things like "with water" or "from leaves" then they are still language. Such composite words are typical for agglutinative languages. Of course it could be not existing language but constructed language. If I understand correctly that's what you are suggesting. That's theoretically possible but I see several problems with it. If you develop your ideas further we could discuss it more. RE: What if it's a "Lossy Encoder" for 15th-Century Latin? - Pointless.. - 27-08-2026 (25-08-2026, 10:11 AM)Rafal Wrote: You are not allowed to view links. Register or Login to view.Quote:my current working hypothesis on this idea is that that the glyph strings are not linguistic Yes, of course, you are right that there is content in some linguistic form. That is, after all, the raison d'être of a book or manuscript. What I am conjecturing is that the script may be mixed with some kind of markers or classifiers, distinct from ordinary diacritics, together with other additions that help guide the interpretation of the underlying linguistic content. So I wouldn't say it i a constructed language, I still think the actual content is in some European language. Unfortunately, teaching commitments leave me with too little time at the moment to explore these ideas as deeply as I would like. I hope that will improve soon. Below is a rough list of topics and possibilities that I have been thinking about. Please do not take them too literally or as fixed claims. Rather, consider them as meaningful tags pointing toward areas or concepts that may be worth exploring.
If anyone is aware of research, discussions, or even speculative work that touches on any of these topics, I would be very grateful for pointers. As always, constructive comments, alternative perspectives, and critical observations are most welcome. RE: What if it's a "Lossy Encoder" for 15th-Century Latin? - Mauro - 31-08-2026 The idea of a lossy encoding has already been proposed in different flavours, mainly because it can lower the characters entropies and create 'words' with a regular structure. I think @quimqu proposed a ciphering method with a many-to-one substitution scheme where part of the entropy was transferred to ancillary 'residuals' which were necessary to fully decode the text (there's a thread somewhere here, sorry for not providing the link, I can't find it now). For myself, I think a possibility is a (inherently lossy) many-to-many substitution scheme, which might possibly lower the character entropies, create words with regular structure and have few repetitions, given the -to-many substitutions, as observed in the VMS. But the hard part is getting this (or any other) scheme to work. RE: What if it's a "Lossy Encoder" for 15th-Century Latin? - Pointless.. - 01-09-2026 (31-08-2026, 04:56 PM)Mauro Wrote: You are not allowed to view links. Register or Login to view.The idea of a lossy encoding has already been proposed in different flavours, mainly because it can lower the characters entropies and create 'words' with a regular structure. I think @quimqu proposed a ciphering method with a many-to-one substitution scheme where part of the entropy was transferred to ancillary 'residuals' which were necessary to fully decode the text (there's a thread somewhere here, sorry for not providing the link, I can't find it now). Thank you, Mauro, for your comments. I can see why a many-to-one mapping could be problematic, although I find Quimqu’s approach interesting and suggestive. I am not yet sure how best to classify my own approach, so here is a brief description of what I am trying to work out. In previous posts, I suggested that reducing Latin to its basic, uninflected roots could create a lightweight, lossy payload. However, once grammatical endings such as -am and -ibus, together with function words such as ad and cum, are removed, the relationships between the remaining elements can become ambiguous. If much of the usual grammar is omitted, how could the text remain readable? One possibility is that rigid positional slots take over part of the role normally performed by syntax. Another is that the receivers, the members of the intended audience, have enough information to recover the text in the correct, unambiguous, form. 1. Slots Instead of Sentences Rather than using free-flowing prose, the system could behave more like a structured form or ledger. Each entry might be divided into implicit but fixed positions: Slot 1, category or condition: headache Slot 2, ingredient or item: hellebore Slot 3, action or preparation: grind in a mortar If the reader knows that the first position identifies the condition, the second the ingredient, and the third the action, many grammatical endings, prepositions, and predictable instructional words no longer need to be written. Position carries some of the relationships that grammar would normally express. The text would therefore be read not as a sequence of ordinary sentences, but as a series of structured entries whose elements are interpreted according to their position. 2. Repetition as Positional Padding This positional requirement might also offer a possible explanation for immediate token repetition, such as daiin daiin. Suppose the system requires a fixed three-slot structure, but a particular entry contains only two pieces of information. Leaving one position empty might cause the following element to be assigned to the wrong slot. The scribe could avoid this by repeating a token or inserting a conventional filler. In that case, the repeated form would not necessarily add new semantic information. It could function as a placeholder, somewhat like a zero in positional notation or a conventional mark in an empty ledger cell. This is just one possible explanation, of course and maybe not even a coherent one as I am still working on it. It would need to be tested by examining where immediate repetition occurs and whether those positions are consistent with the proposed slot structure. 3. A Frequent Core and a Long Hapax Tail A slot-based system might also help explain the manuscript’s large number of forms that occur only once. Some slots could use a snall and frequently repeated set of category or structural markers. These would form the recurring core of the token distribution. Other slots could remain open for highly specific information, such as proper names, uncommon ingredients, local plant varieties, quantities, or locations. These variable slots could produce a long tail of rare or unique forms. In simplified terms: Fixed slots draw from a small set of recurring structural or category markers. Variable slots accommodate specific information that may occur only once or a few times. Such a system could (?) produce a strongly uneven frequency distribution without requiring every space delimited glyph string to function as an ordinary word. 4. Low Character Entropy A rigid positional system could also contribute to low character entropy. If particular positions allow only a limited range of glyphs or glyph combinations, the next character becomes more predictable. Meaning would then be carried not only by the glyph strings themselves, but also by their position, their relationship to neighbouring elements, and perhaps by additional markers. A relatively small and tightly constrained inventory of symbols could therefore participate in a much larger number of meaningful combinations. Positional structure would not automatically explain low character entropy, but it offers a possible mechanism that could be examined statistically. A Structured-Data Interpretation If the manuscript is approached as a kind of a positional ledger rather than as running prose, several of its unusual features may begin to look connected: low character entropy, immediate repetition, weak conventional grammar, positional regularities, and a large number of unique forms. Under this hypothesis, these features would not necessarily be unrelated anomalies or cryptographic tricks. They might instead reflect a structured notation system in which position replaces part of the grammar, placeholders preserve alignment, and open slots accommodate highly specific information. This remains a working hypothesis rather than a conclusion. Its value will depend on whether a consistent positional structure can be identified across the manuscript and whether the model produces predictions that can be tested independently. Hope the description makes sense even if the approach does not, at least not yet. RE: What if it's a "Lossy Encoder" for 15th-Century Latin? - ololololo - 01-09-2026 If I’m reasoning correctly, you end up with something akin to a book cipher, where the connections between words are indicated by morphemes. The only question is whether you’ll be able to gather a sufficient set of morphemes to create a readable text… |