(31-08-2026, 04:56 PM)Mauro Wrote: You are not allowed to view links. Register or Login to view.The idea of a lossy encoding has already been proposed in different flavours, mainly because it can lower the characters entropies and create 'words' with a regular structure. I think @quimqu proposed a ciphering method with a many-to-one substitution scheme where part of the entropy was transferred to ancillary 'residuals' which were necessary to fully decode the text (there's a thread somewhere here, sorry for not providing the link, I can't find it now).
For myself, I think a possibility is a (inherently lossy) many-to-many substitution scheme, which might possibly lower the character entropies, create words with regular structure and have few repetitions, given the -to-many substitutions, as observed in the VMS. But the hard part is getting this (or any other) scheme to work.
Thank you, Mauro, for your comments. I can see why a many-to-one mapping could be problematic, although I find Quimqu’s approach interesting and suggestive.
I am not yet sure how best to classify my own approach, so here is a brief description of what I am trying to work out.
In previous posts, I suggested that reducing Latin to its basic, uninflected roots could create a lightweight, lossy payload. However, once grammatical endings such as -am and -ibus, together with function words such as ad and cum, are removed, the relationships between the remaining elements can become ambiguous.
If much of the usual grammar is omitted, how could the text remain readable? One possibility is that rigid positional slots take over part of the role normally performed by syntax. Another is that the receivers, the members of the intended audience, have enough information to recover the text in the correct, unambiguous, form.
1. Slots Instead of Sentences
Rather than using free-flowing prose, the system could behave more like a structured form or ledger. Each entry might be divided into implicit but fixed positions:
Slot 1, category or condition: headache
Slot 2, ingredient or item: hellebore
Slot 3, action or preparation: grind in a mortar
If the reader knows that the first position identifies the condition, the second the ingredient, and the third the action, many grammatical endings, prepositions, and predictable instructional words no longer need to be written. Position carries some of the relationships that grammar would normally express.
The text would therefore be read not as a sequence of ordinary sentences, but as a series of structured entries whose elements are interpreted according to their position.
2. Repetition as Positional Padding
This positional requirement might also offer a possible explanation for immediate token repetition, such as daiin daiin.
Suppose the system requires a fixed three-slot structure, but a particular entry contains only two pieces of information. Leaving one position empty might cause the following element to be assigned to the wrong slot. The scribe could avoid this by repeating a token or inserting a conventional filler.
In that case, the repeated form would not necessarily add new semantic information. It could function as a placeholder, somewhat like a zero in positional notation or a conventional mark in an empty ledger cell.
This is just one possible explanation, of course and maybe not even a coherent one as I am still working on it. It would need to be tested by examining where immediate repetition occurs and whether those positions are consistent with the proposed slot structure.
3. A Frequent Core and a Long Hapax Tail
A slot-based system might also help explain the manuscript’s large number of forms that occur only once.
Some slots could use a snall and frequently repeated set of category or structural markers. These would form the recurring core of the token distribution. Other slots could remain open for highly specific information, such as proper names, uncommon ingredients, local plant varieties, quantities, or locations. These variable slots could produce a long tail of rare or unique forms.
In simplified terms:
Fixed slots draw from a small set of recurring structural or category markers.
Variable slots accommodate specific information that may occur only once or a few times.
Such a system could (?) produce a strongly uneven frequency distribution without requiring every space delimited glyph string to function as an ordinary word.
4. Low Character Entropy
A rigid positional system could also contribute to low character entropy. If particular positions allow only a limited range of glyphs or glyph combinations, the next character becomes more predictable.
Meaning would then be carried not only by the glyph strings themselves, but also by their position, their relationship to neighbouring elements, and perhaps by additional markers. A relatively small and tightly constrained inventory of symbols could therefore participate in a much larger number of meaningful combinations.
Positional structure would not automatically explain low character entropy, but it offers a possible mechanism that could be examined statistically.
A Structured-Data Interpretation
If the manuscript is approached as a kind of a positional ledger rather than as running prose, several of its unusual features may begin to look connected: low character entropy, immediate repetition, weak conventional grammar, positional regularities, and a large number of unique forms.
Under this hypothesis, these features would not necessarily be unrelated anomalies or cryptographic tricks. They might instead reflect a structured notation system in which position replaces part of the grammar, placeholders preserve alignment, and open slots accommodate highly specific information.
This remains a working hypothesis rather than a conclusion. Its value will depend on whether a consistent positional structure can be identified across the manuscript and whether the model produces predictions that can be tested independently.
Hope the description makes sense even if the approach does not, at least not yet.