kckluge > 21-05-2026, 07:56 AM
(20-05-2026, 11:26 PM)nablator Wrote: You are not allowed to view links. Register or Login to view.(20-05-2026, 06:22 PM)kckluge Wrote: You are not allowed to view links. Register or Login to view.All of which is getting into the weeds. The point is that *if* Voynichese were a cipher that breaks up words into smaller chunks, the process that breaks them up is unlikely to be syllabification (and likely isn't deterministic in general?) due to the extremely low TTR that results.
The TTR of Voynichese can be reduced as much as you need... by simplifications, equivalences, re-spacing. However the babble-like sequences of similar words would not produced plausible Latin (or Chinese).
Jorge_Stolfi > 21-05-2026, 10:02 AM
(20-05-2026, 03:43 PM)nablator Wrote: You are not allowed to view links. Register or Login to view.I estimate it to 800-1000 in long texts (without many exotic words), more than double what Mandarin has (~400).
Jorge_Stolfi > 21-05-2026, 10:33 AM
(20-05-2026, 03:42 PM)Stefan Wirtz_2 Wrote: You are not allowed to view links. Register or Login to view.(19-05-2026, 08:25 PM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.[..]There is no proof for some of them being just „slight variants“ or even an „abbreviation“ for daiin or anything else.
The Voynichese daiin (and sometimes slight variants like dain, kaiin and laiin, and the abbreviation dam)
ReneZ > 21-05-2026, 12:53 PM
(21-05-2026, 10:33 AM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.But I think I do have good proof that dair, dain, kaiin and laiin can be variants of daiin. It seems you stll don't accept that evidence, but I hope it will eventually be irrefutable; I am working on that.
And I have also an explanation for those particular variations: the glyphs k, d and l (and, separately, r and in) would look very similar if written in a "cursive" handwriting, such as the Author may have used in the draft.
Jorge_Stolfi > 21-05-2026, 02:28 PM
(21-05-2026, 07:56 AM)kckluge Wrote: You are not allowed to view links. Register or Login to view.the distribution of the number of words between instances of words with a length-normalized edit distance below some threshold [...] is an incredibly significant statistical signature of the Voynich text that any theory about how the text is generated needs to reproduce.
Quote:It's for Stolfi to address that issue -- quantitatively -- in the context of explaining/defending his theory. [...] The "explanation" would be to show that the same anomalies that you see in the SPS are spelling system, and incidence of errors.
Quote:...I'm acutely aware that this is the "Chinese" theory thread (and not, for instance, the (AFAIK non-existant) "verbose cipher" theory thread), bur moderators and readers please bear with me because what follows will, in fact, wrap back around to being on topic...
Quote:3) results in a critique of Stolfi's approach here
Quote:All of which leads to wrapping this back around to discussing Stolfi's theory [...]. It's not *impossible* that Stolfi has stumbled into a solution with his approach, but [...] I think it is far more likely that he has fallen prey to the siren call of the "crib" that has lured so many mariners sailing on the treacherous seas of the Voynich Mss to their doom.
Quote: What he *should* be doing -- again, IMHO -- is taking texts in one or more SE Asian languages (dealer's choice), assigning some scheme for representing them with Voynich glyphs, simulating whatever "confused ignorant scribe" error processes he thinks are there, and then showing that you wind up with something that -- quantitatively, and for all the "greatest hits" properties (including "babble-like sequences of similar words") -- looks like Voynichese.
eggyk > 21-05-2026, 03:40 PM
nablator > 21-05-2026, 06:53 PM
(21-05-2026, 02:28 PM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.Well, you have that text above to play with. Do you see "babble-like sequences" in it?
nablator > 21-05-2026, 08:04 PM
(21-05-2026, 06:53 PM)nablator Wrote: You are not allowed to view links. Register or Login to view.How to build a babble detector? Not sure. Some compression algorithm run on all (short) substrings?
kckluge > 22-05-2026, 03:54 AM
(21-05-2026, 02:28 PM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.(21-05-2026, 07:56 AM)kckluge Wrote: You are not allowed to view links. Register or Login to view.the distribution of the number of words between instances of words with a length-normalized edit distance below some threshold [...] is an incredibly significant statistical signature of the Voynich text that any theory about how the text is generated needs to reproduce.
Can you be more specific?
Quote:Quote:It's for Stolfi to address that issue -- quantitatively -- in the context of explaining/defending his theory. [...] The "explanation" would be to show that the same anomalies that you see in the SPS are spelling system, and incidence of errors.
Can you see those anomalies in You are not allowed to view links. Register or Login to view.?
Quote:Quote:...I'm acutely aware that this is the "Chinese" theory thread (and not, for instance, the (AFAIK non-existant) "verbose cipher" theory thread), bur moderators and readers please bear with me because what follows will, in fact, wrap back around to being on topic...
I don't mind. Besides, any argument you present for why the "Chinese" Origin theory (COT) is wrong will be in-topic.
Quote:Quote:3) results in a critique of Stolfi's approach here
First, note that the type-to-token ratio (TTR) of a text in any language usually depends on the text size. You are not allowed to view links. Register or Login to view. (which may be just a consequence of Zipf's law) says that the size L of a a text's lexicon (number of "word types") is related to size N of the text (number of tokens) like L ≈ K sqrt(N), where K may depend on the language, style, topic, etc.
So, when comparing TTRs of different texts, it is absolutely important to trim them to the same number of tokens. Or instead compute the apparent constant K = L/sqrt(N) instead of the ratio L/N.
Quote:Second, spelling and transcription errors can inflate L while having little effect on N. Mandarin uses only 1300 syllables (~18%) out You are not allowed to view links. Register or Login to view. by the known phonetic constraints of the language. That means that changing one phoneme in a random syllable will, with high probability, increase the lexicon size L -- even if the change respects the phonetic constraints.
Thus, when comparing lexicon sizes, it may be prudent to exclude word types that occur only once or twice. Or maybe even more, depending on the rate of errors.
Quote:Third, one must not forget that statistics -- word and character frequencies, correlations, Zipf's and Heap's law, etc -- are a property of a text, or a specific collection of texts -- not of a language. There is no such thing as "the frequency of 'e' in English", "the most common word in Latin", "the Heap K of Mandarin". Even if the dialect and spelling are fixed, those statistics are greatly affected by topic, style, nature of the text, cost of vellum, etc. One can have a meaningful and grammatically correct text in English with 50'000 tokens and only 100 word types -- that does not use the word "the" even once.
Quote:Quote:All of which leads to wrapping this back around to discussing Stolfi's theory [...]. It's not *impossible* that Stolfi has stumbled into a solution with his approach, but [...] I think it is far more likely that he has fallen prey to the siren call of the "crib" that has lured so many mariners sailing on the treacherous seas of the Voynich Mss to their doom.
Well, what can I say? I think that the matches I have found are quantitatively categorical.
Quote:Note that my solution is very different from most other proposals. I don't know what is the language and the encoding, but I claim that I have identified a specific plaintext (the SBJ) whose structure matches the SPS much better than could be expected by chance. Specifically, I have a small dictionary (the "cribs") such that those words occur in the SBJ with spacings that are very closely proportional to the spacings of the corresponding words in the SPS.
Quote:Quote: What he *should* be doing -- again, IMHO -- is taking texts in one or more SE Asian languages (dealer's choice), assigning some scheme for representing them with Voynich glyphs, simulating whatever "confused ignorant scribe" error processes he thinks are there, and then showing that you wind up with something that -- quantitatively, and for all the "greatest hits" properties (including "babble-like sequences of similar words") -- looks like Voynichese.
Well, you have that text above to play with. Do you see "babble-like sequences" in it?
Quote:All the best, --stolfi
Jorge_Stolfi > 22-05-2026, 04:22 AM
(20-05-2026, 01:35 AM)JoeyB Wrote: You are not allowed to view links. Register or Login to view.FIRST: which SBJ or ZHB version is the right comparator. And SECOND, whether the source text could have existed in the 'right form' around 1400. And I definitely don't have anything useful to add there.
Quote:BUT, the other issue seems to be more testable: if this is a positional-distance hypothesis, can the method we use to match also recover the rooster/f105v32-38 pair when we run the whole thing blind? Which files shouldI grab for that? EG, compare all SPS paragraphs against all SBJ/ZHB entries without preselecting the rooster pair,
Quote:I started looking at and playing with files in the ic.unicamp....Notes/077 folder but before I go too far down the road, if the method is bad I don't want to keep going, and if there are files that are final versions or authoritative I'd want to use those.