![]() |
|
A mathematical approach to double words, conditional logic and the missing pages - Printable Version +- The Voynich Ninja (https://www.voynich.ninja) +-- Forum: Voynich Research (https://www.voynich.ninja/forum-27.html) +--- Forum: Theories & Solutions (https://www.voynich.ninja/forum-58.html) +---- Forum: The Slop Bucket (https://www.voynich.ninja/forum-59.html) +---- Thread: A mathematical approach to double words, conditional logic and the missing pages (/thread-5861.html) |
RE: A mathematical approach to double words, conditional logic and the missing pages - dashstofsk - 02-07-2026 (02-07-2026, 03:17 PM)Vuk88 Wrote: You are not allowed to view links. Register or Login to view.the existence of a syntactic suppression rule. The expected frequency of word repeats is going to be biased by the affinities that word prefixes have for the suffices of previous words. For instance in quire 20 the number of occurrences where a word starting q follows a word ending y is nearly twice what would be expected if the words were distributed randomly. This has the effect of raising the expectation of q__y word repeats. Also, words starting a have a liking for words ending r ( ~3.3 times the expected number in quire 20 ) and this raises the count of ar repeats. In quire 20 the number of occurrences of a word starting a and following a word ending iin is about 2/3 of what would be expected. This then lowers the expectation of aiin repeats. To try to justify word repeats you might have to attempt also to give some justification for suffix-prefix affinities such as these. RE: A mathematical approach to double words, conditional logic and the missing pages - Vuk88 - 02-07-2026 (01-07-2026, 08:50 PM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.(30-06-2026, 03:46 PM)Vuk88 Wrote: You are not allowed to view links. Register or Login to view.I specifically focused on one of the most famous statistical anomalies: the "double repetitions" (like chol chol). ... These double words aren't isolated or random; they are strictly bound to specific associated words. Dear Professor Stolfi, First of all, I want to thank you for your work, because if today I can run these algorithms (which were unthinkable until a few years ago) it is thanks to the foresight of your work. Without that fundamental dataset, none of this would be possible. Your observation is exactly the thought that constantly grips me during testing. To be honest, when I saw the precision of these spikes for the first time, my first thought was: "there must be an artifact or a bug somewhere". I tried to make the data as clean as possible before feeding it to the algorithm. I took the standard Takahashi file and aggressively stripped it down... I removed the transcription notes you inserted, the uncertain markers like exclamation marks or dashes, and converted dots and commas to spaces. The script isolates (or at least I hope it does) only the pure EVA roots, and the doublet count is a very basic string comparison between adjacent words. Since you are literally the architect of this interlinear archive and know its formatting secrets better than anyone else, please allow me to ask if you believe there might be edge cases in the transcription markers that my cleaning does not intercept? On an empirical level, I performed a visual cross-check on the high-resolution facsimile (examining several samples, for example f108r) and the doublets counted by my parser find an exact physical correspondence in the ink written on the parchment... but I am absolutely open to the idea of having overlooked some structural anomaly in the source data of which I am unaware. Therefore, if you have any methodological suggestions or tests to propose to further falsify this result, I would be honored to listen. P.S. To directly answer your previous request to show the actual text around those boundaries, I have attached a visual document that cross-references the graph spikes with the raw lines extracted directly from your LSI interlinear archive. For each spike, the doublet identified by my script is placed alongside the original transcription line, so you can verify the correspondence at a glance. Respectfully, Alfredo RE: A mathematical approach to double words, conditional logic and the missing pages - asteckley - 02-07-2026 (02-07-2026, 05:33 PM)Vuk88 Wrote: You are not allowed to view links. Register or Login to view.I read your comments regarding the physical permutation test. In reality, setting up the analysis with those rules would result in a structural error... applying a physical shuffle to a Bag-of-Words model triggers a macroscopic homogenization bias, and it is obvious to expect results that would artificially favor the algorithm even more.I'm not really sure what rules you are referring to since I didn't really present any rules (and I don't see that Mauro did either) nor did I suggest to do anything like this latest processing that you described. In any case, none of this can be commented on when you haven't defined what you are using for "accuracy". And the whole approach for your randomized data testing (and the use of Y-Scrambling for the current purposes) remains -- to me-- questionable. I’ll just recap my assessment of all this so far:
RE: A mathematical approach to double words, conditional logic and the missing pages - Mauro - 02-07-2026 (02-07-2026, 05:33 PM)Vuk88 Wrote: You are not allowed to view links. Register or Login to view.I read your comments regarding the physical permutation test... Thank you. Well, I think I can tentatively accept that a set of words tend to co-occur on the same page as word doublets, while another set of words tends to steer clear from those pages. I'm not fully conviced yet, but enough that I think your observation deserves consideration, so I surely look forward to see how your work proceeds. Regarding what this may mean, once fully confirmed, I'd be cautious to say for now. Does 'aiin' inhibits the appearance of doublets, or do doublets inhibit the appearance of 'aiin'? Or something entirely different from triggers/inhibitors causes this behaviour? I hope we'll come to know! RE: A mathematical approach to double words, conditional logic and the missing pages - Jorge_Stolfi - 03-07-2026 (02-07-2026, 08:13 PM)Vuk88 Wrote: You are not allowed to view links. Register or Login to view.Thank you for your work, because if today I can run these algorithms (which were unthinkable until a few years ago) it is thanks to the foresight of your work. Without that fundamental dataset, none of this would be possible. Thanks for the compliment, but my contributions to the interlinear transcription file were minor. Most of the merit should go to Gabriel Landini (who recently passed away) and René Zandbergen (who kept maintaining the file after Gabriel and I left the Voynich scene). Quote:do you believe there might be edge cases in the transcription markers that my cleaning does not intercept? If you are using Rene's current version of the transcription file, he has a very detailed description of the format at his website. The only thing that I would add to that spec is my conviction the manuscript itself (not just the transcription file) has many errors or quirks that could hide doublets. In particular, I am convinced that m is an abbreviation, most likely for iin; that the ending -ir is a scribal error for -iin (not -in), and -iir is an error for -iiin; that Ih, ITh, IKh are just malformed versions of Ch, CTh, CKh; and that the position of the plume on Sh (on the first e, on the second e, or midway) is not significant. I would "fix" these errors before any analysis. There must be also many bogus or omitted word spaces, d replaced by k, confusion between r and s, a and y, ain and aiin etc.; but at present I do not know how to detect and correct these other "errors". However, be aware that the possibility of error is disputed by many, and anyway the effect of such corrections on the number of doublets would not be great. Quote:On an empirical level, I performed a visual cross-check on the high-resolution facsimile (examining several samples, for example f108r) and the doublets counted by my parser find an exact physical correspondence in the ink written on the parchment... I went to look for doublets on my version of the transcription file (which has many small differences from Rene's version). I did NOT do the "error fixes" above, except map Ih ITh, IKh to Ch, CTh, CKh. I counted 292 doublets on 66 different pages. Most pages had 3 or fewer doublets. The exceptions were: 6 <f75r> Page 1 of Bio 6 <f75v> Page 2 of Bio 7 <f76r> Page 3 of Bio(?) Text only 6 <f77r> Page 5 of Bio 4 <f78r> Page 7 of Bio 5 <f80v> Page 12 of Bio 4 <f85v2> Big fold-out 7 <f103v> Starred Parags 6 <f104v> Starred Parags 6 <f107r> Starred Parags 5 <f107v> Starred Parags 7 <f108r> Starred Parags 11 <f108v> Starred Parags 7 <f111r> Starred Parags 4 <f112r> Starred Parags 7 <f116r> Starred Parags + Unknown These counts more or less match the plot on your PDF file, but I do not see signs of the peaks at section boundaries that were so prominent in the You are not allowed to view links. Register or Login to view.. But that was density of doublets per page, not count. So perhaps there was something wrong with the denominator (number of words in the page)? From the above counts it seems that Starred and Bio do have a higher number of doublets per page than the other sections. But they also have more words per page. On the other hand they seem rather uniform with respect to this statistic. The higher numbers for f108r, f108v, and You are not allowed to view links. Register or Login to view. can be explained by the fact that, on each of those pages, the Scribe mashed 10-12 parags in a single giant parag, thus boosting the number of words on those pages. The higher number for 116r may have something to do with the fact that the last half of that page is a block of 20 lines of dense text with no stars. I should have counted also the number of words per page and compute the densities of doublets, not just counts. But I don't have the time now, sorry. Doublets are common in monosyllabic languages compared to other languages. These are examples from a Chinese medical book: 太山山谷 tài shān shān gǔ. 瘀血血瘕欲死 yū xiě xiě jiǎ yù sǐ 劳极洒洒 láo jí xiǎn xiǎn 伤寒寒热 shāng hán hán rè 明目目痛 míng mù mù tòng 洗洗酸痛 xiǎn xiǎn suān tòng ... Some of these are just accidents - a compound term that ends in a syllable followed by another compound that start with the same syllable. Like 太山山谷 = "high mountains and mountain valleys". Some are indeed compounds with two identical syllables. Like 洗洗 xiǎn xiǎn = "chills". In non-monosyllabic languages, one could look instead for repeated syllables within or across words. English seems to be rather scarce on those, while they seem to be more common in Romance. If the doublets in the VMS are mostly two-syllable compounds, like xiǎn xiǎn, then you finding would be that, within a page, certain words are positively correlated with those words, while other words are negatively correlated with them. Which would not be unusual in a natiral language. All the best, --stolfi RE: A mathematical approach to double words, conditional logic and the missing pages - Vuk88 - 07-07-2026 @asteckley first of all, please accept my apologies for mixing the reply to you and to @Mauro in my previous post. I want to thank you for your critical analysis, which prompted me to reopen the data and subject the analysis to much stricter stress tests. I discovered that you were absolutely right on several methodological critiques regarding the way I presented the data, even though the underlying phenomenon seems to robustly survive classical statistical testing. To answer your first question about accuracy…I had not clearly defined the target. The algorithm was set up as a binary classification at the page level—that is, predicting whether the page will contain at least one doublet or not, based on the rest of the vocabulary on that page. The accuracy in cross-validation is 60.7%. I completely agree with you on the weakness of the Y-Scrambling test because in a text structured by topics like the VMS, scrambling the labels makes it obvious that the real data will result in an extreme outlier. This demonstrates a non-random correlation (p-value 0.049), but it is certainly not enough to define such a correlation as a true "syntactic rule". I also agree with your observation regarding the Bar Chart and t-SNE. The bar chart shows a continuous probability gradient, and describing them as "two separate subsets" was an oversimplification in my interpretation. Similarly, the t-SNE map, as you noted, groups pages based on overall vocabulary similarity, but it does not prove any direct causal link with doublets…so…in future revisions of the paper, the use of these tools will be scaled back to a purely exploratory role. However, I welcomed with great interest your methodological suggestion: to verify whether the triggers exist using exclusively classical statistics. I therefore analyzed the entire vocabulary using Fisher's Exact Test. To avoid the multiple comparisons problem and the appearance of false positives in the vast VMS vocabulary, I applied a rigorous Bonferroni Correction. By testing the 450 main words of the manuscript (those present in at least 10 pages), the significance threshold (alpha) to consider a word a potential "trigger" dropped from 0.05 to 0.00011. In the absence of a real correlation, the probability of finding false positives below this extreme threshold is almost zero. Instead, the analysis isolated 28 different words that overcome this statistical barrier. The raw data based on Takahashi's transcription (filtering out punctuation) show a rather clear trend: qoteedy: Appears in 38 pages. Of these, 38 contain at least one doublet (100%). In pages without qoteedy, the probability of finding doublets drops to 52.7%. (P-Value < 0.000001). olkeey: Appears in 26 pages. 26 contain doublets (100%) against a baseline of 55.6%. (P-Value 0.000001). (Incidentally, the main triggers present p-values so infinitesimal — e.g. < 0.000001 — that they would survive the Bonferroni Correction even if it were calculated on the entire unfiltered vocabulary of 8,000 words). To rule out that this correlation is a statistical artifact, I subjected these results to five specific "stress tests". Syntactic Proximity: This could have been a random co-occurrence on the same page. By measuring the physical distance, the median distance between the main triggers and the generated doublet is only 24-25 words. The trigger and the doublet are systematically found within the same paragraph, suggesting a strongly local dependency. Dialect Control (Simpson's Paradox): I hypothesized that the triggers were simply vocabulary words of "Dialect B", and that scribe B naturally had a tendency to write more doublets. I therefore isolated the analysis exclusively to Dialect B pages. The result does not change: within Dialect B alone, the absence of qoteedy generates doublets in 51% of cases, while its presence brings doublets to 100%. The correlation occurs independently of the scribe's overall style. Geographical Distribution: You rightly noted that there are anomalous folios with extreme densities of doublets (e.g., f108). I verified the spatial distribution of the triggers to ensure they were not confined there. The word qoteedy is found 84.2% of the time outside the anomalous folios. The rule is not a geographical "glitch" localized to a few pages. Typology of Triggered Doublets: I verified what is being duplicated. The absolute most frequent doublet in the entire manuscript is daiin daiin. However, when qoteedy is present, the daiin daiin doublet is almost totally suppressed. In its place, the trigger forces the duplication of very different words or, frequently, of isolated bigrams like ol ol and ar ar. (Considering the hypothesis that 'ol' and 'ar' may represent counting sequences, this could imply an "operative" function for these words, though I do not want to go too far into interpretations). Independent Validation (Cross-Transcription): To avert the risk that the correlation depended on a human error by Takahashi in transcribing ligatures, I repeated the Fisher Test by downloading the ZL3b transcription by René Zandbergen. Even in this file, the main triggers survive clearly: olkeey and olchedy remain predictive at 100%, while qoteedy settles at 97.5%. Your critique was invaluable in making me recalibrate my terms. Defining an overall accuracy of 61% (from the ML model) as an "iron rule" was methodologically reckless. However, by isolating the data with univariate statistics, I believe it clearly emerges that there is a subset of vocabulary that radically—and syntactically locally—alters the probability of observing specific duplicated structures. I thank you again for pushing me to use more rigorous verification tools. I would be glad to know your opinion on these proximity and distribution analyses, and what other tests I could possibly run next. Best regards, Alfredo RE: A mathematical approach to double words, conditional logic and the missing pages - Jorge_Stolfi - 07-07-2026 (07-07-2026, 01:02 AM)Vuk88 Wrote: You are not allowed to view links. Register or Login to view.Typology of Triggered Doublets: I verified what is being duplicated. The absolute most frequent doublet in the entire manuscript is daiin daiin. However, when qoteedy is present, the daiin daiin doublet is almost totally suppressed. In its place, the trigger forces the duplication of very different words or, frequently, of isolated bigrams like ol ol and ar ar. (Considering the hypothesis that 'ol' and 'ar' may represent counting sequences, this could imply an "operative" function for these words, though I do not want to go too far into interpretations). This is an important point. Consider the possibility that Voynichese is a natural language with some encoding that maps each distinct word type to the same distinct word type (or word type tuple, e.g. by splitting the word into syllables or morphemes). Then a "doublet" is not an interesting category for analysis. Imagine analyzing the text of an English newspaper by looking at words that have trigraph doublets, of the form xABCABCy. You may get some intriguing results -- like "trigraph doublets occur in only 1% of the pages, but they occur in 90% of the pages where the words 'voice' or 'river' occurs" -- but it would be very hard to figure out what is going on. Because you would be studying the distributions of "Mississippi" and "chiuhaua" and "singing" and "alfalfa" etc. added together. Thus I think you should
For the Bio section, on the other hand, a more natural unit may be a group of a fixed number k of consecutive tokens. Or even not use presence/absence in a unit but instead measure word-word correlation as a function of distance, measured in tokens. All the best, --stolfi RE: A mathematical approach to double words, conditional logic and the missing pages - dashstofsk - 07-07-2026 (07-07-2026, 01:02 AM)Vuk88 Wrote: You are not allowed to view links. Register or Login to view.The absolute most frequent doublet in the entire manuscript is daiin daiin. No. It is chol. Moreover the count of daiin daiin in the language B pages is roughly what would be expected if words were distributed randomly. daiin occurs 479 time in language A pages and 313 in language B pages. There are 11461 word pairs in A and 24056 word pairs in B. Expected doubles in A = 479 x 478 / 11461 = 20. In B, 313 x 312 / 24056 = 4. Actual numbers are 11 and 3. I cannot see that this is taking you anywhere. These numbers are too small to come to any statistical conclusion. You also need to take into account that many of these repeat words are more common in language A or language B. And that language B pages have more words and that they have longer sentences. Also that the main language B sections, quires 13 and 20, have different frequency of repeats and seem to have a slightly different corpus of words. The text in the sections of the manuscript is not uniform and seems to defy generalisation. RE: A mathematical approach to double words, conditional logic and the missing pages - Aga Tentakulus - 07-07-2026 Under normal circumstances, a book has only one author. But here it is different. When one person writes in the singular and another in the plural, things change even if the meaning is essentially the same. “It is like this”, “They are like this”. "es ist so" "sie sind so". RE: A mathematical approach to double words, conditional logic and the missing pages - nablator - 07-07-2026 (07-07-2026, 07:57 AM)dashstofsk Wrote: You are not allowed to view links. Register or Login to view.No. It is chol. Is it in V101? Similar counts in RF1b-er (words extracted by ivtt -x7): 22 chol 20 qokeedy 14 qokedy 13 ar 12 daiin 11 chedy 11 ol 11 qokeey ... In VT0e-n: 22 chol 20 daiin 19 qokeedy 14 qokedy 12 qokeey 11 ar 11 chedy 9 ol 8 dy 8 shedy If doublets across line breaks are not counted: RF1b-er: 20 chol 16 qokeedy 13 qokedy 12 ar 11 ol 9 chedy 9 qokeey 8 daiin 8 shedy 7 or ... VT0e-n: 20 chol 15 qokeedy 14 daiin 14 qokedy 10 ar 10 qokeey 9 ol 8 chedy 8 shedy 6 dy ... I'd like to know which transliteration has more daiin than chol. |