The Voynich Ninja

Full Version: A mathematical approach to double words, conditional logic and the missing pages
You're currently viewing a stripped down version of our content. View the full version with proper formatting.
Pages: 1 2 3 4 5 6
(02-07-2026, 03:17 PM)Vuk88 Wrote: You are not allowed to view links. Register or Login to view.the existence of a syntactic suppression rule.

The expected frequency of word repeats is going to be biased by the affinities that word prefixes have for the suffices of previous words.

For instance in quire 20 the number of occurrences where a word starting  q follows a word ending  y is nearly twice what would be expected if the words were distributed randomly.  This has the effect of raising the expectation of  q__y  word repeats.  Also, words starting  a have a liking for words ending r ( ~3.3 times the expected number in quire 20 ) and this raises the count of  ar repeats.  In quire 20 the number of occurrences of a word starting  a and following a word ending  iin is about 2/3 of what would be expected.  This then lowers the expectation of  aiin repeats.

To try to justify word repeats you might have to attempt also to give some justification for suffix-prefix affinities such as these.
(01-07-2026, 08:50 PM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.
(30-06-2026, 03:46 PM)Vuk88 Wrote: You are not allowed to view links. Register or Login to view.I specifically focused on one of the most famous statistical anomalies: the "double repetitions" (like chol chol). ... These double words aren't isolated or random; they are strictly bound to specific associated words.

That is a very interesting observation! (Even though it is not evidence that the text is not plain language.)  

The legend on this figure says that the frequency of doublets spikes at transtions between sections:

That would be another interesting observation... However, the precise coincidences, the absence of spikes elsewhere,  and the uniform size of the spikes make me suspect that they are artifacts of the way the data is provided or processed.  Are you sure that this is not the case?  

Can you show the actual text around those boundaries where one can see the doublets?

All the best, --stolfi

Dear Professor Stolfi,
First of all, I want to thank you for your work, because if today I can run these algorithms (which were unthinkable until a few years ago) it is thanks to the foresight of your work. Without that fundamental dataset, none of this would be possible.
Your observation is exactly the thought that constantly grips me during testing. To be honest, when I saw the precision of these spikes for the first time, my first thought was: "there must be an artifact or a bug somewhere".
I tried to make the data as clean as possible before feeding it to the algorithm. I took the standard Takahashi file and aggressively stripped it down... I removed the transcription notes you inserted, the uncertain markers like exclamation marks or dashes, and converted dots and commas to spaces. The script isolates (or at least I hope it does) only the pure EVA roots, and the doublet count is a very basic string comparison between adjacent words.
Since you are literally the architect of this interlinear archive and know its formatting secrets better than anyone else, please allow me to ask if you believe there might be edge cases in the transcription markers that my cleaning does not intercept?
On an empirical level, I performed a visual cross-check on the high-resolution facsimile (examining several samples, for example f108r) and the doublets counted by my parser find an exact physical correspondence in the ink written on the parchment... but I am absolutely open to the idea of having overlooked some structural anomaly in the source data of which I am unaware. Therefore, if you have any methodological suggestions or tests to propose to further falsify this result, I would be honored to listen.

P.S. To directly answer your previous request to show the actual text around those boundaries, I have attached a visual document that cross-references the graph spikes with the raw lines extracted directly from your LSI interlinear archive. For each spike, the doublet identified by my script is placed alongside the original transcription line, so you can verify the correspondence at a glance.
Respectfully,
Alfredo
(02-07-2026, 05:33 PM)Vuk88 Wrote: You are not allowed to view links. Register or Login to view.I read your comments regarding the physical permutation test. In reality, setting up the analysis with those rules would result in a structural error... applying a physical shuffle to a Bag-of-Words model triggers a macroscopic homogenization bias, and it is obvious to expect results that would artificially favor the algorithm even more.
I'm not really sure what rules you are referring to since I didn't really present any rules (and I don't see that Mauro did either) nor did I suggest to do anything like this latest processing that you described.

In any case, none of this can be commented on when you haven't defined what you are using for "accuracy". And the whole approach for your randomized data testing (and the use of Y-Scrambling for the current purposes) remains -- to me-- questionable.

I’ll just recap my assessment of all this so far:
  • You have said that you found two sets of words that strongly correlate with word doublets: “triggers,” which have a positive correlation, and “inhibitors,” which have a negative correlation. You have also suggested that these two sets effectively form two identifiable subsets or clusters within the overall lexicon. But I have not yet seen evidence that demonstrates this.
  • Your bar chart appears to show only a simple, continuous trend of coefficients across the full set of words. It does not seem to show two distinct groups or clusters. Likewise, your cluster plot does not appear to support the claim about triggers and inhibitors, and it is not clear what data the plot is actually presenting.
  • Your observations about the t-SNE results—and the extent to which they show some clustering—suggest only that there is some kind of structure in the data. That, by itself, is not surprising. But the nature of that structure cannot really be identified by simply applying t-SNE. At most, the result shows what a black-box dimensional reduction looks like. And the visualization you showed does not clearly indicate any relationship to word doublets, at least not one that you have explained.
  • So far, I do not see anything that connects those dimensional-reduction results to word doublets, or really to any other specific attribute of the text—unless there is some special processing of the data that you have not described.
  • Your randomization tests also remain suspicious to me. What you have explained about them so far has not resolved that concern. It is still unclear what exactly you are measuring as “accuracy.” So far, your explanations of the randomization tests --and the expanded analysis you last described-- seem consistent with the expectation that the original-data case would be an extreme outlier in the distribution shown in the violin plot.
  • The only clear discovery you have presented—which, if correct, is indeed a surprising one—is that the density of word doublets is four to five times greater on five particular folio pages, and that those pages may also be unusual in their semantic position within the overall book. But there does not appear to be any role for your proposed trigger and inhibitor word clusters in that particular result.
  • Confirming and analyzing the spikes in doublet density does not require special algorithms such as t-SNE or cluster identification, or any machine learning techniques for that matter. It only requires good old-fashioned statistical analysis to determine whether those spikes are unlikely to occur by chance. Similarly, if some words have a strong proximal affinity for, or aversion to, word doublets—that is, if there really are triggers or inhibitors—then that too can be tested using good old-fashioned statistical techniques.
I am not saying with certainty that your analysis does not support some of your conclusions. I am saying only that, if it does, then you have not yet described the details needed to justify drawing those conclusions.
(02-07-2026, 05:33 PM)Vuk88 Wrote: You are not allowed to view links. Register or Login to view.I read your comments regarding the physical permutation test... 

Thank you. Well, I think I can tentatively accept that a set of words tend to co-occur on the same page as word doublets, while another set of words tends to steer clear from those pages. I'm not fully conviced yet, but enough that I think your observation deserves consideration, so I surely look forward to see how your work proceeds.

Regarding what this may mean, once fully confirmed, I'd be cautious to say for now. Does 'aiin' inhibits the appearance of doublets, or do doublets inhibit the appearance of 'aiin'? Or something entirely different from triggers/inhibitors causes this behaviour? I hope we'll come to know!
(02-07-2026, 08:13 PM)Vuk88 Wrote: You are not allowed to view links. Register or Login to view.Thank you for your work, because if today I can run these algorithms (which were unthinkable until a few years ago) it is thanks to the foresight of your work. Without that fundamental dataset, none of this would be possible.

Thanks for the compliment, but my contributions to the interlinear transcription file were minor.  Most of the merit should go to Gabriel Landini (who recently passed away) and René Zandbergen (who kept maintaining the file after Gabriel and I left the Voynich scene).

Quote:do you believe there might be edge cases in the transcription markers that my cleaning does not intercept?

If you are using Rene's current version of the transcription file, he has a very detailed description of the format at his website.  

The only thing that I would add to that spec is my conviction the manuscript itself (not just the transcription file) has many errors or quirks that could hide doublets.  In particular, I am convinced that m is an abbreviation, most likely for iin; that the ending -ir is a scribal error for -iin (not -in), and -iir is an error for -iiin; that Ih, ITh, IKh are just malformed versions of Ch, CTh, CKh; and that the position of the plume on Sh (on the first e, on the second e, or midway) is not significant.  I would "fix" these errors before any analysis.

There must be also many bogus or omitted word spaces, d replaced by k,  confusion between r and s, a and y, ain and aiin etc.; but at present I do not know how to detect and correct these other "errors".

However, be aware that the possibility of error is disputed by many, and anyway the effect of such corrections on the number of doublets would not be great.

Quote:On an empirical level, I performed a visual cross-check on the high-resolution facsimile (examining several samples, for example f108r) and the doublets counted by my parser find an exact physical correspondence in the ink written on the parchment...

I went to look for doublets on my version of the transcription file (which has many small differences from Rene's version).  I did NOT do the "error fixes" above, except map Ih ITh, IKh to Ch, CTh, CKh. 

I counted 292 doublets on 66 different pages.  Most pages had 3 or fewer doublets. The exceptions were:

      6 <f75r>  Page  1 of Bio
      6 <f75v>  Page  2 of Bio
      7 <f76r>  Page  3 of Bio(?) Text only
      6 <f77r>  Page  5 of Bio
      4 <f78r>  Page  7 of Bio
      5 <f80v>  Page 12 of Bio

      4 <f85v2> Big fold-out

      7 <f103v> Starred Parags
      6 <f104v> Starred Parags
      6 <f107r> Starred Parags
      5 <f107v> Starred Parags
      7 <f108r> Starred Parags
     11 <f108v> Starred Parags
      7 <f111r> Starred Parags
      4 <f112r> Starred Parags
      7 <f116r> Starred Parags + Unknown

These counts more or less match the plot on your PDF file, but I do not see signs of the peaks at section boundaries that were so prominent in the You are not allowed to view links. Register or Login to view..  But that was density of doublets per page, not count. So perhaps there was something wrong with the denominator (number of words in the page)?

From the above counts it seems that Starred and Bio do have a higher number of doublets per page than the other sections.  But they also have more words per page.   On the other hand they seem rather uniform with respect to this statistic.

The higher numbers for f108r, f108v, and You are not allowed to view links. Register or Login to view. can be explained by the fact that, on each of those pages, the Scribe mashed 10-12 parags in a single giant parag, thus boosting the number of words on those pages.

The higher number for 116r may have something to do with the fact that the last half of that page is a block of 20 lines of dense text with no stars.

I should have counted also the number of words per page and compute the densities of doublets, not just counts.  But I don't have the time now, sorry.

Doublets are common in monosyllabic languages compared to other languages.  These are examples from a Chinese medical book:

山山谷    tài shān shān gǔ.
血血瘕欲死   yū xiě xiě jiǎ yù sǐ
劳极洒洒   láo jí xiǎn xiǎn
寒寒热   shāng hán hán rè
目目痛   míng mù mù tòng
洗洗酸痛   xiǎn xiǎn suān tòng
...
Some of these are just accidents - a compound term that ends in a syllable followed by another compound that start with the same syllable.  Like 太山山谷 = "high mountains and mountain valleys".  Some are indeed compounds with two identical syllables. Like 洗洗 xiǎn xiǎn = "chills".

In non-monosyllabic languages, one could look instead for repeated syllables within or across words.  English seems to be rather scarce on those, while they seem to be more common in Romance.

If the doublets in the VMS are mostly two-syllable compounds, like xiǎn xiǎn, then you finding would be that, within a page, certain words are positively correlated with those words, while other words are negatively correlated with them.  Which would not be unusual in a natiral language.

All the best, --stolfi
@asteckley 

first of all, please accept my apologies for mixing the reply to you and to @Mauro in my previous post. I want to thank you for your critical analysis, which prompted me to reopen the data and subject the analysis to much stricter stress tests. I discovered that you were absolutely right on several methodological critiques regarding the way I presented the data, even though the underlying phenomenon seems to robustly survive classical statistical testing.
 
To answer your first question about accuracy…I had not clearly defined the target. The algorithm was set up as a binary classification at the page level—that is, predicting whether the page will contain at least one doublet or not, based on the rest of the vocabulary on that page. The accuracy in cross-validation is 60.7%. I completely agree with you on the weakness of the Y-Scrambling test because in a text structured by topics like the VMS, scrambling the labels makes it obvious that the real data will result in an extreme outlier. This demonstrates a non-random correlation (p-value 0.049), but it is certainly not enough to define such a correlation as a true "syntactic rule".
I also agree with your observation regarding the Bar Chart and t-SNE. The bar chart shows a continuous probability gradient, and describing them as "two separate subsets" was an oversimplification in my interpretation. Similarly, the t-SNE map, as you noted, groups pages based on overall vocabulary similarity, but it does not prove any direct causal link with doublets…so…in future revisions of the paper, the use of these tools will be scaled back to a purely exploratory role.
 
However, I welcomed with great interest your methodological suggestion: to verify whether the triggers exist using exclusively classical statistics.
I therefore analyzed the entire vocabulary using Fisher's Exact Test. To avoid the multiple comparisons problem and the appearance of false positives in the vast VMS vocabulary, I applied a rigorous Bonferroni Correction.
By testing the 450 main words of the manuscript (those present in at least 10 pages), the significance threshold (alpha) to consider a word a potential "trigger" dropped from 0.05 to 0.00011. In the absence of a real correlation, the probability of finding false positives below this extreme threshold is almost zero. Instead, the analysis isolated 28 different words that overcome this statistical barrier.
 
The raw data based on Takahashi's transcription (filtering out punctuation) show a rather clear trend:
 
qoteedy: Appears in 38 pages. 
Of these, 38 contain at least one doublet (100%). 
In pages without qoteedy, the probability of finding doublets drops to 52.7%. (P-Value < 0.000001).

olkeey: Appears in 26 pages. 26 contain doublets (100%) against a baseline of 55.6%. (P-Value 0.000001).

(Incidentally, the main triggers present p-values so infinitesimal — e.g. < 0.000001 — that they would survive the Bonferroni Correction even if it were calculated on the entire unfiltered vocabulary of 8,000 words).
 
To rule out that this correlation is a statistical artifact, I subjected these results to five specific "stress tests".

Syntactic Proximity:  This could have been a random co-occurrence on the same page. By measuring the physical distance, the median distance between the main triggers and the generated doublet is only 24-25 words. The trigger and the doublet are systematically found within the same paragraph, suggesting a strongly local dependency.

Dialect Control (Simpson's Paradox): I hypothesized that the triggers were simply vocabulary words of "Dialect B", and that scribe B naturally had a tendency to write more doublets. I therefore isolated the analysis exclusively to Dialect B pages. The result does not change: within Dialect B alone, the absence of qoteedy generates doublets in 51% of cases, while its presence brings doublets to 100%. The correlation occurs independently of the scribe's overall style.

Geographical Distribution: You rightly noted that there are anomalous folios with extreme densities of doublets (e.g., f108). I verified the spatial distribution of the triggers to ensure they were not confined there. The word qoteedy is found 84.2% of the time outside the anomalous folios. The rule is not a geographical "glitch" localized to a few pages.

Typology of Triggered Doublets: I verified what is being duplicated. The absolute most frequent doublet in the entire manuscript is daiin daiin. However, when qoteedy is present, the daiin daiin doublet is almost totally suppressed. In its place, the trigger forces the duplication of very different words or, frequently, of isolated bigrams like ol ol and ar ar. (Considering the hypothesis that 'ol' and 'ar' may represent counting sequences, this could imply an "operative" function for these words, though I do not want to go too far into interpretations).

Independent Validation (Cross-Transcription): To avert the risk that the correlation depended on a human error by Takahashi in transcribing ligatures, I repeated the Fisher Test by downloading the ZL3b transcription by René Zandbergen. Even in this file, the main triggers survive clearly: olkeey and olchedy remain predictive at 100%, while qoteedy settles at 97.5%.

Your critique was invaluable in making me recalibrate my terms. Defining an overall accuracy of 61% (from the ML model) as an "iron rule" was methodologically reckless. However, by isolating the data with univariate statistics, I believe it clearly emerges that there is a subset of vocabulary that radically—and syntactically locally—alters the probability of observing specific duplicated structures.
 
I thank you again for pushing me to use more rigorous verification tools. I would be glad to know your opinion on these proximity and distribution analyses, and what other tests I could possibly run next.
 
Best regards,
Alfredo
(07-07-2026, 01:02 AM)Vuk88 Wrote: You are not allowed to view links. Register or Login to view.Typology of Triggered Doublets: I verified what is being duplicated. The absolute most frequent doublet in the entire manuscript is daiin daiin. However, when qoteedy is present, the daiin daiin doublet is almost totally suppressed. In its place, the trigger forces the duplication of very different words or, frequently, of isolated bigrams like ol ol and ar ar. (Considering the hypothesis that 'ol' and 'ar' may represent counting sequences, this could imply an "operative" function for these words, though I do not want to go too far into interpretations).

This is an important point.  Consider the possibility that Voynichese is a natural language with some encoding that maps each distinct word type to the same distinct word type (or word type tuple, e.g. by splitting the word into syllables or morphemes).  Then a "doublet" is not an interesting category for analysis.  Imagine analyzing the text of an English newspaper by looking at words that have trigraph doublets, of the form xABCABCy.  You may get some intriguing results -- like "trigraph doublets occur in only 1% of the pages, but they occur in 90% of the pages where the words 'voice' or 'river' occurs" -- but it would be very hard to figure out what is going on.  Because you would be studying the distributions of "Mississippi" and "chiuhaua" and "singing" and "alfalfa" etc. added together.  

Thus I think you should 
  • focus on specific doublets, like "daiin daiin", rather than "doublets" in general
  • analyze only one homogeneous section at a time, like Herbal-A or Bio
  • use a more natural geographic unit than page
  • study distribution within the units besides across units
About the last two points: in the Herbal section, for example, the natural unit is probably the paragraph, not the page -- because the paragraphs are expected to be largely independent but with non-trivial internal correlations; and we expect that the probability occurrence of certain words and phrases will vary depending on position along the paragraph. 

For the Bio section, on the other hand, a more natural unit may be a group of a fixed number k of consecutive tokens.  Or even not use presence/absence in a unit but instead measure word-word correlation as a function of distance, measured in tokens.
 
All the best, --stolfi
(07-07-2026, 01:02 AM)Vuk88 Wrote: You are not allowed to view links. Register or Login to view.The absolute most frequent doublet in the entire manuscript is daiin daiin.

No. It is chol.

[attachment=16375]

Moreover the count of daiin daiin in the language B pages is roughly what would be expected if words were distributed randomly. daiin occurs 479 time in language A pages and 313 in language B pages. There are 11461 word pairs in A and 24056 word pairs in B. Expected doubles in A = 479 x 478 / 11461 = 20. In B, 313 x 312 / 24056 = 4. Actual numbers are 11 and 3.

I cannot see that this is taking you anywhere. These numbers are too small to come to any statistical conclusion. You also need to take into account that many of these repeat words are more common in language A or language B. And that language B pages have more words and that they have longer sentences. Also that the main language B sections, quires 13 and 20, have different frequency of repeats and seem to have a slightly different corpus of words.

The text in the sections of the manuscript is not uniform and seems to defy generalisation.
Under normal circumstances, a book has only one author. But here it is different.
When one person writes in the singular and another in the plural, things change even if the meaning is essentially the same.
“It is like this”, “They are like this”.
"es ist so" "sie sind so".
(07-07-2026, 07:57 AM)dashstofsk Wrote: You are not allowed to view links. Register or Login to view.No. It is chol.

Is it in V101?

Similar counts in RF1b-er (words extracted by ivtt -x7):
22 chol
20 qokeedy
14 qokedy
13 ar
12 daiin
11 chedy
11 ol
11 qokeey
...

In VT0e-n:
22 chol
20 daiin
19 qokeedy
14 qokedy
12 qokeey
11 ar
11 chedy
9 ol
8 dy
8 shedy

If doublets across line breaks are not counted:

RF1b-er:
20 chol
16 qokeedy
13 qokedy
12 ar
11 ol
9 chedy
9 qokeey
8 daiin
8 shedy
7 or
...

VT0e-n:
20 chol
15 qokeedy
14 daiin
14 qokedy
10 ar
10 qokeey
9 ol
8 chedy
8 shedy
6 dy
...

I'd like to know which transliteration has more daiin than chol.
Pages: 1 2 3 4 5 6