The Voynich Ninja

Full Version: A mathematical approach to double words, conditional logic and the missing pages
You're currently viewing a stripped down version of our content. View the full version with proper formatting.
Pages: 1 2 3 4 5 6
(30-06-2026, 03:46 PM)Vuk88 Wrote: You are not allowed to view links. Register or Login to view.I specifically focused on one of the most famous statistical anomalies: the "double repetitions" (like chol chol). ... These double words aren't isolated or random; they are strictly bound to specific associated words.

That is a very interesting observation! (Even though it is not evidence that the text is not plain language.)  

The legend on this figure says that the frequency of doublets spikes at transtions between sections:
[attachment=16281]
That would be another interesting observation... However, the precise coincidences, the absence of spikes elsewhere,  and the uniform size of the spikes make me suspect that they are artifacts of the way the data is provided or processed.  Are you sure that this is not the case?  

Can you show the actual text around those boundaries where one can see the doublets?

All the best, --stolfi
(01-07-2026, 08:50 PM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.... However, the precise coincidences, the absence of spikes elsewhere,  and the uniform size of the spikes make me suspect that they are artifacts of the way the data is provided or processed.  Are you sure that this is not the case?  

I agree. In fact, that was the first thing that crossed my mind as well, because this is exactly the kind of surprising result that one can get excited about, only to later discover that it was caused by some simple coding bug that seemed unimaginable at the time.

So it definitely calls for extra checks and quality-assurance testing. Vuk88, don't be too surprised—or embarrassed—if that turns out to be the case.

I look forward to seeing more results.
(01-07-2026, 03:17 PM)asteckley Wrote: You are not allowed to view links. Register or Login to view.(I apologize that I have not fully digested other people's comments and your replies. From what I read there, it seemed clearer if I responded directly to your initial post and your zenodo paper only.)
What you have done is very interesting and possibly valuable work, but I am unclear on several aspects of it.
Hi Andrew,
First of all, thanks for your careful reading and for giving me the opportunity to clarify the methodology, since the brevity of the initial post didn't really allow me to explain everything systematically.
I'll try to answer your questions.
Quote:First, if I understand correctly, your first chart shows the words with the highest correlation — that is, the unique words (aka word types) most likely to occur as doublet sequences — and the words with the lowest correlation — that is, the words least likely to occur that way. [...]
If so, then the complete bar graph would presumably show a continuous trend of words, with correlations decreasing gradually from 0.933 for “sy” down to 0.665 for “chees.” In other words, your chart has clipped out the long middle range of a continuous trend related to how often unique words occur in doublet form -- no indication of clustering per se.
To begin with, as you correctly pointed out, the logistic regression produces a continuous probability spectrum (ranging from a high probability of generating doublets to a high probability of inhibition). The bar chart deliberately clips the middle range to highlight the extremes, namely the highest positive and negative coefficients. Actually, the goal of this chart wasn't to prove "clustering" per se, but rather to show that the doublet phenomenon is not uniformly distributed across the vocabulary. Specific words act as heavy "syntactic wrappers" (red), while others act as strong "inhibitors" (blue).
Quote:The coloring of red and blue in that cluster plot is presumably assigned simply by the sign of the z-direction value. [...]
So is this chart not just showing the positive vs negative sign of an unidentifiable metric in the low-dimension reduction from the t-SNE algorithm?
Or did you instead use a seperate clustering algorithm on the points in the t-SNE output space in order to define two clusters?
Regarding how the clusters were defined and what the colors represent, I can confirm that the red and blue colors aren't the result of a clustering algorithm applied to the t-SNE coordinates. They are predefined labels imported directly from the logistic regression (Red = Main Trigger Words; Blue = Main Inhibitors). The t-SNE algorithm was used exclusively to reduce a high-dimensional contextual vector space (Word2Vec/TF-IDF) into a 3D space. The fact that the pre-labeled Red and Blue words naturally separate into distinct regions in the t-SNE space demonstrates that the words triggering doublets and those inhibiting them do not share the same contextual environments. They belong to different syntactic and semantic classes.
Quote:With the Monte Carlo simulations and the violin plot, again, if I understand your description correctly, the predictive accuracy concerns whether one can predict whether a given word belongs to one cluster rather than another. Is that correct?
[...]
But again, I am not clear how you conclude that this structure has anything to do specifically with word doublets, as opposed to other structuring effects, (such as the self-citation patterns described by Timm or the differing vocabularies found in Currier languages A and B).
As for the tests conducted, I can tell you that in each of the 1,000 permutations, the order of the words in the manuscript was completely randomized. However, I strictly preserved the exact global vocabulary, the frequencies of every single word, and the word lengths, while also standardizing the page lengths. I then ran the exact same predictive pipeline (Logistic Regression -> Accuracy Score) on the shuffled text just to see if the model could still predict the appearance of doublets based on the surrounding vocabulary. The predictive accuracy collapsed to about 50%, so pure randomness.
This proves that the doublet phenomenon is not an artifact of Zipf's Law, global word frequency, or macro-scale vocabulary differences like Currier A and B. It is a strictly local, non-random syntactic rule based on word adjacency. Timm's "self-citation" hypothesis would struggle to explain why specific "trigger" words manage to reliably predict the onset of a doublet sequence across the entire corpus.
To eliminate any doubt that these results might be polluted by transcription noise, I should also clarify the preprocessing methodology. Before running any algorithm, I applied a strict purification filter to the entire dataset. Any word containing uncertainty markers, non-alphabetic characters, or anomalous punctuation in the Takahashi transcription (e.g., `dar!o!m`) was cleaned to extract only its pure root (`darom`). So the mathematical models worked exclusively on the real linguistic data, eliminating at the source the risk of statistical "hallucinations" caused by transcription noise.
Additionally, returning to your point about the Currier A and B dialects, I verified that the syntactic rule I discovered ("Trigger -> Doublet") fits perfectly across both dialects. Even though the Herbal section (Currier A) uses a completely different set of nouns compared to the Recipes section (Currier B), the exact same "Trigger" words are used to close data blocks in both sections. It is undeniable that this observation raises serious doubts about the "language simulation device" hypothesis, since it's hard to imagine how a random generator or a medieval hoaxer could maintain such a rigid mathematical rule intact...
Quote:As I understand it, you have plotted the density of word doublets per folio page and compared that to the thematic assignments that have been observed through folio sequence of the manuscript. And you have found that five particular folios stand out by having four to five times as many word doublets as all the other folios. Those five pages each happen to either be the terminal page of a thematic section, or (at least in the last of the five cases) to precede a section of missing folios. Is that correct?
[...]
(By the way, the work about to be published by Colin Layfield and Lisa Fagin-Davis might be worth reviewing when it becomes available soon, to see if it affects the folio sequence positioning part of your finding.)
Finally, you noticed that the last analysis (the massive spike of doublets on terminal folios, like f108) seems a bit disconnected from the previous machine-learning charts. You are absolutely right, it is indeed a structural observation of the manuscript as a whole. However, it perfectly closes the circle of this research.
While the machine-learning models explained how the mathematical rule works at the single-line level (the "trigger" words generating doublets), the analysis of the terminal folios explains why the author (or authors) created it. The fact that this specific mathematical formula is so heavily concentrated precisely at the end of thematic sections proves that these double words act as actual "End of File" or "End of Database" markers, used to close out blocks of information.
Thanks for pointing out the upcoming work by Colin Layfield and Lisa Fagin-Davis. Since the codicological sequence of the quires is highly debated, any new data on the original binding order will be crucial to see if these syntactic "spikes" align even more perfectly with the original transitions between the fascicles.
Thanks for your meticulous observations.
Alfredo
(01-07-2026, 05:33 PM)Mauro Wrote: You are not allowed to view links. Register or Login to view.
(30-06-2026, 10:58 PM)Vuk88 Wrote: You are not allowed to view links. Register or Login to view.The Logistic Regression doesn't use an n-gram model, it uses a Bag-of-Words approach at the page level. So it's not looking at physical adjacency, but rather page-level co-occurrence. It found that words like shedy and chees heavily turn on or off on the exact same pages where the doubles appear.
Thanks, it's clearer now.

(30-06-2026, 10:58 PM)Vuk88 Wrote: You are not allowed to view links. Register or Login to view.On the sample size issue with the 36 chees vs 298 doubles, I had the exact same concern about statistical chance. That's precisely why I ran a Monte Carlo Permutation Test with 1000 iterations. The empirical p-value came out to p=0.0010, and the mathematical distance between the permuted random distribution and the true manuscript data is what confirms this isn't just a random fluke.

Question: did the MonteCarlo permutation reshuffle all the words, or did it keep the original doublets and reshuffled the rest? And, did it keep the page structure of the original (the number of words on each page)? I think I would be convinced your findings are not a statistical fluke with a permutation test which:

- keeps the same doublets as the original in the same positions (that is to say, on the same page as the original, given your algorithm works page by page)
- keeps the same page structure (number of words on each page) as the original
- avoids creating new doublets when reshuffling words
- and finally reshuffles words section by section (Balneo, Herbal A, Herbal B..). I say this because there are big differences in vocabulary between the sections and mixing everything together will tend by itself to dilute any statistical signal.

In any case, nice work  Smile

Hi Mauro,
To answer your question directly...No, I strictly avoided physically reshuffling the words in the manuscript. Diluting the Currier sections or creating random doublets by chance are exactly the structural pitfalls I wanted to avoid. I invite you to read my recent reply to Andrew regarding the general standardization methodology, but here is how I bypassed the structural problem in the permutation test.
As you likely know, to avoid altering the manuscript's physical structure, I relied on Y-Scrambling (Target Permutation) rather than a physical text scramble. The text of the Voynich, its structure, and the exact positions of the words were left 100% untouched. 
I simply assigned a mathematical label ('1' or '0') to the text blocks depending on whether they actually contained a doublet or not. 
During the 1,000 permutations, the algorithm randomly shuffled only these mathematical labels, forcing the model to look for a syntactic rule on fake targets while the underlying text remained perfectly intact.
Thanks to this approach, the Currier vocabularies were never mixed, and the physical page structure was never altered.
I hope this clears up your doubts about the structural integrity of the test.

Thanks Mauro for the help you are giving me. 
What I really need right now is exactly this: to stress-test the algorithms as much as possible before moving on to the next steps of the analysis.
Alfredo
Your concept of “trigger” words was not coming across very clearly before. That may be more my fault than yours, so let me see if I can get a few things straight.
As I understand it, you are treating a word that occurs anywhere on a page containing word doublets as a “trigger” if that word correlates highly with such pages. But the trigger word can appear anywhere on the page: before the doublet, after the doublet, between different doublets, and so on.

Also, there is no particular correlation between the trigger word and the specific words that form the doublet. So, for example, if “sy” appears on a page, it predicts that at least one doublet, or perhaps many doublets, are likely to occur on that same page. But it does not predict which particular doublet or doublets will appear. It could be a doublet of any word.

And if a word correlates negatively with pages containing doublets, then it is an “anti-trigger,” so to speak — what you call an inhibitor. In other words, it predicts the absence of doublets on the page where it occurs.

Regarding your cluster plot: you used the t-SNE algorithm to place the words as points in a three-dimensional space. Then you independently colored them according to whether they were positive predictors of doublets, which you call triggers, or negative predictors, which you call inhibitors.
Presumably, you did that by assigning any word with a coefficient greater than 0 as a trigger, and any word with a coefficient less than 0 as an inhibitor. Or did you use thresholds — for example, one color for words with coefficient greater than 0.5, and another for words with coefficient less than -0.5?

I also do not understand why “sy” appears while none of the other labelled words in that plot even show up in the bar chart of extreme correlation words. I assume the blue “sy” (Instead of red) is just a switch in color choices, but I still cannot understand the labelling on the cluster plot.

In any case, most of the individual points on your cluster plot are not visible, because they have largely been replaced by continuous contoured color shading. So it is not clear whether the intensity of that shading represents the value of the z-dimension or the density of points.

If the shading is based on the trigger/inhibitor assignment I described above, then with or without non-zero thresholds, what I am seeing is simply that the three reduced dimensions chosen by the t-SNE algorithm have some general relationship to the coincidence of a word with doublets occurring on the same page. But that relationship could be as simple as the page having or not having word doublets at all; it is not clear that the reduced dimensions are even capturing the existence of the  trigger or inhibitor words. Also, regardless of whether the shading reflects coefficient value or point density, any clustering that is present in the plot seems to form three clusters about as strongly as it forms two.

In short, the connection between the results of the clustering plot and the coefficients that signal triggers and inhibitors, as illustrated in the bar chart, remains totally unspecified.

(02-07-2026, 02:56 AM)Vuk88 Wrote: You are not allowed to view links. Register or Login to view.As for the tests conducted, I can tell you that in each of the 1,000 permutations, the order of the words in the manuscript was completely randomized. However, I strictly preserved the exact global vocabulary, the frequencies of every single word, and the word lengths, while also standardizing the page lengths. I then ran the exact same predictive pipeline (Logistic Regression -> Accuracy Score) on the shuffled text just to see if the model could still predict the appearance of doublets based on the surrounding vocabulary. The predictive accuracy collapsed to about 50%, so pure randomness.
Can you clarify this further?

By “the same predictive pipeline,” do you mean that in each of the 1,000 trials, “sy” — and in fact all the words — effectively became neutral triggers, with no real predictive power one way or the other?
Or do you mean that in each of the 1,000 trials, all the words were analyzed completely afresh to find a new ordering and range of logistic regression coefficients, and then the predictive accuracy was recomputed using those trial-specific trigger words? In that case, are you saying that no words showed meaningful predictive power, with all coefficients effectively close to zero? (And, of course, as I think someone pointed out earlier in the thread, you must keep the doublets together during the shuffling procdes, as if they were a single word.)

Obviously, under the first scenario, we should not be at all surprised that the “Actual Voynich Data” appears in the extreme tail of the distribution shown in the violin plot.

By the way, I did find that your Zenodo paper goes into a bit more detail than your initial post at the top of this thread, though not by very much. 
I assume you have good answers to many of the questions I have raised here. If so, I think it would be well worth explaining them in your formal research paper. If the conclusions you have stated are correct, then you have a good story to tell.

But to be candid, I remain dubious. I still feel that your major finding is primarily what you have shown in your last “Cross Analysis” plot. Assuming that result is not eventually explained away by a coding bug, I think it could be analyzed much further than you have taken it so far.

For example: are the trigger words and inhibitors correlated with particular word doublets? That question alone could produce a different assignment of triggers versus inhibitors. Are the trigger words correlated with the number of word doublets on a page? What about the proximity of the trigger words to the predicted doublets within individual pages?

If I can find some time, I will try to reproduce your analysis.

Regards,
--Andrew
(02-07-2026, 03:51 AM)Vuk88 Wrote: You are not allowed to view links. Register or Login to view.I simply assigned a mathematical label ('1' or '0') to the text blocks depending on whether they actually contained a doublet or not. 
During the 1,000 permutations, the algorithm randomly shuffled only these mathematical labels, forcing the model to look for a syntactic rule on fake targets while the underlying text remained perfectly intact.

Sorry -- I just saw this part of your previous answers.  If I understand what you did correctly, then your violin result is probably not surprising at all. 

In any randomly selected document (of which the VMS is one such document), you can form a bar chart of logistic regression coefficients where all the words can be ordered from whatever the highest value is to whatever the lowest value is. In the case of the VMS you got a range of 0.933 down to -0.665. (Again, from that chart, there is no clustering apparent -- just a continual trend from some positive value down to some negative value andmore or less centered on zero.) The only question is whether that range that is spread across the value zero is larger than expected by chance or smaller.  (Determining that would require you going one step further to perform a statistical test of significance on the range or variance of the coefficients.)  

By randomly assigning the mathematical label on the text blocks, you come up with a different set of coefficients, but the range of those coefficients is not necessarily any different than in the subject case. (That's what the statistical test of significance would determine.)

It is important to understand just how you came up with the "predictive accuracy" in your violin chart.  Is it the predictive accuracy of your top triggers (and bottom inhibitors), or some averaging of the accuracies of all the words as predictors, or what? If you are considering, for example, the predictive accuracy of just the most powerful trigger plus that of the most powerful inhibitor, then finding the "Actual Voynich Data" located in the extreme tail might be meaningful because your randomization technique is then testing if the VMS case is indeed unusual or not.

Regards,
--Andrew
You used a clever method, but I rather agree with asteckley here (You are not allowed to view links. Register or Login to view.).

I'd feel much more comfortable with a reshuffling test done along the lines of my You are not allowed to view links. Register or Login to view.. If after the reshufflings the plots come out significantly different from the original, then that would be something.
I too have difficulty understanding your claim.

But also when you are doing an analysis on repeat words then it would help to enumerate the repeats. I have done this for you, for quires 13 and 20. See attached.

In particular you will see:
  • In quire 13 the frequency of word pairs that are repeats is 1.1%. For quire 20 it is 0.6%. Roughly a halving.
  • In quire 13  ol is the top word ( 299, 4.2% of words ). The number of  ol ol repeats is 12 and this is nearly the same as the expected number if words were distributed randomly ( 299 x 4.2% = 13 ).  qokedy repeats 11 times, 4 would be expected.  qokeedy, observed 9, expected 3.
  • In quire 20  aiin is the top word ( 322, 2.73% ) but it does not repeat. 9 would be expected.  qokeedy, observed 11, expected 1.  qokeey, observed 9, expected 2.  ar, observed 8, expected 4.

We do know that there is a difference in the words and language of the sections of the manuscript and that this will probably affect the frequency of repeats. In particular quire 13 has a smaller corpus of words [ see You are not allowed to view links. Register or Login to view. ]. Moreover, the number of repeats over and above what could be expected is so low that I doubt that any analysis will lead to any definite conclusion.
(02-07-2026, 01:12 PM)dashstofsk Wrote: You are not allowed to view links. Register or Login to view.I too have difficulty understanding your claim.

But also when you are doing an analysis on repeat words then it would help to enumerate the repeats. I have done this for you, for quires 13 and 20. See attached.

In particular you will see:
  • In quire 13 the frequency of word pairs that are repeats is 1.1%. For quire 20 it is 0.6%. Roughly a halving.
  • In quire 13  ol is the top word ( 299, 4.2% of words ). The number of  ol ol repeats is 12 and this is nearly the same as the expected number if words were distributed randomly ( 299 x 4.2% = 13 ).  qokedy repeats 11 times, 4 would be expected.  qokeedy, observed 9, expected 3.
  • In quire 20  aiin is the top word ( 322, 2.73% ) but it does not repeat. 9 would be expected.  qokeedy, observed 11, expected 1.  qokeey, observed 9, expected 2.  ar, observed 8, expected 4.

We do know that there is a difference in the words and language of the sections of the manuscript and that this will probably affect the frequency of repeats. In particular quire 13 has a smaller corpus of words [ see You are not allowed to view links. Register or Login to view. ]. Moreover, the number of repeats over and above what could be expected is so low that I doubt that any analysis will lead to any definite conclusion.

Hi @dashstofsk,
Thank you for the interest shown in my work and for the detailed PDF. 
The expected vs. observed frequency calculations in your document are mathematically correct, however, there is a fundamental misunderstanding regarding what the Machine Learning model is actually measuring. Your analysis counts self-repeats (how often ol is immediately followed by ol). My Logistic Regression model does not measure self-repeats.
As @asteckley accurately noted (in post #15), the algorithm operates at the page level. It measures whether the presence of a specific word anywhere on a page correlates with the appearance of any doublet on that same page.
When the model identifies aiin as an "Inhibitor", it means that on pages where aiin is heavily used, the author (or authors) actively suppressed the use of doublets entirely, conversely, "Triggers" act as page-level syntactic markers, predicting that a doublet sequence is likely to occur nearby.
Interestingly, your own data highlights this underlying syntax even at the self-repeat level. You noted that qokedy self-repeats 11 times instead of the expected 4. Furthermore, your data shows that aiin (with 322 occurrences) should statistically produce 9 self-repeats, but it produces 0.
If doublets were pure random statistical noise, aiin would have formed its 9 expected repeats. Its total absence further proves the existence of a syntactic suppression rule.
While your analysis shows that self-repeats of frequent words like ol are likely statistical noise, this does not affect my central correlation that the surrounding page vocabulary mathematically predicts the doublet phenomenon.
To understand exactly why certain words trigger this behavior, I am designing a second phase of the study to drop down to a sub-lexical level. But before jumping to hasty conclusions or introducing new variables, I need to further stress-test the underlying logic upon which the current algorithms were built, and that takes time.
Regards,
Alfredo
(02-07-2026, 08:46 AM)Mauro Wrote: You are not allowed to view links. Register or Login to view.You used a clever method, but I rather agree with asteckley here (You are not allowed to view links. Register or Login to view.).

I'd feel much more comfortable with a reshuffling test done along the lines of my You are not allowed to view links. Register or Login to view.. If after the reshufflings the plots come out significantly different from the original, then that would be something.

@asteckley and @Mauro.
I read your comments regarding the physical permutation test. In reality, setting up the analysis with those rules would result in a structural error... applying a physical shuffle to a Bag-of-Words model triggers a macroscopic homogenization bias, and it is obvious to expect results that would artificially favor the algorithm even more.
Out of pure scientific curiosity and to see exactly how the algorithm would run, I tried running it with your rules. As expected I got exactly the result I imagined, the algorithm proved me even more right by returning an accuracy of 63.3% compared to the 54% of the real text. But accepting this result would be a mistake for a very specific reason…If you physically shuffle the words in an entire section you completely destroy the natural variance of the manuscript. The original text is clustered, the author grouped specific words in specific places. When we scramble them we create a homogeneous soup.
 I empirically measured this phenomenon and the vocabulary variance between pages drops by 40% after the shuffle. To explain it simply, let's think about the initial part of the Voynich where there are many drawings and very little text. If we physically shuffle the words of that section, we take the words from the dense pages and spread them over those with the drawings. We end up homogenizing the text, making the empty pages equal to the text-rich ones, irreparably altering the sophisticated and rigid structure of the manuscript's distribution. By doing so, we would end up feeding the model artificial and too “perfect” data... The algorithm has an incredibly easy job when it works with homogeneous data, and that is why the accuracy increases, but it is a trick, a bias.
 It was precisely to avoid this error the reason why the Y-Scrambling algorithm I used is better for this task. Instead of touching the text, Y-Scrambling leaves the complex manuscript 100% intact exactly as it was written. It only randomizes the target labels, namely the doublets. This forces the model to work on the real and unaltered text, taking into maximum consideration the variance and complexity of the manuscript that the physical test instead destroys.
To guarantee maximum transparency, I will write down the logical steps with which I fed the algorithm for the physical shuffling test so that you can verify its setup. Logical Pipeline of the Physical Shuffling Test :
  1. Section Isolation: 
    The manuscript was divided into thematic quires (Herbal, Astro, Balneo, Pharm, Recipes). Words never crossed the section boundaries.
  2. Extraction and Locking: 
    For each section, the length of each page and the exact positions of the original doublets were locked. All other single words were inserted into a section shuffling pool.
  3. Shuffling: 
    The pool of single words was randomly shuffled for 1000 iterations.
  4. Reconstruction: 
    The pages were reconstructed by filling the empty spaces with the words from the shuffled pool. I added a rule to prevent the creation of new accidental doublets. If a word drawn from the pool was identical to the one just inserted, the algorithm discarded it and drew a different one.
  5. Evaluation: 
    The shuffled pages were evaluated by the same Logistic Regression model. I thank you for suggesting this test. Observing how the model reacts to homogenized data was very instructive for understanding the underlying structure of the manuscript. If you believe there are logical errors in this pipeline or that the algorithm needs modifications, I am absolutely ready to listen and discuss them.Alfredo
Pages: 1 2 3 4 5 6