The Voynich Ninja

Full Version: Independent Analysis Identifying Languages Beyond Currier A/ B, compared to René Zand
You're currently viewing a stripped down version of our content. View the full version with proper formatting.
Pages: 1 2
Long time lurker but first time poster here. Professionally, I am a data scientist and for fun I enjoy digging into the VM.

First off, I'd like to start out by saying I have nothing but respect for RZ and the work he has done with respect to Voynich. I frequently use many of the tools he has created, and while I disagree with a few of his classifications and methods, we arrived at many of the same conclusions. I'm posting this mainly because his page on the subject asks for independent review of his parameters. <3 <3

I recently conducted my own analysis to understand if the VM contains any 'languages' beyond Currier A and B, and compared my results against those from You are not allowed to view links. Register or Login to view.(note languages here does not necessarily imply fully different languages, I'm simply using Currier's terminology).


My main motivation for undertaking this analysis is the decryption attack I’m pursuing. I have a substantial historical corpus from which I can obtain reliable botanical cribs in different languages, but the greatest barrier to applying this attack to the Voynich manuscript is the shortage of shared, botanically enriched candidates across Herbal A and Herbal B compared with the historical training works. Understanding the differences between Herbal A and Herbal B is therefore one of my top priorities.

I have two main concerns with RZ's approach. The first is that reliance on a few bigram rules could result in a false continuous gradient, that largely disappears when other features are included. The second is that the selected bigrams could completely miss large outliers, for example, as is the case with f58. While i can see the utility of using bigram segmentations as a language classifier, at least for my purposes, that type of segmentation doesn't capture enough of the underlying differences between the different language clusters.

Method

The approach i took was different than that of RZ.

I represented each page by its within-word transition counts and derived conditional transition probabilities. G² log-likelihood-ratio statistics were used to quantify feature contrasts and construct descriptive axes. Overall profile differences were measured using context-weighted Jensen–Shannon divergence, with clustering cross-checked using a Hellinger representation.

I tested candidate manuscript populations against their within-sample variation and whole-sheet comparisons, and examined whether similar groupings emerged when pages were clustered without supplying those population labels. My final results came from this combination of tests, it was not simply the output of one clustering algorithm. Classification rules similar to those supplied by RZ could be obtained from this process, but my goal was more focused around identifying the populations as opposed to developing ways to classify them.

This method produces segmentations that broadly are in agreement with RZ, although there are some meaningful differences that I'll discuss below. The fact that two different statistical approaches recover much of the same structure supports the theory that the Currier A / B classifications are insufficient to describe the variation between the text across the manuscript.


Results 

My segments are below. The percentages below express JS divergence relative to the early Herbal A–Herbal B split, set to 100%.


1. Early Herbal A
  RZ designation: A
  Divergence from Herbal B: 100%

2. Late Herbal A + Pharma
  RZ designation: Ae
  Divergence from early Herbal A: 50%; from Herbal B: 80%

3a. Herbal B¹
    RZ designation: B
    Divergence from early Herbal A: 100%

3b. Stars / Recipes
    RZ designation: broadly Bs
    Divergence from early Herbal A: 115%; from Herbal B: 35%

4. Biological / Balneological
  RZ designation: Bb; You are not allowed to view links. Register or Login to view. receives Cb at page level
  Divergence from early Herbal A: 155–160%; from Herbal B: 60%

5. Celestial / hand 4²
  RZ designation: overlaps C/Ce
  Divergence from early Herbal A: 105–125%; from Herbal B: 60–80%

6. f58
  RZ designation: A
  Divergence from early Herbal A: 100%; from Herbal B: 75%


¹ Herbal B hand 2 supplies the comparison baseline.
² The range covers astronomical, cosmological and zodiac circle/radial text separately.

Some notes:

(3a+3b) are not sufficiently different to clear the bar for an individual 'language', but represent differences significant enough that they should not be considered the same and may confound some types of analysis

f58, as a single folio, does not enough evidence to say it is fully its own language, but its divergence from the rest of the text is extreme, far more than any other page in the MS, as I will explain more below.

Compared To RZ

RZ separates languages into A, B, or C using a single bigram, 'ed'. While this is indeed an extremely simple and efficient way of segmenting Currier A vs B, relying on it as the primary classification mechanism leads to some misleading and possibly erroneous conclusions like 'C is intermediate to A and B'.  A page can be intermediate on the ed diagnostic while differing substantially on other construction patterns. An intermediate diagnostic value does not establish that some of the text is an intermediate form of Herbal A and Herbal B.

Ce as an independent dialect, in my opinion, is not supported. My threshold is a between group divergence of at least 2x, and a whole sheet permutation test with a p-value <= .01. Ce is 1.18x and .80, so it is not close, at least in my analysis. Certainly you can segment it with the rule eo > 16.7%, but this alone does not establish such a distinction is the result of different languages. The difference between approximately 22% and 12% could reflect two distinct settings, a continuous gradient, differing mixtures of constructions, or section-specific vocabulary. Group averages alone cannot distinguish those possibilities, especially with the limited text for that section.

Similarly, the single page in Cb, f80r, is not sufficiently different from other bio text to warrant its own dialect.  You are not allowed to view links. Register or Login to view. remains firmly within the larger Bio profile (RZ does acknowledge the evidence for this section is lower).

The classifications of f58 and You are not allowed to view links. Register or Login to view. are examples of the limitations of this bigram threshold approach. f58 gets an A designation despite its large departure from ordinary Herbal A, while You are not allowed to view links. Register or Login to view. receives Cb at page level while remaining broadly Bio-like. There are, in fact, are several such outliers that are roughly about as divergent as f80r: 1r, f66r, f86v5 and f99v.

The outlier score below is calculated as JS divergence from the nearest page on another physical sheet divided by the JS divergence between the page’s own odd/even text-line halves. Below 1, another page is closer to it than its own two halves are to each other. Around 1: its distance from another page is comparable to its internal half-to-half difference. Above 1: even its closest external neighbor is farther away than that internal difference. This is a rough benchmark, it does not automatically establish a page as an 'outlier'.

Outlier ranking among 187 eligible prose pages:

- f58r: rank 1; score 1.763
- f58v: rank 2; score 1.673
- f86v5: rank 3; score 1.314
- f66r: rank 4; score 1.222
- f80r: rank 11; score 1.070
- f1r: rank 22; score 0.969
- f99v: rank 25; score 0.942


In my opinion, f58 is the only folio that is an obvious 'outlier', although the You are not allowed to view links. Register or Login to view. score does intrigue me. If only we had the other missing folios from this quire.

All of the above aside, i want to be clear, I do agree with most of RZ's proposed expansions to Currier A/B.  In general, the amount of text divergence across sections exceeds that of other historical works and needs further analysis and explanation. For my attack, I would distinguish at least four main groups, with celestial material and f58 bringing the working inventory to six.
Several AI red flags in this thread, but maybe be research is sound?
(Yesterday, 11:13 AM)Koen G Wrote: You are not allowed to view links. Register or Login to view.Several AI red flags in this thread, but maybe be research is sound?
The conclusions are still interesting. I think it’s worth waiting a little while Smile
(Yesterday, 11:13 AM)Koen G Wrote: You are not allowed to view links. Register or Login to view.Several AI red flags in this thread, but maybe be research is sound?

I definitely understand the risks of accepting AI work uncritically.

I've used AI with help assembling historical reference works, as well as executing some of the tests, but AI did not generate the methods or conclusions. I'm happy to discuss those in more detail, i know it can come across like a lot of technical jargon.

The TLDR is i wanted to use a different method than PCA over bigrams, for the reason i said in my post. 

JS divergence, in my optinion, is the appropriate tool for this type of analysis, regardless of whether or not AI is used, and i think using the transition matrix rather than bigrams captures most of the bigram information while introducing additional context. Hellinger distance is a useful crosscheck as well.


I was able to back out some descriptive bigram-like rules since i submitted this post originally, using the above classifications provide.

you can plug these into voynichese yourself and pretty easily see the impact. Some of them I haven't seen utilized as discriminatory features, because they aren't as useful in the classic A/B split.


In particular, using these splits, 'cho' is actually the most useful bigram for discriminating (yes i know currier identified chol and chor as discriminatory words, but i haven't seen 'cho' specifically discussed, and definitly not as a better discriminator than ED.  After CHO, it is ED, DY, EO, then AL.

ED , EO, and very minimal AO are used in RZ's splits, but not the others. 

Happy to answer any questions!
(Yesterday, 07:01 PM)npcompl33t Wrote: You are not allowed to view links. Register or Login to view.I definitely understand the risks of accepting AI work uncritically.
I've used AI with help assembling historical reference works, as well as executing some of the tests, but AI did not generate the methods or conclusions. I'm happy to discuss those in more detail, i know it can come across like a lot of technical jargon.

Applying a double standard will not yield good results, in my opinion.
(Yesterday, 07:01 PM)npcompl33t Wrote: You are not allowed to view links. Register or Login to view.but AI did not generate the methods or conclusions
This is obviously not the case! The text above laying out the conclusions is one of the more flagrant examples of AI generated text we've seen on this site since the new rules, and a good reminder that people consistently underestimate (if not outright lie) about AI involvement in their proposals
(Yesterday, 09:07 PM)rikforto Wrote: You are not allowed to view links. Register or Login to view.The text above laying out the conclusions is one of the more flagrant examples of AI generated text

No it isn't: 0% AI according to copyleaks and gptzero detectors.
(Yesterday, 09:07 PM)rikforto Wrote: You are not allowed to view links. Register or Login to view.
(Yesterday, 07:01 PM)npcompl33t Wrote: You are not allowed to view links. Register or Login to view.but AI did not generate the methods or conclusions
This is obviously not the case! The text above laying out the conclusions is one of the more flagrant examples of AI generated text we've seen on this site since the new rules, and a good reminder that people consistently underestimate (if not outright lie) about AI involvement in their proposals

Lol i wrote that myself, i almost left in all the spelling mistakes for this exact reason.
It doesn't read too AI to me, which is why I wondered if anyone is able to comment on the research.
(17-09-2026, 01:01 AM)npcompl33t Wrote: You are not allowed to view links. Register or Login to view.The second is that the selected bigrams could completely miss large outliers, for example, as is the case with f58.

The frequency of "al" alone should set f58 apart: You are not allowed to view links. Register or Login to view.

Nick Pelling wrote about the peculiarities of f58: You are not allowed to view links. Register or Login to view.
Pages: 1 2