Long time lurker but first time poster here. Professionally, I am a data scientist and for fun I enjoy digging into the VM.
First off, I'd like to start out by saying I have nothing but respect for RZ and the work he has done with respect to Voynich. I frequently use many of the tools he has created, and while I disagree with a few of his classifications and methods, we arrived at many of the same conclusions. I'm posting this mainly because his page on the subject asks for independent review of his parameters. <3 <3
I recently conducted my own analysis to understand if the VM contains any 'languages' beyond Currier A and B, and compared my results against those from You are not allowed to view links.
Register or
Login to view.(note languages here does not necessarily imply fully different languages, I'm simply using Currier's terminology).
My main motivation for undertaking this analysis is the decryption attack I’m pursuing. I have a substantial historical corpus from which I can obtain reliable botanical cribs in different languages, but the greatest barrier to applying this attack to the Voynich manuscript is the shortage of shared, botanically enriched candidates across Herbal A and Herbal B compared with the historical training works. Understanding the differences between Herbal A and Herbal B is therefore one of my top priorities.
I have two main concerns with RZ's approach. The first is that reliance on a few bigram rules could result in a false continuous gradient, that largely disappears when other features are included. The second is that the selected bigrams could completely miss large outliers, for example, as is the case with f58. While i can see the utility of using bigram segmentations as a language classifier, at least for my purposes, that type of segmentation doesn't capture enough of the underlying differences between the different language clusters.
Method
The approach i took was different than that of RZ.
I represented each page by its within-word transition counts and derived conditional transition probabilities. G² log-likelihood-ratio statistics were used to quantify feature contrasts and construct descriptive axes. Overall profile differences were measured using context-weighted Jensen–Shannon divergence, with clustering cross-checked using a Hellinger representation.
I tested candidate manuscript populations against their within-sample variation and whole-sheet comparisons, and examined whether similar groupings emerged when pages were clustered without supplying those population labels. My final results came from this combination of tests, it was not simply the output of one clustering algorithm. Classification rules similar to those supplied by RZ could be obtained from this process, but my goal was more focused around identifying the populations as opposed to developing ways to classify them.
This method produces segmentations that broadly are in agreement with RZ, although there are some meaningful differences that I'll discuss below. The fact that two different statistical approaches recover much of the same structure supports the theory that the Currier A / B classifications are insufficient to describe the variation between the text across the manuscript.
Results
My segments are below. The percentages below express JS divergence relative to the early Herbal A–Herbal B split, set to 100%.
1. Early Herbal A
RZ designation: A
Divergence from Herbal B: 100%
2. Late Herbal A + Pharma
RZ designation: Ae
Divergence from early Herbal A: 50%; from Herbal B: 80%
3a. Herbal B¹
RZ designation: B
Divergence from early Herbal A: 100%
3b. Stars / Recipes
RZ designation: broadly Bs
Divergence from early Herbal A: 115%; from Herbal B: 35%
4. Biological / Balneological
RZ designation: Bb; You are not allowed to view links.
Register or
Login to view. receives Cb at page level
Divergence from early Herbal A: 155–160%; from Herbal B: 60%
5. Celestial / hand 4²
RZ designation: overlaps C/Ce
Divergence from early Herbal A: 105–125%; from Herbal B: 60–80%
6. f58
RZ designation: A
Divergence from early Herbal A: 100%; from Herbal B: 75%
¹ Herbal B hand 2 supplies the comparison baseline.
² The range covers astronomical, cosmological and zodiac circle/radial text separately.
Some notes:
(3a+3b) are not sufficiently different to clear the bar for an individual 'language', but represent differences significant enough that they should not be considered the same and may confound some types of analysis
f58, as a single folio, does not enough evidence to say it is fully its own language, but its divergence from the rest of the text is extreme, far more than any other page in the MS, as I will explain more below.
Compared To RZ
RZ separates languages into A, B, or C using a single bigram, 'ed'. While this is indeed an extremely simple and efficient way of segmenting Currier A vs B, relying on it as the primary classification mechanism leads to some misleading and possibly erroneous conclusions like 'C is intermediate to A and B'. A page can be intermediate on the ed diagnostic while differing substantially on other construction patterns. An intermediate diagnostic value does not establish that some of the text is an intermediate form of Herbal A and Herbal B.
Ce as an independent dialect, in my opinion, is not supported. My threshold is a between group divergence of at least 2x, and a whole sheet permutation test with a p-value <= .01. Ce is 1.18x and .80, so it is not close, at least in my analysis. Certainly you can segment it with the rule eo > 16.7%, but this alone does not establish such a distinction is the result of different languages. The difference between approximately 22% and 12% could reflect two distinct settings, a continuous gradient, differing mixtures of constructions, or section-specific vocabulary. Group averages alone cannot distinguish those possibilities, especially with the limited text for that section.
Similarly, the single page in Cb, f80r, is not sufficiently different from other bio text to warrant its own dialect. You are not allowed to view links.
Register or
Login to view. remains firmly within the larger Bio profile (RZ does acknowledge the evidence for this section is lower).
The classifications of f58 and You are not allowed to view links.
Register or
Login to view. are examples of the limitations of this bigram threshold approach. f58 gets an A designation despite its large departure from ordinary Herbal A, while You are not allowed to view links.
Register or
Login to view. receives Cb at page level while remaining broadly Bio-like. There are, in fact, are several such outliers that are roughly about as divergent as f80r: 1r, f66r, f86v5 and f99v.
The outlier score below is calculated as JS divergence from the nearest page on another physical sheet divided by the JS divergence between the page’s own odd/even text-line halves. Below 1, another page is closer to it than its own two halves are to each other. Around 1: its distance from another page is comparable to its internal half-to-half difference. Above 1: even its closest external neighbor is farther away than that internal difference. This is a rough benchmark, it does not automatically establish a page as an 'outlier'.
Outlier ranking among 187 eligible prose pages:
- f58r: rank 1; score 1.763
- f58v: rank 2; score 1.673
- f86v5: rank 3; score 1.314
- f66r: rank 4; score 1.222
- f80r: rank 11; score 1.070
- f1r: rank 22; score 0.969
- f99v: rank 25; score 0.942
In my opinion, f58 is the only folio that is an obvious 'outlier', although the You are not allowed to view links.
Register or
Login to view. score does intrigue me. If only we had the other missing folios from this quire.
All of the above aside, i want to be clear, I do agree with most of RZ's proposed expansions to Currier A/B. In general, the amount of text divergence across sections exceeds that of other historical works and needs further analysis and explanation. For my attack, I would distinguish at least four main groups, with celestial material and f58 bringing the working inventory to six.