Eg. Root:
1) find the labels representing root
2) find paragraphs that only mention those labels and not others
3) find words in those paragraphs that are unique to those paragraphs regarding roots
4) cross reference those words found in the paragraphs of other categories and narrow it down like that.
Please let me know what you think and if you can help me at all
There is a word that I suspected to mean "root" that appears only in 2 pharma pages that display many roots. Let's see if your guess is the same as mine.

If you can figure out which word means root, I'll give you £10...
(17-09-2026, 10:45 AM)DjCatac Wrote: You are not allowed to view links. Register or Login to view.Eg. Root:
1) find the labels representing root
2) find paragraphs that only mention those labels and not others
3) find words in those paragraphs that are unique to those paragraphs regarding roots
4) cross reference those words found in the paragraphs of other categories and narrow it down like that.
Please let me know what you think and if you can help me at all
Hello DjCatac!
I think that's a good idea!
How many labels have you analyzed so far?
I have actually been using this approach myself, and unfortunately there are multiple confounding things that end up popping up. I can share my list of candidate terms and the analysis i've done so far if that is helpful.
I assembled several historical reference texts to test against, and sure enough, these all show significant enrichment of botanical terms like root, seed, flower, etc. Root in particular tends to occur in about 50% of entries in comparable herbal mansucripts. Seed can end up having different spellings, but is also common. Flower is less so, but does appear to be more likely on pages that show a flower.
I have tested training a model, based on some established blind translation research, trained against historical works, to identify these terms. while overall absolute accuracy is lopw, it is capable of recovering some of them based on context and frequency analysis alone.
The issue is that when i try to apply this to voynich, the certainty in classifications drops to almost 0.
The first major barrier to this analysis is that there aren't really enough candidates that overlap between Herbal A and Herbal B, with the frequencies you would expect, while also not appearing in non-botanical sections. Frankly, this is a huge issue -- its possible that the word for 'root' is not the same between Herbal A and Herbal B, but in general, the lack of this shared botanical specific vocabulary between the two is extremely concerning.
The second is that 'root' and 'seed' tend to also occur in anatomical, recipe, and biological sections in my historical comparisons. Root is frequently used to describe anatomy, and seed is also often used in reproductive contexts. The issue is that the actual subject matter of the two largest non-botanical sections is not clear. THe bio/balneological section and the stars section could very well use these terms, which would make discriminating them quite difficult.
The top candidates for most botanical categories are: ckhy, chckhy, dam, and oky , however none of these really occur as frequently as i would expect for root, and they appear as likely terms for virtually all botanical concepts.
ykchy is a candidates for flower and seed
yty, and ytey for fruit
This list is not comprehensive -- i can provide more detailed information if that is helpful.
I think this is definitely an avenue worth pursuing, just the major confounding factor is the lack of expected overlap between A and B with the frequencies i'd expect for these terms, based on my historical comparisons.
(Yesterday, 07:24 PM)npcompl33t Wrote: You are not allowed to view links. Register or Login to view.
I have tested training a model, based on some established blind translation research, trained against historical works, to identify these terms. while overall absolute accuracy is lopw, it is capable of recovering some of them based on context and frequency analysis alone.
That seems very different to me from your first message.
My apologies, I mistook the author of the message.
I started with an approach like the author.
Essentially, the idea is once you identify one word as root, that removes it from a pool of words. You then can assign the next word, or swap words to try to maximize a scoring function.
There is actually some You are not allowed to view links.
Register or
Login to view. on this type of 'blind' machine translation. Essentially the same idea, as the author, scaled up to all words.
In my historical-text benchmarks, this approach puts the correct botanical meaning among its ten proposed candidates roughly 40% of the time. First-choice accuracy is much lower, only
3/54, or 6%. So it is useful for narrowing down a shortlist of potential cribs, less good at actually identifying specific ones.
Again the issue is that the overall certainty in the results from this method completely falls apart when used on voynich, combined with the large differences between Herbal A and Herbal B. None of the historical works i was analyzing contain differences as large as Herbal A and B, at least not within the same topic.
(Yesterday, 07:54 PM)npcompl33t Wrote: You are not allowed to view links. Register or Login to view.Again the issue is that the overall certainty in the results from this method completely falls apart when used on voynich, combined with the large differences between Herbal A and Herbal B. None of the historical works i was analyzing contain differences as large as Herbal A and B, at least not within the same topic.
Is there any reason not to test Herbal A and Herbal B separately?
No, that is ultimately what I did. I was hoping that I could identify candidate words in both and perhaps establish some sort of correspondence.
There are plausible candidates in both separately. One issue is that A lacks a suitable non-botanical comparison, so it can be difficult to distinguish botanical from non-botanical words.
For example, the closest word to matching the 'root' profile in my historical reference works is cthol, however, this word is actually one of the markers Currier used to distinguish Currier A itself, so it could just be a common non-botanical currier A word. Cthor, the 2nd runner up, has the same initial cth used by Currier to distinguish A.
Notably, these have almost no appearance in B, and no other words really match the frequency profile of 'root' in the historical comparisons i have. In the historical texts, root is dramatically enriched against unrelated topics, over 5x in Dioscorides botanical work, nearly 18x in medical remedy works. This should not be a difficult word to spot, if the historical reference works are any example.
Some B leads that don't occur in A are: Okam, kar, kchdy, chdy, and ykedy.
You can immediately see the problem though, a word like 'root' should be extremely common in botanical sections vs non, yet the only word that really matches this profile only really occurs in A, and is itself considered a marker of the A language. No obvious correspondence exists between the B candidates and the A candidates.
This is ultimately what led me to start digging into the different 'Languages' in the manuscript, which i discuss in the other thread i posted. The idea is sound, root should be present and fairly obvious to pick out, i just can't seem to find a good candidate for it, and not from a lack of trying.
(17-09-2026, 03:48 PM)nablator Wrote: You are not allowed to view links. Register or Login to view.There is a word that I suspected to mean "root" that appears only in 2 pharma pages that display many roots. Let's see if your guess is the same as mine. 
I wonder: if 10 people agreed to submit the word they think is "root" to an independent person, how many would pick the same word? Maybe we should try it.