The Voynich Ninja

Full Version: A mathematical approach to double words, conditional logic and the missing pages
You're currently viewing a stripped down version of our content. View the full version with proper formatting.
Pages: 1 2 3 4 5 6
(07-07-2026, 07:57 AM)dashstofsk Wrote: You are not allowed to view links. Register or Login to view.
(07-07-2026, 01:02 AM)Vuk88 Wrote: You are not allowed to view links. Register or Login to view.The absolute most frequent doublet in the entire manuscript is daiin daiin.
No. It is chol.

In my transcription of the Starred Parags section ("Quire 20"), these are the string doublets and their frequencies, ignoring spaces and line breaks but not parag breaks:
Code:
    68 ar:ar
    28 ra:ra
    19 al:al
    19 ol:ol
    15 or:or
    11 chedy:chedy
    11 yqokee:yqokee
      9 ain:ain
      7 aiinot:aiinot
      7 okeey:okeey
      6 edyqok:edyqok
      6 qokeedy:qokeedy
      5 eyqok:eyqok
      5 lo:lo
      4 aiin:aiin
      4 ch:ch
      4 daiin:daiin
      4 eyqoke:eyqoke
      4 ro:ro
      4 yqoke:yqoke
      3 air:air
      3 alot:alot
      3 dy:dy
      3 kal:kal
      3 ok:ok
      3 okain:okain
      3 okedy:okedy
      3 ot:ot
      3 otar:otar
      3 oteey:oteey
      2 aro:aro
      2 chedyqok:chedyqok
      2 chol:chol
      2 chor:chor
      2 dyqoke:dyqoke
      2 edyot:edyot
      2 edyqoke:edyqoke
      2 eed:eed
      2 eeyqok:eeyqok
      2 eeyqoke:eeyqoke
      2 eolsh:eolsh
      2 inota:inota
      2 keol:keol
      2 okaiin:okaiin
      2 okaiinch:okaiinch
      2 okeedyq:okeedyq
      2 olkee:olkee
      2 olqok:olqok
      2 orsh:orsh
      2 os:os
      2 otain:otain
      2 qopchedy:qopchedy
      2 yche:yche
      2 yched:yched
      2 yd:yd
      1 aiinpol:aiinpol
      1 aiinqok:aiinqok
      1 ainok:ainok
      1 ainot:ainot
      1 alch:alch
      1 ald:ald
      1 alshe:alshe
      1 am:am
      1 anot:anot
      1 arch:arch
      1 arok:arok
      1 aror:aror
      1 chal:chal
      1 chdy:chdy
      1 che:che
      1 ched:ched
      1 chedyol:chedyol
      1 chedyqot:chedyqot
      1 cheeo:cheeo
      1 cheo:cheo
      1 cheol:cheol
      1 cheos:cheos
      1 chey:chey
      1 cheyqo:cheyqo
      1 chy:chy
      1 da:da
      1 dain:dain
      1 dair:dair
      1 dal:dal
      1 dyche:dyche
      1 dyke:dyke
      1 dyl:dyl
      1 dyote:dyote
      1 dyqo:dyqo
      1 edyche:edyche
      1 edyok:edyok
      1 edyqopch:edyqopch
      1 ee:ee
      1 eedylk:eedylk
      1 eedyqo:eedyqo
      1 eedyqok:eedyqok
      1 eeodyqo:eeodyqo
      1 eeodyqok:eeodyqok
      1 eeylk:eeylk
      1 eeyok:eeyok
      1 eeyqo:eeyqo
      1 eeyqot:eeyqot
      1 eollkch:eollkch
      1 eysh:eysh
      1 hals:hals
      1 inda:inda
      1 inopdai:inopdai
      1 irai:irai
      1 kaiinch:kaiinch
      1 kaiino:kaiino
      1 kaiinqo:kaiinqo
      1 kaiiny:kaiiny
      1 kedyo:kedyo
      1 keedyl:keedyl
      1 keeey:keeey
      1 keey:keey
      1 ko:ko
      1 lara:lara
      1 lch:lch
      1 lchedy:lchedy
      1 lcheedy:lcheedy
      1 lcheey:lcheey
      1 lda:lda
      1 lkeedy:lkeedy
      1 lkeeed:lkeeed
      1 lkeey:lkeey
      1 mo:mo
      1 odaiinch:odaiinch
      1 odyot:odyot
      1 ofchedyq:ofchedyq
      1 okainq:okainq
      1 okal:okal
      1 okar:okar
      1 okcheyq:okcheyq
      1 oke:oke
      1 okeedy:okeedy
      1 okeeyq:okeeyq
      1 okey:okey
      1 olaiin:olaiin
      1 olch:olch
      1 olche:olche
      1 olchedy:olchedy
      1 opchyq:opchyq
      1 oq:oq
      1 otair:otair
      1 otal:otal
      1 otedy:otedy
      1 otee:otee
      1 oteedy:oteedy
      1 otey:otey
      1 pche:pche
      1 pchedy:pchedy
      1 qo:qo
      1 qokal:qokal
      1 qokedy:qokedy
      1 qokeey:qokeey
      1 qotcheo:qotcheo
      1 qoted:qoted
      1 qotedy:qotedy
      1 qoteedy:qoteedy
      1 raii:raii
      1 ralo:ralo
      1 rcheo:rcheo
      1 rcho:rcho
      1 rota:rota
      1 shdar:shdar
      1 shedy:shedy
      1 sheyl:sheyl
      1 shol:shol
      1 soe:soe
      1 teeyo:teeyo
      1 to:to
      1 ylche:ylche
      1 ylched:ylched
      1 ylk:ylk
      1 ylked:ylked
      1 ylkeed:ylkeed
      1 ylt:ylt
      1 yo:yo
      1 yok:yok
      1 yoteed:yoteed
      1 yqok:yqok
      1 yqoked:yqoked
      1 yqokeed:yqokeed
      1 yqoklche:yqoklche
      1 ysh:ysh
      1 yshe:yshe
      1 ysheod:ysheod
This list has all duplicated strings of length 2 to 8.  I did not find any duplicated strings with length between 9 and 15, and I did not look for longer ones.

Some of these duplicated strings are duplicated words, but some are created by accidental juxtaposition of words with certain suffixes and prefixes.  The most common word doublets seem to be ar.ar(68), al.al(19), ol.ol(19), or.or(15). chedy.chedy(11), okeey.okeey(7), qokeedy.qokeedy(6), daiin.daiin(4), okain.okain(3), okedy.okedy(3), otar.otar(3), oteey.oteey(3); however these counts may include some accidental cases. 

The counts for those strings, duplicated or not (again, ignoring spaces and line breaks but not parag breaks) are ar(1104), al(938), ol(1159), or(472), chedy(545), okeey(262),  qokeedy(134), daiin(306), okain(179), okedy(105), otar(110), oteey(88).

The file has N=55886 EVA letters ignoring spaces. Thus the probability of (say) chedy occurring at any random spot is a bit less than p = 545/N or ~0.975%. If the probability did not depend on what came before, we would expect about 545*p  = 545^2/N or ~5 occurrences of chedychedy; thus the 11 actual occurrences are mildly notable.  In the case of qokeedy, we would expect ~0.3 occurrences of quokeedyqokeedy, that is, none; so the 6 actual cases are quite notable. 

All the best, --stolfi
Jorge_Stolfi Wrote:This is an important point. Consider the possibility that Voynichese is a natural language with some encoding that maps each distinct word type to the same distinct word type [...] Then a "doublet" is not an interesting category for analysis.
[...]
Thus I think you should
- focus on specific doublets, like "daiin daiin", rather than "doublets" in general

Dear Professor Stolfi,  
I would like to reply by combining the extraordinary insights you provided in both of your messages (posts #25 and #27), offering you the empirical data I have just calculated.

  1. (Post #25) Finding your orthographic speculations very interesting, I mathematically verified their impact. I modified my algorithm so that it would apply to the LSI corpus (36,364 words cleaned of notes and markers) a restrictive set of substitutions based exclusively on your indications: global substitution of m with iin, of Ih/ITh/IKh to Ch/CTh/CKh, and conditional substitution of the endings -ir and -iir to -iin and -iiin (only at the end of a word).
  Judge the result for yourself:
  • A full 206 Hapax Legomena (rare words appearing only once in the corpus) vanished, assimilated into more common roots; essentially, you have provided an important key to drastically reducing the vocabulary's entropy, cleaning it from transliteration errors!
  • The total number of doublets rose from 301 to 328 (revealing hidden repetitions like daiir daiin).
  • The syntactic correlation remained nothing short of "robust"... the qoteedy trigger continued to predict the presence of doublets in 38 cases out of 38.
(To definitively rule out the hypothesis that this precision was a statistical artifact linked to high frequency, I ran a randomization test by randomly shuffling the order of all 36,364 tokens. In the randomized text, the predictive capacity of qoteedy plummeted from 100% to 39.3%, proving that the pairing in the real text is intentional and positional).

  1. The "Mississippi / Singing" analogy (Post #27) Your analogy regarding the English trigraph doublets is, in my opinion, methodologically rigorous. Indeed, treating all doublets as a single indistinguishable bucket risks creating a statistical "soup". It is for this reason that I have just conducted a disaggregated analysis specifically on the repetition typologies for the top 5 triggers of the manuscript.
Globally, the most frequent doublet in the book is daiin daiin (9.8% of all doublets). But look what happens to the specific distribution when the text is "activated" by the triggers:
  • In the presence of qoteedy: daiin daiin suffers a drastic collapse (2.7%). In its place, the repetition qokeedy qokeedy (11%) and the short syllable ol ol (4.8%) emerge anomalously.
  • In the presence of olkeey: daiin daiin is almost totally inhibited (1.0%). In its place, ol ol dominates (7.1%).
  • In the presence of qokeedy: daiin daiin drops to 3.4%. The syllable ar ar emerges (4.5%).
  • In the presence of qokedy: daiin daiin halves to 4.7%.
  • In the presence of shedy: daiin daiin drops to 4.2%.
This "extreme level of active suppression" exactly corroborates your thesis: we are not looking at a generic propensity to duplicate words, but rather we are observing triggers that "call" only specific roots or targeted monosyllables (inhibiting the other, more common options). Whether it is a monosyllabic language or operative commands, the conditional rigidity of this syntactic structure appears amply supported by the evidence.

Lastly, I enthusiastically accept your methodological recommendations for the next steps. Abandoning the "page" as a geographical unit to test the algorithm on single paragraphs (in the Herbal section) or on sliding windows of k-tokens (in the Bio section) to measure the correlation based on distance will be the focus of my next analysis session.
Thank you again for the time dedicated to this research.
With profound esteem,
Alfredo
(07-07-2026, 11:00 AM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.In my transcription of the Starred Parags section ("Quire 20"), these are the string doublets and their frequencies, ignoring spaces and line breaks but not parag breaks: [...]

I re-counted the string doublets in the Starred Parags section after "correcting" m -> iin,  ir -> iin, hh -> he.  The only significant changes were aiinaiin rising from (4) to (29), and daiindaiin rising from (4) to (10).

The corrected file has N = 56817 EVA letters (including '?'). The string daiin occurs 415 times, and aiin occurs 1774 times (including 415 as part of daiin). By the same reasoning as before, the string doublet aiinaiin should occur ~55 times, so the actual count (29) is a bit too low. The doublet daiindaiin should occur only ~3 times, so the actual count (10) is somewhat notable.

All the best, --stolfi
(07-07-2026, 11:01 AM)Vuk88 Wrote: You are not allowed to view links. Register or Login to view.This "extreme level of active suppression" exactly corroborates your thesis:

Oh I see. No point talking to a bot agreeing with whatever you say. Also notoriously unable to count.
(07-07-2026, 01:24 PM)nablator Wrote: You are not allowed to view links. Register or Login to view.
(07-07-2026, 11:01 AM)Vuk88 Wrote: You are not allowed to view links. Register or Login to view.This "extreme level of active suppression" exactly corroborates your thesis:

Oh I see. No point talking to a bot agreeing with whatever you say. Also notoriously unable to count.

If the formal tone of my post gave the impression of an AI-generated text, it is simply because I use translation and formatting tools to ensure my English is clear. 
If I had tried to translate from Italian to English myself, I would have definitely made mistakes, just as it happened to me before. Therefore, I have no intention of apologizing for using tools to make myself understood.
Furthermore, regarding the accusation of being "unable to count", the frequencies and data I provided are not guessed by a "chatbot". They are the direct output of Python scripts processing the raw LSI transcription.
If you believe the counts are wrong, I invite you to write your own script to parse the Quire 20 text (applying Stolfi's m->iin corrections) and calculate the exact token distance between qoteedy and qokeedy qokeedy.
I am solely interested in continuing to explore and structure my underlying hypothesis, thanks to the feedback of those who are actually following the reasoning. If you have any scientific observations or counter-arguments regarding the data, I am listening. Otherwise, I will proceed on my own path.
(07-07-2026, 02:27 PM)Vuk88 Wrote: You are not allowed to view links. Register or Login to view.If I had tried to translate from Italian to English myself, I would have definitely made mistakes, just as it happened to me before.
Translating the written text is certainly no problem. That’s what I do, after all.
In my case, it’s German to English. I use Deeple. But sometimes I’m not sure about it either. If the site translates it back into German, it’s not right either. Sometimes I have to ask myself, who on earth wrote such nonsense?
The translator gets the meaning of the translation wrong.
(07-07-2026, 02:27 PM)Vuk88 Wrote: You are not allowed to view links. Register or Login to view.If the formal tone of my post gave the impression of an AI-generated text, it is simply because I use translation and formatting tools to ensure my English is clear. 
If I had tried to translate from Italian to English myself, I would have definitely made mistakes, just as it happened to me before. Therefore, I have no intention of apologizing for using tools to make myself understood.

You need not apologize for it.
I noticed early on that your postings were clearly composed with some AI assistance. But it was also very obvious that your posts were not simply AI generation, and that you were likely using it for translation purposes. And to make your responses clearer and more organized. 

That being said, you should try to prune out some of the excessive complimentary language a bit; it comes across as ingratiating and a sign of AI insincerity. (I expect, however, that doing that is probably difficult if  English is not your native language.)
(07-07-2026, 02:27 PM)Vuk88 Wrote: You are not allowed to view links. Register or Login to view.If the formal tone of my post gave the impression of an AI-generated text, it is simply because I use translation and formatting tools to ensure my English is clear. 
If I had tried to translate from Italian to English myself, I would have definitely made mistakes, just as it happened to me before. Therefore, I have no intention of apologizing for using tools to make myself understood.

We have a prohibition on AI-assistance.  There should be a banner on the main page before you sign up.  If you go to The Slop Bucket sub-forum, you will see that we regularly ban posts for being AI generated.  

We accept that AI can be used as a translation tool, but it should not be allowed to rewrite your posts.  When it offers to help with "formatting", that is how slop can creep in, and it also results in the kind of language that makes people think they are talking to a chatbot with you as an intermediary.  Neither is good for discussion.

So if you're using a chatbot/LLM like ChatGPT, Claude, or Gemini for your posts here, please use it for strict translation (I presume from Italian to English) of the words you have written, and don't permit it to alter your text any more than that.  We would rather have your unformatted words rather than a polished LLM version of them.
(07-07-2026, 03:18 PM)tavie Wrote: You are not allowed to view links. Register or Login to view.
(07-07-2026, 02:27 PM)Vuk88 Wrote: You are not allowed to view links. Register or Login to view.If the formal tone of my post gave the impression of an AI-generated text, it is simply because I use translation and formatting tools to ensure my English is clear. 
If I had tried to translate from Italian to English myself, I would have definitely made mistakes, just as it happened to me before. Therefore, I have no intention of apologizing for using tools to make myself understood.

We have a prohibition on AI-assistance.  There should be a banner on the main page before you sign up.  If you go to The Slop Bucket sub-forum, you will see that we regularly ban posts for being AI generated.  

We accept that AI can be used as a translation tool, but it should not be allowed to rewrite your posts.  When it offers to help with "formatting", that is how slop can creep in, and it also results in the kind of language that makes people think they are talking to a chatbot with you as an intermediary.  Neither is good for discussion.

So if you're using a chatbot/LLM like ChatGPT, Claude, or Gemini for your posts here, please use it for strict translation (I presume from Italian to English) of the words you have written, and don't permit it to alter your text any more than that.  We would rather have your unformatted words rather than a polished LLM version of them.

 Thanks for the warning. From now on I will limit the use of AI strictly to translating parts of the text where a literal translation could lead to misunderstandings. I will make sure it does not rewrite or format my posts anymore.
Why don't you use an online translator like deepl.com? You won't have any problems with it.
Pages: 1 2 3 4 5 6