I'm curious how the generator theories can account for the positional tendencies people here keep finding, and I didn’t want to blow up someone else’s thread about their generator. I'm not arguing against a generator, I just can't picture one rule producing all of them, since they seem to work in different ways.
Starting right at the front of the VMS, the two paragraphs of folio f2r:
What's been noticed by the pros here:
1. Tavie's vertical impact: a line's opener avoids the glyph directly above it, and it's directional (o→q is common, q→o isn't; o-o is avoided but s-s is fine). You can watch it in the opener column, they keep changing all the way down.
2. Anton's opener order (from "Regaining the lost order"): openers advance through a set sequence, k y d ch o q s sh. Anton reads You are not allowed to view links. Register or Login to view. as k-d-q in both paragraphs (which you can see), and says paras tend to open with a gallows. The first paragraph here it runs k→d→q→[c]→o→s→sh.
3. Top row: gallows p/f cluster on the first line of a paragraph (Tavie, Currier, Feaster). Here that's p in ypchol/ypchaiin (line 1) and f in fodan (line 8).
4. Feaster's drift: big picture within a line, sh→ch, qo→o, k→t shift rightward, and word makeup drifts as you go down the paragraph.
5. Line ends: the last word tends to differ from mid words (Currier).
So my question: how does a generator that writes left to right, one line at a time, produce these, especially the vertical one? What ties right drift to opener order to being allergic to verticality… maybe it’s separate habits, coincidence, astral phenomena…
// FWIW here is the rabbit hole of other threads I didn’t want to blow up with this question, and sources for positionality such I’ve been sorting through on this. This started with me just privately trying to aggregate rules and tendencies to harmonize over a nice pour of whiskey:
A. Other positional tendencies I left off the image:
benches (ch/sh) are rare as line-openers but spike later: You are not allowed to view links. Register or Login to view.
first words of lines run longer, second words shorter, Vogt 2012 (plain handwriting effect?): You are not allowed to view links. Register or Login to view.
long-range word correlations look like real topic structure, Montemurro & Zanette 2013: You are not allowed to view links. Register or Login to view.
q-start and m-end lines are anti-correlated (MarcoP and pfeaster in the vertical-impact thread below)
B. Sources for what's on the image:
vertical impact, Tavie: You are not allowed to view links. Register or Login to view.
opener order, Anton: You are not allowed to view links. Register or Login to view.
drift, Feaster 2022: You are not allowed to view links. Register or Login to view.
line as a functional unit, Currier 1976
C. Related threads:
ch/sh interchangeability You are not allowed to view links. Register or Login to view.
copy/mutate ledger You are not allowed to view links. Register or Login to view.
hoax/generator debate You are not allowed to view links. Register or Login to view.
generators: Timm & Schinner 2019 (self-citation) You are not allowed to view links. Register or Login to view.
Cardan grille, Rugg 2004 You are not allowed to view links. Register or Login to view.
ReneZ 2021 write-up You are not allowed to view links. Register or Login to view.
Schinner 2007, lots here and the stochastic generator observation You are not allowed to view links. Register or Login to view.
I posted this in my thread, but it might be worth posting it here as well.
I don't have any statistical research on hand, so I can't justify it with scripts or tables. I've determined all of this, let's say, "by eye".
Instead of using linguistic concepts such as "prefix," "root," "suffix," or models like "Crust-Mantle-Core," I decided to try to represent positional patterns in terms of character ratios (identifying which characters are "bigger" than others and assigning weights accordingly).
Here's what I have (sorry for the crooked drawing):
The symbols are grouped conditionally. The largest symbol here is o (ignoring q), and the smallest is y (it appears 23 times after n, which is smaller than even a conventional unit!)
The unit group is special in that it is difficult to define internal relationships. I can assume that r is greater than i based on 18 examples where r precedes i. d is less than e, because I believe that a = ei. But if we consider a as a separate symbol, it will be in the same group.
Some remarkable ratios:
1). ch's more than a dozens of them, but fewer gallows. Also, ch's more than o.
2). x is more than a dozens
3). In the category with the symbol q, I can enter c, which has the property of standing in front of gallows (I mean cases like ct, not cth)
4). It is difficult to determine the weight of s, but I assume that it is the largest in the dozens group, as well as more than ch.
I don't know what good it will do, but I did it Well, at the very least, it perfectly demonstrates the peculiarity of the Voynichese letters. Since they can easily be assigned weights, maybe they're not letters at all, but something else... Maybe they're numbers... Maybe even You are not allowed to view links. Register or Login to view.?
Hi everyone,
I'm Alfredo. My background is actually in social science data analysis, but like many of you, the Voynich Manuscript has been a long-time obsession of mine. I recently had an intuition regarding its structure that I really wanted to test out.
I started wondering if the manuscript is less like standard prose and more like a highly structured technical manual or data ledger. I know that's not a brand-new concept, but I wanted to see if I could find hard statistical proof of its internal grammar. Building on the Latent Semantic Analysis (LSA) research by Bowern, Layfield, and Davis (which the mods here kindly shared!), I decided to run some computational models on the EVA transliteration to hunt for deterministic rules.
I specifically focused on one of the most famous statistical anomalies: the "double repetitions" (like chol chol). While it's common to dismiss these as dittography or scribal copying errors, I hypothesized something different. What if, in a highly structured document with zero punctuation, these consecutive repetitions actually act as mechanical syntactic markers? Essentially, logic gates or section delimiters.
To test this, I applied Machine Learning algorithms (Logistic Regression and LSA) to analyze the words that heavily co-occur on the same pages as these doubles, looking for a mathematical correlation.
The dependency turned out to be massive. These double words aren't isolated or random; they are strictly bound to specific associated words. I've attached a clustering graph below (voynich_syntax_constellation_en.png) where you can see a distinct separation.
The blue and red dots represent two entirely separate grammatical families. The t-SNE algorithm segregated the vocabulary into two distinct functional 'clouds', showing a strong structural dependency. This rigid separation is the fingerprint of a highly structured system that rigorously distinguishes between different functional categories of words
To find out exactly which words do this, I looked at the Logistic Regression coefficients. In the attached bar chart logistic_bar_chart_en.png, you can see these specific "syntactic triggers" isolated. The red bars highlight the exact words that heavily co-occur on the exact same pages as the double sequences, while the blue bars show the ones that mathematically inhibit them (they almost never appear together).
To prove this wasn't just a coincidence, I ran a Monte Carlo Permutation Test with 1,000 iterations. Basically, I had the script completely scramble the text 1,000 times. This destroys the original word order, but keeps the raw word counts exactly the same. As you can see in the second attached graph (permutation_violin_en.png) , the randomly shuffled text completely failed to recreate the pattern (p-value = 0.0010). This proves the manuscript follows rigid conditional rules.
But here is the part with the biggest codicological implications. By mapping the spatial density of these syntactic markers across the whole manuscript (see the attached graph: voynich_all_missing_density.png) , a striking coincidence emerges: the absolute highest peaks of density are concentrated exactly on the pages right next to the codicological gaps—the 14 missing folios (f12, f59-f64, f74, f91-f98, f109-f110). see the attached graph: voynich_all_missing_density.png
A quick note on the attached graph (lsa_vs_syntax.png): This chart visually summarizes the intersection between semantic and syntactic data. The X-axis represents the manuscript folios (timeline from f1 to f116). The colored background blocks show the different thematic sections identified by the LSA algorithm (Herbal, Astronomical, etc.), separated by vertical dashed lines indicating where the topic abruptly changes. The red line represents my data: the density of the double words. As you can clearly see, almost every time there is a semantic transition (a change in topic), there is a massive red spike in syntactic double words. They act as physical boundaries between sections.
For example, one of the highest peaks in the entire book is at folio f108, right before the extraction of the final folios. This strongly supports the historical hypothesis that those removed pages contained the conversion matrices or "Tabulae" required to decode the frequent logical transitions in those sections.
I recently published the full paper, the dataset, and the Python framework I used for the analysis. You can check everything out here: You are not allowed to view links. Register or Login to view.
I would love for the experts here to take a close look, test the Python framework, and offer some constructive criticism. I'm really curious to hear your thoughts on how we might use this syntactic mapping moving forward.
Thanks for your time!
P.S. For those who prefer a less technical read, I have also attached a simplified, non-academic summary of the theory below in PDF format.
You are not allowed to view links. Register or Login to view.
Results strongly indicate that the text was generated via simple methodology (order-3 markov chain).
Also this shows that there is zero measurable correlation between text and the imagery.
Also text written by different "hands" all matches statistical fingerprint, which strongly suggests use of shared methodology for text generation.
Does it prove anything? Guess not, however result seem to leave little to imagination.
All the code necessary for replicating results, detailed findings and vl dataset is in the repo.
"Note: Proposed "solutions" will not be accepted." If the proposed solutions are not accepted, then what is the purpose of the conference?
This policy actually says :"We want to study the Voynich Manuscript, but we do not actually want anyone to solve it." It protects the status quo and ensures that the mystery will continue indefinitely. Refer my thread: You are not allowed to view links. Register or Login to view. Isn't it time on testing evidence, not just maintaining the mystery?
First, my current working hypothesis: The VMS is a compendium of midwifery knowledge compiled during the turmoils within and around the Knights Hospitaller and the rise of the Aragonese, and written in a specifically for this purpose designed and constructed script - on Malta, in the first decades of the 15th. century. Please note, that I honestly doubt this is the ultimative solution - but I assume it's a good pivot point for further investigations.
My main and leading assumptions are as follows:
1. The VMS is (to me at least, obviously) a compendium of midwifery knowledge. No secret society behind it, no cryptography, no deeper metaphorical setup constructing deeper enlighment only to share with the rulers who are worth to know in the sight of God. And not necessarily funded in far-away countries. All just things, your ordinary Jane Doe Midwife should know. This does not automatically rule out metaphorical comparisons or complex systems we nowadays would consider superstition in a given detail, but focussing eg. on the illustrations I would tend to say: if you see a bathing female nude, I assume that Occam's razor surely suggest a bathing nude in first place and not some half-divine figure from forthcoming apocalypses, the power of Qi or similar. Let's keep it simple.
2. Books about midwifery are rare, though midwifery was daily practice and common to use. This is knowledge that for some very good reasons usually was handed over personally from one generation of commited women to the next, normally "on the job", and for centuries, normally strictly orally. And even this had to be done with a lot of precautionary measures. Seeing it written down is a very rare thing. This may have, BTW, led to some blind spot within the one or the other VMS theory. An usual book is an unlikely book. Around 1400~1430 AD (I take this slot as a given) Malta already had at least one vivid hospital, most likely two, and like in Akkon, Rhodos, Sigena or other places, Mulieres Sti. Johannis and their staff (mostly local woman) were active here, too, and as far historians tell us, scientifically leading. But the loss of estates in the Holy Land led to a severe crisis within the Knights Hospitaller and their related enterprises. The Order concentrated their activities to Rhodos. And so (long before the order relocated to Malta), the future of the care work on Malta was unclear. A relocation for nuns, often originatrd from european countries and sometimes of noble birth, was always possible. One for the local midwifes wasn't, at least not for all of them. So there might have been the decision to keep these treasures alive. Even in case the chain of support would not break on Malta, it would have been handy to have a compendium that might be used for teaching lessons to successors. And yes, Malta, is franky spoken, a vague idea. Several places would work without big differences. But, alas, everyone needs a place to start.
3. But how to maintain this knowledge? By writing it down. I assume that a team of several people helped to write down the oral explanations of the local Maltese midwives. Hence, several scribes. Hence, several ducts. But they had to master further challenges: there was no script for (early) Maltese. And most probably, the team had no shared language with those. Maybe the midwives even had different dialects. So they found an obvious solution by setting up a script. This included a few decisons: the simplified radically what they hear into a small group of consonants. Maybe they did the same with the vowels, maybe they skipped them directly. E.g. 13 different fricatives as modern Arabian has, can easily be stripped together to a few. Same with plosives. Nasals, Liquids - one each is enough. I already tested it with classical Arabian, and generally it works - I will soon present here my first steps as well in detail. Of course, with lower Shannon entropy come more homonyms. But the root structure of any semitic language is full of redundance. Nevertheless, some peculiarities of VMS may root in corrective measures to avoid homonyms - e.g.: clear word borders, positional markers, ligatures, and maybe, additionally, vowels or matres. With only a small subset of letters needed they could work on a system for quick writing and, as I assume, integrated this as well. Again, I see no way at the moment, to reduce the possible phonems to graphems in a way that means easy undoing it the other way round. Maybe for ever. But for reading a language one already knows this not necessary. It's just a bummer we don't know and maybe never will.
4. Beside the work on further evidences that this might have been the key to the alphabet and maybe additional estimations there, I'm working on a second corpus lingustic topic, and this is to analyze, whether it is at least not totally unlikely the VMS was written (or better dictated) in a semitic language. I adopted a method here from pattern recognition and I named it "fake lemmata". This includes the setup of word lists (including lists for hapax and dislegomena) and calculate entropy and collocations. Whenever an observed pattern turns out to be significant (fake lemma) it is replaced in a derivated corpus (shadow corpus) as a dedicated token. This is done iteratively. For semitic languages it is expected, that the collocations would rise in distinctiveness and the distribution will be more stable. This comes with several stages of root lexica, word lengthes, TTR, change of Zipf curves, entropy and more Information to analyse. This is, of course, not only done for VMS but also for Maltese, Siculo-arabian, classic Arabian, Hebrew and several addional comparison corpora. The amount of operations is then the measure to compare, like, but not the same as the Levenshtein distance.
Let me please remind at least myself in addition, that every additional assumption makes the overall approach more unlikely - even if makes the approach at first glance additionally plausible. We all tend to such cognitive bias, as we know from Kahneman and Tversky (and Slovic, if I recall it correctly).
So now, challenge me, roast me. As sooner I know where my false assumptions are, the sooner I can adopt my approach.
I've always suspected that some diagrams are enlarged versions of other diagrams.
So far, I've found one interesting match on two diagrams with f68v.
Let's go through the indirect indicators:
1). The colors are the same (only blue)
2). There is a small star in the center
3). The second diagram has one more line.
4). The left diagram has 8 "rays" (I'm referring to those blue areas), while the right diagram has 16.
5).There is no text above the right diagram. If this is a variant of the left diagram, then there is no need to describe it again.
In general, it seems that the right diagram is the left diagram "enlarged" by two times.
They have a few common words: oteey, chey, dal, and dal is in approximately the same position.
A less obvious example is f69r-f69v, but here there is a decrease.
Both diagrams have a star in the center, but the number of rays is different (f69r - 46, You are not allowed to view links. Register or Login to view. - 31).
This word is not on Voynichese, obviously, was left by the author/scribe himself.
It is in a circle with the Voynichese words. The first letter (which is clearly a Latin s) and the third letter (maybe an n) are what distinguish this word. The word is also highlighted by the fact that there are connections between the letters (marked in dark red in the image below), and the distance between them is quite small.
I tried to draw the letters that I could see here:
At least this word has 4 letters, the final one is not identified for sure. The word looks similar to the Bavarian "Suna" (the sun), which is consistent with the sun that is supposedly drawn nearby. However, it is unclear why u has a tail...
If it's a contraction, we won't be able to decipher it.
I don't have any statistical research of my own, but that hasn't stopped me from replacing the ch/sh and k/t pairs. Interestingly, they seem to be interchangeable in many words.
Most words with ch have an equivalent with sh: chedy(501) - shedy(427) chey(344) - shey(283) cheody(89) - sheody(50) chckhy(140)-shckhy(60) chdar(20)-shdar(9) chodaiin(44) - shodaiin(23)
With the exception of chckhy and chdar, you can see that the frequency of sh words is slightly lower than that of ch words, but the difference is quite small.
The situation is similar with k/t: kedy(44) - tedy(42) okedy(118) - otedy(115) okey(64) - otey(57) okeody(37) - oteody(39) okeos(14) - oteos(29) okeor(22) - oteor(12)
The difference is also insignificant, sometimes only a few units.
The reason for this pattern is partly positional equality (the symbols are in the same position), but then the question is what caused it.
By the way, there is no such pattern between f and p, and these letters do not have words that can be compared (well, or I didn't find them).
After spending considerable time in Voynich Ninja forum, I have come to a consideration somhow uncomfortable.
I believe that there is considerable obstacle to solving the manuscript, that goes beyond the text itself.
The response to new theories always seems to be immediate scepticism rather than genuine engagement. Has the community become so attached to the pursuit of the answer, that it no long wants the answer itself? "The journey is its own reward!"- but is that sentiment blocking the destination?