Posted by: Bluetoes101 - 13-12-2025, 01:12 AM - Forum: Imagery
- No Replies
I was looking at "Commentary on the Apocalypse" this week and noted these weird "petal stars".
I wondered if there is any difference in meaning, or if they just had 2 ways of doing stars?
Either way it reminded me of the "stars" on You are not allowed to view links. Register or Login to view.
You are not allowed to view links. Register or Login to view.
Hey folks,
I've been binging Koen's videos on Youtube over the last week, great stuff!
Obviously that means I'm a Voynich novice, but I did use computational linguistics during my PhD at the MPI in Nijmegen, Netherlands, so I couldn't help but dig my fingers into the data :)
I'm not claiming any novelty but I haven't seen the different analysis steps put together in one place so I figured I might as well publish it here. (However, I do think in the end, I have some interesting results that I didn't see anywhere else... but more about this below and in the next post.)
I put the data, scripts and a small analysis report on a dedicated Github repository, if somebody wants to have a deeper look: You are not allowed to view links. Register or Login to view.
My main idea for this round is to perform an TF-IDF analysis. This is a statistical tool where you count how often words occur over the whole text, vs the individual text segments (pages). With this method, one can (approximately) distinguish "content words" from "function words". Content words are those that are specific to a particular topic, like "Voynich" or "Quantum", while function words are words that show up everywhere, like "the", "of", "and", etc. I think it would be tantalizing to produce a list of Voynich words, where we can guess, from section and illustration cues, what they might mean, given where they show up. (Although I don't think it would bring us closer to deciphering the text, it would be fun.)
Down the line, maybe I have time to produce a visual tool where people can explore how words cluster in certain portions of the text. Not quite as fancy as the amazing tool on voynichese.com but in the same spirit.
I'm currently getting into working with LLMs (building them, not talking to Chat GPT) and I am very curious if one can use these tools to identify semantic clusters of Voynich words. Tbd.
I obviously haven't read everything there is on Voynich, but I did my best to go through voynich.nu and Bowern and Lindemann (2020) as well as the latest and pinned posts on this forum, to get a base understanding what's commonly known and what's currently under discussion. I'm looking forward to learn more from you.
I want to start by stating my base preconceptions/assumptions when I went into the analysis, as well as some questions that maybe you can help me with.
Assumptions:
1. The text is real in the sense that somebody in the 15th century wrote something down to communicate information to somebody else.
2. The transcription is reasonably good and conveys the textual content of the VMS to an overwhelming degree, so we can base an analysis on it.
3. The words are words in the sense that they can, through translation, combination, compression, augmentation by auxiliary information, or some other process, be rendered into a language that someone at some time spoke. If there is a cipher, it did not jumble words by moving word boundaries or similar shenanigans.
4. Letters are only meaningful with regards to the manuscript itself. They cannot be identified in a one-to-one manner with any language.
5. The manuscript was written by several scribes/authors, possibly at different times, possibly without knowing each other. The known separations are Currier A and B as well as the 5 Hands (Davis, Lisa Fagin. 2020).
6. There is no hope of me ever decrypting the text by myself since I have none of the necessary skills to actually understand any language that the authors spoke, even less the manuscript itself.
Questions:
1. The whole analysis is based on IT2a-n.txt from You are not allowed to view links. Register or Login to view. Is this the correct choice? As far as I understand, it's a version of the TT transcription, but I don't know what the state of the art is. I noticed that the transcription used on voynichese.com is different and in some places more complete, but I don't know.
2. How is the interaction between "fan groups" like this one and the academic community? From what I saw in the videos, there is a fairly good collaboration, but I still wonder. I know that for some topics that garner so much public interest, there can be a lot of tension.
3. I had trouble finding a "definitive" distinction of pages into Currier A and B. I don't know if that's because it is not fully defined, if there is disagreement or if I just didn't look at the right places.
Base results:
Before we get to the good stuff, I want to post the base analysis as a sanity check but also so they are all in one spot. As I said, these are all well known but it's good for me to see the data myself, so maybe also for others.
I split up all the analysis steps by Currier A and B. Going in, I did not have any idea how close both languages are. My initial assumption was actually that they are as different as German and Latin. These stats helped me understand it better.
[EDIT: I made a mistake in my Currier A/B separation for these plots. Corrected plots in my reply on page 2.]
1. The word length plots for Currier A and B with the distributions for 4 reference languages. (I just chose 4 languages that were easily accessible to me.)
2. The known Zipf distribution of word frequencies with reference languages
3. Bigram heatmap.
This one was interesting to me because it shows a very close correspondence between Currier A and B. I expected a much bigger variation.
As a reference looked at the bigram statistics of the reference languages and one can see that they vary much much more from each other, compared to the currier A and B.
4. Word start and end bigrams/trigrams
The bigrams and word-initial trigrams did not show that much irregularity but the word-end trigrams clearly shows the famous -edy ending for Currier B
What did surprise me is that the -edy ending is also among the most common endings in the Currier A script. From what I read and saw, I assumed that it is almost exclusive to the Currier B. Does that mean that I (a) simply misunderstood or (b) chose the wrong page split between currier A and B?
My current split is defined by the code below. Input is very welcome.
I'll leave it at that for now. I'm curious how the interaction in the community here works and if I'll hear from anybody. I'm still preparing the plots for the TF-IDF analysis, as I said, I actually think they are quite interesting. I will add them as reply when I'm done.
Until then, cheerio,
Marvin.
Last February, I made a breakthrough in my research on the Voynich manuscript and was able to identify its original language. It has been a long journey toward complete translations, but now I can translate entire sentences from the work. I base my translations directly on grammar, using meanings drawn from a dictionary.
Today I wanted to test how well an AI can analyze a text without any background information. To do this, I took a short passage from my own translation and gave it to the AI in complete isolation. I did not reveal which text it came from, what language it was originally written in, or what era it belonged to. The AI was only given the words and their basic meanings.
The purpose of the experiment was to see whether an AI could draw conclusions based solely on the text itself – without any clues about its origin.
The result was surprisingly successful. The AI was able to determine which language group the text belonged to, even though it knew nothing about its background. It didn’t need the original manuscript, the context, or even the knowledge that it was a translation. A single text fragment was enough.
The experiment showed that AI can recognize linguistic and cultural features even when the text is completely disconnected from its original environment. At the same time, it confirmed that my translation is not arbitrary, but based on a correct interpretation of the source text.
Reading Rene Zandbergen's blog, a few things jumped to my mind. In You are not allowed to view links. Register or Login to view., he retraces the most probable evolution of the manuscript in these terms :
Quote:The points that have been presented in relation to the order of production of the MS may now be summarised.
The MS was produced in a bifolio-per-bifolio manner, with the drawing outlines inked first, followed by the inking of the text;
The quire numbers were added before the folio numbers;
The page order has been disturbed, and this happened before both sets of numbers were added;
The painting was done before the present binding;
The quire and folio numbers were added before the present binding;
Some of the painting appears to have been done after the folio numbers were added;
Twelve of the fourteen missing folios were lost after the folio numbers were added, but before the present binding.
This leads to the following tentative reconstruction:
All bifolios of the MS were prepared: the drawing outlines and the text were added in ink;
Sometime after this, the planned order of the bifolios was disturbed. The bifolios were stacked anew in an incorrect order (implying that the person who did this was not the original author) but the set was still complete. (The interesting task of identifying the original page order has not been completed, and has mainly been driven by Nick Pelling);
The quires were numbered first, the MS may have been bound, and the folios were numbered after that. (This initial binding is not necessary but would explain the inconsistency of the quire and folio numbers of quire 9);
At this point, the book had all folios including the now missing ones, and was not painted, or only partially painted. Folio 42 would not have been painted yet;
The MS was disassembled and painted (or the partial painting completed). Six bifolios were lost or removed at this point;
Shortly after the painting, the MS was rebound in the same order, but with the six bifolios missing. Folios 12 and 74 would have still been there. Especially the blue paint transferred on opposite pages;
Folios 12 and 74 were cut out sometime later
We know that the quire numbers were added before the folio numbers, and that this indicates the presence of a first binding, or at least that the manuscript was prepared for binding (the same page mentions earlier that the marks on q9 only show a preparation for binding but no trace of finishing it at this point).
Is there ANY reason that this first preparation for binding might prepare quires consisting of only one bifolio, and that it woud put those single-bifolio quires at any other point than the edges of the finished product ? If there is not, it indicates that q16 and q18 were composed of 4 bifoliae each, like most of the others, but that 3 of those had disappeared by the time the folio numbers were added (Looking further, I suppose it is likely that those pages were foldouts, like q14 through q19 have in abundance, but this doesn't prevent the existence of more missing unnumbered bifoliae)
Then, for which reason would the first preparator prepare uneven quires (I can see two of them, but none can apply to q8 : either a clear semantic/stylistic link, which explains q20 but q8's remaining bifoliae aren't clearly semantically tied, or the physical unwieldyness of long quires with foldouts, which explains q14 through 19 but can't explain LONGER than usual quires) ? q8 is longer than all other quires (except q20, with its very different text layout than the rest), and as long as q13, which is stylistically coherent, but f57, 58, 65 and 66 are quite different to each other, and they aren't even consistent recto to verso (f57r and You are not allowed to view links. Register or Login to view. can be in the same section, but they are clearly different from the pair f57v-f66r). Currier finds both Language A and Language B in this quire, and the images look to belong in different sections, which indicates one of the following :
There is a hidden semantic connection justifying to join together bifoliae like that, and the quirer understood the language (very unlikely, as Lisa Fagin Davis' work tends to suggest that quiring itself was a misunderstanding of the book, which should have stayed as a collection of loose leaves, or should have been quired as a thick pile of singulions)
The missing bifoliae contain drawings and text bridging the gap (possible, but unhelpful)
Q8 was from the start a patchwork quire, gathering everything that doesn't fit (this indicates that q16 and q18 were bigger than a sigular bifolio without foldouts each, as else they could have been joined into q8 and the resulting quire would still not have been thicker than q20, which by its existence, shows that quires this big are practical ; it doesn't explain, though, why it would have been numbered this low, rather than being put at the end)
All quires were initially this big and we shouldn't read into 8's length (not really realistic, as it means 7 bifoliae are missing, one in each of the first 7 quires; the most probable outcome would have been to have unequal quires at the start)
The most probable outcome, for me and for now, is the proposition 3 : q16 and q18 were longer than one standard bifolio each, but all the unnumbered ones were lost between quiring and foliating. I still don't have a good idea of why the quirer would create distinct-length quires in the middle of the book rather than counting the extra leaves at the end of the quiring process, but that might be tied to the process itself, in which case I'd love an idea
Rather than the usual type of steganography, based on the placement of words on a page, the VMs has been equipped with a system based on the placement of extended text segments. Rather than 'hiding" words on a page, the VMs "hides" text segments in circular diagrams. Rather than a grille to find the hidden words, the VMs used patterned markers to designate selected text segments. A grille is too obvious and potentially lost. Patterned markers are more subtle, they might be irrelevant, and they stay in place. A functional structure was created, but was it activated?
Examples of marked text segments occur in the cosmos, zodiac, rosettes, etc. Markers occur in a number of variations. While some are elaborate and obvious, do single lines or blank spaces also constitute 'markers'?
The problem is that designated text is still Voynichese writing that cannot be read. However, the advantage is that this provides specific segments of text to compare and contrast with statistical investigations.
As to whether the artist actually recognized this steganographic technique and chose to provide verification in the VMs illustrations, the answer is provided in VMs White Aries where the text marker joins the blue-striped tub.
2 minutes of hard work writing the prompt mentioning morphemes and the scuola medica salernitana.
Then I clicked the "generate essay" button.
The LLM cites references, not bad:
Quote:Timm (2014, 2016) demonstrates that the majority of Voynichese word to‑
kens can be organized into multidimensional grids, each centered on a core
morpheme. For example, the “daiin” series contains variations such as
“daiin,” “aiin,” “dain,” “ain,” and so forth, with substitutions and
deletions governed by a set of morphophonemic rules:
・ “in,” “iin,” and “iiin” are interchangeable (e.g., “daiin” ↔
“dain”).
・ “ch” and “sh” substitute for each other (“chedy” ↔ “shedy”).
・ “o,” “a,” and “y” interchange in certain positions (“ol” ↔
“al” ↔ “yl”).
・ “k,” “t,” “p,” and “f” are variant forms, often substituting in
medial or terminal positions.
A little bit vague about the decryption of the recipes, but nice try:
Quote:For instance, a typical recipe paragraph in the VMS may be parsed into a
sequence: [verb morpheme] + [plant name morpheme] + [preparation mor‑
pheme] + [application morpheme], with variations reflecting standard me‑
dieval medical formulae. The frequent recurrence of such patterns, com‑
bined with the alignment of morphemes to known medical vocabulary, sup‑
ports the claim of successful decryption.
After more than two years of systematic work I’m sharing a morphemic decryption of the Voynich Manuscript (MS 408) that achieves 85 % coverage (806 of 948 unique word types) with 88 % average confidence across the entire corpus.
Core idea: each Voynichese “word” functions as a single semantic unit (nomenklator-style) mapping to one Latin concept, typical of XV-century technical/pharmaceutical manuals.
Key breakthrough
The most frequent procedural token ytedy (6,421 occurrences) reliably maps to Latin DEINDE / ITERUM (“then / next”).
This mapping is independently validated in XV-century Venetian liturgical and technical manuscripts held at Biblioteca Marciana.
Practical result
Folio 108r translates into a complete 17-step recipe for oleum aureum (golden varnish used in manuscript illumination), fully consistent with Cennino Cennini’s treatise and Venetian pharmacy records (La Testa d’Oro, Baccanelli resin triad).
Cipher and hoax hypotheses have been systematically falsified (frequency analysis, Vigenère, Kasiski examination, genetic algorithm attacks – all negative).
Everything is fully public and reproducible:
• GitHub repository – complete Python code + dataset
You are not allowed to view links. Register or Login to view.
• DOI (concept)
10.5281/zenodo.17617392
• Full dataset (41,912 words, 119,278 morphemes)
You are not allowed to view links. Register or Login to view.
• Academic paper (7 pages)
You are not allowed to view links. Register or Login to view.
Attached:
1. Executive Summary with all statistics
2. Title page + abstract
3. Heat-map of morpheme co-occurrence
4. Example translation of folio 108r
I’m very open to scrutiny and independent verification – just clone the repo and run the scripts.
Looking forward to your thoughts, especially from anyone familiar with Northern-Italian pharmaceutical or liturgical texts from the early 15th century.