The Voynich Ninja

Full Version: Rethinking the Glyph Inventory/Tokenization: Are the Gallows Compositional?
You're currently viewing a stripped down version of our content. View the full version with proper formatting.
I've been manually re-tokenizing selected VMS passages using a compositional glyph model, and in doing so I've found myself questioning some of the conventional assumptions underlying EVA and other attempts to tokenize the manuscript. I don't mean that as a blanket criticism of those systems—they were designed for useful transcription—but I think some of their assumptions about what constitutes an individual glyph may need to be re-evaluated. For that reason, I am not going to use the standard EVA “lettering” here and will instead rely on descriptors.

The main difference I've noticed is what I'm calling the c-shaped joining crossbars. Other researchers have generally treated the four tall, looped glyphs, with or without crossbars, as a collection of eight distinct glyphs. But when I look closely at the manuscript, there are many examples where the left- and right-hand crossbar elements do not fully connect or continue through the upright. Those left- and right-hand c-and-bar-shaped crossbar elements can precede and follow one of the four upright characters, or, less commonly, only precede or only follow one. The point is that these combinations should not necessarily be thought of as atomic characters, but rather as compound structures.

They also appear to combine with one another, forming a separate double-c with a top bar that is often treated as a unique glyph, and they may or may not carry the small diacritical hook. This makes me wonder whether we're looking at independent components rather than eight atomic glyphs.

I would therefore remove the four crossed looped tall characters, the crossbar double-c, and the crossbar-double c with the overhead hook from the atomic character list, and instead introduce left- and right-hand crossbar modifiers with the constrained rules discussed above.

I also question some of the other positionality constraints others have raised and have noted some things of interest.
I've visually checked the manuscript quite extensively with regard to the figure-4-looking character that is often said to occur independently of the following o-like character. I haven't been able to find many convincing examples. The cases I can identify reliably seem to be degraded, overwritten, or partially formed text. That isn't proof that genuine exceptions don't exist, but so far the ambiguous examples aren't convincing to me. There are some standalone occurrences in the marginalia, where the figure-4 form appears by itself, but I'm not sure those should be treated the same way as the running text.

So far, the figure-4 seems to combine with and modify the o-shape to create an initiator complex. I have not found a convincing instance of that character appearing inside a word. Except for some standalone marginalia, I have found the figure-4 shape to be consistently followed by the o-shape in the running text. This appears to be a hard positional rule. I have examined claims of an “i,” “e,” or “c” following the figure-4, but the regions are smudgy and unclear, so I am not convinced. If someone can clearly point out exceptions, I would certainly like to know, so that this can accurately be classified as “almost always” instead of “always.”

The same goes for what I now consider to be the d-shaped terminator mark. It only appears at the end of a word and never inside. Many researchers have already noted it to be positionally restricted. I find it consistently word-final. However, it can be combined with the c-shaped forms and with the undotted-i-like minima in short clusters. From a reasonably large catalog, I can say with some confidence that the c-shaped form occurs once, twice, or three times before the terminator.

Likewise, and here is the interesting part, the minima commonly appears before the terminator once or twice, with three occurrences being rarer; I haven't found convincing examples of longer runs. The minima, however, combines in pen stroke with the d-shaped terminator so closely together that the resulting shape can resemble an “n,” an “m,” or even two “u” or “w” forms. I suspect these should be treated as a carefully constrained terminal-sequence rule.

I therefore propose that the m- and n-like shapes be removed from the token list rather than automatically treated as separate complex glyphs. Similar to the figure-4 being reclassified as a special-case initiator, the d-shaped terminator should instead be treated as a positionally restricted component.

There are softer rules, such as the figure-9-shaped character. It is heavily positionally biased toward the end of words, but can appear much less often at the beginning of words and even less often in the middle. There are others too, but they do not seem to fundamentally affect the core symbol model I am proposing here.

For example, I also find that the small diacritical hook frequently appears above the left and right crossbar components when they are conjoined. It can occur in other positions as well, although much less frequently, and there seems to be some positional structure to its distribution that I haven't worked out yet. For now, I would still classify it as a distinct modifier/component rather than simply part of a particular glyph.

The oft-referred-to “picnic table” character appears, in my examination, to be a standalone character that does not occur inside running words. I would consider it something like a rare pictogram that sits outside the coded text.
My aim isn't to claim that these observations establish a particular interpretation of the manuscript. Rather, I'm wondering whether a more explicitly compositional tokenization might produce a cleaner representation of the underlying system, and whether that could affect subsequent cryptanalytic analysis.

Final note: I have also begun cataloguing the distribution of the four upright forms at paragraph openings and other paragraph breaks. Their distribution is strikingly non-uniform, even apart from the occasional illuminated flourishes at these locations. For example, in the plant section from f1r–f57r, 206 of 219 identifiable paragraphs (94%) begin with a gallows; among those, the double-looped P accounts for 85, T for 66, K for 40, and F for 15. In Q20/Stars, the distribution among distinct paragraph-initial gallows words is similarly skewed: P 58%, T 26%, K 9%, F 7%. I don't yet know what this distribution means, but I think it is important structural information that should be preserved rather than collapsed into a single “gallows” category. I could not find any other clear character-level exceptions, aside from some difficult-to-interpret elaborate crossbar looping.
The positional effect is partly due to CVS. Some symbols cannot be combined with others because they have different shapes.
Also, in VMS, you can find an exception for everything, even for the rule that I wrote. But they’re not that significant...

I am trying to find explanations for such positional anomalies from the perspective of Roman numerals.
I am doing something similar, just yesterday i started making my own transcription of quire 13. One line done so far, oh boy! But as I go I am describing the many ways I see each glyph, including whether they could be seen as constructions, such as the thing that looks like an a but that looks like a attached to an and whether they touch another or stand alone, etc. I want such notes in my transcription to cover all possibilities rather than rely on eva, even in my one line so far there is at least one place I would change the eva transcription from an r to an s or vice versa, can't remember offhand.

I forget which one I was comparing to, but I plan to go back and compare all the ones I can find, in case there are other considerations to add.
@ololololo what does CVS stand for? I tried to look it up in the forum and on the internet but the highest hits are multiple curriculum vitae and some kind of pharmacy.
(05-09-2026, 01:23 PM)Linda Wrote: You are not allowed to view links. Register or Login to view.I am doing something similar, just yesterday i started making my own transcription of quire 13. One line done so far, oh boy! But as I go I am describing the many ways I see each glyph, including whether they could be seen as constructions, such as the thing that looks like an a but that looks like a attached to an and whether they touch another or stand alone, etc. I want such notes in my transcription to cover all possibilities rather than rely on eva, even in my one line so far there is at least one place I would change the eva transcription from an r to an s or vice versa, can't remember offhand.

I forget which one I was comparing to, but I plan to go back and compare all the ones I can find, in case there are other considerations to add.
@ololololo what does CVS stand for? I tried to look it up in the forum and on the internet but the highest hits are multiple curriculum vitae and some kind of pharmacy.
You are not allowed to view links. Register or Login to view.

For me, CLS is associated with a certain rule, according to which in most words, characters of type Curve come first, followed by characters of type Line, or the word consists only of Curve elements. Example: daiin (rule 1), chedy (rule 2).

And I was wrong — it’s CLS, not CVS Smile
bablackthorne
bablackthorneYou can find many examples of deviations from standard character spelling in the thread. You are not allowed to view links. Register or Login to view.
You can find statistics on character combinations (in particular, the character following "4") using the program: You are not allowed to view links. Register or Login to view.  The text even contains the combination 4040 at the beginning of a word.
Ultimately, you will (most likely) arrive at my transcription in elementary strokes. You are not allowed to view links. Register or Login to view.