Rethinking the Glyph Inventory/Tokenization: Are the Gallows Compositional?
bablackthorne > 01-09-2026, 03:27 PM
I've been manually re-tokenizing selected VMS passages using a compositional glyph model, and in doing so I've found myself questioning some of the conventional assumptions underlying EVA and other attempts to tokenize the manuscript. I don't mean that as a blanket criticism of those systems—they were designed for useful transcription—but I think some of their assumptions about what constitutes an individual glyph may need to be re-evaluated. For that reason, I am not going to use the standard EVA “lettering” here and will instead rely on descriptors.
The main difference I've noticed is what I'm calling the c-shaped joining crossbars. Other researchers have generally treated the four tall, looped glyphs, with or without crossbars, as a collection of eight distinct glyphs. But when I look closely at the manuscript, there are many examples where the left- and right-hand crossbar elements do not fully connect or continue through the upright. Those left- and right-hand c-and-bar-shaped crossbar elements can precede and follow one of the four upright characters, or, less commonly, only precede or only follow one. The point is that these combinations should not necessarily be thought of as atomic characters, but rather as compound structures.
They also appear to combine with one another, forming a separate double-c with a top bar that is often treated as a unique glyph, and they may or may not carry the small diacritical hook. This makes me wonder whether we're looking at independent components rather than eight atomic glyphs.
I would therefore remove the four crossed looped tall characters, the crossbar double-c, and the crossbar-double c with the overhead hook from the atomic character list, and instead introduce left- and right-hand crossbar modifiers with the constrained rules discussed above.
I also question some of the other positionality constraints others have raised and have noted some things of interest.
I've visually checked the manuscript quite extensively with regard to the figure-4-looking character that is often said to occur independently of the following o-like character. I haven't been able to find many convincing examples. The cases I can identify reliably seem to be degraded, overwritten, or partially formed text. That isn't proof that genuine exceptions don't exist, but so far the ambiguous examples aren't convincing to me. There are some standalone occurrences in the marginalia, where the figure-4 form appears by itself, but I'm not sure those should be treated the same way as the running text.
So far, the figure-4 seems to combine with and modify the o-shape to create an initiator complex. I have not found a convincing instance of that character appearing inside a word. Except for some standalone marginalia, I have found the figure-4 shape to be consistently followed by the o-shape in the running text. This appears to be a hard positional rule. I have examined claims of an “i,” “e,” or “c” following the figure-4, but the regions are smudgy and unclear, so I am not convinced. If someone can clearly point out exceptions, I would certainly like to know, so that this can accurately be classified as “almost always” instead of “always.”
The same goes for what I now consider to be the d-shaped terminator mark. It only appears at the end of a word and never inside. Many researchers have already noted it to be positionally restricted. I find it consistently word-final. However, it can be combined with the c-shaped forms and with the undotted-i-like minima in short clusters. From a reasonably large catalog, I can say with some confidence that the c-shaped form occurs once, twice, or three times before the terminator.
Likewise, and here is the interesting part, the minima commonly appears before the terminator once or twice, with three occurrences being rarer; I haven't found convincing examples of longer runs. The minima, however, combines in pen stroke with the d-shaped terminator so closely together that the resulting shape can resemble an “n,” an “m,” or even two “u” or “w” forms. I suspect these should be treated as a carefully constrained terminal-sequence rule.
I therefore propose that the m- and n-like shapes be removed from the token list rather than automatically treated as separate complex glyphs. Similar to the figure-4 being reclassified as a special-case initiator, the d-shaped terminator should instead be treated as a positionally restricted component.
There are softer rules, such as the figure-9-shaped character. It is heavily positionally biased toward the end of words, but can appear much less often at the beginning of words and even less often in the middle. There are others too, but they do not seem to fundamentally affect the core symbol model I am proposing here.
For example, I also find that the small diacritical hook frequently appears above the left and right crossbar components when they are conjoined. It can occur in other positions as well, although much less frequently, and there seems to be some positional structure to its distribution that I haven't worked out yet. For now, I would still classify it as a distinct modifier/component rather than simply part of a particular glyph.
The oft-referred-to “picnic table” character appears, in my examination, to be a standalone character that does not occur inside running words. I would consider it something like a rare pictogram that sits outside the coded text.
My aim isn't to claim that these observations establish a particular interpretation of the manuscript. Rather, I'm wondering whether a more explicitly compositional tokenization might produce a cleaner representation of the underlying system, and whether that could affect subsequent cryptanalytic analysis.
Final note: I have also begun cataloguing the distribution of the four upright forms at paragraph openings and other paragraph breaks. Their distribution is strikingly non-uniform, even apart from the occasional illuminated flourishes at these locations. For example, in the plant section from f1r–f57r, 206 of 219 identifiable paragraphs (94%) begin with a gallows; among those, the double-looped P accounts for 85, T for 66, K for 40, and F for 15. In Q20/Stars, the distribution among distinct paragraph-initial gallows words is similarly skewed: P 58%, T 26%, K 9%, F 7%. I don't yet know what this distribution means, but I think it is important structural information that should be preserved rather than collapsed into a single “gallows” category. I could not find any other clear character-level exceptions, aside from some difficult-to-interpret elaborate crossbar looping.