Part 1 can be found here: You are not allowed to view links.
to view.
I know there's a lot of data in this post and I tried to summarize as best as I could. I tried to break it down into sections to make it more readable.
I'll start with the full Voynich ledger which shows the legal letter transitions for the entire corpus, including gallows.
| Prefix | Core midfix | Rare midfix | Suffix |
| a | l, r, i | a, o, d, ch, sh, e, y, s, m, n, x, c, f, k, p, t | d, l, sh, y, s, r, m, n, g, i, p, t |
| o | a, o, d, l, ch, sh, e, s, r, i | y, q, h, m, g, x, c, f, k, p, t | o, d, l, sh, e, y, s, r, m, n, g, x, c, f, k, p, t |
| d | a, o, ch, sh, e | d, l, y, s, r, x, c, i, f, k, p, t | a, o, d, l, sh, e, y, s, r, m, g |
| l | a, o, d, ch, sh | l, e, y, s, r, q, m, g, x, c, i, f, k, p, t | a, o, d, l, sh, e, y, s, r, m, g, k, p |
| ch | a, o, d, e | l, ch, sh, y, s, r, h, x, c, i, f, k, p, t | a, o, d, l, e, y, s, r, m, g, x, k |
| sh | a, o, d, ch, e | l, sh, y, s, h, x, c, f, k, t | a, o, d, l, e, y, s, r, x |
| e | a, o, d, e | l, ch, sh, y, s, r, h, m, g, c, i, f, k, p, t | a, o, d, l, ch, e, y, s, r, m, n, g, f, k, p, t |
| y | a, d, ch, sh | o, l, e, s, r, q, h, c, i, f, k, p, t | o, d, l, y, s, r, m, k, t |
| s | a, o, ch, sh, e | d, y, s, q, c, i, f, k, p, t | a, o, d, l, e, y, s, r, m, n |
| r | a, o, ch | d, l, sh, e, y, r, x, c, i, f, k, p, t | a, o, d, l, ch, e, y, s, r, m, g, k |
| q | o | a, l, ch, e, y, s, c, f, k, p, t | o |
| h | — | a, o, d, ch, e, y, s, h, c, i, k | a, o, d, l, y, s, r, h, g |
| m | — | a, o, d, ch, sh | o, d, y, m, g |
| n | — | a, o, d, r | o, d, e, y, m, n |
| g | — | a, o, e | l, y |
| x | — | a, o | y |
| c | s | o, sh, c, f, k, p, t | o, y, s, f, t |
| i | d, r, i | a, o, l, ch, sh, e, s, h, m, n, c, f, k, p, t | o, d, l, y, s, r, m, n, x, i, p, t |
| f | — | a, o, d, ch, sh, e, y, s, h, c | o, y |
| k | — | a, o, d, l, ch, sh, e, y, s, r, h, c, i | a, o, d, l, ch, sh, y, h, g |
| p | — | a, o, d, ch, sh, e, y, s, h, c, i | a, o, ch, y |
| t | — | a, o, d, l, ch, sh, e, y, s, h, g, c, i, k | a, o, l, ch, sh, e, y, s, h |
| Prefix | Core midfix | Rare midfix | Suffix |
| a | l, r, i | a, o, d, ch, sh, e, y, s, m, n, x, c | d, l, sh, y, s, r, m, n, g, i |
| o | a, o, d, l, ch, sh, e, s, r, i | y, q, h, m, g, x, c | a, o, d, l, ch, sh, e, y, s, r, m, n, g, x, c |
| d | a, o, ch, sh, e | d, l, y, s, r, x, c, i | a, o, d, l, sh, e, y, s, r, m, g |
| l | a, o, d, ch, sh | l, e, y, s, r, q, m, g, x, c, i | a, o, d, l, ch, sh, e, y, s, r, m, g |
| ch | a, o, d, e | l, ch, sh, y, s, r, h, x, c, i | a, o, d, l, ch, e, y, s, r, h, m, g, x, c |
| sh | a, o, d, ch, e | l, sh, y, s, h, x, c | a, o, d, l, e, y, s, r, x |
| e | a, o, d, e | l, ch, sh, y, s, r, h, m, g, c, i | a, o, d, l, ch, e, y, s, r, m, n, g |
| y | a, d, ch, sh | o, l, e, y, s, r, q, h, c, i | a, o, d, l, ch, sh, y, s, r, m |
| s | a, o, ch, sh, e | d, y, s, q, c, i | a, o, d, l, e, y, s, r, m, n |
| r | a, o, ch | d, l, sh, e, y, r, x, c, i | a, o, d, l, ch, e, y, s, r, m, g |
| q | o | a, l, ch, e, y, s, h | o, e |
| h | — | a, o, d, ch, e, y, h | d, y, g |
| m | — | a, o, d, ch, sh | o, d, y, m, g |
| n | — | a, o, d, r | o, d, e, y, m, n |
| g | — | a, o, e | l, y |
| x | — | a, o | y |
| c | s | a, o, ch, sh, e | a, o, y, s |
| i | d, r, i | a, o, l, ch, sh, e, s, h, m, n | a, o, d, l, y, s, r, m, n, x, i |
The above ledgers are for the full corpus. But each scribe had their own subset of this ledger so it became much easier to work with.
Gallows and hapax words do use the core transitions from the ledger, but they also expand the ledger. So, if we make a simple assumption, that the scribe begins with just the ledger for the core words and adds in transitions as they're needed, Scribe 1 starts with this very simple ledger.
While there isn't enough room to show a ledger for each scribe, this heatmap shows how many scribes use each transition. The scribes clearly share a large common set of transitions, but their ledgers are not identical. Each scribe also uses some transitions that the others do not.
The gallows-stripping decision was not arbitrary. It came from several independent tests that all pointed in the same direction. And I'm very likely reproducing some work of others that I'm unaware of, my apologies if I am.
95.68% of all stripped gallows tokens were either an exact earlier word or only one edit away from an earlier word.
With gallows intact, the ledger has 288 midfix transitions and 186 suffix transitions. After stripping f, k, p and t, that drops to 187 midfix transitions and 148 suffix transitions. More importantly, the core and rare transition inventories don't change at all. The core remains 53 midfix and 40 suffix transitions, and the rare remains 31 midfix and 23 suffix transitions.
Stripping can create a new transition where a gallows used to sit between two letters, but across the entire ledger it adds only four new ordered pairs: C→A, C→CH, C→E and Q→H. By comparison, leaving the gallows intact requires 185 gallows-specific midfix and suffix transitions.
So this is what I mean when I say stripping the gallows exposes the underlying ledger. Almost all of the extra transition machinery comes from the gallows themselves. Remove them and nearly all of it disappears while the core and rare ledger remains unchanged.
If you look at the ledger above with gallows included, you may notice that the midfix and suffix transitions for K & T and P & F are very similar. So I asked, if I took all gallows words and swapped the gallows with it's 'pair' how many ledger legal words would that create.
The result is about as close to interchangeable as you could expect. Out of 17,664 possible sibling swaps, 17,636 remain legal, or 99.84%.
It gets even stronger when we look at position. Every initial swap works, 2,377/2,377. Internal swaps work 13,914/13,919 times, or 99.96%. Almost all of the failures are near the end of the word, where suffix restrictions come into play.
And this isn't just something the ledger allows. ZLZB and TTLI independently disagree on a single gallows 113 times. 102 of those 113 disagreements, 90.27%, are exactly these sibling swaps: 85 K/T and 17 P/F. They occur in both directions, K→T and T→K, P→F and F→P.
So the ledger says K/T and P/F can almost always occupy the same structural position, and when the two transcriptions disagree over a single gallows, the disagreement is overwhelmingly K/T or P/F.
One other oddity I noticed in the gallows words is initial i. Across all five scribes there isn't a single retained core word beginning with i. The only six i-initial words are all gallows-bearing. I don't know what that means, but it is another indication that gallows constructions have some peculiarities of their own.
Now, I do personally tend to think of them as decorative but that's just my opinion and it was convenient for me to think of them as such. But, that doesn't mean gallows are meaningless. It doesn't even tell us whether the gallows were added before or after the rest of the word. And while my use of the word lipstick may be a tad dramatic, there are probably a bunch more tests that should be run before Estée Lauder could claim a trademark.
Scribes 4 and 5 do have a limited vocabulary compared to the other scribes so that should be noted. I won't go into extreme depth on these comparisons but suffice it to say that while the scribes did share a large common transition set, they were not identical.
This compares the pooled all-scribes ledger with the five individual scribe ledgers. The numbers show how many scribes use each transition. The orange P marks transitions that enter the core + rare inventory only when the scribes are pooled together. Most of the pooled ledger is therefore made up of transitions already present in one or more individual core + rare ledgers, with only a small number changing status through pooling.
This compares the actual core + rare transition inventories of each scribe using Jaccard similarity. Scribes 1–4 are fairly close to one another, with Scribe 2 and Scribe 3 the most similar at 0.75. Scribe 5 is the clear outlier, sharing only about 0.40–0.53 of its transition inventory with the others. So the scribes are clearly using the same general transition system, but not identical copies of the same ledger.
Hapax look more irregular than repeated words, but most of that irregularity appears to be local to the individual scribe.
Across the five scribes there are 1,817 nongallows scribe-specific hapax. Only 664, or 36.54%, are completely covered by that scribe's own core ledger. But when those same hapax are tested against the pooled all-scribes core, 1,297 of 1,817, or 71.38%, are completely legal. So in most cases the "new" structure in a hapax isn't actually new to the manuscript. It is simply something that particular scribe only used once.
There is also a very nice control using gallows words. If you split stripped gallows words according to whether they occur once or repeatedly:
One-off gallows words behave almost exactly like ordinary hapax, while repeated gallows words behave much more like core words. That suggests rarity is doing more of the work than whether the word contains a gallows.
Again, I have a good bit more data on this but this is the core of what I'm seeing.
In the final post on this topic, I'll provide the repo for my python files so you can reproduce any of these findings with Takahashi or Zandbergen/Landini transcriptions.