24-08-2026, 04:57 AM
So, some of you may remember my work where I created a 'statistical' generator using what I called a ledger. You can find it here: You are not allowed to view links. Register or Login to view.
Well, I have discovered why it didn't exactly inspire anyone. My logic was way off. I have since been improving on my methods and digging deeper into the possibilities of the ledger and I believe I have something that works MUCH better. Before I bury you in data I'll have to state that I have a LOT of information to impart so it will take several posts to explain all of this. I'll start with the big news first then the explanations.
Before I go any further, I want to give credit where it is due. This work follows and expands on the work of Torsten Timm and Andreas Schinner. Their self-citation model proposes that Voynich words were produced by copying existing words and modifying them. This is the foundation of what I have been calling “copy and mutate.” So if you're hoping for a decoding, TLDR; it's not here.
Positionless ledger
This simple ledger is capable of creating and verifying any word in the Voynich, written by any scribe, with the exception of gallows words and hapax. I have intentionally removed those two categories in order to hopefully simplify and explain how this ledger works.
The first thing you'll notice in that table is likely that I have combined CH and SH into one atomic unit. There's also reasons for that which I'll explain later.
How to use the ledger:
Let's assume you want to create the word daiin.
The ledger contains everything needed to reproduce:
That is around 36% of the total Voynich tokens.
Prefix Stacking.
Consider these Voynich words
All four appear in the manuscript.
What immediately stands out is that they share the same ending: aiin. The difference is what has been added in front of it:
This is what I mean by prefix stacking. Instead of creating five completely unrelated words, the system appears to reuse the same base while building a longer structure to its left.
The longest word can be written as:
The shorter words reuse parts of that same chain. This gives us transitions such as:
If you look how the ledger works, each time you pick a new midfix, you go back to the prefix column. It is essentially stacking prefixes in front of each new midfix you choose.
The ledger does not need a separate rule for every complete word. The same transitions are reused as prefixes are added to familiar words. This does not prove that the words were created in this exact order. It shows that the words form a clear family built around the same ending, with progressively larger prefixes.
And for those of you who suspect that the Voynich has Hebrew roots, I have good news: the ledger also works from right to left.
Start with a suffix. Find a prefix row that permits that suffix, then continue working backward by adding permitted midfixes until you reach the prefix where you want to stop.
For example:
You have now built daiin backward. Every word covered by this ledger can be reconstructed in this direction. Naturally, this proves absolutely nothing about Hebrew. But the joke was sitting there, and I was not going to waste it.
Another example of prefix stacking
Consider this family:
All four words appear in the manuscript. None is a hapax or gallows word. At each step, one new atom is placed in front of the previous word:
In ledger form:
The complete chain requires only:
Every one of those transitions appears in the core ledger. This is a particularly clear example of prefix stacking. The word chedy remains intact while L, then O, then D are stacked in front of it.
Thats is exactly the kind of nested word family that prefix stacking would produce and the word families everyone keeps seeing in the Voynich.
Missing Letters?
You may have noticed some dashes in the ledger and some letters with no midfix or suffix and, that is correct. H for example. It has no midfix as it never appears as a prefix. And it has no suffix as it's never the letter before the suffix. Now you would think that choldy would give it a midfix of o. However, H never appears without a C or an H in front of it (keep in mind, no gallows and no hapax tokens). But, it's included because in the full ledger, which I'll describe in the next post, that is going to change.
Which brings me to the next section:
Why I treat CH and SH as atoms
In my earlier ledgers, I treated every transcription character separately. That meant CH was represented as C followed by H, and SH was represented as S followed by H.
For example:
This created a problem in my original table. Every time another prefix was stacked onto the word, C and H were pushed one position farther into the midfix:
I therefore had to keep track of first midfix, second midfix, third midfix and so forth. The ledger became larger and more complicated as words became longer. Eventually I realized that much of this apparent midfix depth was being created by the way I had divided CH and SH. I was splitting units that behaved more cleanly when kept together.
I changed the atomization so that CH and SH are each handled as one ledger atom:
Now the same word family can be described with a few reusable transitions:
The ledger no longer needs to know whether CH is in the first, second, third or fourth midfix position. It only needs to know which atoms can follow CH and which atoms can lead into it.
SH works the same way.
This does not mean that C, S and H have been removed. They remain separate atoms when they occur separately. Only the exact sequences CH and SH are kept together. “Atom” here simply means that the ledger handles the sequence as one unit. The important result is that once CH and SH are treated this way, the positional midfix columns are no longer necessary. The same information can be represented by a much smaller positionless ledger. What had looked like complicated midfix depth was, in large part, the same units being pushed farther to the right by added prefixes.
Eliminating the first argument: Wouldn't any language look like this?
Yes, a positionless ledger can be constructed from any language. For comparison, I chose Finnish. Finnish is an agglutinative language: it commonly builds words by joining smaller elements together. That makes it a useful control for the kind of word-building process I am proposing for Voynichese.
I used the first 12,564 non-hapax Finnish tokens, the same number of tokens represented by the Voynich ledger above.
So yes, Finnish can be placed into the same kind of ledger. But it requires far more word types and far more transitions to describe the same number of tokens. The ledger format itself is not unique to the Voynich Manuscript. What is unusual is how much of the Voynich vocabulary can be described by such a small set of reusable transitions.
Finnish positionless ledger: Seitsemän veljestä
The hook to read the next post: Gallows are lipstick!
And, I'll also present the full ledger, explain why gallows words and hapax words are omitted from the ledger above and hopefully, give some convincing evidence as to why they were omitted.
Thanks. I have donned my flame retardant clothing, fire away.
Well, I have discovered why it didn't exactly inspire anyone. My logic was way off. I have since been improving on my methods and digging deeper into the possibilities of the ledger and I believe I have something that works MUCH better. Before I bury you in data I'll have to state that I have a LOT of information to impart so it will take several posts to explain all of this. I'll start with the big news first then the explanations.
Before I go any further, I want to give credit where it is due. This work follows and expands on the work of Torsten Timm and Andreas Schinner. Their self-citation model proposes that Voynich words were produced by copying existing words and modifying them. This is the foundation of what I have been calling “copy and mutate.” So if you're hoping for a decoding, TLDR; it's not here.
Positionless ledger
| Prefix | Core midfix | Rare midfix | Suffix |
| A | L, R, I | D, M | L, S, R, M, N |
| O | A, O, D, L, CH, SH, E, S, R, I | — | D, L, Y, S, R, M, G |
| D | A, O, CH, SH, E | D, L, Y | O, L, Y, R |
| L | A, O, D, CH, SH | L, E, Y, S | O, D, L, Y, S, R, G |
| CH | A, O, D, E | L, CH, SH, S | A, O, D, L, E, Y, S, R, G |
| SH | A, O, D, CH, E | L, S | O, D, L, E, Y, S, R |
| E | A, O, D, E | CH, Y, S | O, D, E, Y, S, R, N, G |
| Y | A, D, CH, SH | O, L, S | S, R |
| S | A, O, CH, SH, E | D | O, E, Y |
| R | A, O, CH | D, SH, E, I | O, L, Y |
| Q | O | E | — |
| H | — | — | — |
| M | — | — | O |
| N | — | D | O, Y |
| G | — | — | — |
| X | — | A | — |
| C | S | — | — |
| I | D, R, I | L, N | L, S, R, M, N |
This simple ledger is capable of creating and verifying any word in the Voynich, written by any scribe, with the exception of gallows words and hapax. I have intentionally removed those two categories in order to hopefully simplify and explain how this ledger works.
The first thing you'll notice in that table is likely that I have combined CH and SH into one atomic unit. There's also reasons for that which I'll explain later.
How to use the ledger:
Let's assume you want to create the word daiin.
- Look at the prefix column for the row that contains D.
- In the core midfix column you'll see you can choose A. You now have D→A
- Now look in the prefix column for the row that contains A.
- Look in the core midfix column and you'll see you can choose I. You now have D→A→I
- Look in the prefix column for the row that contains I.
- Look in the core midfix column and you'll see that you can also choose I as a midfix. You now have D→A→I→I
- You'd again look in the I prefix row and in the suffix column you can choose N
- You now have the word D→A→I→I→N
The ledger contains everything needed to reproduce:
- 393 core word types
- 412 rare word types
- 805 total word types
- 12,564 manuscript tokens
That is around 36% of the total Voynich tokens.
Prefix Stacking.
Consider these Voynich words
- aiin
- daiin
- odaiin
- chodaiin
All four appear in the manuscript.
What immediately stands out is that they share the same ending: aiin. The difference is what has been added in front of it:
- aiin
- d + aiin
- od + aiin
- chod + aiin
This is what I mean by prefix stacking. Instead of creating five completely unrelated words, the system appears to reuse the same base while building a longer structure to its left.
The longest word can be written as:
Code:
CH → O → D → A → I → I → NThe shorter words reuse parts of that same chain. This gives us transitions such as:
- CH→O
- O→D
- D→A
- A→I
- I→I
- I→N
If you look how the ledger works, each time you pick a new midfix, you go back to the prefix column. It is essentially stacking prefixes in front of each new midfix you choose.
The ledger does not need a separate rule for every complete word. The same transitions are reused as prefixes are added to familiar words. This does not prove that the words were created in this exact order. It shows that the words form a clear family built around the same ending, with progressively larger prefixes.
And for those of you who suspect that the Voynich has Hebrew roots, I have good news: the ledger also works from right to left.
Start with a suffix. Find a prefix row that permits that suffix, then continue working backward by adding permitted midfixes until you reach the prefix where you want to stop.
For example:
- N
- I → N
- I → I → N
- A → I → I → N
- D → A → I → I → N
You have now built daiin backward. Every word covered by this ledger can be reconstructed in this direction. Naturally, this proves absolutely nothing about Hebrew. But the joke was sitting there, and I was not going to waste it.
Another example of prefix stacking
Consider this family:
- chedy 500 occurrences
- lchedy 108 occurrences
- olchedy 34 occurrences
- dolchedy 3 occurrences
All four words appear in the manuscript. None is a hapax or gallows word. At each step, one new atom is placed in front of the previous word:
- chedy
- L + chedy = lchedy
- O + lchedy = olchedy
- D + olchedy = dolchedy
In ledger form:
- CH → E → D → Y
- L → CH → E → D → Y
- O → L → CH → E → D → Y
- D → O → L → CH → E → D → Y
The complete chain requires only:
- D→O
- O→L
- L→CH
- CH→E
- E→D
- D→Y
Every one of those transitions appears in the core ledger. This is a particularly clear example of prefix stacking. The word chedy remains intact while L, then O, then D are stacked in front of it.
Thats is exactly the kind of nested word family that prefix stacking would produce and the word families everyone keeps seeing in the Voynich.
Missing Letters?
You may have noticed some dashes in the ledger and some letters with no midfix or suffix and, that is correct. H for example. It has no midfix as it never appears as a prefix. And it has no suffix as it's never the letter before the suffix. Now you would think that choldy would give it a midfix of o. However, H never appears without a C or an H in front of it (keep in mind, no gallows and no hapax tokens). But, it's included because in the full ledger, which I'll describe in the next post, that is going to change.
Which brings me to the next section:
Why I treat CH and SH as atoms
In my earlier ledgers, I treated every transcription character separately. That meant CH was represented as C followed by H, and SH was represented as S followed by H.
For example:
- chedy = C → H → E → D → Y
- lchedy = L → C → H → E → D → Y
- olchedy = O → L → C → H → E → D → Y
- dolchedy = D → O → L → C → H → E → D → Y
This created a problem in my original table. Every time another prefix was stacked onto the word, C and H were pushed one position farther into the midfix:
- chedy H in the first midfix position
- lchedy H in the second midfix position
- olchedy H in the third midfix position
- dolchedy H in the fourth midfix position
I therefore had to keep track of first midfix, second midfix, third midfix and so forth. The ledger became larger and more complicated as words became longer. Eventually I realized that much of this apparent midfix depth was being created by the way I had divided CH and SH. I was splitting units that behaved more cleanly when kept together.
I changed the atomization so that CH and SH are each handled as one ledger atom:
- chedy = CH → E → D → Y
- lchedy = L → CH → E → D → Y
- olchedy = O → L → CH → E → D → Y
- dolchedy = D → O → L → CH → E → D → Y
Now the same word family can be described with a few reusable transitions:
- D→O
- O→L
- L→CH
- CH→E
- E→D
- D→Y
The ledger no longer needs to know whether CH is in the first, second, third or fourth midfix position. It only needs to know which atoms can follow CH and which atoms can lead into it.
SH works the same way.
This does not mean that C, S and H have been removed. They remain separate atoms when they occur separately. Only the exact sequences CH and SH are kept together. “Atom” here simply means that the ledger handles the sequence as one unit. The important result is that once CH and SH are treated this way, the positional midfix columns are no longer necessary. The same information can be represented by a much smaller positionless ledger. What had looked like complicated midfix depth was, in large part, the same units being pushed farther to the right by added prefixes.
Eliminating the first argument: Wouldn't any language look like this?
Yes, a positionless ledger can be constructed from any language. For comparison, I chose Finnish. Finnish is an agglutinative language: it commonly builds words by joining smaller elements together. That makes it a useful control for the kind of word-building process I am proposing for Voynichese.
I used the first 12,564 non-hapax Finnish tokens, the same number of tokens represented by the Voynich ledger above.
| Ledger | Word types | Core midfix transitions | Rare midfix transitions | Suffix transitions |
| Voynich | 805 | 53 | 31 | 63 |
| Finnish | 3,919 | 277 | 20 | 140 |
So yes, Finnish can be placed into the same kind of ledger. But it requires far more word types and far more transitions to describe the same number of tokens. The ledger format itself is not unique to the Voynich Manuscript. What is unusual is how much of the Voynich vocabulary can be described by such a small set of reusable transitions.
Finnish positionless ledger: Seitsemän veljestä
| Prefix | Core midfix | Rare midfix | Suffix |
| A | A, D, E, H, I, J, K, L, M, N, P, R, S, T, U, V | — | A, I, N, R, S, T |
| D | A, E, I, O, Ä | — | A, E, O, U, Ä |
| E | A, D, E, H, I, K, L, M, N, O, P, R, S, T, U, V, Y, Ä | J | A, E, H, I, N, R, S, T, Ä |
| F | R | L, Ö | — |
| G | A, I, O | E | — |
| H | A, D, E, H, I, J, K, L, M, N, O, T, U, V, Y, Ä, Ö | — | A, E, I, O, T, U, Ä |
| I | A, D, E, H, I, J, K, L, M, N, O, P, R, S, T, U, V, Ä | — | A, E, H, I, N, O, S, T, Ä |
| J | A, E, I, O, U, Y, Ä | — | A, U, Ä |
| K | A, E, I, K, M, O, R, S, U, Y, Ä, Ö | N | A, E, I, K, O, S, U, Y, Ä, Ö |
| L | A, E, H, I, J, K, L, M, O, P, S, T, U, V, Y, Ä, Ö | — | A, E, I, L, O, U, Y, Ä |
| M | A, E, I, M, O, P, S, U, Y, Ä, Ö | — | A, E, H, I, O, U, Ä |
| N | A, E, G, H, I, K, L, N, O, P, S, T, U, Y, Ä | J, M, V | A, E, I, O, Ä |
| O | A, D, E, H, I, J, K, L, M, N, O, P, R, S, T, U, V | — | A, H, I, N, O, S, T |
| P | A, E, I, O, P, S, U, Y, Ä, Ö | L, R | A, I, O, U, Ä |
| R | A, E, H, I, J, K, M, N, O, P, R, S, T, U, V, Y, Ä | Ö | A, I, O, S, Y, Ä |
| S | A, E, I, K, O, P, S, T, U, V, Y, Ä, Ö | M, N | A, E, I, O, S, T, U, Y, Ä |
| T | A, E, H, I, K, O, P, R, S, T, U, Y, Ä, Ö | V | A, E, I, O, U, Y, Ä, Ö |
| U | A, D, E, H, I, J, K, L, M, N, O, P, R, S, T, U, V | — | A, E, I, N, O, S, T, U |
| V | A, E, I, O, U, Ä | Y, Ö | A, E, I, O, U, Ä, Ö |
| Y | D, E, H, I, K, L, M, N, P, R, S, T, V, Y, Ä, Ö | — | I, N, S, T, Y, Ä, Ö |
| Ä | D, E, H, I, J, K, L, M, N, P, R, S, T, V, Y, Ä | — | E, H, I, N, R, S, T, Y, Ä |
| Ö | H, I, J, K, M, N, R, S, T, Y, Ö | D, L, P, V | I, N, S, T, Ä |
The hook to read the next post: Gallows are lipstick!
And, I'll also present the full ledger, explain why gallows words and hapax words are omitted from the ledger above and hopefully, give some convincing evidence as to why they were omitted.
Thanks. I have donned my flame retardant clothing, fire away.
).