![]() |
|
[Article] A Glyph Is Not a Letter, a Token Is Not a Word, a Space Is Not a Space - Printable Version +- The Voynich Ninja (https://www.voynich.ninja) +-- Forum: Voynich Research (https://www.voynich.ninja/forum-27.html) +--- Forum: News (https://www.voynich.ninja/forum-25.html) +--- Thread: [Article] A Glyph Is Not a Letter, a Token Is Not a Word, a Space Is Not a Space (/thread-6034.html) |
A Glyph Is Not a Letter, a Token Is Not a Word, a Space Is Not a Space - pfeaster - 20-08-2026 I haven't had a chance to read this recent (August 2026) article by Liudmila Rozanova and Alexander Temerev at all closely yet, but it looks interesting: "A Glyph Is Not a Letter, a Token Is Not a Word, a Space Is Not a Space: What the Units of Voynichese Are Not" You are not allowed to view links. Register or Login to view. - Patrick RE: A Glyph Is Not a Letter, a Token Is Not a Word, a Space Is Not a Space - JoJo_Jost - 20-08-2026 As I said … that basically confirms the VBM rules RE: A Glyph Is Not a Letter, a Token Is Not a Word, a Space Is Not a Space - Koen G - 20-08-2026 So if I understand one if their conclusions correctly, the final letter of a word (I'll keep saying it) is abnormally good at predicting the first letter of the next word. We would expect this property for whole words, like "you" is often followed by "are", which the VM does not do. So the VM displays an abnormally high predictability between edge glyphs, and near-zero token-to-token predictability. (I know people have researched this in the past, but I haven't learned much about it yet.) So I wonder: can we figure out if edge glyph phenomena are caused by selection, or by alteration? Will a certain end glyph attract an existing word with a specific initial glyph (selection) or will it change the initial glyph of the "planned" next word (alteration)? Can this even be calculated? RE: A Glyph Is Not a Letter, a Token Is Not a Word, a Space Is Not a Space - bi3mw - 21-08-2026 === Global edge coupling === H(head_glyph) = 3.0905 bits H(head_glyph | prev_end) = 2.9070 bits I(head_glyph ; prev_end) = 0.1834 bits === Per-stem alteration test (stem_len=1) === 4775 distinct stems total; 536 testable (>= 5 occurrences, >= 2 distinct head glyphs). stem — The common word stem by which the tokens were grouped. n — The number of occurrences of this stem in the corpus. H_real — How predictable the initial glyph actually is, given the previous final glyph. H_null_mean — How predictable the initial glyph would be, on average, if there were no real correlation (random reference value). p(<=real) — The probability that chance alone would yield a value as low as or lower than the actual value. z — How much the true value deviates from the random mean, measured in standard deviations. Top25: stem n H_real H_null_mean p(<=real) z r 683 1.018 1.082 0.0000 -9.37 kaiin 273 1.253 1.371 0.0000 -5.65 aiin 848 1.417 1.457 0.0000 -4.89 ain 233 1.704 1.801 0.0000 -4.01 ol 309 2.270 2.379 0.0000 -5.34 ar 420 1.748 1.835 0.0000 -5.26 l 741 1.055 1.136 0.0000 -10.98 ly 82 0.603 0.873 0.0000 -8.51 cheey 33 1.704 2.174 0.0000 -4.34 eedy 77 1.370 1.577 0.0000 -3.80 kain 159 1.212 1.338 0.0005 -4.36 dy 73 1.060 1.306 0.0005 -4.54 tey 70 0.418 0.572 0.0005 -4.03 keedy 160 1.235 1.373 0.0005 -5.46 chedy 186 2.078 2.191 0.0005 -3.76 eody 29 0.588 0.929 0.0015 -3.49 m 89 0.663 0.795 0.0020 -3.49 chy 80 2.125 2.325 0.0025 -3.11 dar 31 0.633 0.943 0.0025 -3.37 eey 84 1.545 1.708 0.0040 -3.16 edy 82 1.294 1.425 0.0055 -2.89 hedy 890 0.974 0.985 0.0065 -3.23 iin 479 0.328 0.351 0.0080 -3.00 okeedy 270 0.033 0.064 0.0100 -3.39 kedy 159 1.175 1.251 0.0105 -2.70 23 / 536 stems show head-glyph predictability from prev_end_glyph at p < 0.01 (permutation test). === Head-glyph distribution by context, top significant stems === Each line in the format after ‘...X’ -> Header1:n1/N, Header2:n2/N, ... shows how often the word stem with a given initial glyph occurs after a token with the final glyph X. The fraction (e.g., 154/244) indicates how often this initial glyph was chosen among all N occurrences in this context. Larger numbers indicate a stronger preference. Top10: stem 'r' (n=683, p=0.0000): after '...r' -> a:154/244, o:89/244, c:1/244 after '...n' -> a:63/115, o:52/115 after '...y' -> o:67/108, a:31/108, l:9/108, v:1/108 after '...l' -> o:57/94, a:36/94, l:1/94 after '...s' -> a:49/70, o:21/70 after '...o' -> a:14/23, o:9/23 after '...d' -> a:9/13, o:4/13 after '...m' -> a:3/5, o:2/5 after '...k' -> o:2/3, a:1/3 after '...p' -> a:2/2 after '...e' -> a:1/1 after '...g' -> a:1/1 after '...h' -> a:1/1 after '...x' -> a:1/1 after '...t' -> a:1/1 after '...a' -> a:1/1 stem 'kaiin' (n=273, p=0.0000): after '...y' -> o:59/95, l:23/95, y:10/95, e:1/95, a:1/95, d:1/95 after '...n' -> o:56/68, y:8/68, l:2/68, q:1/68, a:1/68 after '...l' -> o:30/43, l:6/43, y:3/43, a:2/43, c:1/43, s:1/43 after '...r' -> o:29/39, y:8/39, l:1/39, a:1/39 after '...o' -> l:9/13, o:4/13 after '...s' -> o:3/7, l:2/7, a:1/7, y:1/7 after '...d' -> o:1/4, a:1/4, s:1/4, y:1/4 after '...m' -> o:1/2, a:1/2 after '...k' -> o:1/1 after '...a' -> l:1/1 stem 'aiin' (n=848, p=0.0000): after '...y' -> d:291/411, r:38/411, s:33/411, t:18/411, k:17/411, l:7/411, o:4/411, q:2/411, p:1/411 after '...l' -> d:161/224, k:29/224, s:14/224, t:9/224, r:8/224, l:2/224, o:1/224 after '...r' -> d:69/84, k:5/84, o:3/84, s:3/84, r:3/84, l:1/84 after '...n' -> d:62/71, s:6/71, e:1/71, j:1/71, r:1/71 after '...o' -> d:20/38, r:6/38, k:5/38, t:3/38, s:3/38, o:1/38 after '...s' -> d:5/8, o:1/8, s:1/8, k:1/8 after '...a' -> d:3/5, r:2/5 after '...d' -> d:4/4 after '...m' -> d:1/1 after '...h' -> d:1/1 after '...t' -> d:1/1 stem 'ain' (n=233, p=0.0000): after '...y' -> d:52/103, s:18/103, k:16/103, r:10/103, t:4/103, l:2/103, o:1/103 after '...l' -> d:32/64, k:18/64, s:7/64, r:3/64, o:1/64, j:1/64, t:1/64, l:1/64 after '...n' -> d:26/32, s:3/32, r:1/32, o:1/32, t:1/32 after '...o' -> d:13/16, r:2/16, t:1/16 after '...r' -> d:9/13, s:2/13, k:2/13 after '...s' -> l:1/1 after '...p' -> d:1/1 after '...h' -> t:1/1 after '...a' -> k:1/1 after '...m' -> o:1/1 stem 'ol' (n=309, p=0.0000): after '...y' -> q:93/202, d:37/202, l:23/202, k:13/202, s:12/202, r:11/202, t:10/202, p:2/202, f:1/202 after '...l' -> d:30/61, q:12/61, k:9/61, t:3/61, l:3/61, s:2/61, r:2/61 after '...o' -> k:3/14, r:3/14, l:2/14, d:2/14, q:2/14, s:1/14, p:1/14 after '...n' -> d:4/14, s:4/14, q:4/14, k:1/14, p:1/14 after '...r' -> k:2/9, d:2/9, t:2/9, s:1/9, q:1/9, l:1/9 after '...s' -> d:2/3, y:1/3 after '...d' -> d:1/2, q:1/2 after '...a' -> m:1/1 after '...e' -> d:1/1 after '...t' -> d:1/1 after '...m' -> q:1/1 stem 'ar' (n=420, p=0.0000): after '...y' -> d:126/209, s:27/209, k:24/209, t:13/209, r:10/209, l:7/209, f:1/209, p:1/209 after '...l' -> d:62/94, k:14/94, s:6/94, t:5/94, r:4/94, x:1/94, e:1/94, p:1/94 after '...n' -> d:30/42, s:5/42, o:4/42, p:1/42, k:1/42, f:1/42 after '...r' -> d:23/41, t:6/41, k:5/41, s:4/41, f:1/41, m:1/41, o:1/41 after '...o' -> d:16/21, s:2/21, o:1/21, r:1/21, l:1/21 after '...s' -> d:2/5, o:2/5, s:1/5 after '...e' -> d:1/3, k:1/3, s:1/3 after '...d' -> d:1/3, s:1/3, r:1/3 after '...k' -> x:1/1 after '...a' -> k:1/1 stem 'l' (n=741, p=0.0000): after '...r' -> a:108/216, o:103/216, d:4/216, y:1/216 after '...n' -> o:123/188, a:62/188, d:2/188, c:1/188 after '...y' -> o:105/129, a:16/129, d:4/129, s:1/129, y:1/129, k:1/129, r:1/129 after '...l' -> o:89/118, a:22/118, d:5/118, t:1/118, k:1/118 after '...s' -> a:31/52, o:21/52 after '...o' -> o:7/14, a:6/14, r:1/14 after '...d' -> a:7/12, o:4/12, d:1/12 after '...m' -> o:4/4 after '...t' -> a:2/2 after '...e' -> o:2/2 after '...g' -> o:1/1 after '...p' -> a:1/1 after '...a' -> o:1/1 after '...h' -> o:1/1 stem 'ly' (n=82, p=0.0000): after '...r' -> a:16/31, o:15/31 after '...y' -> o:16/16 after '...n' -> o:12/15, a:3/15 after '...l' -> o:10/10 after '...s' -> a:5/8, o:3/8 after '...d' -> d:1/1 after '...o' -> o:1/1 stem 'cheey' (n=33, p=0.0000): after '...y' -> l:9/17, r:4/17, t:2/17, d:1/17, y:1/17 after '...o' -> p:2/7, t:2/7, k:2/7, f:1/7 after '...r' -> f:2/5, o:1/5, k:1/5, y:1/5 after '...l' -> t:1/2, r:1/2 after '...s' -> f:1/1 after '...n' -> o:1/1 stem 'eedy' (n=77, p=0.0000): after '...l' -> k:24/34, o:4/34, e:3/34, t:3/34 after '...y' -> k:16/32, t:11/32, d:3/32, e:1/32, q:1/32 after '...r' -> e:3/5, o:1/5, k:1/5 after '...e' -> k:2/2 after '...s' -> d:1/1 after '...n' -> o:1/1 after '...o' -> k:1/1 after '...d' -> k:1/1 There really is a connection. The last letter of a word influences which letter the next word starts with. This confirms what the paper found. But this connection is not a fixed rule, more like a tendency. After a certain letter, a particular starting letter becomes clearly more likely (often 70-80% of the time). Reliable cases Stem Context Glyph Fraction Percent aiin after '...n' d 62/71 87.3% l after '...y' o 105/129 81.4% ain after '...n' d 26/32 81.3% ain after '...o' d 13/16 81.3% aiin after '...r' d 69/84 82.1% kaiin after '...n' o 56/68 82.4% ly after '...n' o 12/15 80.0% ar after '...o' d 16/21 76.2% l after '...l' o 89/118 75.4% kaiin after '...r' o 29/39 74.4% aiin after '...l' d 161/224 71.9% ar after '...n' d 30/42 71.4% aiin after '...y' d 291/411 70.8% eedy after '...l' k 24/34 70.6% r after '...s' a 49/70 70.0% kaiin after '...l' o 30/43 69.8% r after '...d' a 9/13 69.2% kaiin after '...o' l 9/13 69.2% ain after '...r' d 9/13 69.2% Edit: By definition, the “stem” is simply “whatever remains after the first glyph is removed” (stem_len=1). For a two-letter token like “ar,” the head is ‘a’ and the stem is “r.” That's why we see the individual letters. Posting the code doesn't make sense with the current forum software. RE: A Glyph Is Not a Letter, a Token Is Not a Word, a Space Is Not a Space - JoJo_Jost - 21-08-2026 Mal ganz abgesehen davon, dass das ja im Prinzip schön länger bekannt ist (die 7 Regeln, die 94,x % alles Spaces beschreiben) ist es genau das was ich auch schon herausgefunden hatte und daraus folgt: Die Glyphen um einen Spaces bilden einen Teil der Chiffre. Und die Logik, aufgrund ihre Häufigkeit und Verteilung, ist, dass es in den meisten Fällen ein Vokal sein muss, der durch diese beiden Buchstaben um eine Space herum gebildet wird- Ich hoffe ich darf das schreiben, ich weiß es ist an der Grenze, weil es ja meine Theorie ist, aber hier geht es um die Struktur und die Logik, nicht um die VBM Theorie als ganzes. RE: A Glyph Is Not a Letter, a Token Is Not a Word, a Space Is Not a Space - Torsten - 21-08-2026 (20-08-2026, 11:54 PM)Koen G Wrote: You are not allowed to view links. Register or Login to view.So the VM displays an abnormally high predictability between edge glyphs, and near-zero token-to-token predictability. (I know people have researched this in the past, but I haven't learned much about it yet.) One analyses of the past is the You are not allowed to view links. Register or Login to view. by Vladimir Sazonov from 2003. RE: A Glyph Is Not a Letter, a Token Is Not a Word, a Space Is Not a Space - Stefan Wirtz_2 - 21-08-2026 They just compared to Latin, Italian, German, French and English — the only proof transported in this „news“ is that VMS does not contain western European (words and characters). „Meta-proof“ above: no one can calculate the Voynich language. It‘s funny to see easteuropean-named scientists producing the huge blind spot of Voynichero again. RE: A Glyph Is Not a Letter, a Token Is Not a Word, a Space Is Not a Space - Rafal - 21-08-2026 Quote:So if I understand one if their conclusions correctly, the final letter of a word (I'll keep saying it) is abnormally good at predicting the first letter of the next word. It's not a new result, for example Emma May Smith discussed it: You are not allowed to view links. Register or Login to view. You are not allowed to view links. Register or Login to view. And yes, it is weird. You cannot predict next word based on a previous word. In natural languages words like "pink" will be more often followed by a word like "flower" than some random word like "travel". In VM it doesn't happen. Yet there are these letter dependencies. I would say in some languages there may be some such dependencies, especially the ones with declination: Domum vetustam, ligneam, parvam video I can see an old, wooden, small house But it's last letter - last letter for a pair of words, not last letter - first letter. I cannot think of last letter - first letter example. For me these are "structured gibberish" supporting data, but feel free to disagree
RE: A Glyph Is Not a Letter, a Token Is Not a Word, a Space Is Not a Space - Koen G - 21-08-2026 Well, for example in my Brabantian dialect, there is such a thing. But since the dialect is only spoken (and written informally, and on the decline), this cannot be tested statistically. Anyway, in my spoken dialect, there are three grammatical genders. The gender determines the ending of articles and adjectives (either null, "-e" or "en"). For example: Masculine: den braven hond (the good dog) Feminine: de brave kat (the good cat) However, in the case of the masculine words, whether or not "n" is added also depends on the starting sound of the next word. You get -n before a, b, d, e, i, o, t, u. Also before "h", which is not pronounced so the vowel takes over. So all of the following words are masculine. The presence of "n" in the article only depends on "edge effects": den aap den boom de citroen den dief den ezel de fazant de gieter den hond den iglo de jacht de kabouter de leeuw de man de neushoorn den ooievaar de paling de ruiter de stengel den tram den uil de vos de wachter de zager The edge effects overrules everything, also when there are adjectives involved. To get back to the dog example, we say: den braven hond de groten hond So the word simply looks at its neighbor's starting letter, regardless of the sequence. Does this explain what we see in the VM? I presume everything in the VM is more extreme. But what it does show is that the principle is not anti-linguistic, even in European languages. RE: A Glyph Is Not a Letter, a Token Is Not a Word, a Space Is Not a Space - nablator - 21-08-2026 (21-08-2026, 12:54 PM)Rafal Wrote: You are not allowed to view links. Register or Login to view.You cannot predict next word based on a previous word. Waiting for Stolfi to post his study on conditional entropy of words. Perfectly normal? My quick check: yes, similar to Latin. Quote:In natural languages words like "pink" will be more often followed by a word like "flower" than some random word like "travel". In VM it doesn't happen. It does: "or aiin" etc. Less than in most languages but word order is far from random. |