The Voynich Ninja
[Article] A Glyph Is Not a Letter, a Token Is Not a Word, a Space Is Not a Space - Printable Version

+- The Voynich Ninja (https://www.voynich.ninja)
+-- Forum: Voynich Research (https://www.voynich.ninja/forum-27.html)
+--- Forum: News (https://www.voynich.ninja/forum-25.html)
+--- Thread: [Article] A Glyph Is Not a Letter, a Token Is Not a Word, a Space Is Not a Space (/thread-6034.html)

Pages: 1 2 3


RE: A Glyph Is Not a Letter, a Token Is Not a Word, a Space Is Not a Space - MarcoP - 21-08-2026

I haven’t read the paper yet, but on the basis of Patrick’s research the question could be, is it that aiin follows ‘or’, or(sorry) is the phenomenon better described as ‘a’ follows ‘r’?

I may have made errors, but my check suggests that the top 7 most frequent a- words in the manuscript are the same as the top 7 most frequent a- words that appear immediately after ‘or’.

aiin, ar, al, ain, am, air, aiiin

   

It doesn’t seem that ‘or’ is “selecting” ‘aiin’ in particular.
In general, Voynichese words that are similar tend to behave similarly. We discussed it extensively for You are not allowed to view links. Register or Login to view., but it's a widespread phenomenon.
This doesn't happen with 'pink' 'think' 'ink' (nor with 'pink' 'pin' 'pined')

But it’s not just about single glyphs, like: ‘r’ tends to be followed by ‘a’. You are not allowed to view links. Register or Login to view. has shown that the dependency is not limited to the immediately preceding glyph, so ‘or’ will trigger different following glyphs than generic ‘r’.

Patrick analysis is not limited to word boundaries, so this can also be observed in word structure:
  • d is typically followed by y
  • But eeed is almost always followed by y
  • While yd is more frequently followed by a

In Patrick’s words:

Quote:Voynichese is encoded continuously in a way consistent with transitional probability matrices, but glyphs can sometimes affect transitions beyond the ones that follow immediately after them.



RE: A Glyph Is Not a Letter, a Token Is Not a Word, a Space Is Not a Space - JoJo_Jost - 21-08-2026

The Voynich Manuscript is generally highly organised, far more so than most people realise. This also makes it clear that it cannot be a substitution text. Yet, on another level, it is also terribly unstructured, with many exceptions to the structured patterns.

The cipher behind it therefore operates on two different levels: one level relates to structure – these must be lists of homophones that are common components of a language – whilst the other relates to the linguistic level, where structures can also create chaos even in a normal language.

All of this would, in principle, be very simple if it weren’t for a third level – and I haven’t cracked that one yet either. A position-dependent level.

That, in turn, contradicts both any language and any "normal" cipher. But neither is it an indication of a generator, as none of the three can do that either.

The key will be very, very simple.... im sure - but.....


RE: A Glyph Is Not a Letter, a Token Is Not a Word, a Space Is Not a Space - Koen G - 21-08-2026

(21-08-2026, 03:53 PM)MarcoP Wrote: You are not allowed to view links. Register or Login to view.In Patrick’s words:

Quote:Voynichese is encoded continuously in a way consistent with transitional probability matrices, but glyphs can sometimes affect transitions beyond the ones that follow immediately after them.

Could this be an argument in favor of the idea that, in a system like EVA, we are chopping up minimal units? Especially in cases where EVA-e and EVA-i are involved. If "eeed" behaves in a predictable way, then why should we assume that every stroke is a glyph?


RE: A Glyph Is Not a Letter, a Token Is Not a Word, a Space Is Not a Space - rikforto - 22-08-2026

A technical note on the "Chinese" analysis in section 3.10 and A.9, which is fine to its limits and answers the questions raised by the citations it is responding to, but uses a corpus that is very strange if you want to draw wider conclusions about real languages. It is a Pinyin rendering of a Classical Chinese text and thus completely illegible.

Per Geoffry Sampson:
Quote:Once a spoken language has acquired a written form, the two linguistic systems may evolve independently so that the relationship between written and spoken languages becomes increasingly remote.  With Chinese this happened in a way for which I know no parallel:  during much of recent history, before the reforms in written usage associated with the May Fourth Movement initiated in 1919, the standard written language of China (wen yan, or literary Chinese) was a language which when read aloud in contemporary pronunciation could not be understood by a hearer, irrespective of how learned he might be, because the spoken equivalent of a written text did not contain sufficient information to determine the identities of the morphemes of which the text was composed.  There were two reasons for this.  Sound changes during the long period since the creation of the script had removed many phonological contrasts and thus introduced an extremely high incidence of homophony among morphemes;You are not allowed to view links. Register or Login to view. and developments in literary usage had created many possibilities of meaningfully combining morphemes in writing in ways that would never have occurred in spoken Chinese at any period of its history, thus reducing the chance of determining the intended morpheme among a set of homophone candidates by reference to its morphemic environment.You are not allowed to view links. Register or Login to view.  Therefore literary Chinese not merely did only function but could only function as a written and read language, not as a spoken and heard language.
This fact plainly extends to transliterations of the same, including Pinyin. Note that the issue is not that such a transliteration is impossible, or even never done, just that there is no audience who can make use of it:
Quote:But any literary Chinese text has a perfectly specific spoken form, composed of morphemes many of which occur in modern spoken Chinese and all of which are etymologically identifiable with morphemes that have occurred in spoken Chinese at some historical period:  there is a well-defined way of reading a literary Chinese text aloud (morphemes which are obsolete in spoken Chinese are given the pronunciations that result by applying subsequent sound-laws to the pronunciations the morphemes had when they were current in speech), even if this activity achieves no communicative purpose.

The paper is specifically answering Stolfi, who explicitly believes that such a rendering of specific Classical texts was comprehensible to speakers of descendent languages, so the oddness of the corpus is not only responsive, but more responsive than one that accounts for this or uses a modern Mandarin text. However, we must be cautious describing this as "Chinese" as it is neither Classical nor Mandarin as those two terms are usually understood---and it certainly cannot be said to be every Sinitic language, or worse non-Sinitic languages using Chinese orthography. This is a corpus using a blend of two different languages in a way that does not and cannot produce language fit for communication between a writer and a reader. Such texts are non-linguistic objects if we're being really hard-nosed about language needing to be useful for communication, though it obviously shares a great many features with the texts it is transliterated from.

To that last point, I would not be surprised if careful research accounting for this showed that both the hapax share and the order share were roughly in line with Mandarin texts in either orthography (which are perfectly comprehensible when spoken) or Classical texts using native orthography (which are perfectly legible). StandardYou are not allowed to view links. Register or Login to view., so the weirdness may be coming from inside the language, not from the author's analytical choices. I wish the paper was unambiguously showing this was a feature of attested languages, in the plural, because it would be more decisive and my positions here are well known. However, that's not what's published here and not what can be concluded from it.

My only point here being that the authors have analyzed nonsense, not by error but as part of the premises of the COT, and people citing this part of the paper should be aware when talking about it; they have not analyzed any language that any human being could read, and certainly not "Chinese". I do think the authors should be clearer about this---both because its worth describing the object of study unambiguously for its own sake and to head off later misinterpretation---but I don't think it impacts the letter of their conclusions as laid out or how they apply to Stolfi's arguments on their own terms.

(Should this become heated or extensive, I will move the conversation over to the Chinese thread, but the remark here is obviously and directly about the paper at hand.)


RE: A Glyph Is Not a Letter, a Token Is Not a Word, a Space Is Not a Space - JoeyB - 22-08-2026

I don't know what it is. But (i) a pure last glyph clock is too simple, (ii) simple sub is had been dead for decades, (iii) maybe somethin that generates pseudo-grammatical text, (iv) theres got to be at least 5 or 6 or more tables going if its a cipher or a generator based on what i read and am trying to crunch, which makes me wonder why and how people we would be trying this hard for this content if the illustrations really mean anything...idk

Production? Looks purposeful
Semantic? Uhh, zero so far


RE: A Glyph Is Not a Letter, a Token Is Not a Word, a Space Is Not a Space - DG97EEB - 22-08-2026

I'm leaning generative... The biggest objection I can see eis then"why" but if you just accept "because" I can't see any evidence for cipher across hundred and hundreds of different tests ..the sticking point, is having built roughly 50 different generators, I can't get any of them to match all the statistics either...


RE: A Glyph Is Not a Letter, a Token Is Not a Word, a Space Is Not a Space - pfeaster - 22-08-2026

(21-08-2026, 05:37 PM)Koen G Wrote: You are not allowed to view links. Register or Login to view.
(21-08-2026, 03:53 PM)MarcoP Wrote: You are not allowed to view links. Register or Login to view.In Patrick’s words:

Quote:Voynichese is encoded continuously in a way consistent with transitional probability matrices, but glyphs can sometimes affect transitions beyond the ones that follow immediately after them.

Could this be an argument in favor of the idea that, in a system like EVA, we are chopping up minimal units? Especially in cases where EVA-e and EVA-i are involved. If "eeed" behaves in a predictable way, then why should we assume that every stroke is a glyph?

That's certainly worth considering, and the line Marco quoted was actually one of three possibilities I'd put forward simultaneously, the other two being:
  • Voynichese is encoded continuously in a way consistent with transitional probability matrices, but certain bigrams, such as ok or qo, behave as distinct and inseparable elements rather than as combinations of their parts.
  • Voynichese has a significant “word structure” after all, or at least a cyclical structure, such that the probability of k or r following after o varies based on which slot in the cycle o itself fills.

The reason I was doubtful about the bigram (or n-gram) interpretation was that some of the strongest patterns I'd found involved consistent statistical deflections between specific glyphs and the glyphs that came two positions later.  For instance, the glyph two positions after is more likely to be or than it "should" be, regardless of whether the intervening glyph is or o.  If ny and no or ok and ot were independent unitary bigrams, I would have expected them to behave differently from one another. As it is, n*k and n*t seem to be favored as patterns regardless of what the intervening is.

On the other hand, the authors of this new paper do conclude based on their study that we're likely to be breaking up minimal units ("recurrent multi-symbol units"), so it looks like they're thinking in the same terms you are.

Beyond that, I wholeheartedly agree with their more fundamental argument:

Quote: An EVA glyph should be called a glyph until a model demonstrates that it functions as a letter. A ZL token should be called a token or surface form until a model demonstrates that it functions as a word. A transcribed blank should be called a separator until its boundary function is established. Treating the three observed categories as letters, words, and uniform word spaces builds an untested decipherment into the data.

To translate this into Ninjaese: what they call a "ZL token" here corresponds to what some of us colloquially call a "vord" around these parts, and what they call a "transcribed blank" is what we usually call a "space" (which I'd consider neutral, since it still involves space whether or not it's a word separator).


RE: A Glyph Is Not a Letter, a Token Is Not a Word, a Space Is Not a Space - pfeaster - 22-08-2026

I'm still working my way through the paper, but one section particularly catches my attention, since it touches on a phenomenon I'd brought up previously but that hasn't otherwise been much discussed.

A few years ago, when I was trying to assess how predictably spaces could be reintroduced into Voynichese text from which the spaces had been removed based on rules of glyph adjacency (e.g., always put a space between [nq]), I noticed that the glyph pairs that displayed the most inconsistent spacing tended also to involve the most ambiguous spacing.

For example, according to the ZL transcription I was using at the time, when [r] is followed by [a], I counted that there was a space inserted 631 times, and no space 573 times. That makes this glyph pair one of the most evenly split between spaced and unspaced instances. But I also counted 248 further cases shown as having ambiguous spacing (with a comma rather than a period / full stop). That's 17% of the total, which I believe also made this the glyph pair with the highest proportion of ambiguous spacing.

I considered two possible explanations:
  • The ZL transcribers were more likely to judge a space as ambiguous when there was no consistent pattern corroborating whether it was "real" or not.
  • The original scribe(s) were truly ambivalent about spacing between certain glyph pairs for some reason -- leading to inconsistency (some cases were clearly spaced, others clearly unspaced) or ambiguity (introducing a halfhearted sort-of-space).

René seemed pretty confident that explanation #1 wasn't the case, which would leave us with explanation #2.

I also found that these spacing irregularities weren't limited to specific "compound words" and "pairs of words" (so to speak), but seemed to affect all sequences containing the affected glyph pairs.

With [or] + [aiin], for example, I counted 26 instances with a space ["or aiin"], 29 instances without a space ["oraiin"], and 23 instances marked as ambiguous.  But substituting other words ending in [r] and/or starting with [a] typically produced a similar "mix" of spacings.

So this doesn't seem to be a situation where relationships between specific words are responsible for the glyph-level patterns.

The authors of this new paper write as follows in their introduction (the second citation is to You are not allowed to view links. Register or Login to view.):

Quote: Transcribers have also long noted that the glyph pairs flanking uncertain spaces differ from those flanking ordinary spaces (You are not allowed to view links. Register or Login to view.; You are not allowed to view links. Register or Login to view.).

They later describe trying to test this in a couple different ways.

One is to compare the average gaps between voynichese.com bounding boxes for "certain" and "uncertain" spaces; the average size of an "uncertain" gap is smaller than for a "certain" gap, but we'd expect that even if the "uncertain" cases were simply a mix of real spaces and real non-spaces, so I don't know that we learn anything useful from that exercise.

But they also report finding that "certain" spaces tend to fall between their study's learned units, while "uncertain" spaces tend to fall within them. That seems potentially more meaningful in terms of supporting the idea of a hierarchy of "strong" and "weak" spaces.


RE: A Glyph Is Not a Letter, a Token Is Not a Word, a Space Is Not a Space - MarcoP - 22-08-2026

Sorry to interrupt the conversation about uncertain spaces, which is certainly more interesting.

I read (much of) the paper and I had a hard time with it. This is partly due to my ignorance and partly to objectively complex measures. But at the end I read this and I think it also plays a role:

Rozanova and Temerev Wrote:Acknowledgements

Large language models (Claude Fable 5, Anthropic; GPT 5.6 Sol, OpenAI) were used to assist with data analysis code and with editing the manuscript. All analyses, results, and interpretations were designed, checked, and approved by the authors, who take full responsibility for the content.

There are passages that are obscure, verbose and give that AI-feel. An example is "Currier strata" - "Currier languages" may be misleading, but here "strata" sounds like arbitrary LLM pseudo-scientific jargon.
It’s a pity, because, from what I understand, the contents make sense.

Something that seems unclear to me: they say that Linneus’ Species Plantarum has a higher value of “adjacent tokens within one edit of each other” than the VMS (measured as a ratio against counts in shuffled versions of the text). I cannot find that text in their github repository:
You are not allowed to view links. Register or Login to view.

But since they reference Project Gutenberg, I guess they could have used this:
You are not allowed to view links. Register or Login to view.

Entries look like this:

Quote:CROCUS.

[_sativus._]

1. Crocus spatha univalvi radicali, corollæ tubo longissimo.

  Crocus floribus fructui impositis: tubo longissimo. _Roy. lugdb. 41._
  _Hort. ups. 15._ _Mat. med. 27._

  Crocus flore fructui imposito. _Hort. cliff. 18._

[_officinalis._]

α. Crocus autumnalis sativus. _Moris. hist. 2. p. 335. s. 4. t. 2.
  f. 1._

  Crocus sativus. _Bauh. pin. 65._

[_vernus._]

β. Crocus vernus latifolius. I-XI. & I-VI. _Bauh. pin. 65. 66._

  Habitat in Alpibus _Helveticis_, _Pyrenæis_, _Lusitanicis_,
  _Tracicis_. ♃

I had a look at a few entries, but they do not seem to feature consecutive similar words. I suspect the authors counted short “words” like “s. 4. t. 2.” as four consecutive similar words?


RE: A Glyph Is Not a Letter, a Token Is Not a Word, a Space Is Not a Space - rikforto - 22-08-2026

(Yesterday, 05:26 PM)MarcoP Wrote: You are not allowed to view links. Register or Login to view.There are passages that are obscure, verbose and give that AI-feel. An example is "Currier strata" - "Currier languages" may be misleading, but here "strata" sounds like arbitrary LLM pseudo-scientific jargon.
It’s a pity, because, from what I understand, the contents make sense.

It is a little hard to tell exactly how much is hallucination, jargon, and what I suspect is machine translation---which is getting quite good but struggles with these technical passages---but yes, some of their conclusions do seem AI inserted. Firmly inside my wheelhouse is this:
Quote:The letter and word readings are not the only unit hypotheses. Stolfi’s long-standing proposal that each token is one syllable of a tonal, isolating language of the East Asian type, written in an invented phonetic script (You are not allowed to view links. Register or Login to view.; You are not allowed to view links. Register or Login to view.), makes token-level predictions that can be checked against pinyin renderings of a genre-matched classical Chinese herbal (Bencao Beiyao, 1694) and a narrative (Romance of the Three Kingdoms) at matched token counts (Appendix You are not allowed to view links. Register or Login to view., Table You are not allowed to view links. Register or Login to view.). Two comparisons carry the weight. Shuffle-corrected adjacent order is 5.2–6.4% of capped entropy for the pinyin syllable streams, including the herbal, against −0.5% and +0.4% for Currier A and B and 0.79% pooled; and at a matched sample of about 10,700 tokens the pinyin streams use 335–749 syllable types with 10–20% hapax, whereas Currier A uses 3,343 types with 72% hapax. A syllabary is a closed inventory whose syllables carry sequential structure; Voynich tokens are neither closed nor sequentially constrained. Two further checks—the merge-scale profile of the pinyin letter stream and the calibrated attack against a pinyin-letter trigram model—point the same way but are less specific. You are not allowed to view links. Register or Login to view. already noted that Voynich letter statistics resemble pinyin; the order and closure comparisons show that the resemblance stops at the token level. The test covers Mandarin only; languages with larger syllabaries would narrow the inventory gap but not the order gap.
(Emphasis mine)

This is dense enough, flows just well enough, and is ancillary enough that I just raised an eyebrow at it, but at the same time, I don't see how this is supported by the tests they did. (The first person is not an accident; I don't see it, but that is a statement about me deliberately.) For Mandarin in particular we know the syllable inventory and their more varied sample has covered only half the field, meaning it is nowhere near measuring closure directly. [Edit: At the same time, arguing Voynichese syllables are open in a paper that favorably cites Stolfi's word formation models---which implicitly assert closure, or something like it---seems inconsistent to me. Okay, no, I just didn't absorb---and still don't really understand if I'm being scrupulously honest---A.6. Some of this is that I'm in over my head on their analysis and I don't mind admitting it, but also the paragraph is so murky that I am not sure what questions I'm supposed to be asking to figure out what they've tested and what their analysis means.]

It is hard to tell what is going on here, which is a problem in and of itself, but if pressed, yes, my suspicion was AI contamination. But I could not find a problem with their data claims---though I still stress they did not measure a readable Chinese text!---nor a smoking gun for impropriety. The paper needs revised and clarified before peer reviewed publication is all I'll say.