MarcoP > 21-08-2026, 03:53 PM
Quote:Voynichese is encoded continuously in a way consistent with transitional probability matrices, but glyphs can sometimes affect transitions beyond the ones that follow immediately after them.
JoJo_Jost > 21-08-2026, 04:44 PM
Koen G > 21-08-2026, 05:37 PM
(21-08-2026, 03:53 PM)MarcoP Wrote: You are not allowed to view links. Register or Login to view.In Patrick’s words:
Quote:Voynichese is encoded continuously in a way consistent with transitional probability matrices, but glyphs can sometimes affect transitions beyond the ones that follow immediately after them.
rikforto > Yesterday, 12:01 AM
Quote:Once a spoken language has acquired a written form, the two linguistic systems may evolve independently so that the relationship between written and spoken languages becomes increasingly remote. With Chinese this happened in a way for which I know no parallel: during much of recent history, before the reforms in written usage associated with the May Fourth Movement initiated in 1919, the standard written language of China (wen yan, or literary Chinese) was a language which when read aloud in contemporary pronunciation could not be understood by a hearer, irrespective of how learned he might be, because the spoken equivalent of a written text did not contain sufficient information to determine the identities of the morphemes of which the text was composed. There were two reasons for this. Sound changes during the long period since the creation of the script had removed many phonological contrasts and thus introduced an extremely high incidence of homophony among morphemes;You are not allowed to view links. Register or Login to view. and developments in literary usage had created many possibilities of meaningfully combining morphemes in writing in ways that would never have occurred in spoken Chinese at any period of its history, thus reducing the chance of determining the intended morpheme among a set of homophone candidates by reference to its morphemic environment.You are not allowed to view links. Register or Login to view. Therefore literary Chinese not merely did only function but could only function as a written and read language, not as a spoken and heard language.This fact plainly extends to transliterations of the same, including Pinyin. Note that the issue is not that such a transliteration is impossible, or even never done, just that there is no audience who can make use of it:
Quote:But any literary Chinese text has a perfectly specific spoken form, composed of morphemes many of which occur in modern spoken Chinese and all of which are etymologically identifiable with morphemes that have occurred in spoken Chinese at some historical period: there is a well-defined way of reading a literary Chinese text aloud (morphemes which are obsolete in spoken Chinese are given the pronunciations that result by applying subsequent sound-laws to the pronunciations the morphemes had when they were current in speech), even if this activity achieves no communicative purpose.
JoeyB > Yesterday, 03:45 AM
DG97EEB > Yesterday, 08:06 AM
pfeaster > Yesterday, 02:14 PM
(21-08-2026, 05:37 PM)Koen G Wrote: You are not allowed to view links. Register or Login to view.(21-08-2026, 03:53 PM)MarcoP Wrote: You are not allowed to view links. Register or Login to view.In Patrick’s words:
Quote:Voynichese is encoded continuously in a way consistent with transitional probability matrices, but glyphs can sometimes affect transitions beyond the ones that follow immediately after them.
Could this be an argument in favor of the idea that, in a system like EVA, we are chopping up minimal units? Especially in cases where EVA-e and EVA-i are involved. If "eeed" behaves in a predictable way, then why should we assume that every stroke is a glyph?
Quote: An EVA glyph should be called a glyph until a model demonstrates that it functions as a letter. A ZL token should be called a token or surface form until a model demonstrates that it functions as a word. A transcribed blank should be called a separator until its boundary function is established. Treating the three observed categories as letters, words, and uniform word spaces builds an untested decipherment into the data.
pfeaster > Yesterday, 03:59 PM
Quote: Transcribers have also long noted that the glyph pairs flanking uncertain spaces differ from those flanking ordinary spaces (You are not allowed to view links. Register or Login to view.; You are not allowed to view links. Register or Login to view.).
MarcoP > Yesterday, 05:26 PM
Rozanova and Temerev Wrote:Acknowledgements
Large language models (Claude Fable 5, Anthropic; GPT 5.6 Sol, OpenAI) were used to assist with data analysis code and with editing the manuscript. All analyses, results, and interpretations were designed, checked, and approved by the authors, who take full responsibility for the content.
Quote:CROCUS.
[_sativus._]
1. Crocus spatha univalvi radicali, corollæ tubo longissimo.
Crocus floribus fructui impositis: tubo longissimo. _Roy. lugdb. 41._
_Hort. ups. 15._ _Mat. med. 27._
Crocus flore fructui imposito. _Hort. cliff. 18._
[_officinalis._]
α. Crocus autumnalis sativus. _Moris. hist. 2. p. 335. s. 4. t. 2.
f. 1._
Crocus sativus. _Bauh. pin. 65._
[_vernus._]
β. Crocus vernus latifolius. I-XI. & I-VI. _Bauh. pin. 65. 66._
Habitat in Alpibus _Helveticis_, _Pyrenæis_, _Lusitanicis_,
_Tracicis_. ♃
rikforto > Yesterday, 05:57 PM
(Yesterday, 05:26 PM)MarcoP Wrote: You are not allowed to view links. Register or Login to view.There are passages that are obscure, verbose and give that AI-feel. An example is "Currier strata" - "Currier languages" may be misleading, but here "strata" sounds like arbitrary LLM pseudo-scientific jargon.
It’s a pity, because, from what I understand, the contents make sense.
Quote:The letter and word readings are not the only unit hypotheses. Stolfi’s long-standing proposal that each token is one syllable of a tonal, isolating language of the East Asian type, written in an invented phonetic script (You are not allowed to view links. Register or Login to view.; You are not allowed to view links. Register or Login to view.), makes token-level predictions that can be checked against pinyin renderings of a genre-matched classical Chinese herbal (Bencao Beiyao, 1694) and a narrative (Romance of the Three Kingdoms) at matched token counts (Appendix You are not allowed to view links. Register or Login to view., Table You are not allowed to view links. Register or Login to view.). Two comparisons carry the weight. Shuffle-corrected adjacent order is 5.2–6.4% of capped entropy for the pinyin syllable streams, including the herbal, against −0.5% and +0.4% for Currier A and B and 0.79% pooled; and at a matched sample of about 10,700 tokens the pinyin streams use 335–749 syllable types with 10–20% hapax, whereas Currier A uses 3,343 types with 72% hapax. A syllabary is a closed inventory whose syllables carry sequential structure; Voynich tokens are neither closed nor sequentially constrained. Two further checks—the merge-scale profile of the pinyin letter stream and the calibrated attack against a pinyin-letter trigram model—point the same way but are less specific. You are not allowed to view links. Register or Login to view. already noted that Voynich letter statistics resemble pinyin; the order and closure comparisons show that the resemblance stops at the token level. The test covers Mandarin only; languages with larger syllabaries would narrow the inventory gap but not the order gap.(Emphasis mine)