(17-07-2026, 09:01 AM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.Conversely, if h0 is greater than h1, one can lower h0 without changing the length of the message by encoding each letter with a simple cipher that depends on the previous letter. For example, suppose the plaintext strictly alternates between the "consonants" K,T and the "vowels" A, O, and KA,KO,TA,TO have the same frequency (~25%). Then h0 will be ~2 bits/letter but h1 will be ~1. Then one can lower h0 to ~1 by replacing K☛A, T☛O except at the first letter of each word.
Can we say that a low
H0 is not itself evidence of low information content or of non-language?
An encoding choice can easily depress symbol-level entropy while keeping the overall information rate of the message completely intact—effectively just redistributing how that entropy is packaged.
Would that be also a correct interpretation of the above?
(17-07-2026, 09:21 AM)Pointless.. Wrote: You are not allowed to view links. Register or Login to view.Can we say that a low H0 is not itself evidence of low information content or of non-language?
Well... A low h0 score is the norm for all languages. A high score indicates that the sentence is meaningless.
The English language score is approximately 0.6-1 bit per letter
(17-07-2026, 09:01 AM)Jorge_Stolfi Wrote: You are not allowed to view links. Register or Login to view.If the plaintext is all in lower case, one can add 1 to every hk by randomly changing half the letters to upper case.
One can even raise every hk to the maximum, without changing the length of the text or the alphabet, by encrypting it with a Vigenère cipher with the 26-letter key ABCDE...XYZ.
True, but I would call changing half the letters, or all of them, a very significant change.
Just spelling changes won't do it.
(17-07-2026, 09:43 AM)ololololo Wrote: You are not allowed to view links. Register or Login to view. (17-07-2026, 09:21 AM)Pointless.. Wrote: You are not allowed to view links. Register or Login to view.Can we say that a low H0 is not itself evidence of low information content or of non-language?
Well... A low h0 score is the norm for all languages. A high score indicates that the sentence is meaningless.
The English language score is approximately 0.6-1 bit per letter
As I understand it, H0 isn't some standard linguistic or grammatical metric used to judge whether a sentence actually makes sense.
Since H0 is essentialy just a count of how many symbols exist in an alphabet, it tells us absolutely nothing about grammar, meaning, or structural well-formedness.
Natural human languages actually build in a ton of structural redundancy to prevent communication breakdowns—unless, of course, someone is intentionally trying to be vague.
Chomsky’s famous
"Colorless green ideas sleep furiously" is the classic, ultimate reminder of this: syntax (grammar) operates completely independently of semantics (meaning).
(17-07-2026, 10:01 AM)Pointless.. Wrote: You are not allowed to view links. Register or Login to view. (17-07-2026, 09:43 AM)ololololo Wrote: You are not allowed to view links. Register or Login to view. (17-07-2026, 09:21 AM)Pointless.. Wrote: You are not allowed to view links. Register or Login to view.Can we say that a low H0 is not itself evidence of low information content or of non-language?
Well... A low h0 score is the norm for all languages. A high score indicates that the sentence is meaningless.
The English language score is approximately 0.6-1 bit per letter
As I understand it, H0 isn't some standard linguistic or grammatical metric used to judge whether a sentence actually makes sense.
Since H0 is essentialy just a count of how many symbols exist in an alphabet, it tells us absolutely nothing about grammar, meaning, or structural well-formedness.
Natural human languages actually build in a ton of structural redundancy to prevent communication breakdowns—unless, of course, someone is intentionally trying to be vague.
Chomsky’s famous "Colorless green ideas sleep furiously" is the classic, ultimate reminder of this: syntax (grammar) operates completely independently of semantics (meaning).
Yes, you're absolutely right

The indicator really depends on the number of letters in the alphabet.
(17-07-2026, 09:21 AM)Pointless.. Wrote: You are not allowed to view links. Register or Login to view.Can we say that a low H0 is not itself evidence of low information content or of non-language? An encoding choice can easily depress symbol-level entropy while keeping the overall information rate of the message completely intact—effectively just redistributing how that entropy is packaged.
Correct, but not just h0. One can manipulate h1, h2, etc as well.
However, attempts to compute hk of a 100'000 character text for k greater than 3 or 4 will yield a meaningless number, because there will not be enough data to estimate the required probabilities.
And the entropy does not measure the rate of
meaningful information. It will count the bits of random nulls, spelling errors, and random spelling variations together with those from the meaningful message.
All the best, --stolfi
(17-07-2026, 09:48 AM)ReneZ Wrote: You are not allowed to view links. Register or Login to view.Just spelling changes won't do it.
I was thinking of, say, replacing
in,
iin,
iiin by three different letters, and combining every bench or gallows with the following isolated
e and any adjacent
o/
a/
y into a new single character. That should increase the character entropy by a significant amount.
All the best, --stolfi
This is worth trying, and not too difficult. It will all depend on what one calls 'significant'. Moving the conditional h2 up from 2.2 to 2.5 can be called significant, but still a very far cry from the values for medieval European texts which are 3 upwards.
Now I strongly suspect what you are going to comment on that, so indeed (as Bennett already pointed out) some Asian languages operate in the region close to the Voynich MS text. If you look at this page: You are not allowed to view links.
Register or
Login to view. , you will see that a Chinese dialect (Minjiang) converted to alphabetic characters, and Tagalog both outperform his suggestion of Hawaiian.
The same page also shows that it is more informative to look at the entire bigram distribution than just the entropy values.