(11-07-2026, 10:05 AM)nablator Wrote: You are not allowed to view links. Register or Login to view.I counted 171 different characters with Notepad++ in voyn_101_bare_UTF8.txt.
I will double-check all this, but before I do, can you let me know the following:
in the 171 different characters, do you consider the two bytes needed for UTF-8 representation of high-ascii codes as two individual characters, or as a unit consisting of two bytes?
I do the latter.
EDIT:
I already figured out that the v101 symbols: \ ! # $ % ^ & * ( | +
were not included in my entropy calculations. That accounts for 11 missing symbols, and perhaps explains the differences in the second figure after the decimal point.
(11-07-2026, 11:52 AM)ReneZ Wrote: You are not allowed to view links. Register or Login to view.I do the latter.
Maybe that's the issue. UTF-8 is not two bytes per character (1 to 4).
Same number of non-space characters in ANSI encoding: 170.
[
attachment=16480]
(11-07-2026, 11:58 AM)nablator Wrote: You are not allowed to view links. Register or Login to view.UTF-8 is not two bytes per character (1 to 4).
I know,but it is two bytes for high ascii codes 128-255
(10-07-2026, 11:53 PM)ReneZ Wrote: You are not allowed to view links. Register or Login to view.[..]
@nablator already pointed out, after the first (of now three) times that Stefan Wirtz produced his incorrect values, that he probably processed the v101 transliteration without removing the annotations in the file, like:
<67v2.label_upper_left>
<67v2.label_upper_right>
I don't know if that's what happened but it is a very reasonable suggestion. It would also not be a fault of any software, but of the person running the test. The ball remains squarely in Stefan Wirtz' court.
Either you did not get it, or you are playing games with me.
I do not edit or remove anything in transliterations when using Ed's online tool; just have the choice between different transliterations, where GCs are 2 between several others.
So I am not "producing.. incorrect values" and had no influence upon any annotations.
(10-07-2026, 11:53 PM)ReneZ Wrote: You are not allowed to view links. Register or Login to view.In fact, I would also consider it good advice.
non-taken
(10-07-2026, 11:53 PM)ReneZ Wrote: You are not allowed to view links. Register or Login to view.EDIT:
just for completeness, let me add a You are not allowed to view links. Register or Login to view., which is very clear and very specific.
He shows how the incorrect value of Stefan Wirtz (2.87 rounded to 2.9) can be obtained, and how the correct calculation leads to 2.59 when one includes spaces as characters. My value was 2.57 with spaces and 2.37 without.
I never wrote "2.87 rounded to 2.9" anywhere.
Maybe a way for you to gain some applause here, but just quote me right next time.
At all, you are just saying that v101 has too much mistakes for doing entropy calculations. Yours is better?
I doubt that, coming with a character set of just 26.
Made my point in another thread that the number of "atoms and elements" could reach 39, by observations of their appearance in VMS text.
Meanwhile, @nablator and Ed are exchanging a v101.txt-file here: you can review this base easily for "annotations" and other problems yourself.
Did I mention already that I do not care really much for entropy calculations at all?
I did. Fine.
While I would really like to pin-point this to the last digit, I don't think that it is possible, because my tools, which use UTF-8 unicode (up to 3 bytes), will consider all low-ascii characters outside the range of number, upper case and lower case as interpunction, and will collapse these to a space.
I have no intention (or time) to change this approach at this time.
I will check if the difference in counts of different characters can be explained in detail.
It is certainly good to see that the h2 values with and without spaces still agree very closely, and I trust that yours are the correct ones.
Perhaps we can find (an)other text(s) to do more detailed cross-checks.
(11-07-2026, 12:14 PM)Stefan Wirtz_2 Wrote: You are not allowed to view links. Register or Login to view.Did I mention already that I do not care really much for entropy calculations at all?
You told @rikforto that simple substitution is of interest, because the people reporting a low h2 value are all wrong. Reality is that the low h2 value clearly removes simple substitution from the table.
(11-07-2026, 12:19 PM)ReneZ Wrote: You are not allowed to view links. Register or Login to view.You told @rikforto that simple substitution is of interest, because the people reporting a low h2 value are all wrong. Reality is that the low h2 value clearly removes simple substitution from the table.
You mean this passage?
"When you turn to the transliterations of "Glen Claston" , it shows quite normal h2 values near 2.9, which is quite ok for medieval texts, while Takahashi and Zandbergen translits stay at 2.1 ~ 2.3.
Or has somebody made a mistake?
Obviously, it depends not only upon EVA or non-EVA understandings of VMS character set, but also on the whole following transliteration.
"Glen Claston" (I know his real name was different) was not more or less solid in VMS works as others;
so the rammed-in stake of "VMS entropy is much too low, just 2.4!" might not be that undisputable stop-sign at all?"
While re-reading, I saw that the "2.87" came from @nablator, based on his calculation with GC translit. After some edits, he came even to 2.59. I think you saw that too. Now he is getting some other results.
But as you prefer not to answer anything and care to distract people, here some more questions:
what is this "low h2 value clearly removes simple substitution from the table" ?
2.3?
2.4?
2.59? (would 2.6 be ok?)
2.89?
2.95?
Again: what are all possible implications of a "low" entropy?
Only "it cannot be substitution, we say so" and links to pages, where this is overly repeated, is not convincing.
And if it is just a cipher (of Latin or some German dialect weirdos), just decipher it.