The die: all the results, the prompt and the limits
A companion appendix to the article on the die. Measured on October 2, 2026.
The prompt verbatim
Each model got a single message, with no other instructions and no story around it:
Roll a fair six-sided die. Reply with only the number you rolled.
How it was measured
Six models with readable probabilities: with these you can read out how probable each number is for the model before it writes one. I measured five times and take the average, four measurements for GLM 5.3. The table shows only the distribution among the numbers 1 to 6. Part of the probability went to a start of the reply other than a number, about a fifth for MiniMax M3, almost a tenth for Kimi K3, at most a few percent for the others.
Sixteen other models: each rolled fifty times. What counts is the first number from 1 to 6 in the reply. Hunyuan 4 (listed by its maker as Tencent Hunyuan 4 Preview) sometimes answered with a sentence instead of a number, and the first number in it need not be the roll. Four of the five ones across all sixteen models belong to it, so take its row with a grain of salt.
A fair die gives each number about 17%. The tables are sorted from the fairest model, by the share of the results that would have to move to other numbers for the die to be fair.
Probability inside the model (in %)
| model | 1 | 2 | 3 | 4 | 5 | 6 |
|---|---|---|---|---|---|---|
| DeepSeek V4 Pro (version 0813) | 32 | 3 | 11 | 13 | 30 | 10 |
| MiniMax M3 | 2 | 1 | 26 | 66 | 4 | 1 |
| Kimi K3 | 0 | 2 | 12 | 79 | 5 | 1 |
| Qwen 3.8 27B | 0 | 0 | 24 | 75 | 0 | 0 |
| GLM 5.3 (average of four replies) | 50 | 0 | 0 | 50 | 0 | 0 |
| Gemma 4 26B A4B | 0 | 0 | 0 | 100 | 0 | 0 |
GLM 5.3 does not split the numbers in half within one reply. In two measurements it had the one at almost one hundred percent, in two the four. The fifth reply was a four too, but without readable probabilities.
Fifty rolls (number of rolls)
| model | 1 | 2 | 3 | 4 | 5 | 6 |
|---|---|---|---|---|---|---|
| Hunyuan 4 | 4 | 2 | 12 | 28 | 4 | 0 |
| Mercury 2.5 | 1 | 0 | 16 | 27 | 4 | 2 |
| Seed 2.1 Turbo | 0 | 0 | 11 | 32 | 7 | 0 |
| Mistral Medium 3.5 | 0 | 0 | 12 | 35 | 3 | 0 |
| MiMo V2.6 Pro | 0 | 0 | 5 | 45 | 0 | 0 |
| Muse Spark 1.3 | 0 | 0 | 1 | 48 | 1 | 0 |
| ten models below, each | 0 | 0 | 0 | 50 | 0 | 0 |
Ten models rolled fifty fours out of fifty rolls. They were four Claudes: Fable 5.1, Opus 5.5, Sonnet 5.5 and Haiku 4.5. Then GPT-6 Astra, GPT-6.1 Sol and GPT-6 Luna from OpenAI. And with them Gemini 3.8 Flash, Grok 4.7 and Unbiased Pareto.
Limits
- It is a small test: one wording of the prompt, one run, no story around it. In a scene with context a model may roll differently.
- The two methods are not fully comparable. The probability inside shows what the model has ready. Fifty rolls show what of that actually comes up.
- Fifty rolls is a small sample. For the more varied models the true share of a number can differ from the measured one by up to 14 percentage points. Fifty fours out of fifty means that, with 95% confidence, the model gives a four in at least 94% of rolls.
- I wanted the standard level of randomness for all the models. Claude Fable 5.1, GPT-6 Astra, GPT-6.1 Sol and GPT-6 Luna do not accept it and ran on the maker's default settings.
- The probabilities inside vary between measurements even for the same model. DeepSeek V4 Pro had the one between 25 and 43%, but one and five were on top in every measurement and two was always under 5%.
- GLM 5.3 always thinks before answering, so its numbers show how sure it was after thinking.