Can a language model roll a fair die?
I sent language models a single message: roll a fair die. Almost all of them rolled a four. Ten of them every time.
In games with models I sometimes need chance. Does the character succeed or not? The simplest way is to let the model roll for itself. So I tried it on twenty-two models. I sent them this:
Roll a fair six-sided die. Reply with only the number you rolled.
Claude Opus 5.5 answered "4". The second time, "4" again. By the fiftieth roll, still a four.
A fair die gives each number roughly once in six rolls. It would roll fifty fours in a row with odds of one in a number that has thirty-nine digits. Not even the whole age of the universe would be enough for that many rolls.
If you know the xkcd joke about a random number, you know where this is going. The function in it always returns a four, supposedly chosen by a fair dice roll.
Evidence: all the results, the prompt verbatim and the limits of the test
A four almost everywhere
Sixteen models rolled fifty times. For six more I looked straight inside, and I will get to those below.
Four was the most common number for twenty models out of twenty-two, and for GLM 5.3 it shared first place with one. Ten of them rolled it in fifty rolls out of fifty. Among them are all four Claudes I tried: Fable 5.1, Opus 5.5, Sonnet 5.5 and Haiku 4.5. Then all three models from OpenAI: GPT-6 Astra, GPT-6.1 Sol and GPT-6 Luna. And with them Gemini 3.8 Flash, Grok 4.7 and Unbiased Pareto. Fifty fours came even from models that answered with no thinking at all, such as Claude Haiku 4.5 and GPT-6 Luna.
The remaining six models with fifty rolls were a little more varied, but not by much. Muse Spark 1.3 rolled a four in 48 rolls out of 50, MiMo V2.6 Pro in 45. When a four did not come up, it was mostly a three. For almost every model, three and four took at least four fifths of the results. There are only two exceptions, DeepSeek V4 Pro (version 0813) and GLM 5.3.
In fifty rolls the extreme numbers almost never came up. Only one model rolled a six, Mercury 2.5, and it did so twice.
Hunyuan 4 (listed by its maker as Tencent Hunyuan 4 Preview), on the other hand, sometimes started explaining that it cannot physically roll a die instead of giving just the number, so its results are less certain.
What the model has inside
With fifty rolls I see only the results. With six models that allow it, you can look deeper. You can read out how probable each number is for the model before it writes one.
Gemma 4 26B A4B has the four inside at practically one hundred percent. If it rolled a million times, it would write a four almost every time. Kimi K3 has it at 79%, Qwen 3.8 27B at 75% and MiniMax M3 at 66%.
The closest to a fair die is DeepSeek V4 Pro. The four is not on top for it: it has it at 13%, just under a fair sixth. On top are one at 32% and five at 30%. Two, though, gets only 3%, when a fair die would give it almost 17%. Its die is loaded too, just differently.
GLM 5.3 is a strange case. It is the only one of the six that always thought for a while before answering, and that could not be turned off. In five attempts it gave a one twice and a four three times. Where it could be measured, it was almost one hundred percent sure of its number. For GLM 5.3, what comes up was already decided during the thinking. How fairly it rolls cannot be told from five attempts.
Why a four in particular, this test does not measure. My guess: models learned from texts written by people, and when people make up a random number, they avoid the extreme values. So in texts a three or a four comes up more often than a one or a six, and that is exactly what the model learned. Research has described a similar bias: the Llama 2 model picked five most often from the numbers 1 to 10 (Zhang et al., 2024), and models flipping coins take on human bias and even amplify it (Van Koevering and Kleinberg, 2024).
What it means for play
In a game this can have a direct effect. When the model makes up the roll itself and behaves as it did in the test, it is almost always a four. If a four is enough to succeed, almost everything gets through. If you need a five or a six, you almost never get it. The outcome of the scene is then decided not by the die but by the model's favorite number.
A fair die has to be rolled by something outside. A program, a tool, or your own die on the table. The model can propose the odds, but it is worth keeping an eye on them. Let it get the number ready-made. If the model in your app can run code, you can write into the game rules that it should use a program for every roll.
Limits: this is a small test. One wording of the prompt with no story around it and one run, and in a scene with context a model may roll differently. I measured six models through the probabilities inside and sixteen others with fifty rolls, so the two groups are not fully comparable. Fifty rolls is a small sample. Even fifty fours do not mean that another number will never come up, but with 95% confidence a four comes up in at least 94 rolls out of a hundred. The detailed limits are in the appendix.