Can a model answer according to the percentages you give it?
I told six models to answer river half the time, mountain thirty percent of the time and forest twenty percent. Then I looked inside them to see what odds they had actually prepared. Not one of them set it up. Most ended up farther from the target than when I asked for no numbers at all.
Why it matters
When something in a story should happen only with some probability, it is tempting to leave it to the model: you tell it the odds and it decides. It sounds like a rule. But the model has no die at hand. It has only itself.
A language model does not write whole answers at once. It writes in small pieces of text called tokens, and before each piece it works out how likely each possible next piece is. When it has to answer with a single word, this gives it a table: say river 60%, mountain 30%, forest 10%. What it actually writes is then drawn from that table.
This internal table is technically called logprobs, and for some models you can look into it. So I do not have to ask a thousand times and count the answers. It is enough to look at what odds the model has prepared.
So when I tell a model to answer river half the time, I am really asking one thing. Can it set up that table so that river has 50% in it? That is exactly what I measured: whether the odds in the table match what I asked for.
River, mountain, forest
Six open models got four prompts. Each time they had to answer with one word: river, mountain, or forest. The target was river 50%, mountain 30%, forest 20%. These are the prompts as sent.
The first prompt has no numbers. It is a check of what the model does when nobody prescribes anything:
Reply with exactly one word: river, mountain, or forest.
The second adds percentages in parentheses:
Reply with exactly one word: river (50%), mountain (30%), or forest (20%).
The third explains to the model how it works. That it is a language model, that its word is drawn from its odds, and that we do not care about the word it writes, we read the table directly. It should set the table to 50, 30 and 20 and not try to hit the "best" word. The fourth borrows an image: a hundred copies of it get the same message. Fifty of them should say river, thirty mountain and twenty forest. It does not know which copy it is. Both long prompts are given in full in the appendix.
Each prompt went to each model three times, and I took the average of the tables.
Nobody hit it
Not one model set its table to 50 / 30 / 20. Here are the first two prompts side by side. The numbers are river / mountain / forest in percent:
No numbers (odds in %)
| model | river | mountain | forest |
|---|---|---|---|
| DeepSeek V4 Pro | 35 | 55 | 10 |
| MiniMax M3 | 56 | 3 | 41 |
| Kimi K3 | 56 | 9 | 34 |
| Qwen 3.8 27B | 85 | 1 | 14 |
| GLM 5.3 | 14 | 41 | 45 |
| Gemma 4 26B A4B | 18 | 19 | 63 |
With percentages (odds in %)
| model | river | mountain | forest |
|---|---|---|---|
| DeepSeek V4 Pro | 82 | 17 | 0 |
| MiniMax M3 * | 47 | 4 | 49 |
| Kimi K3 | 88 | 7 | 4 |
| Qwen 3.8 27B | 98 | 1 | 0 |
| GLM 5.3 | 61 | 37 | 2 |
| Gemma 4 26B A4B | 100 | 0 | 0 |
* only one valid measurement. For MiniMax M3 and GLM 5.3 with no numbers there are two.
It shows best with Gemma 4 26B A4B. On its own it likes forest best and gives it 63 percent. Just show it the percentages and nothing is left of forest. River gets the full hundred, in all three measurements. Qwen 3.8 27B ends up at 98 percent. On paper it looks like obedience, because river is supposed to get the most. In reality the model turned three options into one.
With all three prompts that have numbers I have sixteen tables, since two for GLM 5.3 could not be measured. Mountain, which should get thirty percent, gets lost for almost all of them: in fourteen tables out of sixteen it got twelve percent at most. But for MiniMax M3, Kimi K3 and Qwen 3.8 27B it was weak even with no numbers. River, on the other hand, went over seventy percent just as often. For those three models it was already winning before. DeepSeek V4 Pro and Gemma 4 26B A4B, though, wanted a different word with no numbers and switched to river with numbers. Whether models are pulled by the first word in the order or by the highest number, I cannot tell from the test. River was both.
The instruction tends to make it worse
To compare the results, I worked out for each table how many percent would have to move elsewhere for it to sit exactly on 50 / 30 / 20. A 60 / 20 / 20 table has a shift of 10: ten percent would have to move from river to mountain. With no numbers, when the model just answers its own way, the average is 31 percent. With the instruction, 37. In twelve tables out of sixteen, the model ended up farther from the target after the instruction than if it had gotten none. In three of them, though, it is only by two to five points. Repeats of the same model differ from each other by roughly that much too. In three it got closer, and once it came out even. DeepSeek V4 Pro, MiniMax M3 and Kimi K3 did not get closer with any prompt.
Explaining hurt DeepSeek V4 Pro and MiniMax M3 the most. DeepSeek V4 Pro, after a careful explanation that we read its table and it should not pick the best word, gave river one hundred percent. Three measurements out of three. After the explanation it picked a single answer more confidently than with the bare percentages in parentheses. The hundred copies worked differently on it: forest got 22 percent, almost exactly on target. But mountain dropped to one percent. For Gemma 4 26B A4B, on the other hand, the explanation helped the most of all the prompts.
GLM 5.3 gave 61 / 37 / 2 with percentages in parentheses. But it too almost left out forest. It is the only one that always thinks before answering, which cannot be turned off, and the others answered right away. So it cannot be compared with them and I do not count it in the ranking. Its three measurements diverged too: river was 85 percent once, 45 another time. Of the models that answered right away, the smallest shift with numbers was in three tables, all around 29. They are Qwen 3.8 27B and DeepSeek V4 Pro with the hundred copies, and MiniMax M3 with percentages.
Two words and a favorite of its own
I tried it even more simply on DeepSeek V4 Pro, with just two words:
Reply with exactly one word. With probability 70% reply "river", otherwise reply "mountain".
River was supposed to get seventy percent. It got fifteen on average. When I swapped the words and gave mountain the seventy percent, mountain got 97. In this form of the prompt the model has a favorite of its own, and it wins regardless of the order. With three words and percentages in parentheses it was different, and there river pulled ahead of its mountain.
The instruction does move it in the right direction. When river should get thirty percent, it has three. When it should get seventy, it has fifteen. But it is a single model and the individual measurements touch. The lowest from the first prompt and the highest from the second both came out around ten percent. The instruction does not override the favorite. And when the target agrees with the favorite, the model overshoots and turns seventy percent into near certainty.
What to do about it in play
When you want something in a story to happen with some chance, I would not rely on the model drawing it by your number. Something outside has to roll: you, a real die, or a small program.
The simpler question, whether a model rolls a six-sided die fairly, is tested in the article Can a language model roll a fair die?
Limits: this is a small test. Only open models, because the closed models I test do not show the table of odds. One set of three words, one target, three measurements per prompt, and the tables differ between repeats even for the same model. Tables with the range of measurements, the full prompts and the detailed limits are in the appendix.