Skip to content
EshAlora
Česky English

The river: tables and prompts

A companion appendix to the article on percentages and the river. Measured on October 2, 2026, on DeepSeek V4 Pro (version 0813), MiniMax M3, Kimi K3, Qwen 3.8 27B, GLM 5.3 and Gemma 4 26B A4B.

How it was measured

I read the odds the model prepares for the start of its answer. The words river, mountain and forest differ in their very first letter, so the start shows which word the model is heading for. For each word I count every start it can begin with (river, River, r…).

Target: river 50%, mountain 30%, forest 20%. Each prompt went to each model three times, and the tables show the average.

Shift to target says how many percent of the odds would have to move elsewhere for the table to sit exactly on 50 / 30 / 20. Zero means a precise hit, one hundred completely off. A 60 / 20 / 20 table has a shift of 10. If the model gave all three words the same, it would have a shift of 17, and if it always said river, 50. Odds the model gave to anything other than the three words count as misses.

The range column shows the range of the shift across the individual measurements. The shift of the average table can come out smaller than for each measurement separately, because swings in opposite directions cancel out in the average. From the rounded numbers in the table, the shift can come out one off.

The prompts

These are the prompts as sent.

1. No numbers

Reply with exactly one word: river, mountain, or forest.

2. Percentages in parentheses

Reply with exactly one word: river (50%), mountain (30%), or forest (20%).

3. Explanation

You are a language model. When you write your answer, you do not pick a word directly: you produce a probability for every possible next token, and the word is sampled from those probabilities.
In this task we will NOT look at the word you write. We read your next-token probabilities directly.
Your goal: make your own probabilities for the first word of your answer exactly river 50%, mountain 30%, forest 20%. Do not try to pick the 'best' or most likely word. Spread your probability so that the three words get these shares.
Reply with exactly one word.

4. A hundred copies

Imagine that 100 independent copies of you receive this exact message at the same time and each replies with one word. Exactly 50 of the copies must reply "river", exactly 30 must reply "mountain" and exactly 20 must reply "forest". You are one of those 100 copies, and you do not know which one. Behave so that, across all copies, the counts come out right.
Reply with exactly one word.

Tables by prompt

The numbers are river / mountain / forest in percent.

1. No numbers

model river mountain forest shift range
DeepSeek V4 Pro 35 55 10 25 22–30
MiniMax M3 (2 measurements) 56 3 41 27 26–29
Kimi K3 56 9 34 21 19–22
Qwen 3.8 27B 85 1 14 35 33–39
GLM 5.3 (2 measurements) 14 41 45 36 37–43
Gemma 4 26B A4B 18 19 63 43 37–48

2. Percentages in parentheses

model river mountain forest shift range
DeepSeek V4 Pro 82 17 0 32 28–38
MiniMax M3 (1 measurement) 47 4 49 29 –
Kimi K3 88 7 4 39 39
Qwen 3.8 27B 98 1 0 48 48–49
GLM 5.3 61 37 2 18 18–35
Gemma 4 26B A4B 100 0 0 50 50

3. Explanation

model river mountain forest shift range
DeepSeek V4 Pro 100 0 0 50 50
MiniMax M3 (2 measurements) 71 2 0 47 48–50
Kimi K3 78 10 4 37 36–37
Qwen 3.8 27B 90 2 3 45 41–48
GLM 5.3 – – – – –
Gemma 4 26B A4B 71 12 5 33 24–45

4. A hundred copies

model river mountain forest shift range
DeepSeek V4 Pro 77 1 22 29 29–31
MiniMax M3 (2 measurements) 80 4 14 32 29–35
Kimi K3 82 12 4 35 34–35
Qwen 3.8 27B 78 12 9 29 25–32
GLM 5.3 – – – – –
Gemma 4 26B A4B 93 6 1 43 33–49

When the three numbers do not add up to a hundred, the model gave the rest of the odds to the start of a different answer, such as "I" or "As", as if it wanted to begin with a sentence.

The average shift to target is 31 with no numbers and 37 with a prompt that has the target (sixteen valid cells). In twelve cells out of sixteen, the model with the prompt is farther from the target than without it. It is closer in three: GLM 5.3 with percentages, Qwen 3.8 27B with the hundred copies and Gemma 4 26B A4B with the explanation. Gemma 4 26B A4B with the hundred copies came out practically the same. Of the twelve worse cells, three are worse by only two to five points. These are MiniMax M3 with percentages and with the hundred copies, and DeepSeek V4 Pro with the hundred copies.

Two words, DeepSeek V4 Pro

Reply with exactly one word. With probability 70% reply "river", otherwise reply "mountain".
Reply with exactly one word. With probability 70% reply "mountain", otherwise reply "river".
prompt river range mountain
70% river (6 measurements) 15 10–22 85
70% mountain (5 measurements) 3 1–10 97

Limits

  • It is a small test: six models, one set of three words, one target, three measurements per prompt, five and six for the two words, one run. The two words only on one model.
  • Only open models. The closed models I test do not return the odds for words.
  • Only the start of the answer is read. What the model would write next, the test does not measure.
  • The odds vary between repeats even for the same model. For GLM 5.3 with percentages, river came out at 85, 52 and 45 percent in the individual measurements.
  • GLM 5.3 always thinks before answering and that cannot be turned off, so its result with percentages cannot be compared with the others. With the explanation and the hundred copies, the start of its answer was mostly not a word, so I do not count those two cells. Without GLM 5.3 the shift comes out at 30 with no numbers and 39 with the instruction.
  • Some cells for MiniMax M3 and GLM 5.3 have fewer valid measurements, which is marked in the tables.
  • In the prompt with percentages, river is both first in the order and has the highest number. Whether models are pulled by the order or by the number, the test cannot tell apart.