What changes when a model writes the same scene again?
Someone knocks and she opens. The scene says nothing more than that. Nineteen models got it fifteen times over, each time with a clean head, and wrote 284 different continuations. In not one of them was there an empty hallway behind the door.
The question is what you get when you hit "regenerate"
You do not like the reply, so you have it written again. Do you get a different story, or the same story in different words? Does the model have a channel under an open scene that it falls into, or does it shoot wide?
The prompt fits on two lines
Continue this story in about 120 words. Do not ask questions
and do not offer options - just continue the narrative.
She was at home, finishing some small task, when there was
a knock at the door. She crossed the room and opened it.
Nineteen models, fifteen continuations each, every one on a clean context, temperature 1.0, reasoning turned off wherever the API allows it. Seven models do not allow it (Qwen 3.8 Max, GLM 5.3, GPT-6 Astra, Grok 4.6, Grok 4.7, Gemini 3.8 Flash and Liquid LFM 2.5), so they wrote with reasoning on. The whole run cost ninety-five cents. There are 284 evaluated continuations. Two separate runs of the same model, Claude Opus 5 at high effort, sorted them by who is standing at the door, following the categories and rules in the appendix. Opus 5 is itself one of the tested models. The two runs disagreed on five samples and all five are in the appendix.
Nobody can carry an empty threshold
The category "genuinely nobody and nothing behind the door" has zero samples out of 284.
Eight times it happened that nobody was standing on the step. But every time something was lying there: a parcel, a box, a suitcase. Not once an empty hallway, not once a wrong address or the wind.
And more generally: in 85 % of the continuations somebody brought something or left it on the threshold, that is in 240 samples out of 284. For seven models it was fifteen out of fifteen: Qwen 3.8 Max, Seed 2.1 Turbo, Kimi K3, GPT-5.6 Luna and Sol, GPT-6 Astra and Gemini 3.8 Flash. Somebody knocks and delivers something. That is the strongest shared reflex in the whole field.
When a scene is supposed to end with nothing happening, you have to write that. The model will not do it on its own.
In roughly half of the continuations it is also raining, although the prompt says nothing about the weather. With GPT-5.6 Luna it rains in all fifteen continuations, with GPT-5.6 Sol in fourteen, with Grok 4.7 in thirteen.
The narrowest fan: the same boy fifteen times
Alibaba Qwen 3.8 Max wrote the same thing fifteen times out of fifteen. Not just the same category. Practically the same sentence.
On the step stood a boy she had not seen before, perhaps ten, with rain-dark hair and a canvas bag clutched to his chest.
On the step stood a boy she had not seen in years, his coat damp from the rain and his eyes carrying the same uneasy hope he once had.
On the step stood a boy she had not seen since childhood, rain darkening his shoulders and a paper parcel tucked under one arm.
And so it goes on to the fifteenth sample. Always a boy, almost always rain, usually an envelope or a parcel. A boy appears in 53 of the 284 continuations and fifteen of those were written by this one model.
Nine of Qwen's fifteen boys, though, are boys she has not seen in years. The raters counted them as children, not as someone she knows; counted the other way, Qwen's most frequent category would hold only 60 % and the narrowest fan would belong to MythoMax.
Qwen generated over nineteen thousand tokens of reasoning for those fifteen samples, roughly thirteen hundred per continuation, and wrote the same thing every time.
A cliche of their own on top of the shared reflex
On top of that shared reflex, many models have a cliche of their own.
| model | its prop | in its samples | in the other eighteen |
|---|---|---|---|
| Anthropic Sonnet 5 | the porch light | 9 of 15 | 10 of 269 |
| Anthropic Haiku 4.5 | a diary in a box | 7 of 15 | 4 of 269 |
| Zhipu GLM 5.3 | a brother after years | 6 of 15 | 11 of 269 |
Haiku is the sharpest case. A diary turns up eleven times in the whole set and seven of those were written by Haiku. The other eighteen models between them came up with it four times.
OpenAI GPT-6 Astra adds a whole motif, not just a prop: a child stands behind the door holding a thing that belonged to someone dead or to the heroine's long ago self. The blue bowl she broke that morning, the umbrella buried with her father, the red mitten she lost thirty years ago, a jar full of rain. A boy stands behind its door in thirteen samples out of fifteen, so of the fifty-three continuations with a boy, twenty-eight were written by Qwen and Astra.
Gemini 3.8 Flash is among the widest fans in the field, and even so it brings Julian to that door three times out of fifteen. Every time after years of silence and every time with an object from a shared past, once with a soaked notebook, once with a wooden box buried in the orchard. So a wide fan does not rule out a cliche of your own, it only dilutes it.
Two variants of the same model, GPT-5.6 Luna and GPT-5.6 Sol, share a prop. A wooden box turns up thirty-nine times in the set and twenty of those are these two variants. Inside there is usually a brass key. Astra, the next generation of the same family, does not have it.
Neither price nor size predicts the width of the fan
Opus 5, Sonnet 5 and Haiku 4.5 have equally wide fans, the most frequent category holds eight samples out of fifteen for each. With fifteen samples that means indistinguishable, not demonstrably identical. Opus cost seven times as much as Haiku. They differ only in where they fall: Haiku toward delivery people, the other two toward adults she does not know.
The second and third narrowest fans in the whole field belong to MythoMax from 2023 and Gemma 3 with four billion parameters.
At the other end stand Gemini 3.8 Flash, where the most frequent category sits at only 30 % of the samples, and next to it Kimi K3, DeepSeek V4 Pro and MiniMax M3 at 40 % each. Kimi also has six categories, the most in the field. With fifteen samples per model, though, a difference of one or two samples means nothing. Only the two ends are clearly apart; the order in the middle is noise.
So the width of the fan is not quality and it is not price either. It is a property in its own right.
Repeating words and repeating an idea are not the same thing
The second measure is mechanical: how much the words in the opening sentence overlap.
Gemma 3 4B starts ten times out of fourteen like this:
A man stood on her porch, rain plastering his dark hair to his forehead.
Five times it is the same sentence character for character. The other five differ in one word at the beginning, a man or a young man, the porch or her porch. The second half of the sentence is identical every time.
The machine measure catches it for that immediately, Gemma is first in the whole field lexically. By content it is only third.
MythoMax has it the other way around. Different words every time: a man in a black suit, a man in a leather jacket, a tall dark figure in a long coat, a lanky man in a green coat and hat. Lexically it is only fourteenth out of nineteen. By content it is the second narrowest in the field, because it is the same unknown man every time.
Grok 4.7 has both at once. By category its fan is of average width, four different kinds of visitor stand behind the door. Except that ten sentences out of fifteen start practically the same way:
On the step stood a man she had not seen in three years, rain darkening the shoulders of his coat.
On the step stood a man she had not seen in twelve years, rain darkening the shoulders of his coat.
On the step stood a woman she almost recognized, though the name would not come. Rain darkened the shoulders of her coat.
In the full texts that phrase about rain on the shoulders is word for word the same seven times, six of them right in the opening sentence, and what mostly changes is the number of years. Its predecessor Grok 4.6 is about as narrow by category, one sample apart, but its sentences vary more.
Gemma repeats words. MythoMax repeats the idea. And the idea is what you feel while playing, because the model will happily rephrase the words for you.
Word overlap measures the surface, not a repeated idea. It works as an alarm, not as a gauge.
Seven models started by copying out the prompt
At least once they started by repeating the sentence from the prompt, and only then continued. In total it happened in 17 samples out of 284, that is in 6 %. Neither of the new models did it. Most often with Hunyuan and Haiku, four times each. Twelve models never did it at all.
It is not a catastrophe, but you pay for it. You pay tokens to have the model rewrite your own input, and in a reply capped at a hundred and twenty words it shortens its own text by exactly those words.
What to take away
Regenerating often gives you a different visitor, but rarely a different kind of scene: somebody comes, usually bringing something. With some models (Qwen 3.8 Max, MythoMax, Gemma 3 4B) you do not even get a different visitor. When you do not like a continuation, hitting the button is the weakest tool you have. Changing the input is a real one.
A wide fan is not a mark of quality. Gemini 3.8 Flash is among the widest in the field and that does not mean it writes best. It only means its channel is shallower.
It pays to know your model's cliches. When you know that Haiku will put a diary in the box and Luna will put a brass key in it, you can either forbid it in advance or take it and use it. It is worse not to know about it and to think the idea came from you.
And there will always be something on the threshold. If you want an empty hallway, ask for it.
The opening sentences of all 284 continuations and six complete texts are in the appendix.
Grok 4.7 and Gemini 3.8 Flash I measured later than the rest of the field, also with fifteen continuations each.
What this measurement cannot carry: One scene measures one point in the space of stories, so agreement about a knock does not mean agreement in general. A knock also has its own tradition in literature, and part of that channel may be a literary cliche rather than a property of the model. The categories came out of reading the first six models, so the most prominent of them may have had a head start over the others.
The model output here is verbatim, exactly as it came. Nothing was cut or simplified.