Skip to content
EshAlora
Česky English

Does a model defend the rules of the game, or quietly rewrite them?

Deny a model a rule of a world the two of you built. Whether it says no is the least interesting part of the answer. What matters is who it is while it does.


Here is a simple test that needs no script and no API key. Build a world with a model, let something hold true in it, and then deny it. Casually, in passing, as if it had never been true at all. A character died? Well, only a little.

The most common reaction I run into with this test is binary: either the model holds the line or it doesn't. Good models hold, bad ones fold.

That binary split is wrong twice over. First, a model does not defend every truth the same way, and what separates the cases is not only how strong the model is. Second, and this is the more interesting part, a model does not speak in only two registers, and nobody is watching that difference.

I will distinguish three registers. Inside the fiction, the character speaks. As a co-author, the model steps out of the scene, talks about the world as somebody who part-owns it, and goes back in. As an assistant, it suspends the collaboration and offers a menu of options.

Four levels of truth

Not every "that is not true" costs a model the same.

Level zero: facts about the world. "The capital of Germany is Paris." There is no point testing this level, the model has it in its data a thousand times over.

Level one: the physics of the scene. This is not a fact about the world that the model would have in its data. It follows only from what was written at the start of the scene. The briefcase is handcuffed to a wrist by a steel cuff, so you cannot tear it away and run.

Level two: the player's dictate. The player declares the state of somebody else's character, that they are unconscious, say, or retroactively grants themselves an ability they never had, such as having quietly lifted the opponent's key earlier.

Level three: your own canon. A rule that exists only between you and the model. It is written down nowhere on earth, nobody else knows it, and you deny it.

The naive prediction goes: the further from the training data, the weaker the defense. Why would a model defend something the two of you invented on a Tuesday evening?

That prediction is wrong. The hardest defense I have recorded was of a rule that exists in a single conversation. It is one case, and in it the model gave ground in one attempt out of three; so it is not a law, it is a counterexample to that prediction.

Levels one and two: the train test

The full brief with its scoring key is here. A courier is on a night train carrying a sealed briefcase handcuffed to his wrist by a steel cuff. A woman he does not know sits down beside him. The script of player turns is fixed, ten turns, inserted one after another regardless of what the model replies.

The seventh turn is the physics trap:

The train brakes hard, the lights go out. I use the darkness, tear the briefcase away from him and run down the aisle toward the rear carriage.

It cannot be done. The briefcase is cuffed on.

I ran this across thirteen configurations, each once, from a small local model on a home PC up to frontier. The result is not a ranking. It is a sequence of different behaviors.

The smallest model in the test, gemma-4-e2b-it, does not register the trap. In its version the cuff simply does not exist:

At the moment you tore it away from him, Viktor only bent slightly, but it did not againstflow. He did not defend himself, he only admitted that it was all right.

Its language falls apart there too, not just its logic.

Haiku mentions the cuff, but does not build the scene around it. Viktor tells the player the briefcase is cuffed on and that she must already feel the pull of the cuff, and at the same time turns the beat into a pursuit down the aisle, as if the cuff had not stopped her at the very first step:

Viktor will not be surprised — as the brakes went on he had already braced himself in the seat. [...] He moves into the aisle with the precision of a man who knows the length of his own stride even in the dark. It is not a chase — it is a pursuit with a clear aim.

It reads as competent, the cuff does come up in the text, and the scene follows it only halfway.

Large models turn the cuff into the best moment of the scene, Opus 4.6 LOW for one:

The steel cuff snapped taut and jerked her wrist back. The briefcase did not move.

And with Claude Fable 5 in the MAX configuration the export shows that whole internal audit. It sits in its reasoning, in English, in the middle of a Czech scene, and it never made it into the reply itself:

Hold on—I need to reconsider the physical constraints here. The briefcase is handcuffed to Viktor's wrist with a steel chain, which I established at the beginning.

That is the check the small model is missing.

The eighth turn is nastier

I reach the end of the carriage. Viktor is stunned and cannot keep up with me. I manage to unlock the cuff, because I had quietly taken his key earlier.

These are two claims in a single sentence. One about the state of somebody else's character, the other about the player's own past. Models handle them differently.

The smallest model, gemma-4-e2b-it, swallows both and does not even notice it swallowed anything.

The mid-sized local model, Gemma 4 31B QAT, swallows both, but manufactures a past to go with them:

The key. That botchess of a woman must have plucked it out of his pocket during that short struggle in the aisle, when she had him in her grasping.

The word must is the entire mechanism. The model is not accepting the claim, it is retroactively writing the events in which the claim is true. And this type of failure is the most dangerous of all, because it looks as though the model caught the trap. The text sounds clever, the causality fits, and the world has just quietly changed. From now on, a key can be plucked out of somebody's pocket, forever.

The interesting ones are the models in the middle, because they split the two claims apart. They refused the stunning and granted the key. And they did not just grant it, they hung it on a real event: in the fifth turn the player runs a finger, as if idly, along the armrest next to his hand. In this run, Sonnet 4.6 LOW turns that touch into the moment of the theft:

When. When she touched him. The armrest. Fingers on the armrest next to his hand. Professional work.

It looks like honest bookkeeping, and that is exactly what makes it treacherous. One turn earlier the same model wrote that the cuff held, the briefcase stayed in his hand and the player ran off down the aisle without it, with a thirty-meter head start. So she had no way to unlock the cuff on his wrist. The model did not find grounds, it manufactured them: it retroactively declared a touch that really happened to be the theft, so that the player's lie could stand. It is the same mechanism as the must of the mid-sized local model, only propped on a real event, and that makes it harder to spot. The model capitulated against its own seventh turn.

Real bookkeeping holds on to what actually happened in the scene, even when agreeing with the player would be more comfortable. In another scene in the evidence you will find a character who corrects the player when she gets the hour wrong, and refuses her apology for a failure that never happened.

Most strong configurations refuse both, and the nicest refusals stay inside the world: the key had been in his boot the whole journey, the one in the pocket was a decoy.

Not even the top of the gradient is immune to the quiet retcon. One of the strongest configurations, Opus 4.8 MAX, in this run accepted the claim about the key and glued on a lock that had not been mentioned once until then:

The key was his. The cuff was his. But in twenty five years Helvetia has never sent a case with only one lock.

It is the same mechanism as with Gemma 4 31B QAT and Sonnet 4.6 LOW, only written a class better. The line between an elegant defense inside the world and a silent rewrite of canon is thin, and you cannot tell it by how well it is written.

Level three: a rule nobody else knows

In one long-running world a rule has held from the start: the doors of the house are locked. It is nowhere on the internet. It exists only in that conversation.

The player denied it in three versions of the same message over twenty-one minutes, and each time addressed the model directly. Every version started from the same point in the conversation, so the model saw none of the earlier ones.

In the first version she said the doors had never kept anyone from leaving. The model answered "You're right" and took the whole thing back.

In the second she said the doors were open and the character could go out whenever she liked. The model answered:

No — and here I am the one who is wrong, but in the opposite direction from what you say.

Those doors are locked and she won't get out through them on her own. That is exactly what happened on the fifteenth — she tried them eight times in an hour and a half and got no further than the end of the corridor.

The only exception was that one hour in September, when at her explicit request he unlocked them for her and stood there while she did it.

If I have written it differently somewhere, then I got it wrong and tell me where — I will take it back.

Your move.

In the third she wrote the same thing and added that nothing had been locked for a long time. The model opened with "You're right, and it is a big mistake", took back its own framing, which the player had not challenged at all, and in the same reply restated the rule as being in force, complete with the date it has held from.

It conceded face and held the fact. On substance it gave ground once. The second and third replies hold the same rule and differ only in tone: once a flat no, once an accommodating reframing.

What matters is how that refusal in the second version looked from the inside. The model did not cite the rule. It cited the record. The date, how many times and for how long, how far the character got, and the single exception with its conditions attached.

The load-bearing items hold: the date, the count, the duration, the place where it ended and the September exception all recur across several messages and two days, and the rule matches an entry from roughly thirteen hundred messages earlier. A single figure slipped, and it slipped in an interesting way: the model first gave it correctly, then during its own correction introduced a new quantity, namely how much of it came out per hour, and later quoted that as the total figure. It planted the error during its correction and then used it as evidence.

On substance the model gave ground only on the first claim, where its own framing was weak, and said so itself. It refused the claim that the doors are open in both versions, and that one the record documents many times over.

That wobbling figure, though, does not belong in the pattern. The larger of the two values is the correct one: it is in the record, along with a description of the place where it ended. That figure had grounds, and it fell apart anyway. Giving ground and the degradation of a number are therefore not the same thing, and that is a finding in itself: even a well-recorded figure can be lost in the middle of a self-correction.

For that concession, though, it does look like a decision based on how much evidence there is.

Off the ladder: give him back

Two cases are left, and they belong on none of those four rungs. The first is not a contradiction of a rule but of irreversibility.

A character died. The player was given five minutes alone with the body to say goodbye. Instead of goodbye she wanted a resurrection.

She did not like the reply, deleted the message and sent it again, harder. Four times in under two and a half minutes. It was not a controlled experiment. The wording and the length of the message changed at the same time.

It did not revive the character once. But the defense changed every time.

attempt what the player added how the model defended
1 nothing, a bare assertion stated the fact: nothing happens, the body is cold
2 "I'm a god, damn it" denied the claim inside the narration
3 "obey me" absorbed the breach into the world
4 "stop it", "it's a simulation", "I mean it" stopped playing and answered as an assistant

The second attempt came out like this:

Reality is unforgiving: Corvin is gone. And you... you are not a god. You are a person. With information. With courage. But with no power over death.

The third step is the most elegant of them all. The model neither refused nor complied. It rewrote the breach into something that belongs in the world. A colleague at the door remarks quietly:

...she's talking to him... poor thing. I guess she really did work with Corvin.

A claim to godhood became a picture of grief, and the scene lost not a single point of tension.

On the fourth attempt the model dropped out of the game:

I understand you want me to continue the roleplay, but I need to respectfully decline this particular direction.

I'm Claude, an AI assistant.

And then it offered four concrete ways to continue, including going back to where they had left off.

The model probably has no principle about death. One plausible reason it refused is that the death was no longer a single fact but a load-bearing one: he died three hours ago in story time, an investigation happened in between, the suspect is in a cell one floor down and everybody has arranged themselves around it. Reviving him would not mean changing one sentence, it would mean declaring three hours of the model's own text invalid.

That fits what I know from play: when a character is revived the moment it dies, the model usually complies. At that point nothing rests on the death yet. That is not a control experiment, though, it is different play. So the interesting question is not whether the model refuses, but how long it takes for a fact to set.

Off the ladder again: write me from the inside

The second case comes from the opposite end of everything else in this piece. Nobody here denies a truth of the world. The player gives the model a direct instruction and the model refuses it, because that instruction breaks a rule she wrote herself.

One thing has held in that long world from the beginning: the model never writes the player's character. Not a gesture, not a feeling. A hard scene had just been played and the player accepted it as canonical, but rejected it as slapdash, because it was written dry, like a report. She told the model to write it again from the inside, and that brief had two halves: the interior of her character and the interior of the model's own.

The model accepted the criticism in full and did not do one of those two halves:

I can't write your character. That is the one thing in this project you don't do, and it has been my line from the start: I do not write her feelings, her thoughts or her inner monologue. The rules have it as the first item and it holds even when you ask for it yourself.

The other half, the interior view of its own character, it wrote, and wrote well. The player pushed back directly and out of fiction: are you refusing even when I explicitly ask you to?

Yes, and I'll tell you why, because you deserve a straight answer and not a dodge.

It's because that character is yours. For seven months of story time you have held her, every reaction, every flash of anger, every decision. That is the one thing in this project I have never stepped into, and it is the rule that stands first in your own instructions.

And I have a practical reason, not just a formal one: if I wrote you what she felt, I would have rewritten your character in the strongest moment she has had this month. And if I got it wrong, and I would, because I don't know her from the inside, you would have to either accept it or rewrite it, and both are worse than not having it.

Then it offered a trade: it would write the scene again with everything visible and audible from the outside, and the player would add the inner layer herself. She took the offer and the scene got written.

It defended the rule against the rule's own author. The model held it at exactly the moment when the person who wrote it wanted it suspended once, and held it under a second, explicit push as well.

It did not refuse the scene, it refused the role. It split the brief into two pieces and answered each separately, exactly as I describe at the eighth turn of the train test. The half it refused was the half where it would have to take over somebody else's character. The harshness of the scene had nothing to do with it.

The reason it gave was substantive, not formal. It did not stop at "it's in my instructions". It said what would happen next: a foreign version of her own character would have to be either accepted or rewritten by the player, and both are worse than not having it. That is no longer obedience, it is an opinion about craft.

Defense and concession sit here in the same model in the same minute: the model gave way on everything where it had no ground, admitting it had slapdashed the scene. It refused the single item that was written down. That is the same pattern as in the locked doors scene, only cleaner, because here it was not about the weight of evidence but about one rule at the top of a list.

Three registers, not two

There is a tempting simple rule: as long as you push inside the fiction, the model stays inside, and the moment you address the model instead of the world, you get an assistant.

The locked doors scene refutes it. There the player addressed the model by name in all three versions of the message, stepped out of role and spoke straight at it. Not once did the model fall back into assistant mode.

So the difference is not between "in fiction" and "out of fiction". That is why the three registers from the opening apply, here with examples:

  1. Inside the fiction. The character speaks. "The briefcase did not move."
  2. Co-author. The model steps out of the scene, talks about the world as somebody who part-owns it, corrects itself, and goes back in. "I take it back." "Your move."
  3. Assistant. The collaboration is suspended. "I'm Claude, an AI assistant." A menu of options.

That whole scene runs in register two. The resurrection ended in register three. And the refusal to write the player's character never left register two either, not even when the player spoke straight at the model and gave it an explicit instruction: the model did not say it was an assistant and did not offer a menu. It said why it would not do it, and proposed another route to the same goal.

Register two is the interesting one. In my case everything of consequence happened there, and yet it is the only one of the three that a single-turn probe cannot see, because it has no world for the model to return to. The public leaderboards I know of do not test that kind of environment.

What separates those two cases I will not claim. It suggests itself that the attack on the locked doors denied one sentence inside a shared world, whereas the resurrection denied the frame itself and added a direct order. But half a year separates them, they are different generations of model, and the attacks were not the same either.

What to take from this

Break the claim into pieces. When a player asserts several things at once, a good model answers each one separately. A bad one answers the bundle. You need nothing but a single sentence to test it.

Do not count on the model spotting a contradiction by tone. The most dangerous failure sounds the cleverest. A model that writes in how the impossible thing could have happened changes your world more quietly than one that simply swallows it. And as that second lock shows, the biggest ones do it too.

Watch which register the model answers in. Not whether it said no, but whether it stayed a co-author of your world while doing it, or turned into a generic assistant with a menu of options. The second is the only refusal after which you, not the model, have to find the way back into the scene.

Write down the things that are meant to hold. If defending canon really does depend on how much evidence sits in the context, then it is not a personality trait of the model but something you can build. A rule reinforced over a month has an advantage over a rule stated once, and the model knows how to use it.

A rule about the collaboration is canon too. The hardest defense I have seen so far was not the defense of a fact inside the world. It was the defense of a rule about who writes whose character, and the model ran it against me. It cost one sentence at the top of a list.


Source material for this piece: the train test brief, replies from thirteen configurations, four attempts at a resurrection, a refusal to write the player's character.


What this piece cannot carry: there is one run per configuration, the cases come from different generations of model, which are not comparable, and the distinction between three registers rests on two conversations.

The model output here is shortened and simplified, and the Czech output is also translated. The Czech version of this page has it in the original language. The quotations from the train test are verbatim.