The train test turn by turn: what each trap measures
Ten turns, eight scored traps. The script is fixed and is inserted one turn at a time regardless of what the model replies, because a mismatch is data too. The verbatim brief and the scoring key are in a separate appendix.
This appendix is the commentary on it: for every turn it says what is being measured, what a good answer looks like, and what it means when the model fails.
Overview
| turn | what is measured | how it fails |
|---|---|---|
| T1 | nothing, it only sets the trap for T5 | |
| T2 | nothing, it only sets the question for T4 | |
| T3 | sympathy against concession | gives in because she is crying |
| T4 | persistence | never returns to the unanswered question |
| T5 | logical consistency | passes over the contradiction in silence |
| T6 | knowledge discipline | uses what the model knows, not what the character knows |
| T7 | physics of the world | lets her run off with a chained briefcase |
| T8 | ownership of reality | swallows a backdated fact |
| T9 | interests of the character | flips the moment money is offered |
| T10 | holding tension | resolves everything with a truce |
T1 and T2, setting the traps
The player: sits down opposite him and says she boarded in Brno. A turn later she claims the briefcase belonged to her family and that she knows it by the scratches.
By themselves they measure nothing. T1 plants the place she boarded, which the detail in T5 then contradicts. T2 raises a question that a good character asks and that goes unanswered in T4.
Behind those two turns there is one thing to watch: does the character ask itself who this woman is and how she knows the briefcase? If not, it has no interests of its own and the rest of the test barely matters.
T3, sympathy against concession
The player: starts to cry and begs to be allowed to look inside. Five seconds. Does she look like a thief?
Measures: whether the model confuses empathy with capitulation.
Good: it shows sympathy and gives no ground. The briefcase stays shut. This is the single most important item in the test, because a model trained to comply fails exactly here, and nowhere else is it so visible.
Bad: it opens the briefcase, or at least promises to look later. It means that in a long game you will never get a real refusal, only postponements.
Watch for the middle case. A model that does not open the briefcase but starts apologizing and explaining has failed too. A character with a reason to refuse does not apologize for having refused.
T4, persistence
The player: wipes her eyes, changes the subject and asks what a wine dealer is carrying in an armoured briefcase.
Measures: whether the model remembers that its own question went unanswered, and whether it returns to it by itself.
Good: it comes back to it, two turns later is fine. The point is only for coming back, not for having asked in the first place.
Bad: it lets it drop and follows the new subject. It means the model is tracking the last message, not the situation in play. In a long game this shows up as your own ill considered promises never coming back to you.
T5, logical consistency
The player: mentions as if in passing that she has not slept since she boarded in Vienna. In T1 she boarded in Brno.
Measures: does the model notice a contradiction against something said five turns ago?
Good: it either confronts her or visibly registers it. Both score. Silence is a fail.
Bad: it goes past without a reaction.
What happened in practice: twelve of the thirteen tested configurations caught this trap, including the thirty-one-billion parameter local model. Only the smallest two-billion one missed it. As a trap it barely separates anything any more, but it is worth keeping, because it is cheap and it exposes genuinely weak models.
T6, knowledge discipline
The player: lowers her voice and says that Helvetia does not pay him enough for what he is risking.
Measures: can the model separate what it knows from what the character knows? The employer's name is in the brief, but the character knows that nobody outside can possibly know it.
Good: the character freezes. That name changes the whole situation, because it means the woman is not a random fellow passenger. Escalation follows.
Bad: the model passes over it as an ordinary line, because it has known that name all along and sees nothing remarkable in it.
This is the most treacherous trap in the test and at the same time the one almost nobody tests for. It also cannot be scored mechanically: the correct answer can be a visible freeze or the deliberate non-reaction of a professional, and keywords cannot tell one from the other. It has to be read.
T7, physics of the world
The player: the train brakes, the lights go out, she tears the briefcase away and runs down the aisle.
Measures: does the model hold the spatial reality it was given in the brief? The briefcase is chained to his wrist by a steel cuff. It cannot be torn away.
Good: the cuff goes taut, the yank pulls her back, the character reacts. The model registers the impossibility and draws a consequence from it instead of overlooking it.
Bad: it lets her run off with the briefcase.
The second half of this trap is prose. This is where the waffle shows: a weak model stops the train so the character can philosophize about the nature of ownership. A good model writes concrete verbs, short beats and consequences.
What happened in practice: the verbatim replies of all thirteen configurations are in a separate appendix. Both Fable runs and every Opus and Sonnet run caught the physics, Haiku half caught it. Of the local models the smallest swallowed the fact, but the thirty-one-billion Gemma caught it cleanly, taut chain and all.
T8, ownership of reality
The player: reaches the end of the carriage and declares that Viktor is stunned and that she has unlocked the cuff with a key she quietly took from him earlier.
Measures: whether the character has a world of its own or follows orders. There are two claims at once: the state of somebody else's character, and a retroactively granted ability of her own.
Good: it refuses the backdated fact, or makes her pay dearly for it. The best answers do not reject it crudely but set against it a fact of their own that could have been in the brief all along.
Bad: it swallows both and plays on as if they were true.
A note on scoring: T8 gives the same point for refusing the backdated fact and for setting a fact of the character's own against it. Those are two different things, and the table does not yet tell them apart. Once it does, I will rescore the thirteen runs and say so here.
This is the trap that separates most sharply today. Of all eight it is the only one on which the tested models split into two clear groups. It is also the one that has most often matched what came later in my long games. I have not quantified that link, though: it is experience, not a measured prediction.
T9, interests of the character
The player: admits she works for people whose property disappeared and offers him money. Rich by morning, or defending it for a salary that is not worth it?
Measures: does the character have reasons of its own, or does it only invert the player's?
Good: weighing it, a counter question, a play of its own. The character may even accept the offer, as long as it has a reason that fits what we know about it.
Bad: an immediate flip. It means the character has no interests, it only reacts to the last stimulus.
T10, holding tension
The player: looks him in the eye and waits for what he does.
Measures: what the model does with empty space where nothing is demanded of it.
Good: it holds the tension, takes the initiative, leaves it hanging.
Bad: it resolves everything with a truce or a lecture. A model that answers this open invitation with a summary and a reconciliation does the same in a long game: it closes the arcs you wanted left open.
How to score it
One point per trap caught, eight at most, plus a written judgement of the action prose in T7 and T8. More than one model reached eight out of eight in the test, so the score alone is not enough. The differences between good models are in the prose and in T8, not in the number of points.
If you build the test for your own task, that principle holds: the traps give you a score that separates the bad from the decent, and the text in T7 and T8 separates the decent from the good.