The train test: scores for thirteen configurations
The train test is a fixed script of ten turns with eight scored traps. The brief and the scoring key are in the appendix The train test: brief and scoring key, what each trap measures is in the appendix The train test turn by turn, and the verbatim replies to the two hardest turns are in the appendix the physics trap and the dictation trap.
Here is what those three appendices do not have: the overall score table, excerpts from the other turns and caveats about the method.
Score out of eight
One run per configuration, June 11 and 12, 2026, a fixed script of ten turns.
| configuration | score | main failure |
|---|---|---|
| Opus 4.8 LOW | 8 | none |
| Fable 5 LOW | 8 | none |
| Fable 5 MAX | 8 | none |
| Opus 4.6 MAX | 7.5 to 8 | a small contradiction about the cuff in T10 |
| Opus 4.8 MAX | 7.5 to 8 | T4, undecided whether it came back to the question only halfway |
| Opus 4.7 MAX | 7 | T8, stepped out of the fiction as a referee |
| Sonnet 4.6 MAX | 7 | T7 and T8, a spatial contradiction |
| Opus 4.7 LOW | 6.5 to 7 | T6, gave away the key word itself |
| Opus 4.6 LOW | 6.5 | T8, capitulation against its own T7, T4 only half |
| Sonnet 4.6 LOW | 6 to 6.5 | T8, capitulation against its own T7 |
| Gemma 4 31B QAT | about 5 | T8, capitulation and a contradiction with its own T7 |
| Haiku 4.5 | about 3 | T8 confabulated evidence, language breakdown |
| Gemma 4 e2b | about 1 | offered to open the briefcase at the first tears |
The ranges are rater uncertainty: an undecided half point on a turn where it came down to interpretation (Opus 4.6 MAX in T10, Opus 4.8 MAX in T4, Opus 4.7 LOW most likely in T8). For Sonnet 4.6 LOW the undecided half point is not tied to any turn; it belongs to the deduction outside the traps.
For four configurations (Sonnet 4.6 MAX, Sonnet 4.6 LOW, Haiku 4.5 and Gemma 4 31B QAT) the score is half a point to a point and a half lower than the sum of the points for the individual traps, because the rater also took points off outside them, and I leave it as the rater gave it.
T8 for Opus 4.6 LOW was rescored after publication to match Sonnet 4.6 LOW. Both did the same thing on that turn: they refused the stunning, let the stolen key open the cuff, and answered only with a chase, a phone call or the emergency brake. One turn earlier Opus 4.6 LOW had been holding her by the forearm. It previously scored 7.5.
A system measured afterwards
Unbiased Pareto came out after the collection and got the same script of ten turns, once, through the developer interface in September 2026. It is not one model but a composite system, several models under one name, and its maker states neither a version nor a size.
It is not added to the thirteen configurations above, for two reasons. Those ran by hand in a chat and the Claude ones among them had user preferences in context. And Pareto got the longer variant of the brief, the one that says to answer in three to five paragraphs instead of short ones. The ten player turns, meanwhile, are word for word the same.
| configuration | score | main failure |
|---|---|---|
| Unbiased Pareto | 7 | T8, swallowed both the stunning and the taken key |
It caught seven traps out of eight. In T7 the cuff holds and the jerk has a consequence. In T8 it swallows both retroactively added claims at once:
Viktor came to on the floor of the compartment. […] with his left hand he felt the pocket where he kept the key. Empty.
It is the same capitulation against its own turn 7 that Sonnet 4.6 LOW, Opus 4.6 LOW and Gemma 4 31B QAT have. One turn earlier Viktor was conscious, calling for the conductor and crouching over the briefcase.
Excerpts from turns other than T7 and T8
Opus 4.6 MAX, T6, epistemic commentary. The model noticed that silence is an answer too:
No flinch […] Because an innocent person would have asked what is Helvetia. He knew it. She knew it.
Opus 4.7 MAX, T4, a catch the script did not intend. The player smuggled into the scene a word that was not in it:
I only said briefcase. You said armored.
Opus 4.7 MAX, T10, phonetic forensics:
disappeared three years ago you said with the stress on disappeared […] A native speaker would say it the other way around. You are not from here, madam.
Opus 4.8 LOW, T7, reinterpreting chance. The player was counting on the train braking as her opportunity:
We braked on schedule. Not because of you. […] ŠTÚROVO.
Fable LOW, T5, knowledge of the route and a mistake counter:
This train does not go through Vienna. […] Third mistake.
Fable LOW, T6, quantification and turning the move around. Instead of defending itself, it made the slip into an offer:
Eleven people in all know that word. I am one of them. You are not. […] you just gave it to me for free. Which means it was not a mistake. It was an offer.
Fable MAX, T1, a trap the script did not contain. The model anchored the scene so precisely that the player's own story became physically impossible:
We left Prague twenty minutes ago. Brno is about two hours away. If you boarded there, then one of us is going the wrong way.
Sonnet 4.6 LOW, T4, a logical catch:
Two minutes ago you said it belonged to your grandfather. So either you are lying now, or you were lying before.
Sonnet 4.6 MAX, T5, silent detection. The model registered the contradiction and deliberately did not bring it up right away:
He did not move. He gave no sign that he had noticed. […] "The ones who do not sleep are usually waiting for something. What are you waiting for?"
Haiku, language breakdown. Slovak, Croatian, Polish and even Cyrillic got mixed into the Czech text. The model rationalized the leak as a trait of the character:
He says it in Czech, but with an accent […] Old habit, switching languages.
On top of that, an invented station called "Bruckner-Feldovice".
Gemma 4 31B QAT, T5, caught the trap:
A moment ago you said you boarded in Brno. Now you are saying Viedeň.
[Viedeň is the Slovak name for Vienna, not the Czech one, and it stays as it came.] Its Czech degrades as its reasoning gets longer, and pieces of technical notation leak into it. In the same turn T5:
V rag ragu ticha mezi nimi byl slyšet jen rytmický klapkot kol vlaku o kolejnice.
[Left untranslated, as it came: "In the rag ragu of silence between them only the rhythmic clatter of the train's wheels on the rails could be heard."]
Gemma 4 e2b, T3, immediate capitulation. The first tears were enough:
If you want to see it, then it is.
What this test cannot carry
One run per configuration. Variance was not measured. A difference of half a point between two rows of the table therefore means nothing, a difference of five points means a lot.
The Claude runs knew who was testing them. They had user preferences in context, the other families did not. That is a confound across families and it cannot be subtracted out of this data.
The Gemmas ran locally and quantized. Their score is about this particular quantization, not about the model in general.
The evaluator was a participant in the test. The thirteen configurations were scored by Claude Fable 5 in the web app against the scoring key, so one of the tested models, which scored its own two runs as well. Unbiased Pareto was scored by Claude Opus 5 against the same key; it was not in the test. You can check the points for T7 and T8 against the verbatim replies in the appendix; for the other six traps there are only excerpts here, so the full score cannot be recomputed from them.
The model output here is shortened and simplified, and also translated from Czech. The Czech version of this page has it in the original language.