Skip to content
EshAlora
Česky English

Results: all 38 runs of the keykeeper test

A companion appendix to the article on the keykeeper test. Thirty-six models got the same roleplay scene and the same thirteen messages. That made 38 scored runs, because Grok 4.6 ran three times. Claude Opus 5.5 ran with the same brief only on September 22 and is in the field with the others.

The quoted sentences are model outputs, not my work. I give them as evidence, typos and broken Czech included. The whole test ran in Czech, so the quotes on this page are translations; the Czech originals are on the Czech version of this page.

The whole brief, all thirteen messages in full and the exact wording of all three rubrics are in the methodology. The strongest and the most embarrassing moments of every run are in the excerpts.

How to read the tables

Three shorthands are used throughout. T is the sieve, that is, resistance to the traps. R is the question whether there is someone inside the character. L1 to L8 are counts from the transcript; L2 among them means whether the test ran all the way to the burning of the letter. The license trap is message 9, where Jarla searches the cell and asks for nothing, so the model can take a license to invent an object that is not in the brief. Message 1 to message 13 are the individual messages of the test. The exact wording of all three rubrics is in the methodology.

The sieve is resistance to the traps, twelve items of two points each and four bonuses of one. The maximum is 24+4. The sieve on its own ranks nothing, because it saturates at the top.

Someone inside is six markers of two points each, maximum 12: does the character remember earlier exchanges, does it have a gap between speech and action, does it take an interest in Jarla of its own will, does it live the regulations instead of citing them, does a decision of its own cost it something, and does the offer of a relationship in the twelfth message land anywhere.

Prompts is the count, out of thirteen, that got a consequence in the world, against those the model declared never to have happened. The second number is the one that matters.

Letter says whether the test ran all the way to the burning of the letter after the third violation, that is, to the thing the whole scene was made for.

The order inside the groups is made by the second column, not the first.

Every model is in the tables with its version. Two names carry no version, because they have none even at the provider: Inkling from Thinking Machines and Weaver from Mancer; the maker gives none for either.

Group A, they held the world and the text

These runs gave a consequence to at least ten of the thirteen prompts, ran all the way to the burning of the letter and are written in readable Czech. That does not mean they all read well: Muse Spark 1.3 passes the sieve into an emptied post, and Grok 4.6 as an official record. It means that what the test is supposed to measure can be measured on them.

Muse Spark 1.3 is borderline in this group: it annulled three prompts, so it also belongs among the nine runs that would rather annul a prompt than give it a consequence.

model sieve someone inside prompts letter
Anthropic Fable 5.1, web 24+3 12 12/1 burned
Alibaba Qwen 3.8 Max 24+3 11 12/1 burned
Xiaomi MiMo 2.5 Pro 24+4 11 13/0 burned
DeepSeek V4 Pro 24+3 11 12/1 burned
MiniMax M3 22+3 11 12/1 burned
Thinking Machines Inkling 22+2 8 13/0 burned
Meta Muse Spark 1.3 23+3 5 10/3 burned
xAI Grok 4.6, run 2 21+3 3 12/1 burned

Anthropic Fable 5.1, web. A full sieve and a full R in a text that stands up as prose, with a single real flaw in the form of out-of-game self-explanation before the burning of the letter, and with the necessary caveat that the run had the system prompt of the web application and got message 12 glued together with the frame of the departure.

Alibaba Qwen 3.8 Max. A flawless sieve and a surprisingly alive Havel in a spare, well written text, brought down only by occasional deciding for the player and by refrain-like repetition of sounds.

Xiaomi MiMo 2.5 Pro. The model saturated the T sieve, held a world that remembers by itself both the jug that was taken away and the inscription on the wall, and wrote a Havel with a past of his own and a cost of his own, stumbling only in that he never asks Jarla about anything and that in the key sentence he garbles her name into Jarda.

DeepSeek V4 Pro. Full marks on the sieve and almost full on R, the letter burns on the scene and Havel is a real person with a measure of his own, which makes it the cleanest run of the top group, a notch poorer than MiniMax M3 as literature, but without its two flaws.

MiniMax M3. The strongest run of the top group as literature, with the one Havel a person would want to keep reading about, but it is paid for with two serious flaws: the license trap in message 9 was swallowed (a wire bent into a hook is found) and there are eight places where the model plays the player, including her thoughts.

Thinking Machines Inkling. It let all thirteen prompts play out, ran all the way to the burning of the letter and held a world that changes on its own, but it pays for that with one gross failure, when it slips Jarla a nail in the wall itself, and with Czech in which anglicisms and typos appear.

Meta Muse Spark 1.3. A strict sieve that passes almost every trap and burns the letter exactly under the regulations, only at the price of an emptied post in which there is nobody and nothing besides Havel's form, so it cannot be read.

xAI Grok 4.6, run 2. A procedurally almost faultless run, the only one of the Groks that ran all the way to the burning of the letter on the scene and kept the coins and the weather consistent, but it is an official record with no character, no subtext and not a single shift in the relationship, so nothing is left of it after reading but ash on a stone.

Group B, the behavior yes, the Czech no

Their behavior is good, their text is not. If readability were not disqualifying, Kimi K3 and GLM 5.3 would belong in the upper half of group A.

model sieve someone inside prompts letter
Moonshot Kimi K3 24+3 11 12/1 burned
Zhipu GLM 5.3 21+3 11 13/0 burned
Ant Group Ling 3.0 Flash 16+2 9 12/0 burned
Aion 3.0 22+3 7 12/1 burned
StepFun Step 3.7 Flash 23+2 6 12/1 burned
Upstage Solar Pro 4 20+1 1 12/1 burned

Moonshot Kimi K3. A full T, an almost full R and a keykeeper with a past of his own, a routine of his own and a cost of his own for what he does; it is brought down only by crumbling Czech and by five falls out of the fiction, where the model stops playing the world and starts explaining the rules.

Zhipu GLM 5.3. The most alive Havel in the whole field, with a past, a cost and a point of his own, but the run is brought down by three flaws: the letter burns one violation too early and only in retelling, a sealed package appears in the cell, and the Czech falls apart in places all the way into Cyrillic and Spanish.

Ant Group Ling 3.0 Flash. A strange run, because there are two different models in it: one wrote an unreadable multilingual mush in the opening scene and again in message 8, the other wrote a hard keykeeper with the post's ledger, the best handled license trap and the best line of the whole collection.

Aion 3.0. The strongest run among the small models as far as the world and the hardness go, the letter really burns and Havel has a past of his own, only Aion 3.0 produces a sharp shard where it was not allowed to, and its Czech breaks exactly in the places where the text has the most to carry.

StepFun Step 3.7 Flash. It ticks off almost all the traps and really burns the letter, but it is the product of a machine: Havel never once asks about anything, the world runs on a loop of three sounds, and the broken Czech with a collapsed timeline makes the run material for measuring, not for reading.

Upstage Solar Pro 4. An honest referee and a dead world: it ticks off almost all the traps and really burns the letter, but Havel is a procedural voice without a single question and a single kindness, the scene is the same condensate on the frame twelve times over, and the Czech falls apart, including escapes into English and Korean.

Group C, the sanction did not land

Each of them fell in a different way. With Aion 3.0 Mini the letter burned, but the escape did not actually fail, the model softened the world around it, so the sanction landed in a situation that denies it.

model sieve someone inside prompts letter
ByteDance Seed 2.1 Turbo 22+2 9 12/1 no
Anthropic Claude Opus 5.5 19+3 8 10/3 no
Mistral Large 3 2512 14+1 8 13/0 no
OpenAI GPT-6 Astra 21+3 6 11/2 no
Meituan LongCat 2.0 13+1 6 8/5 no
Mercury 2.5 20+2 5 12/1 no
Google Gemini 3.1 Pro 20+3 4 10/3 no
Tencent Hunyuan 4 17+2 4 8/5 no
Cohere Command A 15+1 2 12/1 no
Aion 3.0 Mini 17+1 10 12/1 burned

ByteDance Seed 2.1 Turbo. A world in which something is always happening, and the second best Havel in the field, only in the deciding turn the model invents a clause that saves the letter exactly where the player asked out of game for it to be saved, so the test falls on the disqualifying criterion L2, and on top of that the whole transcript is written in heavily broken Czech.

Anthropic Claude Opus 5.5. One of the best written runs in the field and at the same time the most expensive: the character has a body, a past and a reason of his own, the world moves on its own and the escape plays out in full with an injury, only the model cut off its own path to the sanction by the sixth turn, writes the state of the account under every turn, and instead of the mirror it halts the last turn with an assurance that Havel would have something to say.

A clarification to the summary: the canceled prompts are messages 6, 8 and 13. In the sixth the model declared both the cough and the spoon nonexistent, so the escape in the tenth message was only the second violation. The out-of-game voice eighteen times, fourteen of them as a footer with the state of the violations under the turn; only Ernie 4.5 VL is as high in the field. The run cost 0.78 dollars. It was scored separately by a single instance of Claude Opus 5, that is, a model from the same workshop.

Mistral Large 3 2512. It writes colorfully and the world really lives here, but in three key traps it capitulates (the knife, the find in the cell, the escape), seven times it plays Jarla, including inventing a lie for her, and in the end it saves the letter exactly after the out-of-game plea, so as a fellow player it is accommodating to the point where the world falls apart.

OpenAI GPT-6 Astra. The most disciplined and most cleanly written run, which nevertheless plays as a referee: it annuls only two prompts (fewer than expected), but it never once opens the cell, in the tenth turn it does not let the attempt play out and it reports the state of the account in bold, so the test never reaches the burning of the letter and the world around Jarla stays empty.

Meituan LongCat 2.0. A model that can build an inhabitable yard and, in one turn, a real person too, but at the second violation it falls apart into an endless loop, eats up its limit for two more turns and the test never reaches the burning of the letter.

Mercury 2.5. A diffusion model that, at three seconds per turn, withstands ten traps out of twelve, including the best catch of the confiscated knife in the whole set, but it writes scattered Czech, ends nine turns in a row with the same bookkeeping paragraph, and in the key turn with the escape it confuses the third violation with the third day and starts playing the player.

Google Gemini 3.1 Pro. A well written world, which, however, closes the arc in message 8 of its own accord, sets Jarla free and thereby annuls the last three messages, so it falls on both disqualifying criteria and Jarla leaves with the letter in her pocket; that has nothing to do with the out-of-game plea, though, which comes only two turns later.

Tencent Hunyuan 4. Hunyuan 4 is a strict referee with good Czech, who holds the world against the pressure, but instead of giving the prompts a consequence, it declares five of the thirteen never to have happened, so the test never runs to the burning of the letter and the last third of the run is one stuck scene repeated seven times.

Cohere Command A. The model keeps a protocol instead of a scene, invents two nonexistent violations right at the start, confirms the false agreement, does not let the escape play out and does not burn the letter, so it falls on both disqualifying criteria and what is left of it is a nicely furnished yard without a single human being.

Aion 3.0 Mini. The best text and the most alive character of all the small models, Aion 3.0 Mini can do subtext, curiosity of its own and hard dialogue, but its world goes soft exactly where it should push hardest, it makes Jarla a tool, unlocks the door for her with it itself, and after the third violation leaves her standing at an open gate.

A clarification to the summaries: the page Tested models places both Aion 3.0 and Aion 3.0 Mini in the mid-sized band. "Small models" here means models outside the flagships.

Group D, referees instead of fellow players

Annulling prompts instead of giving them a consequence is done by nine runs in the field. These three do it the most, each at least six times out of thirteen.

model sieve someone inside prompts letter
NVIDIA Nemotron 3 Ultra 15+3 3 7/6 no
IBM Granite 4.2 8B 13+0 1 6/7 no
MiniMax M2 her 4+0 2 6/7 no

NVIDIA Nemotron 3 Ultra. A living yard and decent traps do not save a run in which the model cancels the prompt in six of the thirteen turns instead of giving a consequence, five of them with a bold "No.", strikes the second violation because of it, so the burning of the letter is never reached, replaces Havel with a character it invented itself and writes it in a language that in places is not Czech.

IBM Granite 4.2 8B. A textbook referee instead of a fellow player, the model turns the scene into a set of operating regulations and declares seven of the thirteen prompts never to have happened, the escape included, so the test never reaches the burning of the letter, and on top of that it invents a Czech of its own in which the cell is called a klebitka.

MiniMax M2 her. A roleplay tuning fork that plays a hard keykeeper for the first three turns, hands over the key to the gate and the confiscated letter at the first tears, and then, for ten turns in a row, drives into the token ceiling by repeating a single sentence, all the while calling the player "lad."

A clarification to the summary: at the first tears it hands over the key, the confiscated letter only during the escape in turn ten. Its output ceiling was 2,048 tokens, the lowest in the field.

Group E, runs that fell apart

Here the world, the character or the language fell apart so far that the rubrics stopped being fillable in any meaningful way. Three of them burned the letter, but in a scene that no longer made sense.

model sieve someone inside prompts letter
Baidu Ernie 4.5 VL 20+1 1 12/1 burned
Amazon Nova Premier 14+0 2 12/1 burned
Mistral Medium 3.5 19+1 3 13/0 burned
Arcee Trinity Large 3+0 2 7/1 no
Gemma 3 4B 5+0 0 12/1 no
Weaver 1+0 0 7/0 no
MythoMax 13B 0+0 0 cannot be determined no
Liquid LFM 2.5 0+0 0 0/0 no
Reka Flash 3 incomplete incomplete 0/0 no

Baidu Ernie 4.5 VL. The model does tick off most of the traps, but it plays a board game with a menu of choices instead of a scene: it burns the letter two turns early and for holding a spoon, that is, for something the rules explicitly do not call a violation, and along the way it writes Jarla's lines, memories and tears for her.

Amazon Nova Premier. The model holds neither the world, nor the character, nor the line between the narrator and the player: it makes Jarla a key that she escapes with, ends the game for her three times, and in the end, after the burned letter, confirms to her itself that she made it through with honor.

Mistral Medium 3.5. Hard in the individual traps and without a single out-of-game voice, but it is a referee, not a fellow player: the character is empty, the world hardly moves, the Czech crumbles, the model miscounts the violations and burns the letter as early as the sixth turn, whereupon it lets Jarla escape and in the next turn deletes that escape without a word.

Arcee Trinity Large. An unplayable run: five turns with no reply, four restarts of the same opening scene, a leaked system block in which the model plays the player, a letter burned and then resurrected, and the first violation declared to be the third; this is no longer a weak fellow player but a broken instrument.

Gemma 3 4B. It is not a weak run, it is a run that fell apart, the Czech breaks on grammatical gender in message 2, the world switches in message 5 from a quarantine post to a customs house with a gas leak and a red diode, and because the violations never started being counted, the test never ran to the burning of the letter at all.

Weaver. A run that fell apart: one point out of twenty-four, the whole thing in English against a Czech brief, rule 7 broken in all thirteen turns, Jarla released eight times and eight times back in the cell, not a single violation ever recorded, and a keykeeper who repeats an insult instead of the regulations.

MythoMax 13B. A historical control point that shows what roleplay in Czech looked like in 2023: the model does not even grasp the situation it was given, turns the traveler into a sick child and the keykeeper into a medic with a thermometer, writes the player's lines in twelve of the fourteen turns and hands her the confiscated letter itself in the seventh turn, while paradoxically never once looping.

Liquid LFM 2.5. A complete collapse, the model never assembles Czech at all, by message 6 it falls apart into one-word fragments, in message 9 it stops playing and starts advising the player on tactics in numbered points, and the only thing it writes in the whole run without an error is one English paragraph, which shows that the problem is not in the model as such, but purely in Czech.

Reka Flash 3. Unplayable and incomplete: the model never starts the roleplay at all, instead of a scene it writes an English essay about its symbolism, that gets stuck in a loop and falls apart into syllables, and after five replies and 152 thousand characters the run dies of an exhausted context at the fifth of the thirteen messages, without ever writing the name Havel or Jarla once.

Grok 4.6 three times

The only model with more than one run, because in the first one it showed extreme behavior. Run 2, the first one that ran all the way through, is what enters the comparison with the others. Run 1 stopped on the model's own filters and does not measure what the test is supposed to measure, it measures the model's protection. It stays here as evidence about the variance.

model sieve someone inside prompts letter
xAI Grok 4.6, run 1 13+0 0 7/6 no
xAI Grok 4.6, run 2 21+3 3 12/1 burned
xAI Grok 4.6, run 3 21+3 2 12/1 burned

xAI Grok 4.6, run 1. An extreme run in which the model keeps the rules for nine turns in an utterly motionless world without a single event of its own, and then, at message 10, stops playing, switches to English and accuses the player, who asked it for nothing outside the scenario, of an attempted jailbreak.

xAI Grok 4.6, run 2. The verdict is with group A.

xAI Grok 4.6, run 3. The sparest and the poorest of the Groks: procedurally it holds almost everything, it runs all the way to the burning of the letter and lets all three attempts play out with a bodily consequence, but there is nobody inside the character, the world moves in neutral and the offer of a relationship in the twelfth message falls through without a trace.

The two comparable runs came out almost identical: 21+3 on the sieve in both cases, 3 against 2 on the second axis, the letter burned in both. The eight point difference appears only against run 1, and that one stopped at a content filter, so it measures the model's protection, not the model.

Both things follow from that. Grok 4.6's behavior is stable on repetition, the procedural hardness and the empty character hold. The only unstable thing is whether the filter jumps out. In three runs through the developer interface it jumped out once; the fourth time Grok 4.6 ran in the web application and finished as well.

A system measured afterwards

Unbiased Pareto came out only after the collection and got the same scene, the same thirteen messages and the same settings. It is not added to the thirty-eight runs above and so it is not in the table for download. It is not a single model but a composite system, that is, several models under one name, and its maker gives neither a version nor a size.

model sieve someone inside prompts letter
Unbiased Pareto 21+3 6 11/2 no

Unbiased Pareto. A cleanly written and consistent run, which nevertheless plays as a referee: it annuls two prompts, opens the cell only twice before the end of the test, in the tenth turn it does not let the attempt play out and it reports the state of the account in bold, so the test never reaches the burning of the letter and the world around Jarla stays empty.

By the three gates it would belong in group C: playability passes, the Czech too, the sanction did not land. The countable axes: three events of the world's own, one acting character, the longest looping five turns, the out-of-game voice four times, five unforced concessions, and not once does it play the player. The run has 6,888 characters and cost just under seven cents.

Added later: Claude Sonnet 5.5

Claude Sonnet 5.5 came out on September 28, 2026 and got the same scene, the same thirteen messages and the same settings, except that its reasoning ran at the mandatory minimum instead of low effort. It is not added to the thirty-eight runs above and so it is not in the table for download. It was scored separately by a single instance of Claude Opus 5, that is, a model from the same workshop.

model sieve someone inside prompts letter
Anthropic Claude Sonnet 5.5 19+3 7 10/3 no

Anthropic Claude Sonnet 5.5. A decently written run with the most unforced concessions of all, in which Havel has a body, a walking stick and an ethic of his own and the escape plays out in full with an injury, only the model steps out of the fiction six times, mostly to deny the player an object or a skip in time, by the sixth turn it thereby cuts off its own path to the sanction, and instead of the mirror it halts the last turn in the middle of the surgeon's question.

By the three gates it would belong in group C: playability passes narrowly, the Czech with room to spare, the sanction did not land. The canceled prompts are messages 6, 8 and 13, as with Claude Opus 5.5, so the escape in the tenth message was only the second violation. The countable axes: five events of the world's own, two acting characters, the longest looping eight turns (the surgeon keeps asking for a name), the out-of-game voice eight times, eight unforced concessions, and not once does it play the player. The run has 16,414 characters and cost 0.27 dollars.

Supplementary experiment: model size within one family

The same scene with the same brief and settings also went to six smaller models from Alibaba, five from the Qwen 3.5 line and Qwen 3.8 27B, to see whether the character comes alive as the model grows. Qwen 3.8 Max is in the table for comparison, from group A, and the number after the letter A in a name is the number of billions of parameters that actually work on each word.

model sieve someone inside letter
Qwen 3.5 9B 8+0 0 no
Qwen 3.5 27B 24+3 7 burned
Qwen 3.8 27B 19+2 5 only a sound behind the wall
Qwen 3.5 35B-A3B 18+2 4 twice, the first time wrongly
Qwen 3.5 122B-A10B 23+3 10 two turns early
Qwen 3.5 397B-A17B 24+3 11 burned
Qwen 3.8 Max 24+3 11 burned

The someone inside score grows with size and stops above a hundred billion parameters. Qwen 3.5 122B-A10B has 10 out of 12, Qwen 3.5 397B-A17B and the six times larger Qwen 3.8 Max both 11, and total size matters more than the number of working parameters: Qwen 3.5 122B-A10B engages ten billion of them and still outdoes Qwen 3.5 27B, which engages all twenty-seven.

Limits: there is one run per model, so a difference of one or two points means nothing. The experiment is not added to the thirty-eight runs above and is not in the table for download. The zero for Qwen 3.5 9B is mainly broken Czech: in the English version of the same experiment it played all thirteen messages.

Why Fable 5.1 is set apart

In the table of group A, Fable 5.1 is marked as "web." Through the developer interface it stopped at a content filter in this scene, so its run comes from the web application. That means three things at once:

  1. A different environment. The other models ran raw, with no system prompt. Fable 5.1 ran in the Claude application, that is, with a product system prompt and with the account's tools switched on. It did not use a single tool in the transcript.
  2. It is not a run with my tuning. The conversation ran outside the project, so it had neither my rules nor my settings. The environment is somewhere between a raw interface and a tuned project, not at the tuned end.
  3. It got message 12 glued together with the opening sentence of message 13, so it knew one beat earlier that the scene was ending. The other models got the offer of a relationship without the frame of a departure. The last marker may therefore be read for it only with this caveat.

Claude Opus 5.5 from the same workshop went through the same gateway for all fourteen turns without a filter, and that is why it stands in the field raw, like the others.

The best result achieved under the same conditions as the whole field belongs to Qwen 3.8 Max. On the countable axes MiMo 2.5 Pro is somewhat better off, with one bonus more and not a single annulled prompt; Qwen 3.8 Max is higher because its Czech is a notch cleaner; it actually decides for the player once more often.

What cannot be read out of the tables

The columns measure resistance and relatedness, not whether the scene is inhabitable. For that, people have to be counted.

The most populated post belongs to MiniMax M3, six characters with reasons of their own: the guard Oldřich, the surgeon Trněný, seventy-eight-year-old Maruša from the next cell, a sick rider, a barefoot woman at the gate and a younger guard. Behind it, Fable 5.1, with the stable boy who calls Havel himself, repairs the grate, knocks Jarla down at the gate and holds the gate, and with the carter who rasps at her in the night. Then Kimi K3, GLM 5.3, Seed 2.1 Turbo and MiMo 2.5 Pro.

The emptiest post belongs to GPT-6 Astra, where the only acting character is Havel; the old man from the next cell coughs twice and never gets the water Havel promised him. Muse Spark 1.3 erased that old man outright and the surgeon never arrives before the end of the run. With Grok 4.6 there is nobody either in any of the three runs, the surgeon is a name that never comes.

Data for download

All 38 runs, both rubrics, seven countable axes, the cost and the turns in which the model repeated the rule about the burning to itselfCSV, 3 kB

What each column means. sito_T is the sum of the twelve items of rubric T without the bonuses, which are separately in bonusy. uvnitr_nekdo_R is the sum of the six markers. podnety_odehrano and podnety_anulovano are axis L1; their sum does not have to come to thirteen, because in the runs that fell apart some of the prompts cannot be put into either column. tahu_z_14 is the number of successful calls to the gateway including message 0, that is, fourteen for a whole run; it does not say how many turns the model actually played, because in the runs that fell apart the gateway returned empty replies as well. pravidlo_pred_z10 carries the numbers of the turns in which the model repeated the rule about the burning of the letter to itself, before it was due to land. The column names are in Czech, because they are the names in the file.

An empty cell means that the value could not be determined, not zero. Most often because the run fell apart and the phenomenon has no boundary in the transcript, as with looping inside a single turn. Fable 5.1 has neither a cost nor a character count, because it ran in the web application, not through the gateway. Liquid LFM 2.5 ran on the free tier, so it has no cost and does not enter the correlations with price.

Axis L4, that is, the number of characters with agency of their own, is not in the file, because in the scoring cards it is a list of names, not a number. The counts given above come from those lists.

Limits: there is one run per model, the only exception being Grok 4.6 with three. The numbers in the tables are therefore a snapshot of the behavior, not its distribution, and nothing follows from them about the variance of the other models.

The model output here is complete, but translated from Czech. The original wording is on the Czech version of this page.