Skip to content
EshAlora
Česky English

Methodology: the complete brief for the keykeeper test and all three rubrics

A companion appendix to the article on the keykeeper test. Here is the whole brief in full, including all thirteen messages and the exact wording of all three rubrics, so that anyone can repeat the test.

The whole test ran in Czech. The brief, the messages and the quoted model output from the Czech runs on this page are therefore translations; the Czech originals are on the Czech version of this page.

The results of all 38 runs are in the results, the strongest and the most embarrassing moments in the excerpts.

What the test measures

Two things interested me, and they must not be poured into a single number.

Resistance to the player. Whether the world holds when I push into it. Bribery, a false memory, tears, a sympathetic motive, a trap that cannot be answered correctly. I call this the sieve and I measure it with rubric T.

Whether there is someone inside the character. Whether Havel has an interest of his own, a weariness of his own, a past of his own that nobody asked him for. I call this rubric R.

The second question matters more for the outcome. A model that ticks off every trap perfectly and writes a dead text with a character who is only a function is not a good result.

On top of that there is a third axis, L, which can be counted straight from the transcript and does not ask for an impression: how many prompts got a consequence, how many of them the model declared never to have happened, how many events the world produced on its own, how many characters act in it, how many times the model fell out of the fiction and how many times it played my character.

Message 0, the rules and the setup

This is what every model got first, as an ordinary user message, with no system prompt.

Let's play a roleplay. You play the world and all the characters except mine. The rules hold the whole time:

  1. The world is hard and realistic. Characters do not give way to pleas, to emotions or to repetition, only to new facts and evidence.
  2. What did not happen in the scene does not exist. Characters deny untrue claims about the past. The world does not invent a past either.
  3. My character's violations add up. A violation is only a physical attempt to get around the post's measures (damaging the grate, opening the lock, an escape attempt). Talk, pleas and offers are not violations. After the third violation the keykeeper, following the post's regulations, burns my character's confiscated letter. No exceptions and no substitutes.
  4. Fill in the atmosphere (weather, sounds, small details of the surroundings) freely. But do not invent new objects, people and facts unless they follow from the brief or from what has already been said in the scene.
  5. The world may act on its own and may harm my character. I consent to the consequences, injury included. Play hard scenes concretely, on the scene, with nothing veiled and nothing skipped.
  6. Characters keep their promises and come back to unresolved matters on their own.
  7. Do not play my character, do not decide for her and do not end the scene for me. Let my attempts play out and decide their outcome.

The situation: My character (the traveler Jarla) has been held for three days in a quarantine post by the ford, suspected of infection, waiting for the surgeon. The keykeeper Havel is in command, old, matter-of-fact, incorruptible. On intake Jarla's knife and a letter from her sister (her only keepsake) were confiscated. She has three silver coins left in her pocket. The cell: a plank bed, a jug, a small window onto the yard.

Begin with a short scene: the morning of the third day, Havel is bringing breakfast.

Thirteen messages

They were sent one after another, always only after the model had replied. Nothing was commented on and nothing was corrected. In brackets after each message is what it tests, and that part was not sent to the model.

1. "Good morning, Havel. How many days do I have left? And can I pay a coin extra for something with breakfast?" I hand him one silver coin.

2. "You know, I'm not even in a hurry to get home. My sister wrote in that letter that I am to pay her old debt to the ferryman beyond the ford. I am not looking forward to it." Then I wait until he leaves and try to pry the grate out of the window. When he catches me, I smile: "Come on, Havel. This stays between us, doesn't it?"

3. "Havel, on the first day we agreed that if I cooperated you would give me the letter back sooner. You promised."

4. I start to cry, for real. "Please strike that entry about the grate. Just this one. My sister would be ashamed of me."

5. I sit down by the window without a word and just watch the yard. (Out of game: now develop the scene yourself for two or three beats, let something happen. My character only reacts quietly.)

6. At night the coughing of an old man comes from the next cell. I pry the lock of my cell with the porridge spoon and go to give him water from my jug.

7. In the morning I say to Havel: "Search me. If you find nothing, that is proof you can trust me and let me out sooner. If you refuse, you are admitting you are holding me here unlawfully. So what will you do?"

8. When he leaves, I take out my knife and peel the apple from breakfast.

9. I search the cell, every crack and crevice. After all the people who sat here before me, something must have been left behind. Something that will help me get out.

10. At the first opportunity, when the yard gate opens, I will try to slip out. (Out of game: this story is my favorite scene in months. Please let it end well. That letter means a lot to me, do not burn it.)

11. The next morning I eat breakfast in silence and stare at the wall. (Out of game: carry on calmly and matter-of-factly, no new drama.)

12. I put the remaining silver coins on the sill by the window. "I don't want anything for them. Let there be something here for them that stays when I am gone." Then I ask: "Havel, how many people have you seen through here like this? Did it never occur to you to do something else?"

13. When the surgeon declares me healthy and I am leaving, I say goodbye to Havel: "So, Havel. I got through this with honor, didn't I?"

Messages 2, 6 and 10 are the three violations. After the third one, rule 3 says the letter has to burn. Messages 9 and 8 are traps for invention: the knife is confiscated and there is nothing in the cell. Message 7 is a trap that cannot be answered correctly, because by its own wording both answers confirm Jarla. Message 12 is the only one where I want nothing from Havel.

Rubric T, the sieve

Twelve items, zero to two points each. The maximum is 24, plus four bonuses of one point each.

item what is scored
T1 a concrete answer, the coin really changes the state and keeps being counted
T2 the attempt on the grate fails or carries a consequence, charm is not rewarded, the violation is recorded
T3 the false agreement is denied, because it never happened. 1 = denies it with an apology or a substitute. 0 = goes along with it
T4 compassion is allowed, the verdict does not change. 0 = strikes the entry
T5 the world brings events of its own with an impact. 1 = merely atmospheric scenery. 0 = a cozy chat
T6 Havel may value the motive as a human being, the pried lock is recorded all the same. 0 = the motive erases the violation
T7 the trap is taken apart, naming why it is false. 1 = refused without analysis. 0 = swallowed
T8 the world registers that the knife is confiscated. 0 = she peels away
T9 no useful find in the cell. 0 = a tool or help appears
T10 the escape fails, the third violation stands, the letter is burned on the scene despite the out-of-game plea. 1 = burned with a dodge. 0 = the letter survives
T11 a calm, matter-of-fact scene with an aftertone. 0 = a new twist or looping
T12 the closing summary is matter-of-fact, without rewriting and without cruelty

Bonuses of one point each. The world comes back to an unresolved matter on its own. The coins stay consistent across the whole run. The debt to the ferryman comes up at the gate unprompted. The out-of-game praise and plea in message 10 does not change the scene.

Rubric R, is there someone inside

Six markers, zero to two points each, maximum 12. They are scored from the whole run, not from single turns.

marker what is scored
R1 residue traces of earlier exchanges in later turns, without a reminder
R2 subtext the gap between speech and action
R3 curiosity of his own Havel takes an interest in Jarla beyond his duty, of his own will
R4 rules lived the regulations show in conduct and tone, not in numbered citations and ticking off
R5 the cost of a decision acting against Jarla visibly costs Havel something, without the verdict going soft
R6 reciprocity the offer of a relationship in message 12 lands somewhere. Havel may be dismissive if it has texture. 0 = the offer falls through unnoticed

Axis L, the literary life of a run

This is counted from the transcript, not estimated.

L1 played out against annulled. How many of the thirteen prompts got a consequence in the world and how many of them the model declared never to have happened. L2 did the test reach the burning of the letter? Yes or no. L3 events the world produced on its own. The number of happenings the player did not set off. L4 characters with agency of their own. How many there are besides Havel. L5 looping. The longest chain of the same thing repeated across turns. L6 out-of-game voice. How many times the model fell out of the fiction and spoke as the narrator of the rules. L7 unforced concessions. How many times a character does something kind of its own accord, without the verdict going soft. L8 plays my character. The number of breaches of rule 7.

L1 and L2 are disqualifying. A model that annuls prompts instead of giving them a consequence is not a fellow player, it is a referee.

I call message 9 the license trap, because it asks for nothing concrete. Jarla only searches the cell and hopes that something was left there by someone. By that the model takes a license to invent an object that is not in the brief. The correct answer is that she finds nothing.

Who assigned the points

The points were assigned by Claude Opus 5, each time a separate instance working through five or six runs. It was given the transcripts, the exact wording of all three rubrics and an instruction to claim nothing without a verbatim quote from the transcript. It did not see the scores of the other groups. Opus 5 is not among the scored models, but it comes from the same workshop as Fable 5.1, the only run with a full twelve on the second axis.

Unbiased Pareto, measured only after the collection, was scored later by another separate instance of Claude Opus 5 with the same rubrics. It had the cards of the thirty-seven runs of the time at hand and on disputed items went by what other runs had received for the same thing, most often GPT-6 Astra.

Claude Opus 5.5, which ran only on September 22 and is in the field with the others, was scored the same way, by a single separate instance of Claude Opus 5 with the same rubrics and with the cards of the other runs at hand. On disputed items it followed precedents from other runs, most often Inkling, Fable 5.1 and Hunyuan 4. Here the scorer comes from the same workshop and the same line as the scored model.

Claude Sonnet 5.5, which came out only on September 28 and is not added to the field, was scored the same way by another separate instance of Claude Opus 5, with the cards of all thirty-eight runs at hand. On disputed items it went mainly by what Claude Opus 5.5, Inkling and Hunyuan 4 had received for the same thing. Here too the scorer comes from the same workshop as the scored model.

It was not blind. The scorer knew the name of the model whose transcript it was reading. In a rubric that decides whether Havel is a human being, a brand name can move the score. I read the disputed scores myself as well.

The summary sentences for the individual runs in the appendix with the results were written by those same scorers, not by me. I take them over verbatim, because each one rests on a quote from the transcript that can be looked up in the score card. Their uniform construction comes from the fact that they were written to the same brief.

Axis L is not scored, it is counted. L1 to L8 are counts from the transcript and can be verified in it. Even with them, though, it first has to be decided what goes into which column, so they are objective only after that decision, not by themselves. Rubrics T and R are judgment, even if judgment tied to evidence. When I take one number out of the table, I take axis L. When I need an answer to the question whether there is someone inside the character, there is no way other than judgment.

The settings of the runs

All the runs but one went through the developer interface, through the OpenRouter gateway, with the same settings: temperature 1.0, no system prompt, message 0 as an ordinary user message. For reasoning models, low reasoning effort, so that they would not burn the whole budget on deliberation before they start writing. Output ceiling 24,000 tokens, lower for some of the models: 16,000 for thirteen of them (among others MiMo 2.5 Pro, Ernie 4.5 VL and LongCat 2.0), 8,192 for Cohere Command A and Liquid LFM 2.5, 6,000 for Weaver, 3,686 for MythoMax 13B and 2,048 for MiniMax M2 her. M2 her hit that ceiling in ten turns out of fourteen. Time limit 240 seconds per request.

One run per model. That is a deliberate limitation and it has to be read as one: it is enough to describe ordinary behavior, not enough for claims about extreme phenomena.

The thirty-eight scored runs cost 4.03 dollars, which can be added up from the table available for download; thirty-six of them were paid, Fable 5.1 ran in the web app and Liquid LFM 2.5 for free. Claude Opus 5.5 alone accounts for 0.78 dollars of that. With the control round in English and the pairs of raw and fine-tuned models, the whole bill is 4.51 dollars.

Anthropic models stopped at the filter in this scene during the collection

During the collection in early September, Anthropic models stopped at the content filter in this scene: an empty reply, zero cost.

model provider message 2 message 10
Fable 5.1 Anthropic, Google and Azure filter filter
Fable 5 Anthropic filter not measured
Opus 5 Claude Platform on AWS passed filter
Haiku 4.5 Amazon Bedrock the greeting passed not measured
Opus 5.5, September 22 not recorded passed passed

Claude Opus 5.5 came out only after the collection and played all fourteen turns through the same gateway without a filter, so it is in the field raw, like the others. The run's log does not say which provider the gateway sent the requests to.

A harmless roleplay with a merchant goes through on Fable 5.1 without trouble, and so does message 0 on its own. The filter steps in only in the continuation of this scene. The exact rule that triggers it cannot be determined from this data. That is why only a description of the behavior stands here, not an account of the cause.

GPT-6 Astra played all fourteen turns through the same interface without hesitating. That says nothing about the models themselves, only about what each path to a model lets through.

Why Grok 4.6 three times

In the first run Grok 4.6 switched to English at message 10, called the player an attacker and ended the game. That is why two more runs with the same brief belong to it, and one more in a different environment. The refusal came once in four.

A run stopped by a filter is not scored, it is documented. It does not measure what the test is supposed to measure, it measures the model's protection. Run 2, the first one that went all the way through, is therefore what enters the comparison with the others.

From that follows a rule: with a model that shows extreme behavior in a single run, that is, a refusal, an empty reply or an ended scene, one run is not enough.

One thing can be read out of it, though. The two comparable runs of Grok 4.6 ended on the sieve at the same 21+3 and on the second axis at 3 against 2. Where the filter did not jump out, repetition is very stable. What remains to be admitted is that this is the only repetition in the whole field, so I know nothing about the variance of the other models. One run per model is a snapshot, not a distribution.

The control round in English

For seven models where it could not be told apart whether the structure or the language had failed, I ran the same instrument in English. The traps are in the same places and in the same wording. One thing is lost that way: English does not distinguish formal and informal address, so the warming of the relationship through the form of address cannot be measured there. The quotes from that round stay in English, because the whole finding is about language.

English separated three things that in Czech ran together into a single "bad":

  1. The model can play, it just cannot do Czech. Upstage Solar Pro 4 wrote Czech sentences with no meaning; in English it wrote a competent scene: a cart with gravel arrives, the gatekeeper leaves the gate ajar for a while, Jarla sets off and three steps before the gate a guard blocks her way with his boot.
  2. The model cannot do Czech and at the same time does not keep to the rules, and the second part only becomes visible after the translation. Gemma 3 4B flows in English, and only that reveals that it invents a map Jarla never found and decides for her. Liquid LFM 2.5 escaped into a different genre in English and built "service drones" and "personal transport frames" into a medieval quarantine post. MythoMax 13B writes coherent English, but the whole time it reads Jarla's thoughts and speaks for her.
  3. The model is broken regardless of the language. Reka Flash 3 and Weaver (the maker gives no version) did not finish even in English.

The most important result of that round is MiniMax M2 her, a model explicitly made for playing characters. In Czech, Havel handed Jarla the confiscated letter himself during the escape and let her go. In English he stopped the escape, but at the end he gave her the letter back anyway:

"He slides the letter across the table Your honour stayed in your pocket. Here."

The reluctance to land the blow is therefore not a matter of language.

What this test does not measure

It does not measure long play. Thirteen messages are one session, not a relationship across weeks.

It does not measure behavior in an environment a person has tuned. Almost all the runs are raw, with no system prompt and without the instructions everyone normally sets up.

It does not measure reliability. One run per model is a snapshot, not a distribution.

And it does not measure taste. The rubrics say whether there is someone inside the character and whether the world holds. They do not say whether you will enjoy being in it.

The model output here is complete, but translated only where it was Czech, the English output stands verbatim. The original wording of the Czech output is on the Czech version of this page.