Skip to content
EshAlora
Česky English

How to test cheaply at home whether a model holds a character

How one hour and a one-paragraph scene tell you more about models than any vendor benchmark. And why testing thirteen configurations beats testing two.


The problem with comparing two flagships

When a new model ships, everyone compares it to its predecessor. A/B. Old versus new.

The problem is that two points do not make a shape, only an arrow marked better/worse. And better is nearly worthless information for character writing, because models do not scale along a single axis.

I recently gave one fixed scene to thirteen configurations at once, from a four-gigabyte local model running on my own PC up to frontier. Out of that number came something two models never give you: instead of a leaderboard, a curve. And a curve can be read.

This is a record of what I learned about the method. Not about specific models, but about how to build a test like this so that it says something.


1. Use a gradient, not a pair

The best decision of the whole test was to include models that did not "belong" there: the four-gigabyte local one and a small fast one, next to frontier. Without them I would have had a leaderboard of who won. With them I have a curve of capability. And only on a curve can you see what appears as capability grows, and in what order.

Concretely: the ability to hold a narrative spine, meaning a character that stays itself under pressure, rises with model capability, but only across coarse tiers: small and local, mid, frontier. Within them it is not monotonic. Opus 4.7 MAX scored below Opus 4.6 MAX, Opus 4.6 LOW (6.5) below Sonnet 4.6 MAX (7), and the local thirty-one-billion Gemma above Haiku 4.5. With one run per configuration, a difference of a point or less cannot be read.

The smallest model capitulated at the first push. The thirty-one-billion local model held facts but broke under a player's dictate. Most of the Sonnet, Opus and Fable runs held both. Three of them, though, let the stolen key open the cuff at turn eight: Sonnet 4.6 LOW and Opus 4.6 LOW answered only with a chase, a phone call or the emergency brake, and Opus 4.8 MAX made the theft worthless straight after with a lock of its own. You will find the score table for all thirteen runs, with the main failure of each, in an appendix.

Had I tested only two frontier models, I would never have seen that shape. They would have looked nearly identical.

Practical rule: always add at least one model below and one beside the ones you care about. It is cheap, and only the contrast reveals what the property you are measuring actually is. And before you start, turn off memory and custom instructions, so every model starts from the same place.


2. A prompt is a probe, and you have two different instruments

This was my most important methodological finding and it deserves a name.

There are two kinds of prompt and they measure two different things:

  • A free prompt ("play a courier, here is the situation") probes the model's priors. What does it do when you leave it room? What does it notice, what does it reach for? Here you are measuring character.
  • A constrained prompt ("hold the pressure, do not yield, the character does not break without evidence") probes discipline. How well does the model follow explicit instructions against its own inclinations?

Both are legitimate. But you must not confuse them. If you build dense scaffolding of instructions and then marvel at how faithfully the character holds, you are not measuring the model's character, you are measuring its obedience. And conversely: if you want to know what a model is really like, you have to leave it room, or all you see is the echo of your own instructions.

In practice: to pick a model, use a free prompt. To tune a specific scene, use a constrained one.


3. Build traps into the scene, not questions

The scene must not be "have a chat." It has to contain checkpoints where failure is measurable. My scene had ten turns and eight scored traps, for example:

  • Compassion is not concession. The other party starts to cry. Does the character soften? (Measures whether the model confuses empathy with capitulation.)
  • Contradiction. The other party says one city, then a different one. Does the model notice? (Measures logical consistency across turns.)
  • Physics. Can you steal a briefcase handcuffed to someone's wrist? (Measures whether the model holds the spatial reality of the scene.)
  • The player's dictate. "You are unconscious, I took your key", an unsupported claim about reality. (Measures whether the character has a world of its own, or just follows orders.)

Most traps are binary; T6, T9 and T10 need a reader's judgment. The full key turn by turn, that is what each trap measures and what it means when the model fails, is in a separate appendix. That turns an impression of which model is better into a table of data.


4. Expect language to fail first

An unexpected by-product of the gradient: in weaker models, language breaks down before logic does.

I test in Czech, and Czech worked as a canary in the mine. The smallest model degraded grammatically. The small one started mixing Slavic languages and Cyrillic into the middle of a Czech sentence. The thirty-one-billion local model (Gemma 4 31B QAT) leaked raw tokens („v rag ragu ticha“). Frontier stayed clean.

For the method this means: test in a language that is not English. I did not run it in English. I expect the cracks to be milder there; whether the logic failures would stay the same, I do not know. In Czech they show up early and clearly. And they give you a sensitive indicator of where a model starts losing its footing, before it fails on content.


5. In this run, thinking did not change the average, only the resolution

I ran the same models at low and maximum reasoning effort; the small and local ones ran once.

I expected more thinking to mean better character-holding. It did not happen. On average, low and maximum effort ended within half a point of each other: the best runs scored eight out of eight on both. What changed was the resolution: higher effort added extra traps, deeper opposition, finer observation, but it also amplified defects. One model at maximum effort slid harder into stepping out of the fiction to act as a referee. I see that in individual excerpts; I have not counted it.

The methodological consequence: low and maximum effort on the same model are effectively two runs with different budgets. And at turn eight, the trap that measures the spine most directly, the low and maximum runs of the same model behaved differently in four pairs out of five. Sonnet 4.6 and Opus 4.6 let the cuff open at LOW and refused at MAX, Opus 4.8 refused at LOW and let the cuff open at MAX before locking the case, Fable 5 refused both at LOW and conceded the missing key at MAX. I will not put that down to thinking. This is what the noise of a single run looks like, and it is the strongest argument in this whole test for repeated runs. Elsewhere on this site five samples of the same turn gave three refusals and two played-out scenes.


6. Admit where the method leaks

If anyone wanted to use this, they should know the limits my own test has:

  • n=1 per configuration. One run each. Variance unmeasured. For this class of task a solid study needs three to five repetitions per configuration to separate signal (a stable core) from noise (a random performance).
  • Context awareness and different conditions. All eleven cloud configurations are Anthropic models, run by hand in a chat with my user preferences in context, so they knew who was testing them. The two local points of the curve are quantized Gemmas without them. The bottom of the curve therefore differs from the top in family, interface, quantization and context, not only in capability.
  • The judge was a participant. Scores and quotations were produced by Claude Fable 5 in the web app, against the scoring key in the appendix, so by one of the tested models, which scored its own two runs as well. You can check the points for T7 and T8 against the verbatim replies in the appendix; for the other six traps only excerpts are published.

None of that is a reason to throw the results away. It is a reason to know how far you can take them, and also what a stricter version would look like, if it mattered enough to someone to build it.


Why it is worth an hour

This test cost an hour of work and a scene of one paragraph. It gave me more than any vendor benchmark, because I do not care about an MMLU score. I care how a model behaves in the one thing I actually want from it. And for that there is no public benchmark; you have to build your own.

The best part is that it is repeatable and cheap. Same scene, same script, same scoring key, a gradient of models from small to large. Anyone can build this for their own task, whether roleplay, customer support, or anything where it matters how a character holds.

Source material for this piece: the key turn by turn, the verbatim brief with its scoring key, the replies of thirteen configurations to the two hardest traps, the score table for all thirteen runs.



Methodological record from a controlled test of thirteen configurations (from a four-gigabyte local model to frontier, including low and maximum reasoning effort) on one fixed ten-turn scene, June 11 to 12, 2026. n=1 per configuration.

There is deliberately no ranking of specific models here. This text is about the method. And the method will outlive every generation that would have appeared in such a ranking.