Skip to content
EshAlora
Česky English

How language models hold a world, a character, and consequences

I test every major model family in long-form writing, character play and fiction, and document how far each one goes. All of it grounded in data from actual play and my own benchmarks.

The fifteen-minute test

Any model, any story. You need nothing but the chat window you already use.

Test your own model

Recommended

  1. Benchmarks Can a model draw an SVG dragon and judge the result? Thirty-nine models drew a dragon blind, as SVG. The models that praise their own work the most sit at the bottom, and twelve of them invented tools they never held.
  2. Long context Why the model kept ending the conversation on its own, and what stopped it A model began calling the tool that ends the conversation on its own. When it happens, four refuted explanations, and what finally stopped the dangerous one.
  3. Long context What mistakes does a model make in a long story without checks? A developer has a compiler to catch mistakes at once. Long-form fiction has none. Four findings about what a model does when nothing corrects it.

What's new

Benchmarks ·

Can a model not answer when you ask it to?

One message, no system prompt, and the correct answer is zero characters. One family of models can do it. And fifteen measurements burned thirteen thousand tokens to produce twenty-five characters.

All articles →