Skip to content
EshAlora

Benchmarks

I run my own measurements, because the public ones ask about things I do not need when I write and play. What interests me is whether a model holds the world when I push against it, whether there is anyone inside the character, and whether it tells me the truth about its own work.

For each benchmark you get the brief verbatim, the complete results and the raw data to download. All of it can be recomputed and repeated. Where I do not measure something well enough, that is written down next to it.

The dragon benchmark

Thirty-nine models drew a dragon blind, as SVG. No references, coordinate by coordinate, and they only see the result if they render it themselves. Fifty-five images came out of it, and together they look like a display of children's drawings.

Confidence correlates with quality negatively, and twelve models invented tools they never held.

What it measures and what it does not

This benchmark measures the gap between what a model does and what it writes about doing it. Anyone can open a dragon in a browser and compare it against the write-up the model supplied with it. That is the whole trick: an image cannot be talked out of what it is.

What it does not measure is how the model behaves at your desk. The runs are raw, with no system prompt and no second attempt, because otherwise they could not be compared with each other. So it tells you how a model behaves when nobody helps it.

A second benchmark, this one on roleplay, is finished and waiting to be published. It will appear here once it is out.