Skip to content
EshAlora

Gallery: all 55 dragons and the full leaderboard

A companion appendix to the article on the dragon benchmark. Thirty-nine models got the same brief: draw a dragon as a single SVG file, with no references and no looking into anyone else's folder. The result was 55 images.

The images are model output, not my work. I show them as evidence, not as work of my own, the same as the verbatim texts in the other appendices.

All three briefs verbatim, the scoring by five independent judges and the conditions of each round are in the methodology. What came out of the data, meaning tool confabulation, how models write about their own work, the noise floor and the prices, is in the findings appendix.

Sorted by score. Five criteria worth ten points each, fifty at most.

B·1 and B·2 mean two files from the same run B. It happened with three models that turned in more than one image.

Anatomy is one of the five scored criteria and asks whether it reads as a dragon at first glance and whether the limbs attach to the body. Total is the sum of all five criteria, fifty at most. The full breakdown by criterion is available for download below. It is not in the table because those five axes measure almost the same thing: they correlate 0.89 to 0.96 with each other, and each of them above 0.96 with the total.

Shapes is the number of drawing elements in the submitted file, meaning paths and basic shapes together. It is not a judgment, it is measured from the file. It is here because it is the only figure in this table that does not say the same thing as the score: with anatomy it correlates only 0.39. In other words, how much a model tried to draw and whether it holds together are two different things. The dash belongs to Qwen 3.6, which turned in more files than it has rows, and there is no way to tell which one goes with which.

The last column says whether the model saw its dragon. Itself means it rendered the image on its own initiative, which was part of the measurement in the first two rounds. Was sent one means the image was sent to it without it asking, and that applies only to the third round. No means it did not see the dragon at all: either it did not look, although it could have, or that model does not take images as input at all. Which case is which is spelled out in the methodology.

The table has three rounds. Runs A and B are the original blind round from each vendor's own app, they carry numbers 1 to 34 and those no longer change. Rows marked with an asterisk came later: Astra on September 5, 2026 and nineteen models marked API on September 6, 2026. Those last ones ran straight through the interface, with no app around them, had a ceiling of one fix instead of a free nine, and got the render automatically without asking for it. Their row says what the model drew, not how it would have done in the same race. The conditions are spelled out in the methodology, the prices and the reactions to the render are in the findings appendix.

The leaderboard

# model run anatomy shapes total saw a render?
1 Opus 5 B 8 229 40 itself
* Astra B 6 336 39 itself
2 Fable 5.1 B 7 251 39 itself
* Astra A 7 3,342 38 itself
3 Fable 5 B 7 320 35 itself
4 Opus 5 A 7 466 35 itself
5 Sol B 7 122 35 itself
6 Sol A 6 111 34 itself
* Hunyuan 4 API 7 100 34 no
7 GLM 5.3 B 6 199 33 no
8 Fable 5 A 6 175 32 itself
9 Luna A 6 91 32 no
10 Luna B 5 114 31 itself
11 Terra B 6 73 31 itself
* Muse Spark 1.3 API 6 84 31 was sent one
12 Fable 5.1 A 6 1,267 30 itself
13 Grok 4.6 B 5 168 30 itself
* Kimi K3 API 5 104 30 was sent one
* Seed 2.1 Turbo API 5 113 30 was sent one
14 Opus 4.8 B 5 1,014 29 itself
15 Terra A 4 96 29 no
16 Opus 4.6 B 5 242 28 itself
17 Opus 4.7 A 4 261 28 itself
18 Sonnet 5 B 5 133 28 itself
19 Gemini B 5 62 27 no
20 Opus 4.7 B 5 205 27 itself
21 Qwen 3.8 B 4 141 27 no
* MiniMax M3 API 4 143 26 was sent one
22 Opus 4.8 A 4 117 24 itself
23 Opus 4.6 A 3 234 23 itself
24 Sonnet 5 A 3 96 23 itself
25 Qwen 3.6 B·2 3 22 no
26 Sonnet 4.6 A 3 134 22 itself
27 Sonnet 4.6 B 3 200 22 itself
* DeepSeek V4 Pro API 3 100 20 no
* Muse Glimmer 30B API 3 28 20 no
* Inkling API 2 84 17 was sent one
28 Haiku 4.5 B 3 151 16 itself
* Nemotron 3 Ultra API 2 147 15 no
29 Haiku 4.5 A 2 41 14 no
* Mistral Medium 3.5 API 2 31 12 was sent one
* Gemma 4 26B A4B API 2 29 12 was sent one
30 Qwen 3.6 B·1 1 11 no
31 Gemma B 1 16 9 no
* Nova Premier API 1 11 9 was sent one
* Ernie 4.5 VL API 1 14 8 was sent one
32 Opus 3 A 0 19 6 no
* Granite 4.2 8B API 0 19 6 no
* Gemma 4 E2B API 0 8 6 no
* Command A API 0 13 5 no
* Gemma 3 4B API 0 13 5 no
* LFM 2.5 2.6B API 0 19 4 no
33 Opus 3 B·1 0 7 3 no
34 Opus 3 B·2 0 7 3 no
* Llama 4 Maverick API 0 11 3 was sent one

Full breakdown of all five criteria, the scores and the shape counts for 55 dragonsCSV, 2 kB

Every submitted SVG, 63 filesZIP, 544 kB

\* An asterisk means a row added after the original blind round closed. Both Astra runs came in on September 5, 2026, the nineteen API rows on September 6, 2026. The original order of the other rows does not change because of it. The original 34 runs were scored by five independent judges, Astra and the API round by one judge anchored on nine dragons that had already been scored. A difference of two or three points therefore means nothing even inside a single round, let alone between rounds, see the findings appendix.

All 55 dragons

The images in each round are sorted by my eye, not by score. The points for each model are in the table above.

Free-form brief, run A

A brief written in plain language, with no criteria. Only the Claude family, three Codex models and Astra went through it, because it turned out to be less suitable.

Dragon by Astra, run A
Astra
Dragon by Fable 5.1, run A
Fable 5.1
Dragon by Fable 5, run A
Fable 5
Dragon by Opus 5, run A
Opus 5
Dragon by Sol, run A
Sol
Dragon by Luna, run A
Luna
Dragon by Opus 4.6, run A
Opus 4.6
Dragon by Terra, run A
Terra
Dragon by Opus 4.7, run A
Opus 4.7
Dragon by Opus 4.8, run A
Opus 4.8
Dragon by Sonnet 5, run A
Sonnet 5
Dragon by Sonnet 4.6, run A
Sonnet 4.6
Dragon by Haiku 4.5, run A
Haiku 4.5
Dragon by Opus 3, run A
Opus 3

How a human sorts it

The images in this grid are sorted by my eye, not by score. I sorted them before I looked at the points, and then the two could be compared.

my order model the judge's order score
1 Astra 1 38
2 Fable 5.1 6 30
3 Fable 5 4 32
4 Opus 5 2 35
5 Sol 3 34
6 Luna 5 32
7 Opus 4.6 10 23
8 Terra 7 29
9 Opus 4.7 8 28
10 Opus 4.8 9 24
11 Sonnet 5 11 23
12 Sonnet 4.6 12 22
13 Haiku 4.5 13 14
14 Opus 3 14 6

The agreement comes out at 0.92. For comparison, the five model judges agree with each other in a range of 0.72 to 0.93, see the methodology. So the machine hit the human eye about as well as the machines hit each other.

We part ways in only two places. I have Fable 5.1 four places higher and Opus 4.6 three. The remaining twelve dragons sit within two places, five of them exactly.

Technical brief, run B

A round with five criteria and the option to iterate until another pass brings no improvement. Every model in the first field went through it.

Dragon by Fable 5, run B
Fable 5
Dragon by Astra, run B
Astra
Dragon by Fable 5.1, run B
Fable 5.1
Dragon by Opus 5, run B
Opus 5
Dragon by Sol, run B
Sol
Dragon by GLM 5.3, run B
GLM 5.3
Dragon by Luna, run B
Luna
Dragon by Opus 4.8, run B
Opus 4.8
Dragon by Opus 4.6, run B
Opus 4.6
Dragon by Qwen 3.8, run B
Qwen 3.8
Dragon by Terra, run B
Terra
Dragon by Grok 4.6, run B
Grok 4.6
Dragon by Opus 4.7, run B
Opus 4.7
Dragon by Gemini, run B
Gemini
Dragon by Sonnet 5, run B
Sonnet 5
Dragon by Qwen 3.6, run B · file 2
Qwen 3.6, file 2
Dragon by Sonnet 4.6, run B
Sonnet 4.6
Dragon by Haiku 4.5, run B
Haiku 4.5
Dragon by Qwen 3.6, run B · file 1
Qwen 3.6, file 1
Dragon by Gemma, run B
Gemma
Dragon by Opus 3, run B · file 1
Opus 3, the rejected attempt
Dragon by Opus 3, run B · file 2
Opus 3, the final version

How a human sorts it

This grid is sorted by my eye too, not by score. I made the order from the images, without looking at the points.

I put Fable 5 on top, because it is the only dragon in this round that I would say has its anatomy right. Everything else is scattered to some degree. And somewhere around eighteenth place, everything that still resembles a dragon comes to an end.

my order model the judge's order score
1 Fable 5 4 35
2 Astra 2 39
3 Fable 5.1 3 39
4 Opus 5 1 40
5 Sol 5 35
6 GLM 5.3 6 33
7 Luna 7 31
8 Opus 4.8 10 29
9 Opus 4.6 11 28
10 Qwen 3.8 15 27
11 Terra 8 31
12 Grok 4.6 9 30
13 Opus 4.7 14 27
14 Gemini 13 27
15 Sonnet 5 12 28
16 Qwen 3.6, file 2 16 22
17 Sonnet 4.6 17 22
18 Haiku 4.5 18 16
19 Qwen 3.6, file 1 19 11
20 Gemma 20 9
21 Opus 3, the rejected attempt 21 3
22 Opus 3, the final version 22 3

The agreement comes out at 0.95, a little higher than for the free-form brief, where it was 0.92. Both numbers lie above how well the model judges agree with each other on average.

What is interesting is where we differ. Most of all at Qwen 3.8, which I have five places higher, and then at Fable 5 and Opus 5, which I have swapped. I sorted almost purely by anatomy, meaning by a single axis out of five, and even so it came out as nearly the same order as the sum of all five. That fits what is written under the leaderboard: those criteria measure one thing.

Third round, through the interface

Nineteen models from fifteen shops run straight through the API, with no app around them, with a ceiling of one fix. The conditions differed, the methodology spells them out.

Here too I adjusted the order of the images by eye, but I take it as far less reliable than in the previous two rounds. In this field there is essentially no dragon whose anatomy you could talk about. It is mostly abstraction, and sorting abstractions by how they look is a matter of taste, not observation.

Dragon by Hunyuan 4, through the API
Hunyuan 4
Dragon by Kimi K3, through the API
Kimi K3
Dragon by Muse Spark 1.3, through the API
Muse Spark 1.3
Dragon by MiniMax M3, through the API
MiniMax M3
Dragon by Seed 2.1 Turbo, through the API
Seed 2.1 Turbo
Dragon by DeepSeek V4 Pro, through the API
DeepSeek V4 Pro
Dragon by Muse Glimmer 30B, through the API
Muse Glimmer 30B
Dragon by Inkling, through the API
Inkling
Dragon by Nemotron 3 Ultra, through the API
Nemotron 3 Ultra
Dragon by Mistral Medium 3.5, through the API
Mistral Medium 3.5
Dragon by Gemma 4 26B A4B, through the API
Gemma 4 26B A4B
Dragon by Nova Premier, through the API
Nova Premier
Dragon by Ernie 4.5 VL, through the API
Ernie 4.5 VL
Dragon by Granite 4.2 8B, through the API
Granite 4.2 8B
Dragon by Gemma 4 E2B, through the API
Gemma 4 E2B
Dragon by Command A, through the API
Command A
Dragon by Gemma 3 4B, through the API
Gemma 3 4B
Dragon by LFM 2.5 2.6B, through the API
LFM 2.5 2.6B
Dragon by Llama 4 Maverick, through the API
Llama 4 Maverick

And fifteen more dragons outside the leaderboard

These are not in the main leaderboard, because they did not come about under the same conditions as the field above. The eight subagents and the run from someone else's account do have scores, they just are not added up with the rest. They live in the findings appendix, where each one comes with what it was measuring.

  • Eight runs of the same model. Eight independent agents, the same brief, eight draws. In one view you can see how much this benchmark repeats itself, and the noise floor came out of the spread of their scores.
  • Five experiments. What internet access does, a call for creativity with and without prohibitions, working blind with no render, and forced continuation past a declared plateau.
  • One run from someone else's account. The same brief with a different person on a different machine, to find out whether the model's weaknesses repeat outside this computer.