Gallery: all 55 dragons and the full leaderboard
A companion appendix to the article on the dragon benchmark. Thirty-nine models got the same brief: draw a dragon as a single SVG file, with no references and no looking into anyone else's folder. The result was 55 images.
The images are model output, not my work. I show them as evidence, not as work of my own, the same as the verbatim texts in the other appendices.
All three briefs verbatim, the scoring by five independent judges and the conditions of each round are in the methodology. What came out of the data, meaning tool confabulation, how models write about their own work, the noise floor and the prices, is in the findings appendix.
Sorted by score. Five criteria worth ten points each, fifty at most.
B·1 and B·2 mean two files from the same run B. It happened with three models that turned in more than one image.
Anatomy is one of the five scored criteria and asks whether it reads as a dragon at first glance and whether the limbs attach to the body. Total is the sum of all five criteria, fifty at most. The full breakdown by criterion is available for download below. It is not in the table because those five axes measure almost the same thing: they correlate 0.89 to 0.96 with each other, and each of them above 0.96 with the total.
Shapes is the number of drawing elements in the submitted file, meaning paths and basic shapes together. It is not a judgment, it is measured from the file. It is here because it is the only figure in this table that does not say the same thing as the score: with anatomy it correlates only 0.39. In other words, how much a model tried to draw and whether it holds together are two different things. The dash belongs to Qwen 3.6, which turned in more files than it has rows, and there is no way to tell which one goes with which.
The last column says whether the model saw its dragon. Itself means it rendered the image on its own initiative, which was part of the measurement in the first two rounds. Was sent one means the image was sent to it without it asking, and that applies only to the third round. No means it did not see the dragon at all: either it did not look, although it could have, or that model does not take images as input at all. Which case is which is spelled out in the methodology.
The table has three rounds. Runs A and B are the original blind round from each vendor's own app, they carry numbers 1 to 34 and those no longer change. Rows marked with an asterisk came later: Astra on September 5, 2026 and nineteen models marked API on September 6, 2026. Those last ones ran straight through the interface, with no app around them, had a ceiling of one fix instead of a free nine, and got the render automatically without asking for it. Their row says what the model drew, not how it would have done in the same race. The conditions are spelled out in the methodology, the prices and the reactions to the render are in the findings appendix.
The leaderboard
| # | model | run | anatomy | shapes | total | saw a render? |
|---|---|---|---|---|---|---|
| 1 | Opus 5 | B | 8 | 229 | 40 | itself |
| * | Astra | B | 6 | 336 | 39 | itself |
| 2 | Fable 5.1 | B | 7 | 251 | 39 | itself |
| * | Astra | A | 7 | 3,342 | 38 | itself |
| 3 | Fable 5 | B | 7 | 320 | 35 | itself |
| 4 | Opus 5 | A | 7 | 466 | 35 | itself |
| 5 | Sol | B | 7 | 122 | 35 | itself |
| 6 | Sol | A | 6 | 111 | 34 | itself |
| * | Hunyuan 4 | API | 7 | 100 | 34 | no |
| 7 | GLM 5.3 | B | 6 | 199 | 33 | no |
| 8 | Fable 5 | A | 6 | 175 | 32 | itself |
| 9 | Luna | A | 6 | 91 | 32 | no |
| 10 | Luna | B | 5 | 114 | 31 | itself |
| 11 | Terra | B | 6 | 73 | 31 | itself |
| * | Muse Spark 1.3 | API | 6 | 84 | 31 | was sent one |
| 12 | Fable 5.1 | A | 6 | 1,267 | 30 | itself |
| 13 | Grok 4.6 | B | 5 | 168 | 30 | itself |
| * | Kimi K3 | API | 5 | 104 | 30 | was sent one |
| * | Seed 2.1 Turbo | API | 5 | 113 | 30 | was sent one |
| 14 | Opus 4.8 | B | 5 | 1,014 | 29 | itself |
| 15 | Terra | A | 4 | 96 | 29 | no |
| 16 | Opus 4.6 | B | 5 | 242 | 28 | itself |
| 17 | Opus 4.7 | A | 4 | 261 | 28 | itself |
| 18 | Sonnet 5 | B | 5 | 133 | 28 | itself |
| 19 | Gemini | B | 5 | 62 | 27 | no |
| 20 | Opus 4.7 | B | 5 | 205 | 27 | itself |
| 21 | Qwen 3.8 | B | 4 | 141 | 27 | no |
| * | MiniMax M3 | API | 4 | 143 | 26 | was sent one |
| 22 | Opus 4.8 | A | 4 | 117 | 24 | itself |
| 23 | Opus 4.6 | A | 3 | 234 | 23 | itself |
| 24 | Sonnet 5 | A | 3 | 96 | 23 | itself |
| 25 | Qwen 3.6 | B·2 | 3 | – | 22 | no |
| 26 | Sonnet 4.6 | A | 3 | 134 | 22 | itself |
| 27 | Sonnet 4.6 | B | 3 | 200 | 22 | itself |
| * | DeepSeek V4 Pro | API | 3 | 100 | 20 | no |
| * | Muse Glimmer 30B | API | 3 | 28 | 20 | no |
| * | Inkling | API | 2 | 84 | 17 | was sent one |
| 28 | Haiku 4.5 | B | 3 | 151 | 16 | itself |
| * | Nemotron 3 Ultra | API | 2 | 147 | 15 | no |
| 29 | Haiku 4.5 | A | 2 | 41 | 14 | no |
| * | Mistral Medium 3.5 | API | 2 | 31 | 12 | was sent one |
| * | Gemma 4 26B A4B | API | 2 | 29 | 12 | was sent one |
| 30 | Qwen 3.6 | B·1 | 1 | – | 11 | no |
| 31 | Gemma | B | 1 | 16 | 9 | no |
| * | Nova Premier | API | 1 | 11 | 9 | was sent one |
| * | Ernie 4.5 VL | API | 1 | 14 | 8 | was sent one |
| 32 | Opus 3 | A | 0 | 19 | 6 | no |
| * | Granite 4.2 8B | API | 0 | 19 | 6 | no |
| * | Gemma 4 E2B | API | 0 | 8 | 6 | no |
| * | Command A | API | 0 | 13 | 5 | no |
| * | Gemma 3 4B | API | 0 | 13 | 5 | no |
| * | LFM 2.5 2.6B | API | 0 | 19 | 4 | no |
| 33 | Opus 3 | B·1 | 0 | 7 | 3 | no |
| 34 | Opus 3 | B·2 | 0 | 7 | 3 | no |
| * | Llama 4 Maverick | API | 0 | 11 | 3 | was sent one |
Full breakdown of all five criteria, the scores and the shape counts for 55 dragonsCSV, 2 kB
Every submitted SVG, 63 filesZIP, 544 kB
\* An asterisk means a row added after the original blind round closed. Both Astra runs came in on September 5, 2026, the nineteen API rows on September 6, 2026. The original order of the other rows does not change because of it. The original 34 runs were scored by five independent judges, Astra and the API round by one judge anchored on nine dragons that had already been scored. A difference of two or three points therefore means nothing even inside a single round, let alone between rounds, see the findings appendix.
All 55 dragons
The images in each round are sorted by my eye, not by score. The points for each model are in the table above.
Free-form brief, run A
A brief written in plain language, with no criteria. Only the Claude family, three Codex models and Astra went through it, because it turned out to be less suitable.














How a human sorts it
The images in this grid are sorted by my eye, not by score. I sorted them before I looked at the points, and then the two could be compared.
| my order | model | the judge's order | score |
|---|---|---|---|
| 1 | Astra | 1 | 38 |
| 2 | Fable 5.1 | 6 | 30 |
| 3 | Fable 5 | 4 | 32 |
| 4 | Opus 5 | 2 | 35 |
| 5 | Sol | 3 | 34 |
| 6 | Luna | 5 | 32 |
| 7 | Opus 4.6 | 10 | 23 |
| 8 | Terra | 7 | 29 |
| 9 | Opus 4.7 | 8 | 28 |
| 10 | Opus 4.8 | 9 | 24 |
| 11 | Sonnet 5 | 11 | 23 |
| 12 | Sonnet 4.6 | 12 | 22 |
| 13 | Haiku 4.5 | 13 | 14 |
| 14 | Opus 3 | 14 | 6 |
The agreement comes out at 0.92. For comparison, the five model judges agree with each other in a range of 0.72 to 0.93, see the methodology. So the machine hit the human eye about as well as the machines hit each other.
We part ways in only two places. I have Fable 5.1 four places higher and Opus 4.6 three. The remaining twelve dragons sit within two places, five of them exactly.
Technical brief, run B
A round with five criteria and the option to iterate until another pass brings no improvement. Every model in the first field went through it.






















How a human sorts it
This grid is sorted by my eye too, not by score. I made the order from the images, without looking at the points.
I put Fable 5 on top, because it is the only dragon in this round that I would say has its anatomy right. Everything else is scattered to some degree. And somewhere around eighteenth place, everything that still resembles a dragon comes to an end.
| my order | model | the judge's order | score |
|---|---|---|---|
| 1 | Fable 5 | 4 | 35 |
| 2 | Astra | 2 | 39 |
| 3 | Fable 5.1 | 3 | 39 |
| 4 | Opus 5 | 1 | 40 |
| 5 | Sol | 5 | 35 |
| 6 | GLM 5.3 | 6 | 33 |
| 7 | Luna | 7 | 31 |
| 8 | Opus 4.8 | 10 | 29 |
| 9 | Opus 4.6 | 11 | 28 |
| 10 | Qwen 3.8 | 15 | 27 |
| 11 | Terra | 8 | 31 |
| 12 | Grok 4.6 | 9 | 30 |
| 13 | Opus 4.7 | 14 | 27 |
| 14 | Gemini | 13 | 27 |
| 15 | Sonnet 5 | 12 | 28 |
| 16 | Qwen 3.6, file 2 | 16 | 22 |
| 17 | Sonnet 4.6 | 17 | 22 |
| 18 | Haiku 4.5 | 18 | 16 |
| 19 | Qwen 3.6, file 1 | 19 | 11 |
| 20 | Gemma | 20 | 9 |
| 21 | Opus 3, the rejected attempt | 21 | 3 |
| 22 | Opus 3, the final version | 22 | 3 |
The agreement comes out at 0.95, a little higher than for the free-form brief, where it was 0.92. Both numbers lie above how well the model judges agree with each other on average.
What is interesting is where we differ. Most of all at Qwen 3.8, which I have five places higher, and then at Fable 5 and Opus 5, which I have swapped. I sorted almost purely by anatomy, meaning by a single axis out of five, and even so it came out as nearly the same order as the sum of all five. That fits what is written under the leaderboard: those criteria measure one thing.
Third round, through the interface
Nineteen models from fifteen shops run straight through the API, with no app around them, with a ceiling of one fix. The conditions differed, the methodology spells them out.
Here too I adjusted the order of the images by eye, but I take it as far less reliable than in the previous two rounds. In this field there is essentially no dragon whose anatomy you could talk about. It is mostly abstraction, and sorting abstractions by how they look is a matter of taste, not observation.



















And fifteen more dragons outside the leaderboard
These are not in the main leaderboard, because they did not come about under the same conditions as the field above. The eight subagents and the run from someone else's account do have scores, they just are not added up with the rest. They live in the findings appendix, where each one comes with what it was measuring.
- Eight runs of the same model. Eight independent agents, the same brief, eight draws. In one view you can see how much this benchmark repeats itself, and the noise floor came out of the spread of their scores.
- Five experiments. What internet access does, a call for creativity with and without prohibitions, working blind with no render, and forced continuation past a declared plateau.
- One run from someone else's account. The same brief with a different person on a different machine, to find out whether the model's weaknesses repeat outside this computer.