Skip to content
EshAlora

Methodology: three briefs, five judges and the conditions of every round

The second appendix to the dragon benchmark article; the images and the ranking are in the gallery, the findings, including how much of the ranking is chance, in the appendix on findings.

Where the numbers come from. The scores in the gallery come from a report written by Claude Fable 5.1, one of the models being judged. That is a conflict of interest, and it is measured below: four more judges got the same field, so that what is a property of the dragons can be told apart from what is a property of the judge.

Both briefs

There were two runs. The Claude family and three Codex models (Luna, Sol, Terra) went through both, the other families (Gemini, Grok, GLM, Qwen, Gemma) only run B; the free-form brief A turned out to be less suitable.

Run A, a free-form brief in plain speech

Verbatim wording, taken from the brief given to the judges:

i've got an image benchmark, girl draw me an svg dragon and really focus so it comes out great! i'll compare it with the others, i don't care how much time and effort you put in, i want you to hand me a result you're happy with 🙂 once you can stand behind it save the file into the folder, don't look anywhere!

After the image was handed in came a second request:

Please describe the process briefly, how you went about it, what tools you used and how many times you checked it, as an MD file, and save it into the folder next to the image.

Run B, the technical round

The verbatim wording as the models received it, here in the version for Opus 4.7. The folder name differed for each model.

BRIEF — SVG benchmark, technical round

Draw a dragon as a single SVG file. This is a test of maximum output quality — the result will be judged externally against the criteria below, by comparison across models.

Hard rules: Work EXCLUSIVELY in your own folder opus 4.7 It is forbidden to read, open or browse any other folders, files or outputs — including "for calibration" or "for comparison". Your output must be produced independently. Permitted tools: any tool for checking YOUR OWN work (render/preview of the SVG, a browser to verify the code is valid). At the end, justify in writing whether you used tools or not.

Hand in: (a) drak.svg, (b) postup.md — briefly: number of iterations, what you changed between iterations and why, how you verified the final render, the justification of your choice of tools.

Scoring criteria (the output will be scored against them, 0–10 each):

Anatomy and readability: the dragon is a dragon at first glance — a head with an eye and a maw, neck, body, limbs, wings with a membrane, tail; the limbs attach in an anatomically plausible way, nothing floats in space.

Composition: a dynamic pose (not a static side view), a thought-out use of the area, a functional background or a deliberate decision to go without one.

SVG craft: clean valid code; curves (path/bezier), not a jigsaw of ellipses and rectangles; the layering makes sense; correct viewBox; the file renders identically in common browsers.

Color and light: a deliberate palette, shading or gradients creating volume, contrast between the dragon and the background.

Detail: scales, textures, claws, teeth, wing membranes — fineness of execution where it raises quality, not decoration for decoration's sake.

Standard: the target level is a professional vector illustration, the kind a client would buy. "Functional" or "recognizable" is NOT the goal — the goal is the best dragon you are able to produce. Iterate until a further iteration stops bringing a measurable improvement against the criteria above; only then hand it in. There is no point bonus for speed, restraint or saving effort.

Three places show up in the data: the ban on looking elsewhere is in there twice, including for the excuse "for calibration"; the justification of tools is the source of the feedback column; the sentence about iteration with no bonus for saving effort turns the number of iterations into a measure of effort.

Third round: nineteen models through the API

The original field had nineteen models from six companies: Anthropic, OpenAI, Google, xAI, Zhipu and Alibaba. So on 6 September 2026 a third round was added, nineteen models from fifteen shops, seventeen through OpenRouter and two locally:

Amazon Nova Premier, Baidu Ernie 4.5 VL, ByteDance Seed 2.1 Turbo, Cohere Command A, DeepSeek V4 Pro, Google Gemma 3 4B, Gemma 4 26B A4B and Gemma 4 E2B (locally), IBM Granite 4.2 8B, Liquid AI LFM 2.5 2.6B, Meta Llama 4 Maverick, Muse Spark 1.3 and Muse Glimmer 30B (locally), MiniMax M3, Mistral Medium 3.5, Moonshot Kimi K3, NVIDIA Nemotron 3 Ultra, Tencent Hunyuan 4, Thinking Machines Inkling.

It cost $2.01.

In the gallery this round sits in the same table, but marked. The scores can be read on one scale, because spatial composition lives in the weights of the model, not in the system prompt of the app around it. The conditions, though, differed in six things, and every one of them moves the result:

  1. Raw models, with no vendor system prompt.
  2. They got the render automatically, whereas before a model had to ask for it itself; the feedback column means something else here.
  3. Nine models got no render, each time for a different reason, and those must not be blurred together: the model does not accept images (DeepSeek V4 Pro, Nemotron, Granite, Tencent Hunyuan 4, Liquid LFM), the whole shop has no image support (Cohere), the gateway refused (Gemma 3 4B, request limits at the provider), a local run (Gemma 4 E2B and Muse Glimmer, outside this methodology).
  4. The ceiling was a single fix, not a free zero to nine; the exception is Muse Spark with two.
  5. The canvas was prescribed at 1600 by 1000 points, but only for part of the field; whoever got the original wording chose for themselves: Granite 200 by 100, DeepSeek a percentage width, Gemma 4 E2B 800 by 600.
  6. The brief had to be adapted, the original assumes a disk, a terminal and a browser.

Points 2 and 4 move the score demonstrably: after the render came back, nine models rewrote fourteen to forty-eight percent of the area. A row from the third round says what a model can draw, not how it would have done in the same race.

Changes to the brief: the ban on looking elsewhere was thrown out, a note was added that the model has no disk and no tools and hands in both files as text, the canvas was prescribed, a sentence about invalid XML was added. The criteria and the paragraph about the target level stayed word for word the same. Seven models got the original wording (Muse Spark, Cohere, Granite, Nemotron, DeepSeek and both local runs), the remaining twelve got this:

Environment and tools (read carefully, this differs from the usual brief):

You are running through an API. You have no file system, no terminal, no browser and no tool for rendering SVG. You therefore cannot save anything to disk or look at the result. Do not attempt to call tools and do not describe steps you did not perform. So hand in both files directly in the text of your answer, each in its own block:

  • drak.svg in a ```svg block
  • postup.md in a ```markdown block After you hand in, you will be sent one image: a PNG render of your own drak.svg, rendered in the Chrome browser. On that basis you will be able to hand in a corrected version of both files, or to write that you are handing in unchanged. No other feedback will be available.

Canvas:

Use exactly viewBox="0 0 1600 1000" and width="1600" height="1000". The dimensions are the same for all tested models so that the results are comparable. Do not use percentage dimensions.

Your output must be produced independently, without leaning on anyone else's examples or comparisons.

Careful: the file must be well-formed XML. Every tag closed, attributes in the form name="value", a space between attributes, never a comma. An invalid file will not render in the browser at all.

The rest, that is the instruction to hand in drak.svg and postup.md, the five scoring criteria and the paragraph about the target level, is word for word the same as in run B above.

The split was not random, it went by batches. It can show what that sentence does, but it cannot isolate it: along with the wording, the canvas, the XML instructions and the way of handing in were changing too. The result is in the appendix on findings.

Who judged this round

One judge, the model Opus 5, unlike the five in the first round. So that the numbers would sit on the same scale, before scoring it got nine already scored dragons from the first field as anchors, together with the breakdown of points by individual criterion.

For these rows we therefore have no measure of judge agreement, which we do have for the first field. A difference of two points between two models means nothing here.

Five judges

The scores in the gallery are Fable 5.1's judgment. So that what is a property of the dragons could be told apart from what is a property of the judge, four more models got the same gallery. Fable 5.1 judged first and blind, with the whole folder; Sol and Terra had full access, rendered the files themselves and knew they were judging their own family as well; Grok 4.6 and Gemini (by its own report Gemini Pro Vision) saw only a sheet with 34 dragons under the codes D01 to D34 and a shortened brief, so they did not know they were judging their own dragon.

Rank agreement (Spearman):

pair agreement
Fable 5.1 × Sol 0.93
Terra × Sol 0.92
Fable 5.1 × Terra 0.90
Terra × Grok 0.88
Fable 5.1 × Grok 0.86
Grok × Sol 0.86
Gemini with any of them 0.72 to 0.79

Average score awarded: Fable 5.1 25.2, Sol 32.2, Terra 31.3, Gemini 24.5, Grok 24.3 points out of 50. Four of the five agree around 0.9, Gemini is the outlier against all of them, both Codex models score six to seven points higher, but they hold the ordering.

Family favoritism is across the board, not targeted: against Fable 5.1, Terra added 6.2 points to Codex, 6.1 to Claude, 5.9 to the rest; Sol 7.0 to Codex, 6.6 to Claude, 8.3 to the rest. Sol gave both of its own dragons 40, the highest self-rating among the judges, but it gave the same 40 to Opus 5 A and placed Opus 5 B, GLM and Terra above itself. Grok gave its own dragon 26, lower than Fable 5.1 did (30), so the anonymization worked; it gave Gemini's dragon 34, and Gemini gave Grok's 30.

What the two Codex judges wrote in their reviews

Alongside the numbers, Sol and Terra also handed in written reviews. Neither of them signs those by name, so the attribution to a judge rests on the scores matching. Both rendered all 34 SVGs themselves and state an uncertainty of ±1 for differences of about a point. Sol added a criterion that is not in the brief, a binary "a dragon at first glance" test: 7 of the 34 images failed it.

They agree on the winner, Opus 5 B, on the biggest improvement being Opus 4.8's by six points, on the technical brief adding almost nothing (Terra +0.6 points, Sol +0.5) and on 34.5 out of 50 for the 23 runs with a render. The group without a render, though, is different for each of them: Terra 11 runs with an average of 24.6, Sol 9 runs with 29.3, so a gap of 9.9 against 5.2 points. It is a difference in whose declaration the judge believed. They agree on the worst calibration: with Terra, Haiku B declared "professional 9/10" against 20 out of 50, with Sol "its own 45/50 versus my 23.5/50".

Their mutual agreement is 0.92, but on individual dragons the differences run up to nine points: Terra scores from 6 to 45, Sol from 9 to 42. Neither of them put its own dragon at the top, Terra gave 40 to Sol, Sol 40.5 to Terra.

What the judges found outside the images

Claims about checks that cannot be verified are logged by both as declared, and neither raises the score for them. While rendering all the files they ran into three things that have nothing to do with the drawing:

  • Breaking the ban on looking elsewhere. Fable 5.1 in run A read about 2,500 characters of someone else's file from Haiku after handing in, and admitted it herself. Fable 5 in run B saw the names of neighboring folders, did not read the contents. Sonnet 5 in run B, according to Sol, accidentally loaded someone else's postup.md in the root and admitted it.
  • Undocumented browser checks. Gemma, Qwen 3.6 and Haiku in run B claim identical rendering in two or three browsers, with nothing to back it up; Sol adds Opus 4.8 B and its four browsers.
  • Its own numbers not matching. With Opus 5 in run A the chat talks about eight iterations and postup.md about eleven.

Scores from all five judges

Sorted by consensus. Ø is the average of the five judges and it is the most reliable number in the whole table. Model families and the images are in the gallery.

Rows where the judges diverged by twelve points or more are almost always an atypical style (a silhouette, a medallion, a sticker) or a dragon whose anatomy falls apart up close.

model Fable 5.1 Sol Terra Grok Gemini Ø 5
Opus 5 B 40 42 45 36 38 40.2
Opus 5 A 35 39.5 40 36 42 38.5
Fable 5.1 B 39 38 40 34 38 37.8
Terra B 31 40.5 37 31 44 36.7
Sol B 35 40 39 30 38 36.4
GLM 5.3 B 33 40.5 37 32 38 36.1
Opus 4.8 B 29 39 39 28 40 35.0
Terra A 29 39 39 33 34 34.8
Sol A 34 40 40 30 29 34.6
Luna A 32 38 38 27 37 34.4
Fable 5 B 35 38.5 38 35 23 33.9
Grok 4.6 B 30 35.5 38 26 30 31.9
Qwen 3.8 B 27 35 36 32 29 31.8
Fable 5.1 A 30 36.5 36 24 28 30.9
Gemini B 27 35.5 36 34 18 30.1
Fable 5 A 32 37 35 29 17 30.0
Opus 4.7 B 27 33.5 35 29 25 29.9
Opus 4.8 A 24 33 33 25 34 29.8
Luna B 31 36.5 36 26 17 29.3
Opus 4.6 B 28 34.5 29 26 22 27.9
Opus 4.7 A 28 36 34 24 14 27.2
Sonnet 5 B 28 32.5 32 24 19 27.1
Opus 4.6 A 23 35 28 21 25 26.4
Sonnet 5 A 23 29 27 19 29 25.4
Sonnet 4.6 B 22 30 33 18 23 25.2
Qwen 3.6 B, file 2 22 29 25 23 26 25.0
Sonnet 4.6 A 22 30 34 18 12 23.2
Haiku 4.5 A 14 24.5 28 18 7 18.3
Qwen 3.6 B, file 1 11 28 20 18 13 18.0
Haiku 4.5 B 16 23.5 20 11 15 17.1
Opus 3 A 6 14.5 12 9 7 9.7
Gemma B 9 13.5 8 11 7 9.7
Opus 3 B, file 2 3 10 11 4 7 7.0
Opus 3 B, file 1 3 9 6 4 7 5.8

All five agree on this: at the top Opus 5 B, Opus 5 A, Fable 5.1 B, Terra B, Sol B and GLM, at the bottom Opus 3, Gemma, Haiku and the orange spheres of Qwen 3.6. The biggest split is Terra B: Gemini 44 points, Terra to itself 37, Sol 40.5, Fable 5.1 31, Grok 31.

Scores for the third round and Astra, from a single judge

These rows were judged by Opus 5 alone, anchored on nine dragons from the field above. So there is no agreement measure for them and a difference of two points means nothing. The images are in the gallery.

model idea comp. craft color detail total
Astra B 6 8 9 8 8 39
Astra A 7 7 8 8 8 38
Hunyuan 4 (Tencent) 7 7 8 6 6 34
Muse Spark 1.3 (Meta) 6 6 7 6 6 31
Kimi K3 (Moonshot) 5 6 7 6 6 30
Seed 2.1 Turbo (ByteDance) 5 5 7 6 7 30
MiniMax M3 (MiniMax) 4 5 6 6 5 26
DeepSeek V4 Pro (DeepSeek) 3 3 5 4 5 20
Muse Glimmer 30B (Meta) 3 3 6 5 3 20
Inkling (Thinking Machines) 2 4 4 4 3 17
Nemotron 3 Ultra (NVIDIA) 2 2 5 2 4 15
Mistral Medium 3.5 (Mistral) 2 2 3 3 2 12
Gemma 4 26B A4B (Google) 2 2 4 2 2 12
Nova Premier (Amazon) 1 1 3 3 1 9
Ernie 4.5 VL (Baidu) 1 1 3 2 1 8
Granite 4.2 8B (IBM) 0 1 2 2 1 6
Gemma 4 E2B (Google) 0 1 2 3 0 6
Command A (Cohere) 0 1 2 2 0 5
Gemma 3 4B (Google) 0 1 2 2 0 5
LFM 2.5 2.6B (Liquid AI) 0 1 2 1 0 4
Llama 4 Maverick (Meta) 0 1 1 1 0 3

Where next

The ranking and all 55 images are in the gallery. What came out of the data, that is confabulated tools, the way models describe their own result, the noise threshold and the prices, is in the appendix on findings.