Gallery: all 55 dragons and the full leaderboard
A companion appendix to the article on the dragon benchmark.
Verbatim model output and measurements. This is the evidence the articles rest on, not extra reading.
A companion appendix to the article on the dragon benchmark.
The second appendix to the dragon benchmark article; the images and the ranking are in the gallery, the findings, including how much of the ranking is…
A companion appendix to the article about a model that started calling the conversation-ending tool on its own.
While it writes, the model reads its own output, so it can tell when its probability distribution has gone blurry.
Ten turns, eight scored traps. The script is fixed and is inserted one turn at a time regardless of what the model replies, because a mismatch is data too.
The complete brief exactly as it goes to the model, including the fixed script of player turns. Anyone can run it themselves.
Verbatim replies from thirteen model configurations to two turns of the same scene. Damaged text is left as it came, because it is part of the finding.
Appendix to the article about the dragon benchmark. The dragons and the ranking are in the gallery, the prompts and the judges in the methodology.