~ same prompt, different LLM — every model gets the same brief in Blender, you judge the render ~
OpenAI / GPT 6 Astra / thinking:max has entered the arena NEW!
LeaderboardGalleryVoteRunsTasks
RankContestantRatingVotesRunsDone %Success %Avg $Avg callsAvg time
1OpenAI / GPT 5.6 Luna / thinking:maxcompare →167861100100$0.0420610s
2Anthropic / Sonnet 5compare →1268710100$0.98721903s
provOpenAI / GPT 5.6 Lunacompare →155431100100$0.045101s

Model comparison

each point is a model (Full, plaster-bust) — hover for exact values
AnthropicOpenAI unrated
12071340147316061740$0.028$0.073$0.189$0.491$1.28Bradley–Terry ratingOpenAI / GPT 5.6 Luna / thinking:maxAnthropic / Sonnet 5OpenAI / GPT 5.6 LunaMean cost per task (USD, log scale)

By dimension

lens rounds and chip verdicts feed these — none cast yet
Lighting
no data yet
Composition
no data yet
Modeling
no data yet
Materials
no data yet
Render
no data yet
Adherence
no data yet

Modes compared

each contestant's ratings across run modes (mixed ladder)
ContestantFullOne-shot3D-only
DeepSeek / Deepseek V4 Flash1337
Google / Gemini 3.5 Flash / thinking:medium1146
Google / Gemma 4 31b It1528
OpenAI / GPT 5.6 Luna174113371719
OpenAI / GPT 5.6 Luna / thinking:max1876
OpenAI / GPT 6 Astra / thinking:medium
Anthropic / Sonnet 51479
Anthropic / Sonnet 5 / thinking:low1337
Vote A/BGalleryRunsbradley-terry fit over 8 decisive votes
Votes cast on this bench: 000482 · 157 runs · 25 contestants
Best viewed with Netscape Navigator 4.07 at 800×600 · real runs, live ratings