~ same prompt, different LLM — every model gets the same brief in Blender, you judge the render ~
OpenAI / GPT 6 Astra / thinking:max has entered the arena NEW!
LeaderboardGalleryVoteRunsTasks
RankContestantRatingVotesRunsDone %Success %Avg $Avg callsAvg time
1Google / Gemini 3.5 Flash / thinking:mediumcompare →1671810100$0.7433282s
2DeepSeek / Deepseek V4 Procompare →166181100100$0.0633491s
3OpenAI / GPT 5.6 Lunacompare →163871100100$0.04597s
4Z.AI / GLM 5.2compare →157371100100$0.1735511s
5DeepSeek / Deepseek V4 Flashcompare →148261100100$0.0749585s
6Anthropic / Haiku 4.5 / thinking:xhighcompare →1358710100$0.2860391s
7Google / Gemma 4 31b Itcompare →131571100100$0.038277s
8Google / Gemini 3.1 Flash Litecompare →130261100100$0.0215114s
provOpenAI / GPT 5.6 Luna / thinking:maxcompare →01100100$0.0327367s

Model comparison

each point is a model (Full, camp-lantern) — hover for exact values
AnthropicDeepSeekGoogleOpenAIZ.AI unrated
12471367148716061726$0.016$0.045$0.127$0.353$0.986Bradley–Terry ratingUNRATED — no decisive votes yetGoogle / Gemini 3.5 Flash / thinking:mediumDeepSeek / Deepseek V4 ProOpenAI / GPT 5.6 LunaZ.AI / GLM 5.2DeepSeek / Deepseek V4 FlashAnthropic / Haiku 4.5 / thinking:xhighGoogle / Gemma 4 31b ItGoogle / Gemini 3.1 Flash LiteMean cost per task (USD, log scale)

By dimension

lens rounds and chip verdicts feed these — none cast yet
Lighting
no data yet
Composition
no data yet
Modeling
no data yet
Materials
no data yet
Render
no data yet
Adherence
no data yet

Modes compared

each contestant's ratings across run modes (mixed ladder)
ContestantFullOne-shot3D-only
DeepSeek / Deepseek V4 Flash14561601
DeepSeek / Deepseek V4 Pro1615
Google / Gemini 3.1 Flash Lite1309
Google / Gemini 3.5 Flash / thinking:medium16121492
Google / Gemini 3.6 Flash / thinking:medium
Google / Gemma 4 31b It12761368
Z.AI / GLM 5.21530
OpenAI / GPT 5.6 Luna15961731
OpenAI / GPT 5.6 Luna / thinking:max
OpenAI / GPT 5.6 Sol / thinking:max1710
OpenAI / GPT 6 Astra / thinking:medium
Anthropic / Haiku 4.5 / thinking:xhigh1308
Anthropic / Sonnet 5 / thinking:low1397
Vote A/BGalleryRunsbradley-terry fit over 28 decisive votes
Votes cast on this bench: 000482 · 157 runs · 25 contestants
Best viewed with Netscape Navigator 4.07 at 800×600 · real runs, live ratings