~ same prompt, different LLM — every model gets the same brief in Blender, you judge the render ~
OpenAI / GPT 6 Astra / thinking:max has entered the arena NEW!
LeaderboardGalleryVoteRunsTasks
RankContestantRatingVotesRunsDone %Success %Avg $Avg callsAvg time
1OpenAI / GPT 5.6 Solcompare →174091100100$0.4510196s
2OpenAI / GPT 5.6 Lunacompare →1676172100100$0.045106s
3OpenAI / GPT 5.6 Luna / thinking:maxcompare →165991100100$0.2816350s
4DeepSeek / Deepseek V4 Flashcompare →150451100100$0.20741175s
5Google / Gemini 3.5 Flash / thinking:mediumcompare →149951100100$2.0252709s
6Z.AI / GLM 5.2compare →1422810100$1.00911836s
7DeepSeek / Deepseek V4 Procompare →140951100100$0.20581488s
8Anthropic / Haiku 4.5 / thinking:xhighcompare →139071100100$0.77971114s
9Google / Gemma 4 31b Itcompare →136861100100$0.036377s
10Google / Gemini 3.1 Flash Litecompare →133251100100$0.011086s

Model comparison

each point is a model (Full, corgi-beach-ball) — hover for exact values
AnthropicDeepSeekGoogleOpenAIZ.AI unrated
12711403153616691801$0.008$0.035$0.154$0.685$3.05Bradley–Terry ratingMean cost per task (USD, log scale)

By dimension

lens rounds and chip verdicts feed these — none cast yet
Lighting
no data yet
Composition
no data yet
Modeling
no data yet
Materials
no data yet
Render
no data yet
Adherence
no data yet

Modes compared

each contestant's ratings across run modes (mixed ladder)
ContestantFullOne-shot3D-only
DeepSeek / Deepseek V4 Flash1535
DeepSeek / Deepseek V4 Pro1483
Google / Gemini 3.1 Flash Lite1356
Google / Gemini 3.5 Flash / thinking:medium15241240
Google / Gemma 4 31b It1378
Z.AI / GLM 5.21448
OpenAI / GPT 5.6 Luna17121430
OpenAI / GPT 5.6 Luna / thinking:max1687
OpenAI / GPT 5.6 Sol1769
OpenAI / GPT 6 Astra / thinking:medium1519
Anthropic / Haiku 4.5 / thinking:xhigh1420
Vote A/BGalleryRunsbradley-terry fit over 38 decisive votes
Votes cast on this bench: 000482 · 157 runs · 25 contestants
Best viewed with Netscape Navigator 4.07 at 800×600 · real runs, live ratings