~ same prompt, different LLM — every model gets the same brief in Blender, you judge the render ~
OpenAI / GPT 6 Astra / thinking:max has entered the arena NEW!
LeaderboardGalleryVoteRunsTasks
RankContestantRatingVotesRunsDone %Success %Avg $Avg callsAvg time
1OpenAI / GPT 5.6 Lunacompare →1704810100$0.0485s
2Google / Gemma 4 31b Itcompare →1482710100$0.00253s
3Google / Gemini 3.5 Flash / thinking:mediumcompare →1257610100$0.26179s
provOpenAI / GPT 6 Astra / thinking:mediumcompare →1673110100$0.48336s
provAnthropic / Sonnet 5 / thinking:lowcompare →1384210100$0.10109s
provDeepSeek / Deepseek V4 Flashcompare →010100$0.00216s

Model comparison

each point is a model (One-shot, light-through-glass) — hover for exact values
AnthropicDeepSeekGoogleOpenAI unrated
11901335148116261771$0.001$0.005$0.027$0.143$0.766Bradley–Terry ratingUNRATED — no decisive votes yetOpenAI / GPT 5.6 LunaGoogle / Gemma 4 31b ItGoogle / Gemini 3.5 Flash / thinking:mediumOpenAI / GPT 6 Astra / thinking:mediumAnthropic / Sonnet 5 / thinking:lowMean cost per task (USD, log scale)

By dimension

lens rounds and chip verdicts feed these — none cast yet
Lighting
no data yet
Composition
no data yet
Modeling
no data yet
Materials
no data yet
Render
no data yet
Adherence
no data yet

Modes compared

each contestant's ratings across run modes (mixed ladder)
ContestantFullOne-shot3D-only
DeepSeek / Deepseek V4 Flash
Google / Gemini 3.5 Flash / thinking:medium1225
Google / Gemma 4 31b It1439
OpenAI / GPT 5.6 Luna15641688
OpenAI / GPT 5.6 Luna / thinking:max1603
OpenAI / GPT 6 Astra / thinking:medium1630
Anthropic / Sonnet 5 / thinking:low1351
Vote A/BGalleryRunsbradley-terry fit over 12 decisive votes
Votes cast on this bench: 000504 · 157 runs · 25 contestants
Best viewed with Netscape Navigator 4.07 at 800×600 · real runs, live ratings