~ same prompt, different LLM — every model gets the same brief in Blender, you judge the render ~
OpenAI / GPT 6 Astra / thinking:max has entered the arena NEW!
LeaderboardGalleryVoteRunsTasks
RankContestantRatingVotesRunsDone %Success %Avg $Avg callsAvg time
1Anthropic / Opus 5 / thinking:mediumcompare →173291100100$1.07543s
2OpenAI / GPT 5.6 Sol / thinking:maxcompare →1636810100$1.431041s
3Anthropic / Sonnet 5 / thinking:lowcompare →160651100100$0.08161s
4OpenAI / GPT 6 Astra / thinking:mediumcompare →1603510100$0.67598s
5Google / Gemini 3.6 Flash / thinking:mediumcompare →1432710100$0.20152s
6DeepSeek / Deepseek V4 Flash 0731 / thinking:maxcompare →1338710100$0.02567s
7DeepSeek / Deepseek V4 Flashcompare →1311610100$0.00125s
8Anthropic / Claude Fable 5.1 / thinking:mediumcompare →118651100100$2.951368s
provGoogle / Gemma 4 31b Itcompare →167311100100$0.00120s
provOpenAI / GPT 5.6 Lunacompare →1483110100$0.0369s
provGoogle / Gemini 3.5 Flash / thinking:mediumcompare →010100$0.24151s

Model comparison

each point is a model (One-shot, abandoned-village) — hover for exact values
AnthropicDeepSeekGoogleOpenAI unrated
11041282145916361814$0.000$0.005$0.049$0.527$5.69Bradley–Terry ratingUNRATED — no decisive votes yetMean cost per task (USD, log scale)

By dimension

lens rounds and chip verdicts feed these — none cast yet
Lighting
no data yet
Composition
no data yet
Modeling
no data yet
Materials
no data yet
Render
no data yet
Adherence
no data yet

Modes compared

each contestant's ratings across run modes (mixed ladder)
ContestantFullOne-shot3D-only
Anthropic / Claude Fable 5.1 / thinking:medium1196
DeepSeek / Deepseek V4 Flash14731315
DeepSeek / Deepseek V4 Flash 0731 / thinking:max1346
DeepSeek / Deepseek V4 Pro1449
Google / Gemini 3.1 Flash Lite1495
Google / Gemini 3.5 Flash / thinking:medium17501155
Google / Gemini 3.6 Flash / thinking:medium1438
Google / Gemma 4 31b It13461537
Z.AI / GLM 5.21688
OpenAI / GPT 5.6 Luna15991325
OpenAI / GPT 5.6 Luna / thinking:max1781
OpenAI / GPT 5.6 Sol / thinking:max1647
OpenAI / GPT 6 Astra / thinking:medium1613
Anthropic / Haiku 4.5 / thinking:xhigh1489
Anthropic / Opus 5 / thinking:medium1740
Anthropic / Sonnet 5 / thinking:low1616
Vote A/BGalleryRunsbradley-terry fit over 27 decisive votes
Votes cast on this bench: 000504 · 157 runs · 25 contestants
Best viewed with Netscape Navigator 4.07 at 800×600 · real runs, live ratings