~ same prompt, different LLM — every model gets the same brief in Blender, you judge the render ~
OpenAI / GPT 6 Astra / thinking:max has entered the arena NEW!
LeaderboardGalleryVoteRunsTasks
RankContestantRatingVotesRunsDone %Success %Avg $Avg callsAvg time
no runs on disk — bench idle

Model comparison

each point is a model (3D-only, camp-lantern) — hover for exact values
unrated
14001450150015501600$0.001$0.006$0.032$0.178$1.00Bradley–Terry ratingMean cost per task (USD, log scale)
No rated models yet — cast decisive votes to plot the field.

By dimension

lens rounds and chip verdicts feed these — none cast yet
Lighting
no data yet
Composition
no data yet
Modeling
no data yet
Materials
no data yet
Render
no data yet
Adherence
no data yet

Modes compared

each contestant's ratings across run modes (mixed ladder)
ContestantFullOne-shot3D-only
DeepSeek / Deepseek V4 Flash14371584
DeepSeek / Deepseek V4 Pro1599
Google / Gemini 3.1 Flash Lite1291
Google / Gemini 3.5 Flash / thinking:medium15931477
Google / Gemini 3.6 Flash / thinking:medium
Google / Gemma 4 31b It12571373
Z.AI / GLM 5.21510
OpenAI / GPT 5.6 Luna15771712
OpenAI / GPT 5.6 Luna / thinking:max
OpenAI / GPT 5.6 Sol / thinking:max1675
OpenAI / GPT 6 Astra / thinking:medium1748
Anthropic / Haiku 4.5 / thinking:xhigh1291
Anthropic / Sonnet 5 / thinking:low1376
Vote A/BGalleryRunsratings appear after first votes — BT fit pending
Votes cast on this bench: 000504 · 157 runs · 25 contestants
Best viewed with Netscape Navigator 4.07 at 800×600 · real runs, live ratings