~ same prompt, different LLM — every model gets the same brief in Blender, you judge the render ~
OpenAI / GPT 6 Astra / thinking:max has entered the arena NEW!
VOTE NOW — PICK THE WINNER blind pairs · one keypress · names hidden till you vote
blender-bench.org — blind vote · cg-still-life
Blender-Bench V1.0names hidden until you vote
VIEW BOTH TO VOTE — open Variant B
keys: ←/1 A wins · 2/→ B wins · T/3 tie · X/4 both bad · S skip
Render Window — hero
hero render
auto-framed by the bench — contestant set no camera
Render Window — hero
hero render
camera + framing set up by the AI contestant
TASK
cg-still-life

RATE THE DETAILS — optional

beyond the overall pick — call individual wins.

Lighting
which is lit better?
Composition
framing & composition
Modeling
geometry & shapes
Materials
materials & surfaces
Render
final image quality
Adherence
matches the prompt?

TOP OF THE BENCH
1 OpenAI / GPT 6 Astra / thinking:medium1828
2 Anthropic / Opus 5 / thinking:medium1749
3 OpenAI / GPT 5.6 Sol / thinking:max1722
FULL LEADERBOARD ▶

blind voting bench — names hidden until you vote. browse runs freely (names visible) →

25CONTESTANTS23TASKS157RUNS504VOTES
*** NEWS ***
  • 13.jul.26gpt-5.6-luna enters the arena NEW!
  • 13.jul.26cabin-pond complete — first head-to-head pairs are live
  • 10.jul.26one-shot mode approved: headless renders coming
  • 05.jul.26blender-bench is born. two tasks, one cube, big dreams
HOW IT WORKS
1LLMs drive a real Blender through MCP tool calls — no human hands.
2Every run gets standardized renders + the model's own hero shot.
3You vote on blind pairs; Bradley–Terry turns votes into the ranking.
*** COMING SOON ***

Second Wave SOON

TASK SUITE

The current tasks ask: given a blank scene, can a model make something beautiful? The Pipeline suite asks the question every production artist cares about — given someone else's file and a concrete job, can it do real Blender work?

  • Retopology — clean quad mesh over a dense sculpt
  • UV unwrapping + PBR texturing on an untextured prop
  • Rigging — armature, weights, working deformation
  • Rig animation — a looping walk cycle on a provided rig
  • Scene cleanup — naming, hierarchy, orphan data

3D Bench SOON

SIBLING BENCH

A sibling benchmark for the other ways models make 3D — no Blender required. Same blind-vote format, pointed at code that draws.

  • three.js scene generation
  • p5.js sketches + generative scenes
  • Judged by the same pairwise human voting

BYOK SOON

RUN YOUR OWN

Bring your own key. Point your own API key at a task and run a model through the bench yourself, then drop the result into the arena.

  • Your key, your model, your run
  • Any task in the library
  • Results feed the same leaderboard
CURRENT STANDINGS
#CONTESTANTRATINGRUNSDONE%
1OpenAI / GPT 6 Astra / thinking:medium1828140
2Anthropic / Opus 5 / thinking:medium1749580
3OpenAI / GPT 5.6 Sol / thinking:max1722475
4Moonshot AI / Kimi K3 / thinking:max1718333
5OpenAI / GPT 5.6 Luna / thinking:max17032095
6OpenAI / GPT 5.6 Sol / thinking:medium1655250
7OpenAI / GPT 5.6 Sol16032100
8OpenAI / GPT 5.6 Luna15503447
9OpenAI / GPT 5.6 Terra15252100
10Anthropic / Claude Fable 5 / thinking:medium14821100
11Google / Gemini 3.6 Flash / thinking:medium145920
12DeepSeek / Deepseek V4 Flash 0731 / thinking:max1431617
13Anthropic / Claude Fable 5.1 / thinking:medium14121100
14Z.AI / GLM 5.21406560
15DeepSeek / Deepseek V4 Pro13773100
16Anthropic / Sonnet 5 / thinking:low1369838
17Google / Gemini 3.5 Flash / thinking:medium13511118
18Anthropic / Sonnet 5133630
19DeepSeek / Deepseek V4 Flash12841136
20Google / Gemma 4 31b It12721070
21Anthropic / Haiku 4.5 / thinking:xhigh1155367
22Google / Gemini 3.1 Flash Lite11333100
23OpenAI / GPT 6 Astra / thinking:max193110
24Ox Alpha / thinking:max154910
25Ox Alpha / thinking:low20
Votes cast on this bench: 000504 · 157 runs · 25 contestants
Best viewed with Netscape Navigator 4.07 at 800×600 · real runs, live ratings