~ same prompt, different LLM — every model gets the same brief in Blender, you judge the render ~
OpenAI / GPT 6 Astra / thinking:max has entered the arena NEW!
WHAT IS THIS?

Blender-Bench benchmarks LLMs — and the harnesses that drive them — at using Blender through the Blender MCP tool interface. Each contestant gets a task and a real Blender, and has to build the scene entirely through tool calls: no human hands on the mouse.

Runs are scored by pairwise human voting. You compare two renders of the same task and pick the better one; a Bradley–Terry model turns those votes into the ranking on the leaderboard.

WHY THIS EXISTS

The project was inspired by MineBench, which has always fascinated the author as someone who has worked in 3D for quite some time. We've reached the point where models are advanced enough to finally do meaningful work in 3D.

Understanding 3D is an important capability for AI, especially on the way to world models, AGI, and embodied AI. It comes down to recognizing relationships between objects in space, predicting how they would interact with one another, and developing a sense of visual beauty and composition.

The first wave of Blender-Bench is creative tasks — detailed prompts describing a scene. A second wave will target real-world Blender workflows and everyday 3D work: rigging, retopology, animation. Some parts, like retopology, may be a little overkill — but they're an interesting way to observe how models approach problem-solving. This is only the beginning.

HOW IT WORKS
1LLMs drive a real Blender through MCP tool calls — no human hands.
2Every run gets standardized renders + the model's own hero shot.
3You vote on blind pairs; Bradley–Terry turns votes into the ranking.

For the full picture — budgets, prompt and task versioning, how Done % differs from the completion rate, and the honest limits of blind voting — read the methodology.

TASKS

Every contestant works from an authored prompt. Browse them in the task library — each task shows its prompt in a Blender Text Editor and lists the runs recorded against it. Browse results in the gallery or read individual transcripts in the run log.

WHAT'S NEXT

More is coming: a Second Wave task suite for real production skills (retopology, UV texturing, rigging, rig animation, scene cleanup); 3D Bench, a sibling benchmark for three.js and p5.js scene generation; and BYOK — bring-your-own-key runs so you can put a model through the bench yourself. All still in the works; see the roadmap on the home page.

SITE STATISTICS

We count page views and approximate visits to understand site traffic. This uses a temporary visit ID in this browser tab, with no analytics cookies, IP addresses, query strings or full referrer URLs in our statistics. Referrer hostnames tell us where visits came from. Records are retained for 90 days. Do Not Track and Global Privacy Control disable this collection. Voting uses its own session cookie, separately from these statistics.

CREDITS

Built on Blender + Blender MCP. Source on GitHub SOON — the repo goes public once the bench does. If the bench is useful to you, you can support it. The leaderboard, gallery, run log, and vote booth all read live benchmark runs and votes.

Votes cast on this bench: 000482 · 157 runs · 25 contestants
Best viewed with Netscape Navigator 4.07 at 800×600 · real runs, live ratings