Leaderboard ANIMALS

An LLM writes code (e.g. Blender) that builds the 3D model.

Human voting is currently focused on the commercial image→3D and text→3D models, so these scores are paused where the votes so far left them. The AI-judge board still ranks this modality.

Ranked by human votes · 340 cast · see the AI-judge board →

Share Post on X
How ranking works

Rank groups models whose 95% bootstrap CIs overlap into the same rank (they are not statistically separable). BT = Bradley–Terry score (Elo-scaled); the bar shows the 95% CI with the point estimate marked. Every board ranks a SINGLE method — scores from different methods come from disconnected match pools and aren't comparable, so there is no cross-method ranking. Elo updates live per vote. Recompute in Admin.

LLM procedural (code-gen)

15 generators
Rank Generator BT score 95% interval BT 227.1–1660.1 Trend Votes Status
1 qwen/qwen3.7-plus code 1523.3
[1097.0, 1660.1]
2 28 more votes → firm
1 openai/gpt-5.6-sol code 1514.5
[1080.2, 1502.6]
2 28 more votes → firm
1 x-ai/grok-4.5 code 1145.6
[1092.2, 1462.6]
2 28 more votes → firm
1 anthropic/claude-sonnet-5 code 1139.4
[1060.0, 1413.7]
3 27 more votes → firm
1 anthropic/claude-opus-4.8 code 1126.5
[861.5, 1374.8]
2 28 more votes → firm
1 deepseek/deepseek-v3.2 code 1112.8
[631.3, 1262.7]
2 28 more votes → firm
1 z-ai/glm-5.2 code 1102.2
[614.9, 1132.7]
2 28 more votes → firm
1 z-ai/glm-4.6v code 1096.2
[642.1, 1250.5]
2 28 more votes → firm
1 minimax/minimax-m3 code 729.2
[595.5, 1141.2]
1 29 more votes → firm
1 moonshotai/kimi-k2.7-code code 726.6
[620.0, 1232.1]
3 27 more votes → firm
1 x-ai/grok-4.20 code 722.9
[553.1, 1141.6]
1 29 more votes → firm
1 openai/gpt-5.1 code 716.3
[467.8, 1140.7]
2 28 more votes → firm
1 deepseek/deepseek-v4-pro code 703.0
[307.3, 1134.2]
2 28 more votes → firm
1 mistralai/mistral-medium-3-5 code 679.5
[248.6, 1152.4]
1 29 more votes → firm
1 meta-llama/llama-4-maverick code 310.1
[227.1, 1146.4]
1 29 more votes → firm

95% credible interval · point estimate · trend · firm a rank backed by 30+ votes; below that, Status counts the votes still needed · click a row for detail

Filters & bias audit

Bias audit — left(A) win rate 0.569 (≈0.50 = unbiased) · tie 0.059 · bad 0.191

VLM judge (Sonnet 4.6, multi-view) — automated LLM-judge rankings by paradigm

Loading automated rankings…