Leaderboard ALL KINGDOMS
An LLM writes code (e.g. Blender) that builds the 3D model.
Human voting is currently focused on the commercial image→3D and text→3D models, so these scores are paused where the votes so far left them. The AI-judge board still ranks this modality.
Ranked by human votes · 340 cast · see the AI-judge board →
How ranking works
Rank groups models whose 95% bootstrap CIs overlap into the same rank (they are not statistically separable). BT = Bradley–Terry score (Elo-scaled); the bar shows the 95% CI with the point estimate marked. Every board ranks a SINGLE method — scores from different methods come from disconnected match pools and aren't comparable, so there is no cross-method ranking. Elo updates live per vote. Recompute in Admin.
LLM procedural (code-gen)
18 generators| Rank | Generator | BT score | 95% interval BT -100.2–2328.9 | Trend | Votes | Status |
|---|---|---|---|---|---|---|
| 1 | openai/gpt-5.6-sol-pro | 1832.5 | [1128.0, 2328.9] | 1 | 29 more votes → firm ► | |
| openai/gpt-5.6-sol | 1430.6 | [1329.8, 1696.6] | 16 | 14 more votes → firm ► | ||
| 1 | x-ai/grok-4.5 | 1416.0 | [1213.9, 1939.3] | 8 | 22 more votes → firm ► | |
| 1 | anthropic/claude-opus-4.8 | 1133.6 | [828.1, 1585.0] | 13 | 17 more votes → firm ► | |
| 1 | google/gemini-3.1-pro-preview | 1125.6 | [866.1, 1629.8] | 9 | 21 more votes → firm ► | |
| 1 | qwen/qwen3.7-plus | 1109.8 | [768.1, 1523.7] | 7 | 23 more votes → firm ► | |
| 1 | anthropic/claude-sonnet-5 | 1076.0 | [732.3, 1406.6] | 14 | 16 more votes → firm ► | |
| 1 | x-ai/grok-4.3 | 1023.1 | [414.2, 1555.8] | 3 | 27 more votes → firm ► | |
| 1 | moonshotai/kimi-k2.7-code | 998.5 | [569.3, 1394.4] | 9 | 21 more votes → firm ► | |
| 2 | z-ai/glm-5.2 | 972.2 | [544.8, 1176.7] | 13 | 17 more votes → firm ► | |
| 1 | openai/gpt-5.1 | 931.8 | [388.4, 1627.5] | 3 | 27 more votes → firm ► | |
| 1 | mistralai/mistral-medium-3-5 | 905.7 | [170.1, 1216.5] | 9 | 21 more votes → firm ► | |
| 1 | x-ai/grok-4.20 | 903.1 | [321.7, 1351.3] | 9 | 21 more votes → firm ► | |
| 1 | minimax/minimax-m3 | 900.0 | [61.8, 1394.7] | 5 | 25 more votes → firm ► | |
| 1 | deepseek/deepseek-v4-pro | 869.2 | [315.3, 1232.5] | 7 | 23 more votes → firm ► | |
| 3 | deepseek/deepseek-v3.2 | 782.0 | [344.8, 983.0] | 9 | 21 more votes → firm ► | |
| 3 | z-ai/glm-4.6v | 736.7 | [134.8, 1024.6] | 8 | 22 more votes → firm ► | |
| 2 | qwen/qwen3.6-plus | 486.5 | [-77.7, 1167.4] | 1 | 29 more votes → firm ► | |
| 3 | meta-llama/llama-4-maverick | 468.2 | [-100.2, 894.8] | 4 | 26 more votes → firm ► |
95% credible interval · point estimate · trend · firm a rank backed by 30+ votes; below that, Status counts the votes still needed · click a row for detail
Filters & bias audit
Bias audit — left(A) win rate 0.569 (≈0.50 = unbiased) · tie 0.059 · bad 0.191
VLM judge (Sonnet 4.6, multi-view) — automated LLM-judge rankings by paradigm
Loading automated rankings…