Leaderboard ALL KINGDOMS
An agent renders, critiques, and revises its own 3D model in a loop.
Human voting is currently focused on the commercial image→3D and text→3D models, so these scores are paused where the votes so far left them. The AI-judge board still ranks this modality.
Ranked by human votes · 340 cast · see the AI-judge board →
How ranking works
Rank groups models whose 95% bootstrap CIs overlap into the same rank (they are not statistically separable). BT = Bradley–Terry score (Elo-scaled); the bar shows the 95% CI with the point estimate marked. Every board ranks a SINGLE method — scores from different methods come from disconnected match pools and aren't comparable, so there is no cross-method ranking. Elo updates live per vote. Recompute in Admin.
Agentic 3D
15 generators| Rank | Generator | BT score | 95% interval BT -53.8–1911.2 | Trend | Votes | Status |
|---|---|---|---|---|---|---|
| 1 | openai/gpt-5.6-sol-pro (agentic) | 1585.5 | [1353.8, 1911.2] | 5 | 25 more votes → firm ► | |
| 1 | openai/gpt-5.6-sol (agentic) | 1274.1 | [1076.5, 1594.5] | 12 | 18 more votes → firm ► | |
| 1 | x-ai/grok-4.5 (agentic) | 1128.6 | [748.2, 1624.4] | 6 | 24 more votes → firm ► | |
| 1 | google/gemini-3.1-pro-preview (agentic) | 1106.4 | [860.3, 1426.3] | 13 | 17 more votes → firm ► | |
| 1 | x-ai/grok-4.20 (agentic) | 1095.4 | [838.7, 1374.5] | 4 | 26 more votes → firm ► | |
| 1 | minimax/minimax-m3 (agentic) | 1090.5 | [584.1, 1595.4] | 5 | 25 more votes → firm ► | |
| 1 | anthropic/claude-opus-4.8 (agentic) | 1065.2 | [659.8, 1367.6] | 13 | 17 more votes → firm ► | |
| 2 | anthropic/claude-sonnet-5 (agentic) | 1057.8 | [677.4, 1346.4] | 10 | 20 more votes → firm ► | |
| 1 | qwen/qwen3.6-plus (agentic) | 1021.8 | [729.7, 1395.9] | 1 | 29 more votes → firm ► | |
| 3 | qwen/qwen3.7-plus (agentic) | 754.6 | [472.1, 953.9] | 6 | 24 more votes → firm ► | |
| 2 | meta-llama/llama-4-maverick (agentic) | 714.1 | [305.5, 1146.8] | 3 | 27 more votes → firm ► | |
| 3 | mistralai/mistral-medium-3-5 (agentic) | 686.9 | [339.6, 929.2] | 5 | 25 more votes → firm ► | |
| 2 | openai/gpt-5.1 (agentic) | 624.0 | [172.3, 1138.4] | 4 | 26 more votes → firm ► | |
| 4 | z-ai/glm-4.6v (agentic) | 605.3 | [240.3, 858.7] | 6 | 24 more votes → firm ► | |
| 2 | moonshotai/kimi-k2.7-code (agentic) | 349.0 | [-53.8, 1153.8] | 3 | 27 more votes → firm ► |
95% credible interval · point estimate · trend · firm a rank backed by 30+ votes; below that, Status counts the votes still needed · click a row for detail
Filters & bias audit
Bias audit — left(A) win rate 0.569 (≈0.50 = unbiased) · tie 0.059 · bad 0.191
VLM judge (Sonnet 4.6, multi-view) — automated LLM-judge rankings by paradigm
Loading automated rankings…