Leaderboard ALL KINGDOMS

An agent renders, critiques, and revises its own 3D model in a loop.

Human voting is currently focused on the commercial image→3D and text→3D models, so these scores are paused where the votes so far left them. The AI-judge board still ranks this modality.

Ranked by human votes · 340 cast · see the AI-judge board →

Share Post on X
How ranking works

Rank groups models whose 95% bootstrap CIs overlap into the same rank (they are not statistically separable). BT = Bradley–Terry score (Elo-scaled); the bar shows the 95% CI with the point estimate marked. Every board ranks a SINGLE method — scores from different methods come from disconnected match pools and aren't comparable, so there is no cross-method ranking. Elo updates live per vote. Recompute in Admin.

Agentic 3D

15 generators
Rank Generator BT score 95% interval BT -53.8–1911.2 Trend Votes Status
1 openai/gpt-5.6-sol-pro (agentic) code 1585.5
[1353.8, 1911.2]
5 25 more votes → firm
1 openai/gpt-5.6-sol (agentic) code 1274.1
[1076.5, 1594.5]
12 18 more votes → firm
1 x-ai/grok-4.5 (agentic) code 1128.6
[748.2, 1624.4]
6 24 more votes → firm
1 google/gemini-3.1-pro-preview (agentic) code 1106.4
[860.3, 1426.3]
13 17 more votes → firm
1 x-ai/grok-4.20 (agentic) code 1095.4
[838.7, 1374.5]
4 26 more votes → firm
1 minimax/minimax-m3 (agentic) code 1090.5
[584.1, 1595.4]
5 25 more votes → firm
1 anthropic/claude-opus-4.8 (agentic) code 1065.2
[659.8, 1367.6]
13 17 more votes → firm
2 anthropic/claude-sonnet-5 (agentic) code 1057.8
[677.4, 1346.4]
10 20 more votes → firm
1 qwen/qwen3.6-plus (agentic) code 1021.8
[729.7, 1395.9]
1 29 more votes → firm
3 qwen/qwen3.7-plus (agentic) code 754.6
[472.1, 953.9]
6 24 more votes → firm
2 meta-llama/llama-4-maverick (agentic) code 714.1
[305.5, 1146.8]
3 27 more votes → firm
3 mistralai/mistral-medium-3-5 (agentic) code 686.9
[339.6, 929.2]
5 25 more votes → firm
2 openai/gpt-5.1 (agentic) code 624.0
[172.3, 1138.4]
4 26 more votes → firm
4 z-ai/glm-4.6v (agentic) code 605.3
[240.3, 858.7]
6 24 more votes → firm
2 moonshotai/kimi-k2.7-code (agentic) code 349.0
[-53.8, 1153.8]
3 27 more votes → firm

95% credible interval · point estimate · trend · firm a rank backed by 30+ votes; below that, Status counts the votes still needed · click a row for detail

Filters & bias audit

Bias audit — left(A) win rate 0.569 (≈0.50 = unbiased) · tie 0.059 · bad 0.191

VLM judge (Sonnet 4.6, multi-view) — automated LLM-judge rankings by paradigm

Loading automated rankings…