Leaderboard ALL KINGDOMS

An LLM writes code (e.g. Blender) that builds the 3D model.

Human voting is currently focused on the commercial image→3D and text→3D models, so these scores are paused where the votes so far left them. The AI-judge board still ranks this modality.

Ranked by human votes · 340 cast · see the AI-judge board →

Share Post on X
How ranking works

Rank groups models whose 95% bootstrap CIs overlap into the same rank (they are not statistically separable). BT = Bradley–Terry score (Elo-scaled); the bar shows the 95% CI with the point estimate marked. Every board ranks a SINGLE method — scores from different methods come from disconnected match pools and aren't comparable, so there is no cross-method ranking. Elo updates live per vote. Recompute in Admin.

LLM procedural (code-gen)

18 generators
Rank Generator BT score 95% interval BT -100.2–2328.9 Trend Votes Status
1 openai/gpt-5.6-sol-pro code 1832.5
[1128.0, 2328.9]
1 29 more votes → firm
openai/gpt-5.6-sol code 1430.6
[1329.8, 1696.6]
16 14 more votes → firm
1 x-ai/grok-4.5 code 1416.0
[1213.9, 1939.3]
8 22 more votes → firm
1 anthropic/claude-opus-4.8 code 1133.6
[828.1, 1585.0]
13 17 more votes → firm
1 google/gemini-3.1-pro-preview code 1125.6
[866.1, 1629.8]
9 21 more votes → firm
1 qwen/qwen3.7-plus code 1109.8
[768.1, 1523.7]
7 23 more votes → firm
1 anthropic/claude-sonnet-5 code 1076.0
[732.3, 1406.6]
14 16 more votes → firm
1 x-ai/grok-4.3 code 1023.1
[414.2, 1555.8]
3 27 more votes → firm
1 moonshotai/kimi-k2.7-code code 998.5
[569.3, 1394.4]
9 21 more votes → firm
2 z-ai/glm-5.2 code 972.2
[544.8, 1176.7]
13 17 more votes → firm
1 openai/gpt-5.1 code 931.8
[388.4, 1627.5]
3 27 more votes → firm
1 mistralai/mistral-medium-3-5 code 905.7
[170.1, 1216.5]
9 21 more votes → firm
1 x-ai/grok-4.20 code 903.1
[321.7, 1351.3]
9 21 more votes → firm
1 minimax/minimax-m3 code 900.0
[61.8, 1394.7]
5 25 more votes → firm
1 deepseek/deepseek-v4-pro code 869.2
[315.3, 1232.5]
7 23 more votes → firm
3 deepseek/deepseek-v3.2 code 782.0
[344.8, 983.0]
9 21 more votes → firm
3 z-ai/glm-4.6v code 736.7
[134.8, 1024.6]
8 22 more votes → firm
2 qwen/qwen3.6-plus code 486.5
[-77.7, 1167.4]
1 29 more votes → firm
3 meta-llama/llama-4-maverick code 468.2
[-100.2, 894.8]
4 26 more votes → firm

95% credible interval · point estimate · trend · firm a rank backed by 30+ votes; below that, Status counts the votes still needed · click a row for detail

Filters & bias audit

Bias audit — left(A) win rate 0.569 (≈0.50 = unbiased) · tie 0.059 · bad 0.191

VLM judge (Sonnet 4.6, multi-view) — automated LLM-judge rankings by paradigm

Loading automated rankings…