Leaderboard ALL KINGDOMS

An LLM writes code (e.g. Blender) that builds the 3D model.

Ranked by human votes · 1614 cast · see the AI-judge board →

Share Post on X
How ranking works

Rank groups models whose 95% bootstrap CIs overlap into the same rank (they are not statistically separable). BT = Bradley–Terry score (Elo-scaled); the bar shows the 95% CI with the point estimate marked. Every board ranks a SINGLE method — scores from different methods come from disconnected match pools and aren't comparable, so there is no cross-method ranking. Elo updates live per vote. Recompute in Admin.

Show all 53 · 31 unrated

LLM procedural (code-gen)

3 generators
Rank Generator BT score 95% interval BT 1005.3–1021.3 Trend Votes Status
1 z-ai/glm-4.6v code 1012.9
[1005.3, 1021.3]
2 28 more votes → firm
1 x-ai/grok-4.5 code 1012.9
[1005.3, 1021.3]
1 29 more votes → firm
1 anthropic/claude-sonnet-5 code 1012.9
[1005.3, 1021.3]
1 29 more votes → firm

95% credible interval · point estimate · trend · firm a rank backed by 30+ votes; below that, Status counts the votes still needed · click a row for detail

Filters & bias audit

Bias audit — left(A) win rate 0.5 (≈0.50 = unbiased) · tie 0.07 · bad 0.162

VLM judge (Sonnet 4.6, multi-view) — automated LLM-judge rankings by paradigm

Loading automated rankings…