Leaderboard PLANTS

An agent renders, critiques, and revises its own 3D model in a loop.

Human voting is currently focused on the commercial image→3D and text→3D models, so these scores are paused where the votes so far left them. The AI-judge board still ranks this modality.

Ranked by human votes · 340 cast · see the AI-judge board →

Share Post on X
How ranking works

Rank groups models whose 95% bootstrap CIs overlap into the same rank (they are not statistically separable). BT = Bradley–Terry score (Elo-scaled); the bar shows the 95% CI with the point estimate marked. Every board ranks a SINGLE method — scores from different methods come from disconnected match pools and aren't comparable, so there is no cross-method ranking. Elo updates live per vote. Recompute in Admin.

Agentic 3D

13 generators
Rank Generator BT score 95% interval BT 53.3–1750.7 Trend Votes Status
1 openai/gpt-5.6-sol (agentic) code 1470.4
[1056.3, 1750.7]
5 25 more votes → firm
1 google/gemini-3.1-pro-preview (agentic) code 1262.0
[976.7, 1600.1]
9 21 more votes → firm
1 anthropic/claude-opus-4.8 (agentic) code 1113.4
[454.5, 1398.1]
6 24 more votes → firm
1 x-ai/grok-4.20 (agentic) code 958.0
[633.6, 1336.6]
2 28 more votes → firm
1 mistralai/mistral-medium-3-5 (agentic) code 888.7
[537.1, 1260.4]
3 27 more votes → firm
1 qwen/qwen3.7-plus (agentic) code 883.7
[169.9, 1490.9]
3 27 more votes → firm
1 anthropic/claude-sonnet-5 (agentic) code 845.6
[565.4, 1269.6]
3 27 more votes → firm
1 openai/gpt-5.1 (agentic) code 845.5
[508.0, 1180.1]
1 29 more votes → firm
1 meta-llama/llama-4-maverick (agentic) code 845.5
[522.4, 1169.5]
1 29 more votes → firm
1 minimax/minimax-m3 (agentic) code 696.8
[53.3, 1163.6]
1 29 more votes → firm
1 moonshotai/kimi-k2.7-code (agentic) code 595.9
[135.3, 1188.3]
2 28 more votes → firm
1 z-ai/glm-4.6v (agentic) code 429.1
[117.4, 1120.4]
3 27 more votes → firm
1 x-ai/grok-4.5 (agentic) code 429.0
[130.5, 1176.9]
1 29 more votes → firm

95% credible interval · point estimate · trend · firm a rank backed by 30+ votes; below that, Status counts the votes still needed · click a row for detail

Filters & bias audit

Bias audit — left(A) win rate 0.569 (≈0.50 = unbiased) · tie 0.059 · bad 0.191

VLM judge (Sonnet 4.6, multi-view) — automated LLM-judge rankings by paradigm

Loading automated rankings…