Leaderboard FUNGI

An LLM writes code (e.g. Blender) that builds the 3D model.

Human voting is currently focused on the commercial image→3D and text→3D models, so these scores are paused where the votes so far left them. The AI-judge board still ranks this modality.

Ranked by human votes · 340 cast · see the AI-judge board →

Share Post on X
How ranking works

Rank groups models whose 95% bootstrap CIs overlap into the same rank (they are not statistically separable). BT = Bradley–Terry score (Elo-scaled); the bar shows the 95% CI with the point estimate marked. Every board ranks a SINGLE method — scores from different methods come from disconnected match pools and aren't comparable, so there is no cross-method ranking. Elo updates live per vote. Recompute in Admin.

LLM procedural (code-gen)

16 generators
Rank Generator BT score 95% interval BT -362.7–2173.4 Trend Votes Status
1 openai/gpt-5.6-sol code 1565.5
[890.2, 1892.8]
7 23 more votes → firm
1 anthropic/claude-opus-4.8 code 1301.8
[516.3, 1887.8]
5 25 more votes → firm
1 moonshotai/kimi-k2.7-code code 1301.8
[347.2, 2173.4]
2 28 more votes → firm
1 x-ai/grok-4.5 code 1301.8
[756.1, 2032.4]
3 27 more votes → firm
1 mistralai/mistral-medium-3-5 code 1301.8
[678.0, 1762.7]
4 26 more votes → firm
1 deepseek/deepseek-v4-pro code 1218.1
[563.8, 1897.6]
2 28 more votes → firm
1 minimax/minimax-m3 code 1038.0
[375.6, 1786.4]
4 26 more votes → firm
1 google/gemini-3.1-pro-preview code 885.2
[52.8, 1431.0]
3 27 more votes → firm
1 z-ai/glm-5.2 code 885.2
[339.6, 1274.0]
1 29 more votes → firm
1 qwen/qwen3.7-plus code 885.2
[216.0, 1495.2]
3 27 more votes → firm
1 x-ai/grok-4.20 code 870.6
[332.7, 1640.7]
6 24 more votes → firm
1 meta-llama/llama-4-maverick code 621.5
[-18.2, 1232.5]
1 29 more votes → firm
1 anthropic/claude-sonnet-5 code 547.2
[-6.2, 1439.1]
5 25 more votes → firm
1 z-ai/glm-4.6v code 468.7
[-362.7, 1184.7]
1 29 more votes → firm
1 qwen/qwen3.6-plus code 454.1
[16.4, 1202.1]
1 29 more votes → firm
1 deepseek/deepseek-v3.2 code 146.7
[-346.5, 1249.3]
4 26 more votes → firm

95% credible interval · point estimate · trend · firm a rank backed by 30+ votes; below that, Status counts the votes still needed · click a row for detail

Filters & bias audit

Bias audit — left(A) win rate 0.569 (≈0.50 = unbiased) · tie 0.059 · bad 0.191

VLM judge (Sonnet 4.6, multi-view) — automated LLM-judge rankings by paradigm

Loading automated rankings…