AI-judge board PLANTS

These ranks come from a VLM judge, not human votes. A vision-language model (Sonnet 4.6, multi-view) casts the ballots on this board — no person voted on them. It is a separate, automated surface: its scores are never mixed into the human leaderboard, and the two can disagree. For the ranking humans voted for, see the human-vote leaderboard →

Agentic 3D

Scores are only comparable within a paradigm.
Rank (UB) Generator BT score Votes
1 openai/gpt-5.6-sol (agentic) 1152.7 20
2 x-ai/grok-4.5 (agentic) 1118.2 16
3 google/gemini-3.1-pro-preview (agentic) 1070.5 14
4 anthropic/claude-opus-4.8 (agentic) 1062.8 14
5 qwen/qwen3.7-plus (agentic) 1013.7 10
6 x-ai/grok-4.20 (agentic) 1003.4 16
7 moonshotai/kimi-k2.7-code (agentic) 983.8 7
8 anthropic/claude-sonnet-5 (agentic) 966.7 8
9 qwen/qwen3.6-plus (agentic) 927.7 3
10 z-ai/glm-4.6v (agentic) 892.2 16
11 meta-llama/llama-4-maverick (agentic) 883.5 12
12 openai/gpt-5.1 (agentic) 870.7 12
13 mistralai/mistral-medium-3-5 (agentic) 723.5 2

Image→3D reconstruction

Scores are only comparable within a paradigm.
Rank (UB) Generator BT score Votes
1 TRELLIS 2 1205.8 18
2 Hunyuan3D 3.1 1155.0 60
3 Meshy 6 1143.5 18
4 Hunyuan3D v3 1084.8 52
5 TRELLIS via fal 1001.3 60
6 Hunyuan3D v2 995.3 40
7 Rodin/Hyper3D 986.3 60
8 InstantMesh 959.4 8
9 TRELLIS via Replicate 929.7 40
10 SAM 3D 911.6 18
11 Pixal3D 909.0 18
12 TripoSR 861.6 56

LLM procedural (code-gen)

Scores are only comparable within a paradigm.
Rank (UB) Generator BT score Votes
1 openai/gpt-5.6-sol-pro 1341.9 4
2 openai/gpt-5.6-sol 1210.1 24
3 google/gemini-3.1-pro-preview 1189.8 36
4 z-ai/glm-5.2 1132.7 34
5 x-ai/grok-4.20 1074.9 18
6 anthropic/claude-opus-4.8 1048.6 36
7 moonshotai/kimi-k2.7-code 1035.0 22
8 anthropic/claude-sonnet-5 1026.4 18
9 deepseek/deepseek-v4-pro 1018.5 14
10 minimax/minimax-m3 997.3 19
11 deepseek/deepseek-v3.2 971.8 22
12 x-ai/grok-4.5 956.3 23
13 qwen/qwen3.6-plus 947.3 23
14 mistralai/mistral-medium-3-5 943.1 24
15 qwen/qwen3.7-plus 932.8 22
16 meta-llama/llama-4-maverick 894.9 20
17 z-ai/glm-4.6v 893.7 20
18 x-ai/grok-4.3 883.3 32
19 openai/gpt-5.1 879.1 33

Text→3D (native)

Scores are only comparable within a paradigm.
Rank (UB) Generator BT score Votes
1 Tripo H3.1 text 1179.2 55
2 Hunyuan3D v3 text 1047.5 55
3 Hunyuan3D 3.1 text 1044.1 57
4 Tripo P1 text 998.4 57
5 Rodin text via fal 935.1 58
6 Rodin text via Replicate 831.5 50

BT = Bradley–Terry score over VLM-judge ballots (multi-view condition), fitted within a single method — scores from different methods come from disconnected match pools and aren't comparable. Votes = judge ballots, not human votes.