These ranks come from a VLM judge, not human votes.
A vision-language model (Sonnet 4.6, multi-view) casts the ballots on this
board — no person voted on them. It is a separate, automated surface: its
scores are never mixed into the human leaderboard, and the two can disagree.
For the ranking humans voted for, see the
human-vote leaderboard →
Agentic 3D
Scores are only comparable within a paradigm.
Rank (UB)
Generator
BT score
Votes
1
openai/gpt-5.6-sol-pro (agentic)
1139.7
20
2
google/gemini-3.1-pro-preview (agentic)
1105.9
46
3
x-ai/grok-4.5 (agentic)
1093.4
49
4
anthropic/claude-opus-4.8 (agentic)
1058.1
38
5
openai/gpt-5.6-sol (agentic)
1053.9
56
6
anthropic/claude-sonnet-5 (agentic)
1002.1
29
7
x-ai/grok-4.20 (agentic)
999.7
35
8
moonshotai/kimi-k2.7-code (agentic)
990.2
25
9
z-ai/glm-4.6v (agentic)
979.5
34
10
minimax/minimax-m3 (agentic)
970.4
10
11
qwen/qwen3.7-plus (agentic)
948.7
32
12
openai/gpt-5.1 (agentic)
919.1
28
13
meta-llama/llama-4-maverick (agentic)
882.7
35
14
mistralai/mistral-medium-3-5 (agentic)
870.2
8
15
qwen/qwen3.6-plus (agentic)
829.0
7
Image→3D reconstruction
Scores are only comparable within a paradigm.
Rank (UB)
Generator
BT score
Votes
1
Hunyuan3D v3
1114.9
93
2
TRELLIS 2
1103.3
40
3
Hunyuan3D 3.1
1102.9
102
4
Meshy 6
1073.7
52
5
Rodin/Hyper3D
1027.0
104
6
TRELLIS via fal
1026.8
104
7
SAM 3D
995.0
42
8
TRELLIS via Replicate
967.4
76
9
InstantMesh
955.2
8
10
Hunyuan3D v2
952.4
67
11
Pixal3D
920.9
50
12
TripoSR
844.4
84
LLM procedural (code-gen)
Scores are only comparable within a paradigm.
Rank (UB)
Generator
BT score
Votes
1
openai/gpt-5.6-sol-pro
1346.5
4
2
openai/gpt-5.6-sol
1163.8
68
3
z-ai/glm-5.2
1110.7
74
4
anthropic/claude-sonnet-5
1096.8
60
5
moonshotai/kimi-k2.7-code
1073.7
62
6
google/gemini-3.1-pro-preview
1065.7
73
7
anthropic/claude-opus-4.8
1052.9
61
8
x-ai/grok-4.5
1043.2
67
9
x-ai/grok-4.20
1025.0
58
10
z-ai/glm-4.6v
1018.9
48
11
deepseek/deepseek-v4-pro
1014.2
48
12
deepseek/deepseek-v3.2
1010.0
62
13
qwen/qwen3.7-plus
975.7
64
14
minimax/minimax-m3
964.6
58
15
mistralai/mistral-medium-3-5
933.2
65
16
qwen/qwen3.6-plus
916.9
56
17
openai/gpt-5.1
898.1
67
18
meta-llama/llama-4-maverick
873.6
57
19
x-ai/grok-4.3
872.2
32
Text→3D (native)
Scores are only comparable within a paradigm.
Rank (UB)
Generator
BT score
Votes
1
Tripo H3.1 text
1143.3
95
2
Hunyuan3D v3 text
1053.9
87
3
Hunyuan3D 3.1 text
1023.5
81
4
Tripo P1 text
996.0
97
5
Rodin text via Replicate
928.0
72
6
Rodin text via fal
923.4
90
7
Meshy v6 text
783.8
2
BT = Bradley–Terry score over VLM-judge ballots (multi-view condition), fitted
within a single method — scores from different methods come from disconnected
match pools and aren't comparable. Votes = judge ballots, not human votes.