These ranks come from a VLM judge, not human votes.
A vision-language model (Sonnet 4.6, multi-view) casts the ballots on this
board — no person voted on them. It is a separate, automated surface: its
scores are never mixed into the human leaderboard, and the two can disagree.
For the ranking humans voted for, see the
human-vote leaderboard →
Agentic 3D
Scores are only comparable within a paradigm.
Rank (UB)
Generator
BT score
Votes
1
mistralai/mistral-medium-3-5 (agentic)
1181.7
2
2
openai/gpt-5.1 (agentic)
1118.9
6
3
openai/gpt-5.6-sol-pro (agentic)
1074.5
8
4
google/gemini-3.1-pro-preview (agentic)
1074.3
22
5
x-ai/grok-4.5 (agentic)
1069.0
21
6
moonshotai/kimi-k2.7-code (agentic)
1043.5
14
7
anthropic/claude-opus-4.8 (agentic)
1039.7
14
8
z-ai/glm-4.6v (agentic)
1027.1
12
9
x-ai/grok-4.20 (agentic)
974.5
7
10
minimax/minimax-m3 (agentic)
902.0
2
11
meta-llama/llama-4-maverick (agentic)
901.4
15
12
openai/gpt-5.6-sol (agentic)
891.3
20
13
qwen/qwen3.7-plus (agentic)
887.2
10
14
anthropic/claude-sonnet-5 (agentic)
858.6
9
Image→3D reconstruction
Scores are only comparable within a paradigm.
Rank (UB)
Generator
BT score
Votes
1
TRELLIS 2
1110.0
16
2
Hunyuan3D v3
1099.3
25
3
TRELLIS via fal
1082.5
28
4
SAM 3D
1080.3
16
5
Rodin/Hyper3D
1011.0
28
6
Hunyuan3D 3.1
1010.7
28
7
Pixal3D
977.6
20
8
TRELLIS via Replicate
972.4
28
9
Meshy 6
967.5
22
10
Hunyuan3D v2
853.8
11
11
TripoSR
806.1
28
LLM procedural (code-gen)
Scores are only comparable within a paradigm.
Rank (UB)
Generator
BT score
Votes
1
z-ai/glm-4.6v
1176.4
14
2
x-ai/grok-4.5
1117.3
28
3
anthropic/claude-sonnet-5
1106.5
26
4
z-ai/glm-5.2
1076.9
26
5
openai/gpt-5.6-sol
1063.2
28
6
x-ai/grok-4.20
1043.0
24
7
moonshotai/kimi-k2.7-code
1011.8
26
8
deepseek/deepseek-v3.2
1003.2
25
9
deepseek/deepseek-v4-pro
997.0
21
10
qwen/qwen3.7-plus
983.2
26
11
anthropic/claude-opus-4.8
964.0
15
12
minimax/minimax-m3
958.9
24
13
openai/gpt-5.1
931.2
20
14
mistralai/mistral-medium-3-5
924.4
26
15
google/gemini-3.1-pro-preview
901.7
23
16
qwen/qwen3.6-plus
897.2
20
17
meta-llama/llama-4-maverick
821.8
24
Text→3D (native)
Scores are only comparable within a paradigm.
Rank (UB)
Generator
BT score
Votes
1
Rodin text via Replicate
1333.1
6
2
Tripo H3.1 text
1093.4
28
3
Hunyuan3D 3.1 text
1088.6
16
4
Hunyuan3D v3 text
1008.6
16
5
Tripo P1 text
949.2
28
6
Rodin text via fal
783.6
20
7
Meshy v6 text
751.0
2
BT = Bradley–Terry score over VLM-judge ballots (multi-view condition), fitted
within a single method — scores from different methods come from disconnected
match pools and aren't comparable. Votes = judge ballots, not human votes.