AI-judge board ALL KINGDOMS

An LLM writes code (e.g. Blender) that builds the 3D model.

These ranks come from a VLM judge, not human votes. A vision-language model (Sonnet 4.6, multi-view) casts the ballots on this board — no person voted on them. It is a separate, automated surface: its scores are never mixed into the human leaderboard, and the two can disagree. For the ranking humans voted for, see the human-vote LLM procedural (code-gen) board →

LLM procedural (code-gen)

Scores are only comparable within a paradigm.
Rank (UB) Generator BT score Votes
1 openai/gpt-5.6-sol-pro 1346.5 4
2 openai/gpt-5.6-sol 1163.8 68
3 z-ai/glm-5.2 1110.7 74
4 anthropic/claude-sonnet-5 1096.8 60
5 moonshotai/kimi-k2.7-code 1073.7 62
6 google/gemini-3.1-pro-preview 1065.7 73
7 anthropic/claude-opus-4.8 1052.9 61
8 x-ai/grok-4.5 1043.2 67
9 x-ai/grok-4.20 1025.0 58
10 z-ai/glm-4.6v 1018.9 48
11 deepseek/deepseek-v4-pro 1014.2 48
12 deepseek/deepseek-v3.2 1010.0 62
13 qwen/qwen3.7-plus 975.7 64
14 minimax/minimax-m3 964.6 58
15 mistralai/mistral-medium-3-5 933.2 65
16 qwen/qwen3.6-plus 916.9 56
17 openai/gpt-5.1 898.1 67
18 meta-llama/llama-4-maverick 873.6 57
19 x-ai/grok-4.3 872.2 32

BT = Bradley–Terry score over VLM-judge ballots (multi-view condition), fitted within a single method — scores from different methods come from disconnected match pools and aren't comparable. Votes = judge ballots, not human votes.