Coverage & governance ALL KINGDOMS

How much evidence backs each ranking — a generator's rank is only as trustworthy as the votes behind it, and confidence firms up as vote counts grow. Published in full so a rank can be read with its uncertainty, not in isolation.

108 Agentic 3D
123 Image→3D reconstruction
185 LLM procedural (code-gen)
99 Text→3D (native)

Generator coverage — 53 generators with arena outputs, by total Mode-A votes

GeneratorTasksOutputsA-votesConfidenceIn arena
Hunyuan3D v3 text model 16 16 26 provisional
Tripo P1 text model 15 15 24 provisional
Rodin text via fal model 14 14 23 provisional
Meshy v6 text model 14 14 22 provisional
Tripo H3.1 text model 16 16 22 provisional
Hunyuan3D 3.1 text model 13 13 20 provisional
Rodin text via Replicate model 11 11 20 provisional
Hunyuan3D v3 model 11 13 19 provisional
SAM 3D model 16 16 17 provisional
Hunyuan3D 3.1 model 11 13 16 provisional
deepseek/deepseek-v3.2 model 14 14 16 provisional
anthropic/claude-opus-4.8 (agentic) model 11 11 15 provisional
Meshy 6 model 11 11 14 provisional
anthropic/claude-sonnet-5 model 13 13 14 provisional
google/gemini-3.1-pro-preview (agentic) model 12 12 14 provisional
openai/gpt-5.6-sol model 14 14 14 provisional
x-ai/grok-4.20 model 14 14 14 provisional
Hunyuan3D v2 model 10 11 13 provisional
Rodin/Hyper3D model 11 13 13 provisional
openai/gpt-5.6-sol (agentic) model 12 12 13 provisional
TRELLIS via Replicate model 11 12 12 provisional
anthropic/claude-opus-4.8 model 12 12 12 provisional
anthropic/claude-sonnet-5 (agentic) model 10 10 12 provisional
z-ai/glm-5.2 model 12 12 12 provisional
Pixal3D model 11 11 11 provisional
mistralai/mistral-medium-3-5 model 12 12 10 provisional
qwen/qwen3.7-plus model 13 13 10 provisional
deepseek/deepseek-v4-pro model 10 10 9 provisional
meta-llama/llama-4-maverick model 6 6 9 provisional
moonshotai/kimi-k2.7-code model 10 10 9 provisional
x-ai/grok-4.5 (agentic) model 9 9 9 provisional
z-ai/glm-4.6v (agentic) model 8 8 9 provisional
TRELLIS via fal model 8 9 8 provisional
google/gemini-3.1-pro-preview model 11 11 8 provisional
qwen/qwen3.7-plus (agentic) model 8 8 8 provisional
x-ai/grok-4.20 (agentic) model 7 7 8 provisional
x-ai/grok-4.5 model 13 13 8 provisional
TRELLIS 2 model 6 6 7 provisional
minimax/minimax-m3 model 8 8 7 provisional
openai/gpt-5.1 (agentic) model 6 6 7 provisional
meta-llama/llama-4-maverick (agentic) model 6 6 6 provisional
minimax/minimax-m3 (agentic) model 6 6 6 provisional
z-ai/glm-4.6v model 8 8 6 provisional
mistralai/mistral-medium-3-5 (agentic) model 4 4 5 provisional
openai/gpt-5.6-sol-pro (agentic) model 4 4 5 provisional
TripoSR model 6 6 4 provisional
moonshotai/kimi-k2.7-code (agentic) model 4 4 4 provisional
qwen/qwen3.6-plus model 4 4 4 provisional
x-ai/grok-4.3 model 4 4 4 provisional
InstantMesh model 2 2 3 provisional
openai/gpt-5.1 model 6 6 3 provisional
openai/gpt-5.6-sol-pro model 1 1 1 provisional
qwen/qwen3.6-plus (agentic) model 1 1 1 provisional

How models enter & compete

  • Entry — every output is added by an administrator or via the public submission queue, then moderated before it enters the arena. There is no private or preferential pre-testing: a model is not quietly trialled and published only if it scores well.
  • Fair matchmaking — pairs are drawn by a sampler that biases toward the least-compared outputs, so coverage spreads evenly rather than concentrating votes on a favoured few. Every model is drawn from the same pool by the same sampler.
  • No silent deprecation — outputs are not removed to flatter the board; the arena has no deprecation mechanism, so what competed stays on the record.
  • Exclusions are principled, not selective — only reference ground-truth scans and untextured geometry-only outputs are held out of the Mode-A perceptual pool (they confound a visual-preference vote); they remain in the objective Mode-B board. Excluded generators are flagged in the table above.

Reading a rank

A generator's rank is only as trustworthy as the votes behind it. Below 30 total Mode-A votes a rank is marked provisional; at or above it, firm. Bradley–Terry confidence intervals on the leaderboard make the same uncertainty visible per generator.

Current limits, stated plainly: the arena is in an internal evaluation phase, so vote volume is low and most ranks read provisional — ranks will firm up as voting scales. Mode-A votes here are internal and not yet from a public pool.

Task coverage — 20 active tasks

"Mode-B" = objective scoring against held-out ground-truth is available for that task.

TaskCategoryTierGeneratorsOutputs A-votesJudge votesMode-BMode-C
Solanum lycopersicum — single-image → 3D reconstruction Plants easy 42 51 32 0
Anas platyrhynchos — single-image → 3D reconstruction Animals hard 43 43 21 0
Glycine max — single-image → 3D reconstruction Plants moderate 39 39 27 0
Rosa — single-image → 3D reconstruction Plants hard 38 38 18 0
Canis lupus familiaris — single-image → 3D reconstruction Animals hard 37 37 21 0
Zea mays — single-image → 3D reconstruction Plants moderate 37 37 28 0
Arabidopsis thaliana — single-image → 3D reconstruction Plants hard 35 35 33 0
Danaus plexippus — single-image → 3D reconstruction Animals hard 35 35 13 0
Carassius auratus — single-image → 3D reconstruction Animals moderate 34 34 19 0
Morchella esculenta — single-image → 3D reconstruction Fungi moderate 29 29 22 0
Amanita muscaria — single-image → 3D reconstruction Fungi moderate 28 28 10 0
Boletus edulis — single-image → 3D reconstruction Fungi easy 26 26 17 0
Lycoperdon perlatum — single-image → 3D reconstruction Fungi easy 26 26 20 0
Hericium erinaceus — single-image → 3D reconstruction Fungi hard 20 20 10 0
Pinus sylvestris — single-image → 3D reconstruction Plants hard 20 20 30 0
Trametes versicolor — single-image → 3D reconstruction Fungi hard 17 17 10 0
Arabidopsis thaliana — botanical plausibility Synthetic Plants hard 0 0 0 0
Pinus sylvestris — botanical plausibility Synthetic Plants hard 0 0 0 0
Solanum lycopersicum — botanical plausibility Synthetic Plants easy 0 0 0 0
Zea mays — botanical plausibility Synthetic Plants moderate 0 0 0 0

Mode-C — botanical-trait accuracy experimental

Generator-level mean botanical accuracy — a 3D model graded against a literature-sourced per-taxon trait rubric by a calibrated VLM trait-checker. Only trait classes that have passed the human↔VLM agreement gate count toward the score. No class has passed the gate yet, so this axis is experimental: shown for inspection, not yet a ranking signal.

No scored outputs yet. Per-output scorecards live at /trait/<id>.

Machine-readable: /api/coverage.json.