MetaBench · Preprint 2026

A Statistical Audit of Physical AI Benchmark Redundancy

Zaruhi Navasardyan1   Hrant Davtyan1

1Metric AI Lab  ·  zaruhi@metric.am   hrant@metric.am

Twelve physical‑AI benchmarks. Four of them carry 78.5% of the signal — the rest largely repeat what you already know.

51models audited,
1B–241B parameters
12benchmarks, from a
registry of 51
0.49mean rank correlation
between benchmarks
22 / 51models move ≥3 places
when duplicates collapse

Physical AI models are evaluated on suites of benchmarks that differ across model reports, leaving the model-by-benchmark matrix sparse and the relationship between benchmarks unmeasured. We construct a matrix of 51 models on 12 physical AI benchmarks, selected from a registry of 51 benchmarks and 152 models by reporting density, combining scores from model cards and benchmark papers with our own evaluation runs under each benchmark's official protocol. We measure how much information the benchmarks share and show quantitative evidence of redundancy. Redundancy affects reported rankings: collapsing the two substitute pairs into single columns moves 22 of 51 models by three or more places under an equally weighted average. We then select benchmarks greedily under a utility combining score dispersion with variance not explained by the already-selected set, and obtain a four-benchmark subset retaining 78.5% of the utility of all 12, on which we fit a Bradley–Terry ranking. The procedure requires only benchmark-level scores with sufficient overlap and is not specific to physical AI.

01

The redundancy map

Spearman rank correlation between every pair of benchmarks, computed on pairwise-complete model scores. Two pairs stand out as near-substitutes: EmbSpatial–CV-Bench (ρ = 0.876) and Where2Place–RefSpatial-Bench (ρ = 0.860). Hover any cell for detail.

Darker blue means the two benchmarks rank models more alike. Benchmarks marked ● are the four retained by the greedy selection.

02

Which benchmarks repeat the others

Each benchmark's mean rank correlation with the other eleven — a direct measure of how much of it is already covered elsewhere. The four retained benchmarks all sit in the least-redundant half, three of them at the very bottom: the selection picked independent signal without being told to.

03

Rankings move when duplicates collapse

Model rank under an equally weighted average of all twelve benchmarks (left) versus the same average after each substitute pair is merged into a single column (right). 22 of 51 models shift by three or more places, the largest by nine — purely from removing double-counted signal, with no new evaluation.

04

How firm is a rank?

Bradley–Terry rank on the four-benchmark core suite, with the span of ranks the same model takes across leave-one-benchmark-out refits. A long bar means the model's standing depends on which benchmark you happen to include.

05

Bradley–Terry ranking

All 51 models, ranked on the retained four-benchmark suite (RefSpatial-Bench, MindCube, VSI-Bench, BLINK).

Download CSV ↓
06

The score matrix

The full model×benchmark matrix behind every figure above. Click a column to sort.

Download CSV ↓