This file maps paper-facing manipulation benchmark audit claims to the lightweight public files in this repository. Claims are marked as public-recomputed when scripts/recompute_claims.py recomputes the number from included files; intentional exclusions are called out where they affect artifact scope.
- LIBERO fixed-instruction shortcut policy: public-recomputed from
shortcut_solvability/results/libero/results.csvandshortcut_solvability/results/libero/best_checkpoint/*/{summary.json,trials.csv}. Headline cells are Spatial495/500 = 99.0%, Object500/500 = 100.0%, Goal494/500 = 98.8%, and Long462/500 = 92.4%. - CALVIN fixed-instruction shortcut policy: public-recomputed from
shortcut_solvability/results/calvin/results.csvandshortcut_solvability/results/calvin/best_checkpoint/*/{summary.json,trials.csv}. Headline complete cells areD->DATC3.123over1000sequences andABCD->DATC3.872over1000sequences.ABC->DATC3.242is included as supporting evidence and remains labeledno_full_official_artifact_manifest.
- LIBERO Goal shared-instance pairwise disagreement: public-recomputed from
statistical_significance/libero_goal_5x5k/policy_success_summary.csv,pairwise_disagreement.csv,policies/*/{policy_summary.json,episodes_combined.csv}, andshared/{libero_config.yaml,init_state_goal_5000_MANIFEST.json}. The released package verifies five policies with5000shared instance IDs each, then recomputes every joined policy pair. The paper-facing value is mean pairwiseD = 0.03528and median0.0354. - Aggregate-data previous-SOTA significance categories: generated from the five released Sam official-protocol previous-SOTA exports under
leaderboards/stat_significance_sam_export_20260522T233929/usingstatistical_significance/code/generate_significance_categories.pyand the cutoff logic instatistical_significance/code/significance_cutoffs.py. The released category tables contain1349comparable rows:497no-improvement,145provably-not-significant,331provably-significant, and376indeterminate. Another212benchmark/track rows are explicitly excluded inexcluded_missing_scores.csvbecause the current score or previous-SOTA score is missing in the source export. Count conversion uses Python nearest-evenround()to match Sam's reference code, and the released CSVs include scaled-count and rounding-residual columns.
- Leaderboard CSV snapshots: copied from
/home/ripl/workspace/leaderboardsat commit725613ff4e10f2725de3ac4ebcdbff28fc39586b. The release includes five primary benchmark citation trackers, four supplementary benchmark trackers, and the five Sam official-protocol previous-SOTA exports. It intentionally excludes raw audit folders, caches, zips, scripts, downloaded papers, and non-CSV artifacts from the source repository.
- SimplerEnv fixed-grid calibration: public-recomputed from
creeping_overfitting/results/simplerenv/fixed_grid_calibration/per_episode_results_all.csvand the summary CSVs. Aggregate policy rates are CogACT-Base561/1152 = 48.70%, SpatialVLA432/1152 = 37.50%, InternVLA-M1705/1152 = 61.20%, X-VLA-WidowX834/1152 = 72.40%, and Dexbotic / DB-MemVLA745/1152 = 64.67%. - SimplerEnv Protocol A-E distribution-overfitting matrix: public-recomputed from
creeping_overfitting/results/simplerenv/distribution_overfitting/per_episode_results_all.csvand the summary CSVs. Aggregate policy rates are CogACT-Base230/2016 = 11.41%, SpatialVLA182/2016 = 9.03%, InternVLA-M1287/2016 = 14.24%, X-VLA-WidowX1012/2016 = 50.20%, and Dexbotic / DB-MemVLA841/2016 = 41.72%. - CALVIN Protocol 1 resampled-pose distribution-overfitting: public-recomputed from
creeping_overfitting/results/calvin/resampled_pose_per_sequence.csv,distribution_overfitting_summary.csv, andcombined_summary.json. ATC drops are X-VLA1.027, GR-10.749, and RoboFlamingo0.498. - CALVIN fresh-sequence sample-overfitting: public-recomputed from
creeping_overfitting/results/calvin/fresh_sequence_per_sequence.csv,fresh_sequence_summary.csv, andcombined_summary.json. Pooled ATC deltas versus matched calibration are X-VLA-0.0145, GR-10.1070, and RoboFlamingo0.0690, where positive means calibration scored higher. Result provenance is verified atEval_Policies_CoRLcommiteba7c0037294557427a7a854c91be56d3f2838ec. - LIBERO Layer 2 fresh-init-state summaries: public-recomputed from
creeping_overfitting/results/libero/fresh_init_state/*.csv,official_calibration/*.csv, andsample_overfitting_summary.csv. Policy-level fresh rates are Spatial Forcing9718/10000 = 97.18%, SimVLA9760/10000 = 97.60%, and Pi05 / LeRobot9741/10000 = 97.41%. Per-episode rollout rows are excluded from this minimal package.
- SimplerEnv WidowX scripted-demo DSD: public-recomputed from
data_source_dependency/results/scripted_widowx/trials.csv,results.csv, andaggregate_summary.json. The result is stack24/24, carrot23/24, spoon21/24, eggplant23/24, and overall91/96 = 94.79%.