Robot Policy Evaluation: Why 90% vs 92% Proves Little
Abstract When evaluating robot control policies, many practitioners draw direct conclusions from simple success‑rate percentages. For instance, given Policy A with 90 % success and Policy B with 92 % success, people frequently claim Policy B performs better. Nevertheless, purely comparing percentage figures without sample size, confidence intervals, paired experimental design and statistical power analysis often produces unreliable judgments. Drawing on Clopper‑Pearson exact confidence intervals, Wilson score intervals, McNemar’s paired testing and hierarchical episode‑within‑task structure, this article lays out a complete practical workflow for robot policy evaluation, covering pre‑experiment planning and post‑hoc result checking. For engineering teams running robot‑simulation benchmarks mixed with LLM‑based agent workloads, an API gateway such as 4sapi can help standardize telemetry collection and multi‑backend request orchestration. 1. The Pitfall: Percentages Without Sample Sizes Lack Evidentiary Weight Statements such as “Policy A achieves 90 % success; Policy B achieves 92 % success” are ubiquitous in robotics papers and technical reports. However, these two numbers alone cannot support the conclusion that Policy B is stronger. Valid interpretation must account for roll‑out count, task composition, random seeds, paired‑group configuration and statistical power. The RoboLab v4 benchmark illustrates this concrete risk. Each policy runs only 10 episodes per task. Under this setup, when a policy reaches a 90 % success rate, its 95 % confidence interval spans approximately 19 percentage points . Even expanding to 100 roll‑outs, the interval width still sits near six percentage points. Authors explicitly classify 10‑episode runs as coarse‑grained indicators and warn that fine‑grained policy comparison remains untrustworthy. This warning generalizes across most high‑cost robot benchmarks: reported numbers may print with high numerical precision, yet real statistical