Robotic manipulation in lab simulation: top model scores 89.4%. Robotic task completion in real household environments: top model scores 12%. Same generation of models. Different benchmarks. Order-of-magnitude gap.
The underlying benchmarks:
RLBench is a software-based simulation environment that tests robotic policies on 18 standard manipulation tasks with 100 demonstrations per task: picking objects, opening drawers, placing items, basic assembly. The benchmark is run in simulation. EquAct leads the leaderboard at 89.4% average success rate, with SAM2Act at 86.8%.
The 2025 BEHAVIOR Challenge tests robotic policies on real household tasks: environments with clutter, unfamiliar object layouts, and tasks that compound multiple manipulation primitives (clean the kitchen, set the table, organise items by category). The top model achieves 12% on a held-out test set evaluation, with the leading-model Q-score metric reflecting the same gap.
The two benchmarks are not measuring the same thing. RLBench measures whether a policy can execute a manipulation primitive when the environment matches the training distribution. BEHAVIOR-1K measures whether a policy can complete an end-to-end household task when the environment includes the long tail of real-world variation that training data did not specifically cover.
The gap between the two benchmarks is the gap between "robotic manipulation that works in conditions resembling the training environment" and "robotic manipulation that works in conditions resembling the deployment environment." For enterprise procurement decisions about robotics, this gap is what determines whether a robotics deployment is economically viable.
For procurement evaluation, three observations.
The first: vendor demonstrations of robotic capability are predominantly performed in conditions that resemble RLBench, not BEHAVIOR-1K. A vendor demo showing a robot completing 90% of pick-and-place tasks in a controlled lab environment is consistent with the data. The same vendor's deployment in a real-world environment may achieve 12–15% task success, and that is also consistent with the data. The demo and the deployment performance are measuring different things, and procurement frameworks that treat the demo as evidence of deployment capability are operating under a misunderstanding the data clearly contradicts.
The second: the gap is not closing as fast as the language-model gap is closing on equivalent benchmark dimensions. Language model benchmarks at the "novel reasoning" frontier (HLE, FrontierMath Tier 4) are climbing at multiple percentage points per year. The robotics gap between lab and deployment is wider, and the rate at which it is closing is not currently calibrated by available data. The procurement implication: robotics capability investments should be planned against a longer time horizon than language model capability investments, even when the underlying model architectures share approaches.
The third: the gap is not uniform across robotics task categories. Pick-and-place tasks with well-defined objects show a smaller lab-to-deployment gap than tasks involving compound manipulation, clutter handling, or sequential reasoning across multiple steps. Procurement decisions about robotics deployment should be category-specific. Which task categories work in deployment now, which work in controlled environments but not in deployment, and which do not yet work in either.
The data observation: there is an order-of-magnitude gap between robotic capability in simulation and robotic capability in real household environments. This gap is the central structural feature of the 2026 robotics landscape. It is not a temporary state that closes within a year. It is the lens through which all robotics procurement decisions need to be made.
For procurement frameworks evaluating robotics vendors, this surfaces a specific question structure. The vendor demonstration is necessary evidence but not sufficient evidence. The procurement question is "can this vendor's robotic capability survive the lab-to-deployment gap on our specific deployment environment?" That question is not answered by RLBench scores, BEHAVIOR-1K scores, or vendor lab demos. It is answered by deployment trials in environments that resemble the procurement context: same lighting, same clutter, same edge cases.
Three concrete actions for procurement teams evaluating robotics.
Run deployment-environment trials before commitment. The lab-to-deployment gap means that vendor selection based on benchmark or demo performance has a high probability of producing under-performing deployments. Pilot deployments in conditions that resemble the procurement context surface failure modes that benchmarks do not.
Plan for a longer evaluation cycle than language-model procurement. Robotics evaluation takes longer than language model evaluation, partly because deployment trials require physical setup time and partly because the failure modes only appear in extended operation. Procurement timelines built on language-model evaluation cycle times will under-allocate evaluation budget for robotics.
Distinguish "robotic capability for controlled environments" from "robotic capability for real environments" in procurement specifications. The first is what most current robotic deployments achieve well. The second is where most enterprise deployment plans place the requirement, and where the lab-to-deployment gap means most plans will not achieve their stated goals on the stated timeline.
The data is clear about what is currently possible in robotics, where the structural gaps are, and what types of procurement decisions are likely to produce successful deployments versus failed ones. The data is in. Whether your procurement frameworks weight lab demos and benchmark scores against real-environment trials at the right ratio is the choice in front of you.
Sources
- Primary: Stanford AI Index 2026, Chapter 2 (Technical Performance) 2.7 — hai.stanford.edu/ai-index/2026
- Lab simulation data: AI Index 2025 own analysis (RLBench, re-published 2026) — 18-task robotic manipulation suite
- Household task data: BEHAVIOR Challenge Leaderboard, 2025 — end-to-end household task evaluation
- Responsible robotics reference: Zhang et al. (2025) — ResponsibleRobotBench evaluation
Discussion