Fast Facts
A new benchmark from Dalian University of Technology tested 12 multimodal AI models on robot-arm tasks. Every model could locate objects at ~100% accuracy. Every model understood what needed to be done at ~99%. Yet the best model completed only 53.93% of whole tasks. Industrial robotics training sims teach perception and reasoning effectively, but the translation from “understanding” to “doing” remains unsolved. For procurement teams, this isn’t an AI research problem — it’s a deployment timeline and budget problem.
Industrial robotics training sims just produced a number that should reshape how buyers evaluate vendor claims. A research team led by Dalian University of Technology released VA-Bench, a benchmark testing whether general-purpose multimodal models can turn what they see into successful robot-arm actions. Across 12 model configurations, the top performer completed 53.93% of tasks. But the same models scored 100% on locating targets and 99.6% on understanding what manipulation was required.
The gap between understanding and doing is now measurable. It is also the gap most procurement teams have been paying for without realizing it.
What the Benchmark Actually Measures
VA-Bench is not a perception test. It is an action test. Before each trial, a model watches a successful demonstration video but receives no object coordinates, expert trajectories, or precise action parameters. It must issue concrete commands: translations in millimeters along each axis, rotation angles, and when to open or close the gripper. A fixed controller executes only what the model specifies.
The benchmark covers 14 manipulation task types across 280 physics-verified scenes, plus held-out geometry and layout variants. Each model condition was run three times.
The results are consistent across model families:
| Metric | Best Model (Qwen3.8-max) | Runner-up (Opus-5) | Third (GPT-5.6-sol) |
|---|---|---|---|
| Whole-task success | 53.93% | 52.86% | 51.55% |
| Subtask completion | 65.9–68.6% | — | — |
| Error detection | 73.6% | — | — |
| Online error correction | 46.7% | — | — |
| Single-arm success | 65.61% | — | — |
| Dual-arm success | 11.11% | — | — |
The authors note the top three sit within 2.38 points and do not claim a reliable ordering.
Why the Error Correction Gap Matters More Than the Success Rate
The most consequential finding is not the 53.93% whole-task success rate. It is the gap between detecting an error and correcting it.
Qwen3.8-max spotted errors in 73.6% of cases. It corrected them online in only 46.7% of cases.
Industrial robotics training sims are good at teaching models to recognize when something is wrong. They are significantly worse at teaching models to recover in real time. In a factory, recognizing a failure and correcting it are the same task. Recognizing it without correcting it is just a logged incident.
This gap also explains why dual-arm tasks fail so catastrophically. Single-arm success reached 65.61%. Dual-arm success collapsed to 11.11%. Coordinating two arms requires assigning objects to each arm and managing their relative positions — a spatial reasoning task that current models handle poorly.
The Active Viewing Advantage
One finding has direct procurement implications: models that chose their own camera viewpoints succeeded 57.50% of the time. When given five fixed views — more information, less control — success dropped to 27.86%. All four models in that test declined without active observation.
The lesson is straightforward. Industrial robotics training sims that give models control over what they look at outperform sims that give models more cameras but no agency. The architecture of the training environment matters more than the volume of sensor data.
What This Means for Procurement Timelines
The DALIAN benchmark does not mean industrial robots are unusable. It means the use cases where they work are narrower than vendor marketing implies.
⚠ Fiction—composite scenario, not a real event: A plant manager evaluates a robotic arm for dual-arm assembly. The vendor’s simulation demo shows the robot completing the task flawlessly. The procurement team approves the pilot. On the production floor, the robot succeeds on single-arm pick-and-place but fails on dual-arm coordination more than 80% of the time. The vendor’s simulation used pre-positioned objects and a single arm. The factory floor requires two arms working together on variable incoming parts. The benchmark would have predicted the failure. The demo hid it.
Global Implications
For manufacturers in emerging markets, the industrial robotics training sims gap creates a specific risk. A facility that cannot afford to run physical pilots before committing capital is more reliant on vendor simulation claims. The DALIAN benchmark provides a counterweight: ask vendors what their dual-arm success rate is on held-out tasks, not their best-case demo.
The benchmark also confirms that simulation is not a substitute for deployment data. Models that performed well in simulation degraded significantly when object shapes, containers, or layouts changed — one reference model dropped from 80% to 47.86% success when only geometry varied.
What to Ask Before the Next Robot Pilot
Three questions separate a defensible deployment from a simulation-fueled gamble:
| Question | Why It Matters |
|---|---|
| What is your dual-arm success rate on held-out tasks? | The benchmark shows a 54-point gap between single-arm and dual-arm performance |
| How does your model perform when object geometry changes? | One reference model lost 32 percentage points when only shapes varied |
| What is your error correction rate, not just your error detection rate? | Detecting a failure without correcting it is not deployment readiness |
Industrial robotics training sims have advanced far enough to teach models what to see. They have not yet advanced far enough to teach models what to do when what they see changes. That gap is the difference between a pilot and a production line.
💡 CreedTec Analyst’s Note — Daniel Ikechukwu
Strategic Impact: The DALIAN VA-Bench quantifies the gap between perception and action in simulation-trained robot policies. Buyers evaluating robotics vendors should treat simulation success rates as an upper bound, not a prediction. The dual-arm failure rate is the single most important number to demand before any pilot approval.
Stop: Accepting vendor simulation demos as evidence of production readiness without held-out task data.
Start: Requiring dual-arm success rates, error correction rates, and geometry-variation performance as conditions of pilot approval.
Watch: Whether the next generation of industrial robotics training sims closes the error correction gap or continues to improve perception while leaving action reliability behind.
ROI Outlook: A pilot that succeeds on single-arm tasks and fails on dual-arm coordination wastes the pilot budget and delays deployment by a full learning cycle. The benchmark gives procurement a framework to avoid that cost before it is incurred.
Sources:
Further Reading:
- The 100x Robot Simulation Speed Claim Has an Asterisk
- Google’s Intrinsic Open-Sources Its Robot Simulation Stack
- Synthetic Robot Training Data: The 40% Threshold Undercutting a Billion-Dollar Race
- Brain Connectome Simulation Just Traded Crypto With 166,700 Neurons
- MIT SceneSmith Attacks the Cost Nobody Talks About in Robot Training
Subscribe to CreedTec’s weekly briefing—robotics training sims, simulation economics, and the procurement signals behind the benchmark headlines.


