Skip to content
Türkçe

Robotics · Research · OpenAI

Robot tests: even the best model completed only 19% of 84 robot tasks, and 63 tasks were solved by no model

Published: 1 sourceTürkçe

Thirty-three researchers from 11 institutions, including HKUST, Stanford and MIT, released RobotWorld on October 7. Its 84 simulated tasks cover manipulation, locomotion, driving and drone control. OpenAI’s GPT-6 Astra scored best, completing just 16 of 84 tasks, or 19 percent. Anthropic’s Claude Opus 5.5 came second with 13 (15.5 percent), Kimi K3 solved two, and DeepSeek V4.1 Flash and Gemini 3.8 Flash solved one each. Even combining all five models, 63 tasks went unsolved. The authors say agents often reach the right pose yet lose the object, correct mistakes too late or believe unfinished tasks are done.

Sources

  1. arXiv, “RobotWorld: Benchmarking Multimodal Agents for Robot Use Across Diverse Tasks and Embodiments”, (arxiv.org)

About this story

This story was posted on Instagram by @jarrus.tech on Oct. 10, 2026.

Spotted an error in this story? [email protected] · Instagram

This story in Turkish: Robot testleri: en iyi model bile 84 robot görevinin sadece %19’unu başardı, 63 görev hiçbir model tarafından çözülemedi

On the same topic