Robot tests: even the best model completed only 19% of 84 robot tasks, and 63 tasks were solved by no model
Thirty-three researchers from 11 institutions, including HKUST, Stanford and MIT, released RobotWorld on October 7. Its 84 simulated tasks cover manipulation, locomotion, driving and drone control. OpenAI’s GPT-6 Astra scored best, completing just 16 of 84 tasks, or 19 percent. Anthropic’s Claude Opus 5.5 came second with 13 (15.5 percent), Kimi K3 solved two, and DeepSeek V4.1 Flash and Gemini 3.8 Flash solved one each. Even combining all five models, 63 tasks went unsolved. The authors say agents often reach the right pose yet lose the object, correct mistakes too late or believe unfinished tasks are done.
Sources
- arXiv, “RobotWorld: Benchmarking Multimodal Agents for Robot Use Across Diverse Tasks and Embodiments”, (arxiv.org)
About this story
This story was posted on Instagram by @jarrus.tech on Oct. 10, 2026.
Spotted an error in this story? [email protected] · Instagram
Short link: thejarrus.com/en/robotworld-test
This story in Turkish: Robot testleri: en iyi model bile 84 robot görevinin sadece %19’unu başardı, 63 görev hiçbir model tarafından çözülemedi
On the same topic
OpenAI’s image model posts a near-perfect score on text-dense images
On UltraText Bench, a dense-text test led by Westlake University, OpenAI’s GPT Image 2 ranked first of 24 configurations with 99.35 out of 100.
A flaw in ChatGPT’s Mac app that could expose private chats was found and patched
Patrick Wardle found CVE-2026-100754, a flaw in ChatGPT’s Mac app that could expose chat history; OpenAI fixed it on September 25 in version 26.924.20706.
Anthropic is running a lab that does biology experiments in the Bay Area
Anthropic runs a Bay Area biology lab where it aims to have Claude direct automated lab equipment under human supervision, with a focus on fundamental biology.