In C5R’s AI-run lab, the best model completed only 45% of the tasks
On September 24 San Francisco-based C5R unveiled Facility-0, a lab built in twelve weeks that combines biology, chemistry and materials science, along with SciUniverse, a benchmark that measures how well models can carry out physical experiments. Frontier models scored between 9.4% and 45.3%, with Anthropic’s Claude Fable 5.1 on top at 45.3%. The failures were recurring: trying to pipette frozen samples, reusing pipette tips across DNA-containing wells, vortexing open plates and misreading noise in spectroscopy data. The takeaway: models know the science but lack an intuition for the physical reality of a lab.
Sources
- C5R, post on X (@c5rcorp), (x.com)
- C5R, “Can frontier models carry out scientific work?” (c5r.net)
About this story
This story was posted on Instagram by @jarrus.tech on Sept. 29, 2026.
Spotted an error in this story? [email protected] · Instagram
Short link: thejarrus.com/en/sciuniverse-benchmark
This story in Turkish: C5R’ın yapay zekanın yönettiği laboratuvarında en iyi model görevlerin yalnızca %45’ini başardı
On the same topic
Robot tests: even the best model completed only 19% of 84 robot tasks, and 63 tasks were solved by no model
In RobotWorld, a benchmark released October 7, the top model, GPT-6 Astra, completed only 16 of 84 simulated robot tasks; no model solved 63 of the tasks.
OpenAI’s image model posts a near-perfect score on text-dense images
On UltraText Bench, a dense-text test led by Westlake University, OpenAI’s GPT Image 2 ranked first of 24 configurations with 99.35 out of 100.
AWS adds Chinese lab Z.ai’s GLM 5.3 model to Amazon Bedrock
AWS added GLM 5.3, a 753-billion-parameter model from Beijing-based Z.ai (formerly Zhipu AI), to Amazon Bedrock for eligible enterprise customers on October 5.