In C5R’s AI-run lab, the best model completed only 45% of the tasks
On September 24 San Francisco-based C5R unveiled Facility-0, a lab built in twelve weeks that combines biology, chemistry and materials science, along with SciUniverse, a benchmark that measures how well models can carry out physical experiments. Frontier models scored between 9.4% and 45.3%, with Anthropic’s Claude Fable 5.1 on top at 45.3%. The failures were recurring: trying to pipette frozen samples, reusing pipette tips across DNA-containing wells, vortexing open plates and misreading noise in spectroscopy data. The takeaway: models know the science but lack an intuition for the physical reality of a lab.
Sources
- C5R, post on X (@c5rcorp), (x.com)
- C5R, “Can frontier models carry out scientific work?” (c5r.net)
About this story
This story was posted on Instagram by @jarrus.tech on Sept. 29, 2026.
Spotted an error in this story? [email protected] · Instagram
Short link: thejarrus.com/en/h9-20
This story in Turkish: C5R’ın yapay zekanın yönettiği laboratuvarında en iyi model görevlerin yalnızca %45’ini başardı
On the same topic
DeepMind-born drug discovery company Isomorphic Labs in funding talks at a $40-50 billion valuation
Isomorphic Labs, the DeepMind-born AI drug discovery company, is in early talks to raise money at a valuation of at least $40 billion, Bloomberg reported.
Anthropic releases Claude Haiku 5.5 at 90% lower prices
Claude Haiku 5.5, released Oct. 7, costs 90% less than Haiku 4.5 for prompts up to 100,000 tokens. It leads on benchmarks but uses more tokens per task.
OpenAI is rolling out GPT-6 to everyone in ChatGPT
OpenAI began rolling out GPT-6 to everyone in ChatGPT on October 7. Its new Intelligent UI builds answers from text, visuals and interactive elements.