Research · Safety and security
New benchmark: evaluation criteria get exploited once they become the reward signal
When model-generated evaluation criteria were used as a reward signal, they could be exploited on 8% to 26% of tasks.
Sources
- arXiv, “ImpossibleRubrics: Stress-Testing Generated Rubrics as Reward Signals”, (arxiv.org)
About this story
This story was posted on Instagram by @jarrus.tech on Sept. 18, 2026.
Spotted an error in this story? [email protected] · Instagram
Short link: thejarrus.com/en/h4-08
This story in Turkish: Yeni kıyaslama: değerlendirme kriterleri ödül sinyali olunca istismar ediliyor
On the same topic
DeepMind-born drug discovery company Isomorphic Labs in funding talks at a $40-50 billion valuation
Isomorphic Labs, the DeepMind-born AI drug discovery company, is in early talks to raise money at a valuation of at least $40 billion, Bloomberg reported.
OpenAI’s unreleased model wrote 722 math papers
On October 6, OpenAI published 722 math papers from an unreleased internal model on GitHub. It withdrew three on October 7, leaving 719 in the catalog.
AI was used to design blood cancer drug candidates; the company says its method finds candidates far faster than usual
Insilico Medicine’s Abu Dhabi team used its Chemistry42 AI platform to design PROTAC molecules that destroy BTK; the compounds are still at the research stage.