Research · Safety and security
New benchmark: evaluation criteria get exploited once they become the reward signal
When model-generated evaluation criteria were used as a reward signal, they could be exploited on 8% to 26% of tasks.
Sources
- arXiv, “ImpossibleRubrics: Stress-Testing Generated Rubrics as Reward Signals”, (arxiv.org)
About this story
This story was posted on Instagram by @jarrus.tech on Sept. 18, 2026.
Spotted an error in this story? [email protected] · Instagram
Short link: thejarrus.com/en/impossible-rubrics
This story in Turkish: Yeni kıyaslama: değerlendirme kriterleri ödül sinyali olunca istismar ediliyor
On the same topic
Robot tests: even the best model completed only 19% of 84 robot tasks, and 63 tasks were solved by no model
In RobotWorld, a benchmark released October 7, the top model, GPT-6 Astra, completed only 16 of 84 simulated robot tasks; no model solved 63 of the tasks.
OpenAI’s image model posts a near-perfect score on text-dense images
On UltraText Bench, a dense-text test led by Westlake University, OpenAI’s GPT Image 2 ranked first of 24 configurations with 99.35 out of 100.
AWS adds Chinese lab Z.ai’s GLM 5.3 model to Amazon Bedrock
AWS added GLM 5.3, a 753-billion-parameter model from Beijing-based Z.ai (formerly Zhipu AI), to Amazon Bedrock for eligible enterprise customers on October 5.