Users claimed Opus 5.5’s performance has dropped; an independent measurement showed it within the normal range of variation
This story was updated after publication (Oct. 9, 2026). Details
Users on X and Reddit claim that Claude Opus 5.5, which Anthropic released on September 22, has gotten worse since launch. On BridgeMind’s NerfBench, which retests models regularly, Opus 5.5 fell to 94.2% of its launch score on October 2 and stood at 96.5% in the latest test on October 4. BridgeMind treats 90% to 110% as normal variance, so this is not evidence of a downgrade. Only five tests have been run so far, though. Another possible explanation: according to Anthropic’s documentation, Opus 5.5 defaults to medium effort, while Opus 5 defaulted to high.
Sources
- Anthropic, “Introducing Claude Opus 5.5”, (anthropic.com)
- BridgeBench (BridgeMind), “Claude Opus 5.5 — Nerf Bench” (bridgebench.ai)
- BridgeMind (X), post on X: Opus 5.5 drops to 94.2% on NerfBench, (x.com)
- Anthropic (Claude Platform Docs), “Claude Opus 5.5 migration guide” (platform.claude.com)
Corrections and updates
- After this story was published, NerfBench added a sixth measurement on October 9: Opus 5.5 stood at 95.7% of its launch score. That is still within the 90% to 110% range BridgeMind treats as normal variance, so the conclusion is unchanged.
About this story
This story was posted on Instagram by @jarrus.tech on Oct. 5, 2026.
Spotted an error in this story? [email protected] · Instagram
Short link: thejarrus.com/en/opus-5-5-nerfbench
This story in Turkish: Kullanıcılar Opus 5.5’in performansının düştüğünü iddia etti, bağımsız ölçüm normal dalgalanma aralığında olduğunu gösterdi
On the same topic
Robot tests: even the best model completed only 19% of 84 robot tasks, and 63 tasks were solved by no model
In RobotWorld, a benchmark released October 7, the top model, GPT-6 Astra, completed only 16 of 84 simulated robot tasks; no model solved 63 of the tasks.
AWS adds Chinese lab Z.ai’s GLM 5.3 model to Amazon Bedrock
AWS added GLM 5.3, a 753-billion-parameter model from Beijing-based Z.ai (formerly Zhipu AI), to Amazon Bedrock for eligible enterprise customers on October 5.
Microsoft releases three new voice models, one transcribing 60 languages in real time
On October 1, Microsoft AI announced MAI-Transcribe-2-Streaming, which transcribes 60 languages in real time, and the 23-language MAI-Voice-2.1 and Flash.