Safety and security · Anthropic · OpenAI
The AI Evaluator Forum’s AEF-1 standard has started appearing in independent assessments
This story was corrected after publication (Oct. 9, 2026). Details
The standard, from the AI Evaluator Forum founded in December 2025, sets out five pillars: sufficient access and resources, minimized conflicts of interest, analytic autonomy, transparent methods and results, and protection of sensitive information. It has already started appearing in independent assessments covering Anthropic, Google, Meta and OpenAI. The counter-reading arrived the same week: more than a hundred experts signed an open letter saying Anthropic and OpenAI need truly independent safety evaluators.
Sources
- AI Evaluator Forum, “AEF-1: Minimum Operating Conditions for Independent Third Party AI Evaluations”, (aievaluatorforum.org)
- AI Evaluator Forum, “Minimum Conditions for Embedding Evaluators”, open letter, (aievaluatorforum.org)
- International Business Times, “More Than 100 AI Experts Sign a Letter Saying that OpenAI & Anthropic Need Independent AI Safety Evaluators”, (ibtimes.com)
Corrections and updates
- The headline first posted on Instagram said xAI, OpenAI and Anthropic cosigned the AEF-1 standard. The AI Evaluator Forum’s page lists no company signatories, and no source shows the companies signing it; the headline has been corrected.
About this story
This story was posted on Instagram by @jarrus.tech on Sept. 21, 2026.
Spotted an error in this story? [email protected] · Instagram
Short link: thejarrus.com/en/aef-1-standard
This story in Turkish: AI Evaluator Forum’un AEF-1 standardı bağımsız değerlendirmelerde kullanılmaya başladı
On the same topic
Robot tests: even the best model completed only 19% of 84 robot tasks, and 63 tasks were solved by no model
In RobotWorld, a benchmark released October 7, the top model, GPT-6 Astra, completed only 16 of 84 simulated robot tasks; no model solved 63 of the tasks.
OpenAI’s image model posts a near-perfect score on text-dense images
On UltraText Bench, a dense-text test led by Westlake University, OpenAI’s GPT Image 2 ranked first of 24 configurations with 99.35 out of 100.
AWS adds Chinese lab Z.ai’s GLM 5.3 model to Amazon Bedrock
AWS added GLM 5.3, a 753-billion-parameter model from Beijing-based Z.ai (formerly Zhipu AI), to Amazon Bedrock for eligible enterprise customers on October 5.