Skip to content
Türkçe

Google · Models · Safety and security

Google introduces Gemini 4 Argon, the first model in its Gemini 4 series

Published: 7 sourcesTürkçe

On September 30, Google introduced the first model in the Gemini 4 series: Gemini 4 Argon. It’s built for coding, enterprise knowledge work and cybersecurity.

It can write up to 1 million tokens in a single response, up from 64,000 in earlier models. That lets it work through long, multi-step problems in one go.

Google shared performance test results for its new flagship model, Gemini 4 Argon.

DeepSWE v1.1 tests AI coding agents on 113 long, multi-step software tasks written from scratch in real open-source projects.

Argon 77.9% · Opus 5.5 74.2% · Astra 74.1%

AutomationBench is a test built by Zapier.

It checks whether AI agents can complete real business workflows in apps like Gmail, Slack and Salesforce while following company rules.

Argon 51.3% · GPT-6 Astra 41.4% · Opus 5.5 42.5%

CWE-bench, from Collinear AI, hands an agent a real codebase and expects it to find and fix a security flaw nobody has pointed out. Argon 68% · Astra 68%

By Google’s count, it leads outright on 12 of the 18 tests it published. Astra or Opus 5.5 still win some coding and science tests. These are Google’s own numbers.

The Vals Index is Vals AI’s independent suite of work tasks in finance, law, tax and coding, weighted by each sector’s share of the US economy. Independent testers paint a strong picture too. It ranks first of 41 models on the Vals Index. Artificial Analysis puts it level with GPT-6 Astra at about 60% of the cost per task at launch pricing. Its hallucination rate is 15%, the lowest among models at that level.

The boldest decision is on cybersecurity. Google is giving Argon to trusted defenders and its own teams with cyber guardrails switched off. The goal is to put its full vulnerability-finding ability in the hands of people patching systems. It scores 85.8% on Google’s own test for finding flaws in real code, up from 71% for its previous cyber model.

Google’s cyber tests

  • Finding flaws: Gemini 4 Argon 85.8%, Gemini 3.8 Flash Cyber 71.0%
  • Wiz pen test: Gemini 4 Argon 70.9%, Gemini 3.8 Flash Cyber 58.2%

According to Bloomberg, some Google staff aren’t convinced. In their view, it does well on tests, less well on real work. Google disputes the report. For now only Fairwind defenders and Google’s own teams can use it, with Ultra subscribers and API customers next. The introductory price is $2 input and $10 output per million tokens.

Who gets it, when, and for how much

  • Now → Fairwind defenders and Google teams
  • Next → Google AI Ultra subscribers and API customers
  • Launch price, per million tokens → $2 input, $10 output

Sources

  1. Google, “Gemini 4 Argon: our next era of frontier intelligence”, (blog.google)
  2. Artificial Analysis, “Gemini 4 Argon: Google is back as one of the top three labs in intelligence achieved”, (artificialanalysis.ai)
  3. Vals AI, “Vals Index Leaderboard and Methodology” (vals.ai)
  4. Vals AI, “Gemini 4 Argon Benchmarks, Cost and Capabilities | Vals AI”, (vals.ai)
  5. Bloomberg, “Google grapples with employee skepticism about new Gemini model”, (bloomberg.com)
  6. Zapier, “AutomationBench: AI Agent Benchmarks | Zapier” (zapier.com)
  7. Collinear AI, “CWE-bench: a cybersecurity benchmark by Collinear AI” (cwe-bench.com)

About this story

This story was posted on Instagram by @jarrus.tech on Oct. 1, 2026.

Spotted an error in this story? [email protected] · Instagram

This story in Turkish: Google, Gemini 4 serisinin ilk modeli Gemini 4 Argon’u tanıttı

Related