MOUNTAIN VIEW

Google announced Gemini 4 Argon on Sept. 30, offering a 1 million token output limit—a 15-fold jump from the 64,000 token ceiling of its Gemini 3 models—and positioning the frontier model squarely at enterprise reasoning workloads in software engineering, legal, finance, and cybersecurity.

Introductory pricing is set at $2 per million input tokens and $10 per million output tokens, half the cost of Anthropic's Claude Opus 5.5. Standard pricing will rise to $4 and $20 per million respectively, with cached inputs discounted 95 percent off the input rate. That pricing structure matters: cached inputs allow enterprises to amortize large context windows—legal documents, code repositories, policy databases—across repeated queries, lowering the marginal cost of complex reasoning.

The rollout is restricted. Gemini 4 Argon is available only to vetted Google Cloud customers, government agencies, and cybersecurity partners through Google's Fairwind Program, plus internal teams. General developers and consumers gain access later, with paid API users and Google AI Ultra subscribers prioritized. Google is also part of the U.S. government's voluntary pre-release access process for frontier models.

Koray Kavukcuoglu, SVP of Google DeepMind, framed Argon as purpose-built for software engineering, enterprise knowledge work, and cyber defense. The model arrives four weeks after Google shipped Gemini 3.8 Flash and Gemini 3.8 Flash Cyber on Sept. 2.

Google published 19 benchmark comparisons. Argon leads on 13 rows, ties one with OpenAI's GPT-6 Astra, and trails on five. In knowledge-work-specific tests, Argon swept: 19.6 percent on Harvey's Legal Agent Benchmark (more than triple the next model), 77.9 percent on DeepSWE v1.1, and 68.9 percent on the Vals Index. On Gray Swan's 2026 Indirect Prompt Injection test, Argon showed a 0.7 percent attack success rate, the lowest of 13 models tested.

Argon trails on FrontierSWE v2, Terminal-Bench 4.0, and OSWorld-2.0. All scores are Google-provided and have not been replicated by independent laboratories.

The competitive calculus hinges on context window economics. A 1 million token output window enables multi-stage reasoning, code generation, and iterative document drafting in a single query—reducing API calls and end-user latency. For enterprises running high-volume knowledge work, the math shifts: fewer round-trips offset higher per-token rates, and caching amortizes context costs. Whether Argon's benchmark leads translate to customer adoption depends on whether enterprises value the reasoning depth over Claude's entrenched position in legal and financial sectors.