HomeArticle

Just now, Google released Gemini 4 Argon, which has a maximum single output limit of 1 million tokens.

机器之心2026-10-01 09:00
The maximum output token limit has been increased from the previous 64,000 to 1 million.

Just now, Google announced the launch of its new-generation cutting-edge model Gemini 4 Argon.

Google positions Argon as a model oriented to complex, long-cycle workflows, with a focus on enterprise knowledge work such as software engineering, law and finance, as well as cybersecurity defense.

Different from the previous approach of opening directly to the public, Argon will adopt a phased release strategy. The model will be first provided to trusted cybersecurity defenders through the Fairwind Program.

In terms of pricing, the initial pricing of the Argon API is $2 per million input Tokens, $10 per million output Tokens, and the price of cached input Tokens is 95% lower than the standard input price.

The single output upper limit is expanded to 1 million Tokens

The most notable change of Gemini 4 Argon is that the output Token upper limit has been increased from the previous 64,000 to 1 million.

The 1 million Tokens here refers to the upper limit of single output, which is two different indicators from the context window. A larger output space means that the model can continuously reason in one task trajectory and generate results of hundreds of thousands or even nearly a million Tokens. This capability is mainly designed for large-scale code migration, in-depth research and multi-step enterprise processes.

Some netizens pointed out that the output upper limit of most cutting-edge models is about 128,000 Tokens. In contrast, Argon's capability of generating 1 million Tokens in a single run is quite rare.

Google stated that Argon scored 77.9% in the DeepSWE v1.1 software engineering benchmark, reaching the current highest level. This benchmark mainly measures the model's ability to handle real, long-term software engineering tasks.

In terms of enterprise knowledge work, Argon ranks first in the Vals Index. This index covers finance, programming, law, taxation and other fields, and is weighted according to the contribution of each industry to the US gross domestic product.

Argon also achieved leading performance in benchmarks such as Vals Finance Agent v2, Harvey Legal Agent Benchmark and Zapier's AutomationBench, with a score of 51.3% in AutomationBench.

Argon also emphasizes visual understanding capabilities. Google stated that the model can analyze professional charts, understand long videos, and take actions based on multiple documents. In the long video understanding benchmark LVBench, Argon scored 91.7%.

Has been integrated into Google's internal engineering processes

Google said that thousands of employees have used Argon in their daily work, with application scenarios including code debugging, algorithm design and large codebase migration.

In the field of quantum computing, Argon helps researchers optimize the space-time resources in quantum algorithms, that is, the product of the number of qubits and the number of gate operations. Google stated that in one test, Argon improved the published benchmark result by 40% in just a few minutes.

In the data center memory optimization project, the Argon agent analyzed the performance monitoring data of the entire network, and automatically identified and implemented optimization schemes. After the relevant scheme was launched, more than 300 TiB of memory has been released, and the total expected savings will reach 500 TiB to 1 PiB.

Argon also participated in the migration of Google's internal C/C++ codebase to Rust, covering core libraries such as re2 and libgav1, and extending to the Fuchsia Zircon kernel with more than 800,000 lines of code. Due to the involvement of critical infrastructure, these migrations still need to go through automated auditing, manual review and simulation testing.

In the open source video decoder libgav1 project, the Argon agent rewrote about 32,000 lines of SIMD code through multiple rounds of performance testing and compiler analysis. The new Rust version runs 2.7 times faster than the original Rust version while maintaining consistent video output.

Cybersecurity capabilities go hand in hand with real-world tests

Cybersecurity is the key field for Argon's first batch of implementation. Google stated that the model can independently discover, verify and fix critical software vulnerabilities.

Cybersecurity company Wiz has used Argon through the "Scan for Good" program to freely check high-risk exposure surfaces for public infrastructure. Google said that Argon once discovered a serious vulnerability affecting medical software used by hospitals around the world, which could expose sensitive personal information, and this risk was not identified by previous cutting-edge models.

In the CWE-bench v1 vulnerability fixing benchmark, Argon tied for first place with a score of 68%.

Google also stated that in internal vulnerability tests covering 20 programming languages, as well as black-box penetration tests without providing source code, Argon outperformed Gemini 3.8 Flash Cyber.

As model capabilities improve, Google is also strengthening security mechanisms simultaneously. Argon will restrict cyberattacks and abuse requests related to chemical, biological, radiological and nuclear materials, and identify potential risks through internal activation monitoring, automated red team testing and manual evaluation.

For indirect prompt injection attacks, Google improves the model's protection capabilities through adversarial training and automated testing. The company has also deployed a monitoring mechanism for model mismatch, tracks the model's reasoning process and action trajectory, suspends task execution when necessary, and isolates high-risk training and benchmark environments.

However, there is still a gap between benchmark scores and real-world work experience.

Bloomberg reported that Google internally still has doubts about the performance of Gemini 4 in key tasks such as coding. People familiar with the matter said that although the model performs well in industry benchmarks, some coding tasks are still difficult to complete stably when employees actually use it.

Whether Argon can deliver on its capabilities of long-cycle reasoning and complex task execution requires more real-scenario verification.

Google stated that the company is participating in the voluntary pre-release process for models promoted by the US government, and the first batch of testers will help Google evaluate the model's performance and further improve security protection measures.

After completing the early tests, Argon will be gradually opened to developers, enterprises and consumers. The first batch of public users will include paid API customers and Google AI Ultra subscribers.

Reference links:

https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/

https://www.bloomberg.com/news/articles/2026-09-30/google-grapples-with-employee-skepticism-about-new-gemini-model?srnd=phx-ai

© THE END 

For reprint authorization, please contact this official account

For contribution or coverage requests: liyazhou@jiqizhixin.com

This article is from WeChat Official Account "Almost Human" (ID: almosthuman2014), Author: AI-focused, published with authorization from 36Kr.