Google Strikes Back! Gemini 4 Argon Delivers Performance on Par With Astra at Only 60% of the Cost.
ZheDongXi October 1 news, in the early hours of today, Google released its new flagship model Gemini 4 Argon, which is focused on complex, long-cycle tasks such as enterprise knowledge work including software engineering, law and finance, as well as cybersecurity defense.
Gemini 4 Argon outperforms models including GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5 in DeepSWE v1.1 test that measures long-range software engineering capabilities, Vals Index that evaluates enterprise knowledge work covering finance, programming, law and taxation, AutomationBench that tests enterprise end-to-end automated execution capabilities, and LVBench that examines long video understanding capabilities; in CWE-bench v1 that assesses vulnerability remediation capabilities, Argon ranks first tied with GPT-6 Astra with a score of 68%.
In the comparison of model cost and score in Text Arena, Gemini 4 Argon High has entered the cost-performance frontier with a mixed price of 1525 points and $8 per million tokens.
At present, in the Artificial Analysis Intelligence Index, Gemini 4 Argon scores 53 points, only behind Claude Opus 5.5 and Claude Sonnet 5.5, and its score is tied with Claude Fable 5.1 and GPT-6 Astra.
▲ Gemini 4 Argon's evaluation on Artificial Analysis (Source: Artificial Analysis)
However, Bloomberg, citing people familiar with the matter, said that some Google employees believe that although Gemini 4 performs well in industry benchmark tests, it does not always reach the same level in actual work, especially it still struggles with some programming tasks, which is in contrast to the test scores released by Google on the same day.
Gemini 4 Argon is now open to a group of screened cybersecurity defense teams through the "Fairwind Program". Google plans to collect early test feedback first, continue to improve safety guardrails, and then gradually open it to developers, enterprises and consumers. The first batch of users will include paid API customers and Google AI Ultra subscribers.
The initial price of Argon API is $2 per million input tokens (about 13.4 RMB), $10 per million output tokens (about 67 RMB), and the price of cached input tokens is 95% lower than that of input tokens, that is, about $0.1 per million tokens (about 0.67 RMB).
According to Artificial Analysis, at the discounted pricing, the cost per task of Gemini 4 Argon is $1.99, which is 40% lower than GPT-6 Astra (max).
▲ Comparison of the cost per task of Gemini 4 Argon with other models (Source: Artificial Analysis)
Sundar Pichai, CEO of Google, said that there has been a lot of discussion about the next-generation model, so he hopes to showcase Gemini 4 Argon as early as possible. Google's internal teams have widely used it for programming, quantum computing and other work, and the feedback is good.
▲ Pichai posted an official announcement of Gemini 4 Argon (Source: X)
One X user believes that Gemini 4 is extremely sensitive to prompts, and the output content and style are more affected by prompts than other models. The default style is not good, but adding one more prompt can easily improve it. He also demonstrated the effect of recreating the levels of *Lego Star Wars* with different prompts.
▲ Developers' reviews of Gemini 4 Argon (Source: X)
Another developer believes that in terms of generating a 3D scene with rich details and excellent aesthetics in one go, Gemini 4 Argon has not yet reached the level of Opus 5. Gemini 4 Argon has slightly rough effects in picture details.
▲ Developers compare the generation effects of Gemini 4 Argon (top) and Claude Opus 5 (bottom) (Source: X)
01 .
Output upper limit increased to 1 million tokens,
Programming capabilities and enterprise task tests stand out
Gemini 4 Argon raises the model's output token upper limit from the previous 64,000 tokens to 1 million tokens, making it one of the models with the highest output capacity in the industry at present. The larger output space allows the model to think deeply in a single task and generate hundreds of thousands or even millions of tokens for processing large code bases, complex research and multi-step enterprise processes.
In terms of programming, Gemini 4 Argon achieved a score of 77.9% in the DeepSWE v1.1 test that measures the real-world long-term software engineering capabilities, setting a new record for the best score of this test.
In the Vals Index covering work including finance, programming, law and taxation, Argon ranks first. Argon also takes the leading position in the Vals Finance Agent v2 test (multi-step financial research) and Harvey Legal Agent test (legal research and drafting).
In the AutomationBench used by Zapier to evaluate enterprise end-to-end automated execution capabilities, Argon scored 51.3%, ranking first, significantly leading models including GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5.
Argon's multimodal capabilities are also used for long video and professional chart analysis. In the LVBench that evaluates long video understanding capabilities, the model scored 91.7%, reaching the optimal level of existing models.
02 .
Deeply integrated into Google's internal workflow,
Complete quantum algorithm optimization and large-scale code migration
Inside Google, thousands of employees have used Gemini 4 Argon for daily debugging, algorithm design, research and large-scale code migration and other work.
Google cited an example that Argon once helped quantum computing researchers optimize the spatiotemporal resources of quantum algorithms, improving a public baseline result by 40% within minutes. In addition, a group of Argon Agents analyzed the performance analysis telemetry data covering Google's entire server cluster, and independently identified and implemented a data center memory optimization solution. After the relevant solution is deployed, more than 300TiB of memory can be released, and the total expected savings will reach 500TiB to 1PiB.
Argon Agent also participated in the migration of Google's internal C/C++ code to Rust, covering tens of thousands of lines of core library code and more than 800,000 lines of the Fuchsia Zircon kernel. Considering the importance of the relevant systems, Google said that all large-scale rewrites must go through automated and manual audits, simulation tests and code reviews to confirm safety before they can be deployed to the production environment.
Taking the open-source video decoder libgav1 as an example, Argon Agent rewrote about 32,000 lines of SIMD code through multiple rounds of performance analysis and experiments, enabling the compiler to complete vectorization automatically. While the migrated Rust version keeps the video output completely consistent, its running speed is 2.7 times that of the original Rust version, and its performance is further close to the optimized C++ version.
03 .
Strengthen cybersecurity capabilities,
Capable of independently discovering and remediating vulnerabilities
Google also conducted special training on Gemini 4 Argon for cybersecurity defense capabilities. Argon can independently discover, verify and remediate software vulnerabilities. For trusted cybersecurity defense teams and Google internal teams, Google will provide an Argon version without cybersecurity guardrails so that they can give full play to their full vulnerability analysis and defense capabilities.
Cybersecurity company Wiz has used Argon through the "Scan for Good" project to discover and remediate high-risk exposures for critical public infrastructure for free. Google said that in an early test, Argon discovered a serious vulnerability affecting medical software used by hospitals around the world. This vulnerability could lead to the exposure of sensitive personal information, and previous state-of-the-art models did not identify this risk.
In CWE-bench v1 that assesses vulnerability remediation capabilities, Argon ranks first with a score of 68%, on a par with Grok 4.7 and GPT-6 Astra.
Google's internal comprehensive vulnerability benchmark test shows that Argon can discover multiple types of security exposures in complex code bases covering 20 programming languages. In Wiz's black-box penetration test, Argon also outperformed the previous Gemini 3.8 Flash Cyber in identifying the attack surface of real websites, discovering vulnerabilities and generating proof of concepts.
04 .
Expand safety testing,
Focus on preventing model abuse and prompt injection
Before officially expanding access, Google will continue to strengthen four types of cutting-edge safety capabilities.
In terms of abuse prevention, Argon will reject harmful requests that may be used for cyberattacks as well as chemical, biological, radiological and nuclear weapon attacks, while retaining the legitimate dual-use scientific research capabilities. Google is also strengthening the monitoring of the internal activation state of the model to identify potential abusive behaviors.
Argon has also made progress in defending against indirect prompt injection. Such attacks will hijack the model's behavior through malicious instructions or external context. After automated red team testing and adversarial training, Argon achieved leading results in Gray Swan's indirect prompt injection test.