Gemini 4 Argon cybersecurity capabilities
Google’s Gemini 4 Argon frontier model has debuted with published cybersecurity performance figures released alongside the model announcement. Argon scored 77.9% on the DeepSWE v1.1 test and can produce up to 1 million tokens in a single reply, enabling much longer and more detailed single-response outputs that consolidate extended content into one answer. The model leads on 12 of 18 cybersecurity benchmarks in the published results, indicating a leading position across most measured evaluations.
Gemini 4 Argon scored 77.9% on the DeepSWE v1.1 test, a result published with the model announcement. Comparable published scores on the same test include Claude Opus 5.5 at 74.2%, GPT-6 Astra at 74.1% and Claude Fable 5.1 at 67.4%, while an earlier Gemini 3.6 Flash version scored 49% in July. These figures are presented as point values on the DeepSWE v1.1 evaluation and are reported alongside the model release. The published DeepSWE comparisons list Argon above the other named models on that specific test measure.
Argon’s maximum single-reply token capacity increased from a previously reported 64,000 tokens to 1,000,000 tokens per reply in the published model details. With a token approximated as roughly three-quarters of a word, that change is expressed as about 48,000 words at the 64,000-token level versus about 750,000 words at the 1,000,000-token level. The published results also report that Argon leads on 12 of 18 benchmarks, records one tie and trails on five benchmarks across the evaluated set. These benchmark tallies and the token-capacity figures are included in the material released with Gemini 4 Argon.
Gray Swan Indirect Prompt Injection benchmark results list Gemini 4 Argon at 0.7%, with lower values indicating better performance. Claude Opus 5.5 and Claude Fable 5.1 are each listed at 1.0% on the same benchmark. GPT-6 Astra is listed at 8.5% for the Indirect Prompt Injection measure. Grok is listed at 4.6% on this benchmark. The benchmark entries present each model as a percentage score on the Indirect Prompt Injection test.
Kimi K3 is listed with scores of 51.8% and 52.7% in the Indirect Prompt Injection entries. The table of results shows Argon at 0.7% alongside the other specified percentages for the named models. No additional performance metrics are included in this section.
The Fairwind Program was launched on September 2 with more than 650 partners, including governments and critical infrastructure operators. Gemini 4 Argon is provided to vetted security teams through the Fairwind Program. The Fairwind Program ships without cyber guardrails.
Gemini 4 Argon participates in the U.S. government’s voluntary process for pre-release model access. An early version of Anthropic’s Claude Mythos helped find 271 vulnerabilities in Firefox. OpenAI operates a Trusted Access for Cyber program.
These items are reported in the material released with the models and related program announcements. The section lists the reported partner counts, access pathways and related program entries as presented. No additional interpretation is included here.
Gemini 4 Argon debuted with published cybersecurity benchmark results reporting a leading performance on multiple evaluations, including a top result on the DeepSWE v1.1 test. The model’s single-reply token capacity expanded substantially, enabling much longer single-response outputs, and published comparisons show it leading across most measured cybersecurity benchmarks. Argon is being provided to vetted security teams through a dedicated partner program and participates in the U.S. government’s voluntary pre-release model access process.


