
Gemini 4 Argon vs Claude Opus 5.5: The Benchmark Split Nobody Expected
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 223 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 125 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1148 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 48 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 103 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 212 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Gemini 4 Argon and Claude Opus 5.5 disagree about who is better, and the disagreement is not noise. On Artificial Analysis's Intelligence Index the two are five points apart in Opus 5.5's favour. On the Vals Index — the GDP-weighted economic-impact board — Gemini 4 Argon sits at number one and Claude Opus 5.5 sits at number three, behind both Argon and Claude Sonnet 5.5. Both boards measured the same two models in the same week. That is the whole article: which one you should pick depends on whether your work looks more like an exam question or a Tuesday at a law firm, and on whether you have API access to Argon at all, which as of today almost nobody does.
Google announced Gemini 4 Argon on 2026-09-30 in a DeepMind post credited to Koray Kavukcuoglu, SVP for Google DeepMind and Chief AI Architect. Anthropic released Claude Opus 5.5 eight days earlier, on 2026-09-22. One of those is a GA release you can call this afternoon. The other is a phased rollout to a named cohort of cyber defenders plus Google's own engineers, with a stated intent to reach "paid API customers and Google AI Ultra subscribers" as soon as guardrails are ready and no date attached to that.
The one-line answer, before the detail
If you need the best measured general reasoning today and you need it on a production key, Claude Opus 5.5 is available on OrcaRouter right now and Gemini 4 Argon is not on any public API. If you are building for finance, legal, tax or coding agents that get scored on economic output per dollar, Argon's numbers are better and its price is a fifth of Opus 5.5's per task on the one board that measures exactly that. If you are in cybersecurity defense for critical infrastructure, Argon has a program you can apply to and the answer might be yes.
Where each number actually comes from
Two index boards, two methodologies, and it is worth being precise about which is which because the two camps will quote past each other for months.
• Artificial Analysis Intelligence Index v4.3 — Claude Opus 5.5 at 57.6 and Gemini 4 Argon at 52.6, both read off Artificial Analysis on 2026-09-30. This is a composite of ten evaluations, deliberately task-general, and it is the board most people cite as "the" ranking.
• Vals Index, updated 2026-09-29 — Gemini 4 Argon first at 68.90% (±0.97), Claude Sonnet 5.5 second at 67.04% (±0.92), Claude Opus 5.5 third at 66.97% (±0.89). Vals weights finance, coding, legal and tax tasks by each sector's share of US GDP, so a model that is mediocre at puzzles but excellent at contract review and spreadsheet work scores well here.
A five-point AA gap is real and it is not small. A two-point Vals gap is also real, and with the confidence intervals printed on the leaderboard Argon's 68.90 ±0.97 against Opus 5.5's 66.97 ±0.89 is a separation of roughly 1.5 standard errors — better than a coin flip, short of overwhelming. Neither board is wrong. They are measuring different jobs.

One configuration caveat, and it cuts against reading the AA gap as a pure model comparison. Artificial Analysis charts Gemini 4 Argon at its high effort setting and Claude Opus 5.5 at "Adaptive Reasoning, Max Effort, Default Fallback". Those are not the same point on the two effort ladders, so the 57.6-versus-52.6 headline is not a like-for-like comparison at matched reasoning budgets. It is the number both pages publish, and it is the number to quote if you are being fair to Anthropic, but it is not a controlled experiment.
Cost per task is where the gap gets uncomfortable
Artificial Analysis also publishes an accounting of what it cost to run its index suite, and here the two models are not close.
• Cost per index task — Gemini 4 Argon $1.99 versus Claude Opus 5.5 $5.98.
• Cost to run the whole index suite — Gemini 4 Argon $2,407 versus Claude Opus 5.5 $8,708.
• List price per million tokens — Gemini 4 Argon $2 in / $10 out (Google's stated introductory rate, with cached input at 95% off, so $0.10) versus Claude Opus 5.5 $4 in / $20 out.
• Throughput — Gemini 4 Argon 48.9 output tokens per second, first token after 2.79 seconds; Claude Opus 5.5 92.2 output tokens per second but a first-token latency of 702 seconds on the same harness, which is the extended-thinking tax showing up in the open.
• Vals cost per run — Gemini 4 Argon $15.68 against Claude Opus 5.5 $32.14 on the same GDP-weighted task set.
There is a wrinkle worth flagging rather than smoothing over: Vals's leaderboard lists Argon's rates as $4 in / $20 out, while Google's announcement states $2 in / $10 out for the introductory period. The most likely reading is that Vals is showing a standard list rate and Google is showing a promotional one, but that is an inference and not a published statement from either side. If your budget depends on the number, ask Google.
What Argon claims and what has been checked

Google's post is dense with figures and every one of them is vendor-reported. None has been independently reproduced that I can find, and the independent boards above measure different things than the ones Google chose to headline. Stated as Google's own claims:
• DeepSWE v1.1 — 77.9%, described as a new state of the art on real-world long-horizon software engineering.
• LVBench long-video understanding — 91.7%, described as state of the art.
• AutomationBench (Zapier's benchmark for end-to-end business tasks) — rank #1 at 51.3%.
• CWE-bench v1 vulnerability remediation — tied for first at 68%, building on Gemini 3.8 Flash Cyber's result on the v0 board.
• Gray Swan Indirect Prompt Injection — described as leading, with no figure published in the post.
• Vals Finance Agent v2 and Harvey's Legal Agent Benchmark — described as leading, likewise without a printed number in the announcement.
The two claim families that do have independent support are the ones worth trusting most: Vals confirms the index position with a figure and an interval, and Artificial Analysis confirms a 52.6 index and the price. The internal-productivity anecdotes — a 40% improvement over a published baseline on a quantum subroutine, more than 300 TiB of memory freed across Google's fleet by Argon agents — are company-internal results with no external test, and should be read as such.
What Opus 5.5 is, if you have been living under a rock
Claude Opus 5.5 is Anthropic's flagship: 1M-token context, 128K maximum output, multimodal input across text, image and file, extended thinking, tool use and structured outputs. It succeeded Claude Opus 5 on 2026-09-22 and the headline of its launch was arguably the price cut — the tier came down rather than up, at $4/$20 with cached reads at $0.20. It is routable on OrcaRouter today at exactly that rate, provider list price passed through with zero markup, and a vendor price change lands on our side the same day it lands on theirs.
On the task-level scores that both models have, the picture is a genuine seesaw rather than a sweep:
• Terminal-Bench 4.0 — Claude Opus 5.5 59.6% versus Gemini 4 Argon 57.1%.
• SciCode — Claude Opus 5.5 66.9% versus Gemini 4 Argon 61.8%.
• Humanity's Last Exam — Claude Opus 5.5 61.4% versus Gemini 4 Argon 57.1%.
• Long-context reasoning — Claude Opus 5.5 84.7% versus Gemini 4 Argon 79.7%.
• CritPt — Claude Opus 5.5 31.7% versus Gemini 4 Argon 27.1%.
• Omniscience — Claude Opus 5.5 46.4 versus Gemini 4 Argon 42.4.
• GDPval — Claude Opus 5.5 1,846 versus Gemini 4 Argon 1,611.
• AutomationBench-AA — Gemini 4 Argon 77.5% versus Claude Opus 5.5 69.5%. This is the one head-to-head task score where Argon clearly wins, and it is also the task family closest to what the Vals Index rewards.
So the pattern is consistent: Opus 5.5 is ahead on the academic-flavoured reasoning ladder, and Argon is ahead on end-to-end task execution. That is not a contradiction between the two boards any more. It is the same finding expressed twice.
The capability nobody is comparing on
Gemini 4 Argon's output ceiling is 1,000,000 tokens in a single trajectory, up from the previous generation's 64K. Claude Opus 5.5 caps output at 128K. Both hold a 1M-token context window, but context and output are different budgets, and this is the largest single specification jump in the announcement.
Whether that matters depends entirely on what you run. A 128K output ceiling is generous for chat, code review, structured extraction and agent steps. It is a hard wall for generating a whole migration patch set, a full audit report, or a long-form document in one pass — the workflows Google chose to illustrate the launch with. If your pipeline chains twenty calls to stay under a cap, Argon's ceiling will save you latency and orchestration code. If it does not, you are paying for headroom you will not use.
Availability is the deciding variable, not quality
This is the part of the comparison that has no benchmark attached and matters more than all of them.
Gemini 4 Argon is not generally available. Google's post says it is rolling out to trusted cyber defenders through the Fairwind Program, that the company is engaged in the US government's voluntary pre-release model access process, and that broader access to developers, enterprises and consumers comes "as soon as possible", starting with paid API customers and Google AI Ultra subscribers. There is no model identifier, no endpoint, no published quota and no date. It is not on our catalogue and it is not on anyone else's public API either, so a pricing page for Argon does not exist yet.
Claude Opus 5.5 ships today. On OrcaRouter it is anthropic/claude-opus-5.5, reached through the same OpenAI-compatible endpoint as the other 200-plus models in the catalogue, with automatic failover across providers when one route degrades, and a routing DSL for cases where you want to send some traffic to Opus 5.5 and the rest elsewhere on the same key. No second contract, no second SDK.

The practical recommendation for the next month is unglamorous: build on Opus 5.5, keep your model name in a config value rather than a string literal, and put a cheap model behind your retry path so a latency spike on a 702-second first-token run does not take your product down with it. That last point is not theoretical — Opus 5.5's first-token latency on AA's harness is eleven minutes, and whether that is a harness artefact or the model genuinely thinking longer than anything before it, your timeout budget should be set by the number, not by the marketing.
What would change this verdict
Three things, in order of likelihood. General availability of Argon with a published price, which would make the cost argument immediately actionable and would test whether $2/$10 survives contact with real traffic. An independent, reproducible run of DeepSWE v1.1 and CWE-bench v1 outside Google, which would either confirm the vendor claims or puncture them. Or an Artificial Analysis configuration of Argon at its top effort setting, which would settle whether the five-point index gap is a capability difference or an effort-settings artefact.
Until one of those lands, the honest summary is that Opus 5.5 is the better-verified model and Argon is the better-priced claim. Which of those you can actually use today is not a close call.
Claude Opus 5.5 is live on OrcaRouter at the vendor's list rate with no markup, alongside the rest of the catalogue — if you want to benchmark it against your own workload rather than against a leaderboard, anthropic/claude-opus-5.5 is the id to call, and swapping in a different model later is a one-line change on the same key.
Compared in this article2
Detected from this article · Benchmarks: Artificial Analysis · updated daily
