Gemini 4 Argon Can Write a Million Tokens. Long Is Not the Goal.
Maestro Brief · Published by Maestro Mojo
2026-10-02
Maestro’s take. Google gave Gemini 4 Argon room to stay on one difficult job far longer than most models. That could help with big migrations and long research tasks. It does not make a million-token answer useful, correct, or cheap to review.
Google announced Argon as a frontier model for long-running software engineering, enterprise work, and defensive security. It is not broadly available yet. Access starts with selected cyber defenders. Google says paid API customers and Google AI Ultra subscribers will come later, but it has not given a date.
TL;DR
- Argon can produce up to one million output tokens in one run, up from Google’s previous 64,000-token limit.
- That is a ceiling, not a target and not a quality score.
- Google says Argon scored 77.9% on its DeepSWE v1.1 coding test and is already helping with large internal code migrations. Those are Google’s claims.
- Artificial Analysis gives Argon a score of 53 on its independent Intelligence Index, ranking it eighth, and found it somewhat more verbose than comparable models.
- Introductory API pricing is $2 per million input tokens and $10 per million output tokens. Google says the standard price will double to $4 and $20.
- Do not route normal coding work to Argon just because it can run longer. Make long work earn the extra time and review.
What does “one million output tokens” mean?
It means Argon can keep generating inside one model run for much longer before it hits an output limit.
For a coding agent, that headroom could matter when the job needs many steps: study a large codebase, make changes, run tests, inspect failures, revise the patch, and repeat.
It does not mean you should ask for a million-token chat response. It does not prove the model remembers every detail, stays on course, or produces production-ready code.
A full one-million-token output would cost $10 at the introductory rate and $20 at the later standard rate, before counting input. The bigger cost may be the human time needed to understand what happened.
Where Argon may help
Google says its engineers are using Argon for debugging, algorithm work, and large C/C++-to-Rust migrations. One described project covers more than 800,000 lines in the Fuchsia Zircon kernel. Google also says those changes still go through automated tests, emulation, manual audits, and review.
That last sentence matters most. Long model runs are useful when the work has checkpoints outside the model.
A reasonable Argon job might be:
Port this isolated library from C++ to Rust in small batches. Preserve the public API. After each batch, run the named compatibility and performance tests. Stop when a test fails twice. Return the changed files, test results, performance difference, and unresolved risks. Do not merge or change callers outside this package.
That is very different from:
Rewrite this old system in Rust and make it better.
The first job gives the model a boundary, tests, stop conditions, and evidence. The second gives it a long runway and no steering.
Do this
- Use Argon for work that genuinely needs a long chain of reasoning or tool use.
- Break a large goal into checkpoints the model cannot mark complete by itself.
- Set your own output and spending limits far below the maximum until the workflow proves it needs more.
- Require tests, diffs, logs, or another machine-checkable result at every major stage.
- Compare cost per accepted result, including review and rework—not just price per token.
- Budget at the later $4/$20 rates. Google has not said how long the introductory price will last.
Do not do this
- Do not treat one million tokens as one million tokens of reliable attention.
- Do not give a long-running agent a vague repo-wide goal.
- Do not accept Google’s benchmark table as proof for your codebase.
- Do not count a generated patch as finished before independent tests and review.
- Do not pick Argon for small fixes, summaries, or routine code generation when a faster model already passes your tests.
Why Maestro users care
Long-running agents fail in two common ways: they run out of room, or they keep running after they have lost the plot.
Argon attacks the first problem. Your harness still has to control the second.
Artificial Analysis measured Argon at 53 on its Intelligence Index and about $1.99 per evaluation task at the introductory price. It also recorded 110 million output tokens across the evaluation, above the 81 million median for comparable models. That is useful early evidence, not a verdict. It suggests the model can be competitive without being especially concise.
For Maestro, the routing rule is simple: a long-output model belongs on long, high-value work with strong tests. It should not become the default just because the number is large.
One thing to try
When Argon becomes available to you, give it one difficult task your current model struggles to finish.
Run the same task with both models. Record:
- Whether the result passes the same tests.
- How much human rework it needs.
- Total runtime.
- Total model cost at Argon’s standard price.
If Argon produces a better accepted result for less total effort, route that kind of job to it. If it only produces more work to read, keep your current model.
The bottom line
Argon’s million-token output limit is useful capacity. It is not a reason to generate a million tokens.
Give long jobs more room only when tests, checkpoints, and a clear stop condition keep that room from turning into drift.
Sources considered
Published October 2, 2026 · Tags: Google, Gemini, AI Models, Coding Agents, Model Costs