Gemini 4 Argon Just Changed the AI War From Chatbots to Digital Workers

Google’s Gemini 4 Argon is built to stay on difficult tasks, use tools, work across enterprise workflows and autonomously modify real systems. The bigger story is not another benchmark fight with OpenAI and Anthropic. It is that all three companies are now competing to build AI that can perform sustained work rather than simply answer questions.

AI#Gemini 4 Argon#Google DeepMind#AI Agents#Agentic AI#Enterprise AI#Digital Workers
By TD Research| · 11 min read4 views
Share
Gemini 4 Argon Control Room
Google’s Gemini 4 Argon is designed for sustained software engineering, enterprise knowledge work and cybersecurity tasks as frontier AI competition shifts toward autonomous agents. Photo: Google DeepMind

Key Takeaways

  • 01Gemini 4 Argon is designed for long-horizon work, with Google expanding its output limit from 64K to 1 million tokens.
  • 02Google is already using Argon agents internally for code migrations, data-center optimisation and research tasks.
  • 03Argon leads several business and knowledge-work evaluations, including AutomationBench, but its 51.29% score also shows that unattended enterprise automation remains far from solved.
  • 04Cybersecurity exposes the central tension: Google says Argon can autonomously find, validate and patch vulnerabilities, yet its strongest cyber configuration is initially limited to trusted defenders.
  • 05OpenAI and Anthropic are moving in the same direction. Dots, GPT-6 Astra, Claude Code and dynamic workflows show the frontier shifting from conversational AI toward increasingly autonomous work execution.

Google’s most interesting Gemini 4 feature is the one most people cannot use

Google launched Gemini 4 Argon on September 30 with the usual ingredients of a frontier-model announcement: stronger benchmarks, lower pricing, improved reasoning and comparisons against OpenAI and Anthropic.

But the most consequential part of the launch is buried elsewhere.

Google is already using Argon agents internally to modify large codebases, optimise data-center memory and work on quantum-computing problems. At the same time, the company is refusing to broadly release Argon’s strongest cybersecurity capabilities because the model can autonomously find, validate and patch critical software vulnerabilities.

That tension says more about where artificial intelligence is going than another leaderboard.

For most of the generative AI boom, the central question was which chatbot produced the best answer. With Argon, OpenAI’s Dots and Anthropic’s increasingly autonomous Claude tooling arriving around the same period, the competitive question is changing.

Which AI can actually be trusted with the work?

“Digital worker” is not Google's product label for Argon. It is a useful description of the direction these systems are moving. The model is designed not merely to generate a response but to remain engaged with a complicated objective, reason across many steps, interact with software, modify things, evaluate the result and continue working.

That is a fundamentally different product from the chatbot that started this AI cycle.

Argon is designed to keep working

The technical feature behind much of Google’s positioning is unusually large.

According to Google’s Gemini 4 Argon announcement, the model's maximum output-token allowance has increased from 64,000 to 1 million tokens.

That detail matters. It is not simply a claim that users can paste a million tokens into Gemini. Google specifically describes the increase as output headroom that allows Argon to reason and generate hundreds of thousands of tokens during a single trajectory.

In practice, the goal is to give the model room to stay with a problem instead of producing an answer and stopping.

That could mean inspecting a software project, identifying dependencies, writing code, running tests, encountering failures, modifying the implementation and trying again. In finance it could mean collecting information from multiple documents, comparing it, performing analysis and producing a finished piece of work. In cybersecurity, it can mean finding a vulnerability and continuing far enough to validate and patch it.

The raw token number will make headlines. The more important change is the task horizon.

AI companies are trying to increase the amount of useful work a model can complete before a human has to take control again.

Google is already testing the idea on itself

The strongest evidence for Google's intended direction does not come from Gemini chat demos.

It comes from Google's engineering infrastructure.

The company says a team of Argon agents analysed profiling telemetry across its data centers and identified memory optimisations that freed more than 300 TiB of memory, with Google estimating eventual savings between 500 TiB and 1 PiB once the changes are fully deployed.

Google also says Argon is being used in large C and C++ to Rust migration projects. Those projects range from libraries containing tens of thousands of lines of code to more than 800,000 lines in the Fuchsia Zircon kernel. The company stresses that changes to critical systems still undergo automated testing, emulation, manual auditing and human review.

One example is particularly revealing. Google says Argon agents worked on the Rust version of its libgav1 video decoder, replacing roughly 32,000 lines of SIMD code through repeated profiling and compiler analysis. Google reports that the resulting implementation ran 2.7 times faster than the previous Rust port while producing identical video output.

Another internal use case involves quantum algorithm optimisation. Google says Argon improved a published baseline by 40% in one experiment. These results are company-reported, not independent validations, but together they make Google's product thesis clear: Argon is being developed for projects rather than prompts.

This distinction matters commercially. A chatbot sells access to intelligence. An agent capable of taking responsibility for a bounded piece of work starts competing with software tooling, outsourcing, professional services and eventually portions of human labour.

The benchmark that should make enterprises pay attention is not the coding benchmark

Google leads with impressive coding numbers. It reports 77.9% on DeepSWE v1.1, a benchmark for long-horizon software-engineering tasks. But Argon's enterprise result may be more consequential.

On Zapier’s AutomationBench, which tests agents performing end-to-end work across realistic business systems, Gemini 4 Argon at high reasoning effort currently scores 51.29%. Claude Sonnet 5.5 follows at 44.75%, Claude Opus 5.5 at 42.47% and GPT-6 Astra at 41.4%. The benchmark spans sales, marketing, operations, support, finance and HR, using 47 real tools and scoring whether the correct changes were actually made in the underlying environment.

That result is simultaneously impressive and sobering.

Argon leads.

It also still fails nearly half of the benchmark.

That is the reality behind the “digital worker” narrative.

The frontier is moving toward autonomous execution, but the evidence does not support simply giving these systems unrestricted access to a company and walking away.

The real enterprise architecture will require permission boundaries, verification, monitoring, audit trails, human escalation and controls over what an agent can change.

Google appears to understand that problem because it has built exactly those concerns into Argon’s rollout.

Google built the worker, then put a security guard around it

Argon's cybersecurity capability is where the story becomes uncomfortable.

Google says the model can autonomously find, validate and patch critical software vulnerabilities. The company is giving selected trusted defenders and its internal security teams versions without its normal cyber guardrails so they can use the model's full defensive capabilities.

That unrestricted version is not being handed to the general public.

Instead, initial access is being distributed through Google DeepMind's Fairwind Program. Google says Fairwind works with more than 650 partners globally, with selected partners receiving Argon access and the ability to use it through CodeMender for vulnerability research and patching.

Cloud-security company Wiz is among the early users. Google says Wiz used Argon through its Scan for Good program and that the model found a critical vulnerability exposing sensitive personal information in healthcare software used by hospitals worldwide. According to Google, previous frontier models had missed it.

The obvious problem is that vulnerability discovery is dual use. The same class of reasoning that helps a defender find an unknown weakness can potentially help an attacker understand one.

So the feature that makes Argon commercially valuable is also part of the reason Google is reluctant to fully release it.

That is an important inversion of the old AI model race. Previously, vendors were mostly worried about what a model might say. Agentic models create another category of risk: what the model might do.

Google is monitoring the AI while the AI works

Google's safety architecture provides another clue about where agent products are heading.

The company says it is specifically strengthening Argon against indirect prompt injection, where malicious content encountered by an agent attempts to manipulate its behaviour. It also says it monitors Argon's reasoning and actions for signs that the system is exceeding the user's intentions and can terminate execution when necessary.

Google says similar monitoring was applied during training, with suspicious activity escalated to a dedicated incident-response team. The company is also hardening isolated sandbox environments for higher-risk training and evaluations.

These safeguards are technically important because an AI worker has a different attack surface from an AI chatbot.

A chatbot can hallucinate something wrong.

An agent connected to a browser, terminal, cloud environment, source-code repository or enterprise application can potentially turn a mistake into an action.

The security layer around the model may therefore become almost as important as the model itself.

But Argon has not won the AI race

Google's launch benchmarks make Argon look formidable, but they also provide reasons to resist a simplistic “Google beats OpenAI and Anthropic” conclusion.

Independent evaluator Vals AI currently ranks Gemini 4 Argon first on its overall Vals Index at 68.90%, ahead of Claude Sonnet 5.5 at 67.04% and Claude Opus 5.5 at 66.97%. Argon also ranks first on Vals' Finance Agent v2.

But the same evaluation suite shows weaknesses. Argon sits fifth on Terminal-Bench 4.0 and seventh out of eight evaluated systems on CUA-bench, a computer-use assessment.

Google's own published comparisons show a similarly mixed picture. Argon scores 55% on FrontierSWE v2 compared with 65.5% for GPT-6 Astra, while Claude Opus 5.5 leads Google's Terminal-Bench 4.0 comparison at 66.4% against Argon's 57.4%. On the OSWorld 2.0 offline computer-use subset, GPT-6 Astra scores 72.6% versus Argon's 69.2%.

So the evidence does not show a model that is simply superior at everything.

It suggests different frontier systems are developing different strengths.

And that is where the competitive story gets more interesting.

OpenAI, Anthropic and Google are converging on the same destination

OpenAI launched Dots one day before Argon. Reuters describes Dots as always-on autonomous agents able to pursue goals across applications, work through tools such as Codex and ChatGPT Work, and interact through systems including Slack and Microsoft Teams.

OpenAI's GPT-6 Astra, meanwhile, is explicitly trained for computer use, browsing, software engineering and professional workflows. OpenAI says Astra can operate interfaces, update CRM records, work with documents and spreadsheets, install software, test it and troubleshoot problems on screen.

Anthropic is pursuing the same transition from another direction.

Claude Code can already accept a GitHub task and continue working remotely after the user leaves, eventually returning a pull request for review. Claude Help Center Anthropic's dynamic workflows go further, allowing Claude to create orchestration scripts and coordinate tens or hundreds of parallel subagents for projects that can run for hours or days.

Anthropic's own research found that its longest Claude Code sessions had already increased from less than 25 minutes of autonomous operation to more than 45 minutes over a three-month period earlier this year.

The companies are using different interfaces, model families and product strategies.

But they are increasingly chasing the same scarce resource:

the amount of work a human is willing to hand over.

That may become a more important metric than chatbot preference.

Argon introduces another battle: the economics of AI labor

Google is also attacking the market through price.

Argon will initially cost $2 per million input tokens and $10 per million output tokens, with cached input discounted by 95%. Google says standard pricing after the introductory period will rise to $4 and $20 respectively.

For comparison, OpenAI currently lists GPT-6 Astra at $10 per million input tokens and $50 per million output tokens through its API. OpenAI Developers Anthropic prices Claude Opus 5.5 at $4 and $20, while Sonnet 5.5 costs $2 and $10.

Raw token prices are not enough to determine which system is cheaper for a finished task. Models can consume radically different amounts of tokens, tool calls and execution time to reach the same result.

That distinction becomes much more important with agents.

Companies will eventually care less about the price of a million tokens and more about the cost of completing a migration, investigating an incident, preparing an analysis, processing a finance workflow or resolving a support case correctly.

The economic unit of AI could shift from tokens consumed to work completed.

The digital worker is arriving before the digital workplace is ready

There is one final complication.

Most companies spent the last several years asking how employees should use AI.

Agentic systems create a harder question: what should AI be allowed to do without an employee?

A model that drafts a legal memo is one thing. A model that researches the case, accesses company records, updates the document repository and sends the resulting material elsewhere requires a completely different operating model.

The same is true in engineering. Suggesting code is different from modifying hundreds of thousands of production lines. Identifying a cybersecurity vulnerability is different from testing that vulnerability against a live system.

Argon makes that distinction difficult to ignore.

Google says broader access will eventually begin with paid Gemini API customers and Google AI Ultra subscribers, but as of October 2, it has not provided a firm date for general availability.

That makes Gemini 4 Argon an unusual flagship launch: Google has announced what the model can do before most people are allowed to do it.

And perhaps that is the real signal.

The frontier AI race is no longer just about who can build the smartest model. OpenAI, Anthropic, and Google are trying to determine how long their systems can operate, how much software they can control, how much professional work they can absorb, and how safely humans can delegate authority to them.

The chatbot war asked who could give the best answer.

The next AI war is about who can finish the work.

Sources & References

  1. Google’s Gemini 4 Argon announcementblog.google
  2. Zapier’s AutomationBenchzapier.com
  3. Fairwind Programdeepmind.google
  4. Vals AI currently ranks Gemini 4 Argon first on its overall Vals Indexvals.ai
  5. Dotstdisrupt.com
  6. Claude Help Centersupport.claude.com
  7. OpenAI Developersdevelopers.openai.com
Filed Under:#Gemini 4 Argon#Google DeepMind#AI Agents#Agentic AI#Enterprise AI#Digital Workers#Google#Google DeepMind#Gemini#Gemini 4 Argon#Digital Workers

About the author

TDisrupt's research desk focused on technology markets, emerging infrastructure, digital assets and industry intelligence.