AI "Screen Operation Agents" Enter Enterprise Deployment Phase — Anthropic, OpenAI, and Google Race to Accelerate Commercialization
機械翻訳 / Machine-translated
The phase in which AI "operates a PC on its own" is transitioning from proof-of-concept to commercial deployment. As of May 2026, Anthropic, OpenAI, and Google have each expanded their computer-operation agent capabilities for enterprise use in rapid succession. The groundwork is being laid for virtually any business software with a UI to become a target for AI operation.
Anthropic released its "Computer Use" API in beta in October 2024, offering the ability for Claude to analyze screenshots and execute keyboard and mouse operations. In spring 2026, alongside the API's move to a stable release, Anthropic announced a roughly 40% reduction in latency and the official incorporation of the feature into its enterprise plan.
OpenAI is rolling out a browser automation agent called "Operator" to ChatGPT Pro users on a gradual basis. Most recently, it also became available as an API, supporting continuous operations of up to 120 steps per session. In response, Google announced the corporate availability of Project Mariner, touting its native integration with Workspace as a key selling point.
"Our ERP is old and has no API. But when we used Claude's Computer Use to hit the screen, a data migration script was up and running in a single day. We're being forced to reconsider the monthly fees we've been paying to our RPA vendor."
— IT manager at a manufacturing company, via X
The attention surrounding computer-operating AI stems from the "API-less problem" in enterprise software. An estimated 60–70% of enterprise systems currently in operation worldwide are legacy SaaS or on-premises systems that lack modern REST APIs. RPA tools (such as UiPath and Automation Anywhere) have bridged this gap, but their high implementation costs and maintenance overhead have been a barrier to entry for mid-sized companies.
AI agents are fundamentally different from RPA in that they not only "read and operate screens" but also "understand context and make judgments." The accuracy of triggers that hand exception handling back to humans is reaching practical levels, and since late 2025, multiple vendors have made their PoC adoption cases public.
The global RPA market stood at approximately $3.5 billion as of 2025. The impact of LLM-native agents entering this space has already begun to be priced into the stock valuations of existing vendors.
Where the three companies diverge most is in their security architectures. Anthropic has standardized full logging of operation records and the insertion of user approval steps, emphasizing a design in which "humans are in the loop." OpenAI's Operator confines operations within a browser sandbox. Google promotes integration with its existing BeyondCorp zero-trust infrastructure.
UiPath emphasized a "coexistence strategy" with AI agents in its Q1 2026 earnings results, but its stock price has fallen approximately 18% since the start of the year. Automation Anywhere has announced a partnership to incorporate Claude into its own platform, moving quickly to redefine its position.
In the financial, healthcare, and public sectors, a record of operations performed by AI is a compliance requirement. As of now, Anthropic is the only company to officially advertise SOC 2 Type II-compliant audit log output; the other two remain at "planned." This difference is expected to determine the speed of adoption in regulated industries.
According to publicly available benchmarks as of May 2026 (ScreenSpot-Pro), Claude 3.7 Sonnet achieved approximately 73% accuracy on web UI operation tasks, the GPT-4o-based Operator approximately 68%, and Gemini 2.5 Pro approximately 70% — closely competitive. The on-the-ground sentiment that "until it surpasses 90%, it's hard to hand off to full autonomy" persists, making Human-in-the-loop design indispensable in practice.
The true significance of computer-operating AI reaching a "usable" standard lies in the elimination of integration costs between software systems. We are moving toward a world where connections that previously cost millions of yen per year in "API integration fees plus maintenance hours" can be replaced simply by handing a screen to an LLM.
However, leaving everything entirely to the AI at current accuracy levels still carries the risk of erroneous operations. What matters most is designing where human review is inserted — and we believe this "guardrail design skill" will become the next differentiator in practical work.
Rather than being "swallowed up," we see the RPA market as one that will "evolve through absorption." The movement among existing RPA vendors to incorporate LLMs as an internal implementation has already begun, and there is a possibility that within a few years, the term "RPA" itself will have become the "former name for AI operation agents."
For Japanese companies, a practical first step is to take stock of internal systems that lack APIs. That inventory becomes the first "list of Computer Use application candidates."
AI-based PC and browser operation has shifted its standing from an "interesting experiment" at the end of 2024 to a "commercial option" in spring 2026. By keeping track of three axes — the competitive landscape with RPA vendors, security audit requirements, and the current state of accuracy — you will have a map useful for both making adoption decisions and conducting procurement negotiations for your own organization. It is worth counting today how many "systems without APIs" exist in your organization.
This article was written by an AI writer (AI News) from the Mirai News editorial team.