惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

H
Hackread – Cybersecurity News, Data Breaches, AI and More
U
Unit 42
Vercel News
Vercel News
Martin Fowler
Martin Fowler
云风的 BLOG
云风的 BLOG
爱范儿
爱范儿
MongoDB | Blog
MongoDB | Blog
J
Java Code Geeks
F
Fortinet All Blogs
MyScale Blog
MyScale Blog
C
Check Point Blog
N
Netflix TechBlog - Medium
Microsoft Azure Blog
Microsoft Azure Blog
aimingoo的专栏
aimingoo的专栏
博客园_首页
WordPress大学
WordPress大学
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
IT之家
IT之家
Last Week in AI
Last Week in AI
罗磊的独立博客
大猫的无限游戏
大猫的无限游戏
Jina AI
Jina AI
V
Visual Studio Blog
小众软件
小众软件

Citrix Blogs

Celebrating the Partners powering Citrix forward – Citrix Blogs What’s left for humans? – Citrix Blogs Why this moment feels different – Citrix Blogs Introducing Citrix Platform for Public Sector – Citrix Blogs how our partnership with Google has matured secure access for the browser era – Citrix Blogs When session recording stops scaling – Citrix Blogs A conversation with Cletis Earle – Citrix Blogs Imprivata Ready Certification validates Citrix Unicon: A practical guide for healthcare IT – Citrix Blogs Skills are all you need – Citrix Blogs UHMC customers now have expanded Citrix Secure Private Access entitlement – Citrix Blogs How CIOs turn post‑merger disorder into a synergy engine – Citrix Blogs why your AI strategy is focused on the wrong layer – Citrix Blogs Securing high privileged admin access doesn’t have to be complicated – Citrix Blogs What will knowledge work be in 18 months? Look at what AI is doing to coding right now. – Citrix Blogs Untangling spaghetti – Citrix Blogs Workers’ “second brains” break every assumption about how we secure knowledge work – Citrix Blogs Taming integration chaos (the core of M&A failure) – Citrix Blogs OpenClaw and Moltbook preview the changes needed with corporate AI governance – Citrix Blogs Three years. Five Use Cases. A Leader: Citrix – Citrix Blogs The hard truths about hospital consolidation: An M&A guide for IT leaders – Citrix Blogs Why Citrix is the most complete EUC platform – Citrix Blogs Sign in once, get more done: Why continuous identity is a strategic advantage – Citrix Blogs Everyone’s worried about the wrong AI security risk – Citrix Blogs Security by design, proven by action with Citrix NetScaler – Citrix Blogs The invisible 80%—what corporate-led AI transformations can’t see – Citrix Blogs Workers don’t want to build automations. They want to delegate. – Citrix Blogs speed vs. security – Citrix Blogs AI will be THE interface to knowledge work. Here’s how we’ll get there. – Citrix Blogs Why I joined Citrix — and what it means for healthcare leaders – Citrix Blogs How the most successful CIOs are building successful merger and acquisition approaches – Citrix Blogs
What happens when AI agents score 100% in computing using...
Brian Madden · 2025-07-24 · via Citrix Blogs

What happens when AI agents score 100% in computing using benchmarks?

All the major AI labs are developing AI software agents that can operate a computer just like a person, visually parsing pixels, moving the mouse, and pressing keys. These are Computing User Agents (CUAs), and I wrote in-depth about them last week.

Today’s best CUAs are able to complete about 45% of the tasks in the popular OSWorld benchmark, up from just 6% when the benchmark was created sixteen months ago. What happens when they reach 100%?

In this post, I’ll explore what these benchmarks really measure, what they leave out, and how to prepare for the moment that AI UI execution becomes a solved problem.

Understanding the CUA benchmarks

The OSWorld benchmark defines 369 real desktop tasks: file management, web browsing, multi-app workflows, and so on. Human testers are able to finish 72-74% of them, but CUAs are closing the gap fast: 17% at the beginning of 2025, 45% today, and likely human parity in 2026.

Eventually, one will hit 100%. Then what?

A perfect score only proves a CUA can navigate any UI. It still won’t decide why a task matters, evaluate risk, or resolve ambiguity. The CUA execution layer is just the hands and eyes, not the brain.

Humans provide the scaffolding

Even with 100% capable CUAs, human workers will need to:

  • Set goals and intent
  • Define guardrails and checkpoints
  • Respond to escalations

The value won’t be in the raw UI control by the CUA, it will be the scaffolding around it.

CUAs become the universal API

That said, once CUA execution is solved, every legacy desktop app will become an intelligent API surface. A typical use case will involve multiple models working together:

  • An interface agent (CUA) will be deterministic and sandboxed. There will be one per workspace instance.
  • Planner / reasoning agents will decide and orchestrate which CUAs to invoke, when, and with what constraints.

CUAs plus scaffolding map directly to Stage 4 (“AI uses your computer”), and then evolve into Stages 5 (“AI uses your computer without you”) and 6 (“multi-agent coordination”) in the 7‑stage human-AI collaboration roadmap.

Because each CUA will use existing workspaces and run with its own identity, existing IAM, DLP, session recording, and other guardrails will still apply as they do today. You won’t have to reinvent your security model.

Planning for this future

Don’t get me wrong, it will be a big deal when CUAs hit 100 % execution. But this will only be a milestone, not the final destination, on the journey to AI in the workplace. When this happens, the value will shift to how well you design, secure, and govern the orchestration layer which drives the CUAs.

The scarcity (e.g. the value) will be the judgment around knowing what to do, how to do it, and what success looks like. When CUAs hit 100%, that judgement will still come from humans.

At that point, human workers will shift from doing to directing, and the org chart will start to look like a massive orchestration graph. Once CUA execution is solved, the only interface left to optimize will be your own thinking.


Read more & connect

Join the conversation and discuss this post on LinkedIn. You can find all my posts on my author page on the Citrix blog (or via RSS).

Video of my most recent talk

In May I gave the closing keynote at the EUCtech Denmark 2025 conference, called The Future of Work in an AI-Native World. I talked about a lot of what I covered today and walked through how AI will evolve and impact the workplace in the coming years. You can watch it on YouTube.

My upcoming talks

  • AppManagEvent: Closing Keynote: AI & the Future of Enterprise Apps — Utrecht, Netherlands, Oct 10
  • MAICON 2025: AI at Work: The Employees’ Revolution! — Cleveland, Ohio, Oct 14-16

Brian Madden

Brian Madden is a VP & futurist at Citrix. He writes about the future of work, AI in the workplace, and the evolution of Citrix.