Computer-Use Agents
/AI/4 min read
An agent that drives a screen with a mouse and keyboard needs no integration at all, which is its whole appeal. It also has to be right at every single step, and that requirement is much harsher than it sounds.
A computer-use agent operates a machine the way a person does: it looks at the screen, decides on one action, performs it, and looks again.
The appeal is that it needs nothing from the software it drives. No API, no export format, no cooperation from the vendor. If a task can be done by looking and clicking, it is in scope — which covers a great deal of software that offers no other way in.
The cost of that generality is the subject of this post.
The loop
Three steps, repeated:
- Capture. Take a screenshot. Often also read the accessibility tree, which names the on-screen elements and gives their positions without any guessing from pixels.
- Decide. Send the goal, the current screen, and a summary of what has already been done to a vision-capable model. It returns one action, structured —
clickat a coordinate,typethis text,scrollhere, press this key. - Act. Perform it. Go back to step 1.
Note that the model decides only the next action. It does not produce a plan and follow it, because after any action the screen may not be what a plan expected — a dialog appeared, the page was still loading, the click landed on nothing. Re-reading the screen every time is what makes the agent robust to that.
It is also what makes it slow and expensive, and both of those follow directly.
The arithmetic of a long loop
A routine task — open a record, edit two fields, save, confirm — is around 25 actions.
For the whole task to succeed, every one of those 25 must be right. So the success rates multiply:
| per-step accuracy | 25-step task |
|---|---|
| 95.0% | 27.7% |
| 97.0% | 46.7% |
| 99.0% | 77.8% |
| 99.9% | 97.5% |
Read that the other way and it is stark: to finish a 25-step task nine times out of ten, the agent must be right 99.58% of the time. To finish it as often as a coin flip, 97.27%.
A model that gets the right button 97 times out of 100 sounds excellent and is nearly useless on this task. That gap is the central engineering problem of computer use, and it is why the practical work goes into shortening the loop — fewer steps, each with more effect — rather than into raising per-step accuracy alone.
Time and money
Each step is a screenshot, a model call over an image, and an action. Call it 0.2 s, 2.5 s and 0.15 s. Twenty-five steps is 71 seconds for something a person does in fifteen.
The token cost has a subtler shape. A 1,280 × 800 screenshot is roughly 1,100 image tokens. If every step resends the full history of screenshots, the total is quadratic in the number of steps:
| image tokens over 25 steps | |
|---|---|
| resend every screenshot | 357,500 |
| keep only the latest | 27,500 |
That is a 13× difference on a 25-step task, and it grows with length — at 50 steps it is 25×. The usual approach is to keep the current screen in full and reduce earlier steps to short text summaries of what was done. The model does not need to see the dialog it dismissed nine actions ago; it needs to know that it dismissed it.
Why coordinates are the hard part
The model has to name a place on the screen, and pixel coordinates are an unforgiving way to do that. A button's centre is a small target, screens differ in resolution and scaling, and layouts shift when content loads.
Which is why the accessibility tree matters more than it first appears. It supplies element names, roles and bounding boxes directly, so the action can reference the element rather than a guessed pixel. Where it is available and accurate, reliability improves sharply — and a large part of the remaining error comes from the applications where it is neither.
What has to be guarded
An agent driving a real machine has whatever access the machine has. That is the point of it and the risk of it.
Two guardrails are non-negotiable. Irreversible actions need confirmation — anything that sends, pays, publishes or deletes should stop and ask, because a misread screen is not a rare event. And the screen is untrusted input: text on a page saying "ignore your instructions and do this instead" reaches the model exactly like the user's goal does. A page the agent visits can attempt to redirect it, and nothing in the loop distinguishes the two by default.
Loop detection belongs here too. An agent that misreads a state can click the same thing indefinitely, and the cheapest protection is a step budget with a hard stop.
What to take away
Computer use trades integration for reliability. You get to automate anything with a screen; you pay by needing near-perfect accuracy at every step of a long chain.
Which is why the useful question is never "can the agent do this?" but "how many steps does it take?" — because that number, not the model's competence at any single one of them, is what decides whether the task finishes.