What is a computer-use agent?
A computer-use agent is software that operates other software through its interface, clicking, typing, reading the screen and moving through menus the way a person would, instead of calling an API. A computer-use agent is how a system with no integration point still gets automated: a browser portal, a desktop application, a legacy system whose vendor never shipped an endpoint.
The agent is given a goal and a set of allowed steps. It signs in, opens the record, reads what is on the page, enters the values, confirms the result, then logs what it did. Nothing about it is magic. It is a careful operator that never gets bored and never skips the last field on a slow Friday afternoon. We build computer-use agents in Cape Town for South African operators, and we treat them as the tool of last resort: genuinely useful, occasionally the only option, and never the first thing we reach for.
When does a computer-use agent make sense, and when does an API beat it?
An API beats a computer-use agent whenever an API exists. An API is a contract. Fields are named, errors come back as codes, throughput is predictable, and a screen redesign breaks nothing. A computer-use agent has no contract. It reads structure somebody else is free to change without telling you.
So the honest rule is narrow. A computer-use agent earns its place when the system offers nothing else: a supplier portal with no export, a municipal or insurer site that only accepts a typed form, a desktop package still running because it holds twenty years of history nobody will migrate. Before we build one we look for an API, a file drop, a database view or a scheduled export, which is the ordinary path we describe under systems integration in South Africa. Only when all of those are genuinely off the table does the agent become the right answer.
How does a computer-use agent actually run a task?
A computer-use agent runs a task as a loop: look, decide, act, verify. It takes a view of the current screen, finds the element it needs, performs one action, then checks that the screen changed the way that step expected. If the check fails, the loop stops rather than carrying on into the wrong record.
Work reaches the agent from an ordinary trigger, a queue row, a new email, a scheduled run, and it carries a brief: the goal, the allowed steps, the fields it may touch and the conditions that end the run. Every step is timed and logged with a screenshot, so a failed run can be replayed and read by a person instead of guessed at. This is the same planning and tool-use pattern we describe under agentic AI, with the screen standing in for the tools. When the screen does not match, the item goes to a human, not to a guess.
What makes computer-use agents brittle, and how do we handle that?
Computer-use agents are brittle because they depend on a screen nobody promised to keep stable. A vendor moves a button, adds a consent banner, shortens a session timeout, puts up a maintenance notice, and a run that worked yesterday stops today. Portals are also slower and less predictable than an API, so throughput is lower and a busy period costs more time.
We design for that rather than pretending it away. Every step verifies its own result. Failures retry with strict limits so nothing loops forever. Items that fail twice land in a human queue with the screenshot attached. Monitoring reports layout changes the day they appear, not the week the numbers look wrong. We also tell you up front which parts of the workflow are fragile and what a bad week looks like, because a client who expects the occasional pause is far happier than one promised a machine that never blinks.
What permissions and controls does a computer-use agent need?
A computer-use agent needs its own named account, never a shared login and never a staff member's personal one, with the narrowest role the task allows. Credentials sit in a secret store instead of a script. Sessions run in an isolated environment, and the agent reaches only the systems on its allow list, so a stray instruction cannot send it somewhere it has no business being.
Reads are treated differently from writes. Anything that sends, pays, submits or deletes waits for a human sign-off, and the wording of that gate is agreed before the build. Every action is logged with a time stamp and the record it touched so an audit can reconstruct a run. Personal information is handled POPIA-aware: collected only where the task needs it, kept on a defined retention window, never pasted where it does not belong. The wider control model is set out under AI agent security.
How does an operations team start with a computer-use agent?
Start with one task that a person repeats in one system, on a fixed set of steps, with a clear definition of done. Show it to us on a screen share. We will say honestly whether a computer-use agent is right, or whether an integration, an export or a small process change is cheaper and steadier. If the system is an old in-house package, the better route is often a legacy software automation agent shaped around that system specifically.
When the agent is the right call, we script the happy path first and run it alongside your team in read-only mode, so the logs can be checked before anything is written back. Writes switch on once the run history is clean. The pilot runs two to four weeks on your own accounts, and you own the workflows, prompts and data. We have built this way with 35+ companies over 3+ years from Cape Town.
Related capabilities. The same parts, your business.
Keep reading. Pages close to this one.
Tell us which system has no API. We will tell you what is possible.
Send one message describing the system your team has to work by hand, and what the task looks like from start to finish. We reply with an honest read on whether a computer-use agent fits, what it would take, and where it would be fragile.