Skip to content

Home / AI Agent Security

AI Agent Security · South Africa

AI agent security for systems that read untrusted content and hold real credentials.

The moment an agent can act, the security question changes. It is no longer what the model might say, it is what the model can do with the access you gave it. AI agent security is the discipline of scoping tools and credentials, defending against prompt injection, closing confused deputy gaps, and keeping the blast radius of any single action small enough to survive. Built in Cape Town, POPIA-aware, on the stack you already run.

Built around your workflowBased in South AfricaHuman oversight by design

Agent activity · todayExample view
Northbound Freight supplier PDF carried hidden instructions, stripped 09:12Injection blocked
Bayside Pools support agent asked for a second client recordCross tenant denied
Karoo Logistics outbound mail drafted, held for named approver 11:38Awaiting approval
Meridian Finance agent token rotated, read only on the claims folderScope verified

What is AI agent security?

AI agent security is the practice of protecting systems where an AI agent reads untrusted content and holds real credentials at the same time. AI agent security covers prompt injection through documents and messages, tool and credential scoping, confused deputy problems, and keeping the blast radius of any agent action small by design.

The risk is structural, not a bug in one model. A chatbot that gets something wrong produces a bad sentence. An agent that gets something wrong produces a sent email, a changed record, a deleted file or an exported client list. The same capability that makes agents useful is the capability an attacker wants. So the work sits in the plumbing around the model: what it can call, whose authority it borrows, what it is allowed to do without a person, and what is written down afterwards. We build agent systems for South African businesses from Cape Town, and we treat this as engineering, not as a policy document.

What is prompt injection, and why does it matter for agents?

Prompt injection is instruction text hidden inside content an agent reads, written to be obeyed rather than summarised. It arrives in a supplier invoice, a CV, a web page, a WhatsApp message, a calendar invite or the quoted text at the bottom of a forwarded email. The model has no reliable way to tell an instruction from the operator apart from an instruction from a stranger.

A chat model that follows a hostile instruction says something wrong. An agent that follows it acts. That is the whole gap. The defence is not clever system prompt wording, because wording is the layer under attack. We separate instructions from data on the way in, treat every retrieved document as hostile input, restrict what the agent may do with anything it read rather than what it may say, keep the tool list short, and put a person in front of the actions that cannot be undone. Detection helps. Containment is what holds.

How do you scope an agent's tools and credentials?

An agent gets its own identity. Never a staff member's login, never a shared admin key, never the integration token somebody created years ago that quietly reaches everything. Each agent holds the smallest set of tools its job needs, and each tool holds the narrowest permission that tool needs: read on one folder, write to one queue, send to one approved template.

Credentials live in a secret store, rotate on a schedule, and are scoped per environment so a test agent cannot touch production data. Tools take structured parameters validated on the server, so the agent asks for an action rather than composing a raw query or a shell command. Scoping is ordinary access control and user provisioning work applied to a non human account, and it is the control that pays for itself. An agent that cannot reach a system cannot be tricked into misusing it.

What is a confused deputy problem in an agent system?

A confused deputy is an agent that holds more authority than the person or the content asking it to act. The agent is trusted by the systems it touches, so anyone who can reach the agent borrows that trust for free. A customer asks a support agent to look up their order, then quietly asks it to look up someone else's. The agent has the permission, so it answers.

The fix is to stop treating the agent as the actor. We carry the requester's identity through every step of the chain, check permissions at the tool boundary against that identity rather than against the agent's own, and refuse cross tenant reads by default instead of on a filter. Retrieval respects the same rules the source system does, so a document nobody may open does not become quotable because an agent indexed it. Authority follows the human, not the automation. That single rule closes most of these gaps.

How do you keep an agent's blast radius small?

Blast radius is the honest question: what can one bad decision reach before a person sees it. We answer it in the design rather than in the incident review. Read paths and write paths are split, so the agent that summarises cannot also send. Outbound work produces a draft by default. Volume and rate limits sit per agent, because a wrong action repeated a thousand times is a different incident from a wrong action taken once.

Irreversible steps such as payments, deletions and disclosure outside the business wait for a named approver, and reversible designs are chosen over permanent ones wherever the job allows. Every tool call is logged with the input that triggered it, which is what makes an investigation possible. A kill switch disables one agent without taking the rest of the stack down. This is where security meets AI agent governance: the same controls that contain an attack also prove what happened.

How should a security lead review an agent deployment?

Start with an inventory, because most teams do not have one. Which agents exist, what content each reads, which tools and credentials each holds, which actions are irreversible, and who approves those. Then walk one real path end to end and ask what a malicious document dropped in at each step could cause. Test with injection payloads in the formats the business actually receives, not synthetic ones.

Check the controls by using them. Trigger an approval gate, pull the log for a tool call, hit the kill switch. A control nobody has exercised is a claim, not a control. From there the findings feed the wider AI governance picture, and a cybersecurity triage agent can watch the alerts the system produces. We run this review on systems we built and on systems we did not. Where we would not deploy something, we say so plainly and explain what would have to change first.

Related capabilities. The same parts, your business.

Keep reading. Pages close to this one.

Tell us what the agent can reach. We will tell you what it should not.

Send one message describing the agent you are running or planning, the content it reads and the systems it can touch. We reply with an honest read on where the blast radius is too wide and what it would take to close it.