Skip to content

Home / Private AI

Private AI · South Africa

Private AI is about the boundary, not the brand of model.

Private AI means running models so that sensitive business data stays inside a boundary the organisation controls: on its own servers, on a dedicated private endpoint, or on a machine with nothing going out. This page defines the category, sets out what genuinely has to stay in-house and what does not, and states the trade-offs in capability, cost and effort plainly. Built and run from Cape Town for South African businesses.

Built around your workflowBased in South AfricaHuman oversight by design

Model routing log · todayExample view
Rondebosch Medical Group patient summary at 07:12, kept on local modelIn boundary
Vermeulen Attorneys matter notes redacted, then routed on-premiseRedacted
Table Bay Underwriters claim file blocked from hosted model at 11:38Blocked
Silverstroom Retail product copy sent to hosted model, non-sensitivePublic class

What is private AI?

Private AI is the practice of running AI models so that sensitive business data stays inside a boundary the organisation controls. Private AI covers models hosted on the company's own servers, models running on a dedicated private endpoint, and models running on a laptop or an office machine with nothing leaving it. The boundary is the point, not the brand of model.

That framing matters because private AI is often sold as a product when it is really a design decision. The same open-weight model is private on a server in a company's own rack and not private at all behind a shared consumer app. What makes a deployment private is a documented answer to three questions: where the data sits, who else can read it, and what is retained afterwards. We build in this space from Cape Town, and it sits next to the wider question of sovereign AI in South Africa, which asks the same thing at a national level.

Why does private AI matter for a regulated business?

Private AI matters because a regulated business has to answer one specific question about every system it runs: where does this record live, who can read it, and how long is it kept. A public API can be a perfectly reasonable answer to that question. The contract, the processing region and the retention terms can all line up, and often do.

It stops being an answer when the records are medical files, legal matter notes, client financials or personnel data governed by rules the business cannot renegotiate. Then the difference between a vendor's stated policy and a boundary the business itself enforces becomes the whole compliance position. Private AI turns the answer into something the business states rather than something it accepts. That is also where this overlaps with AI governance under POPIA: the boundary is only worth anything if there is evidence of where each request actually went.

What actually has to stay in-house, and what does not?

Less has to stay in-house than most teams assume, and the split is worth doing carefully rather than by instinct. Raw identifiers, medical and legal detail, unredacted personnel files and anything covered by a contractual no-third-party clause belong inside the boundary. Public marketing copy, published policy text, product descriptions and already-anonymous aggregates usually do not.

Most real work sits between those two, which is why the useful question is which fields cross the line, not which department the work came from. A claims summary may be sensitive only because of three identifier fields. Strip those and the remaining text is ordinary prose. Classifying at the field level is what keeps a private deployment small enough to actually run. We do this mapping before any model is chosen, because a business that treats everything as restricted ends up with a slow system nobody uses, and one that treats nothing as restricted has no boundary at all.

What are the honest trade-offs of running private or local models?

Private and local models cost capability, money and attention. Saying otherwise is how deployments end up abandoned six months in. The strongest reasoning models are not the ones a business runs on its own hardware, so long document analysis and difficult multi-step reasoning come out noticeably weaker on a local model than on a frontier hosted one. Short, well-scoped tasks hold up far better.

There is an operational cost too. Someone owns the hardware, the patching, the backups and the quiet failures at month end. That someone is either an internal person with the time or a partner with a support arrangement, and if it is neither, the deployment decays. The gain is a boundary that can be described precisely and evidenced in an audit, plus predictable behaviour that does not change when a vendor updates a model. For regulated work that trade is usually worth making. For general drafting it often is not.

How do private and public models work together in one system?

They work together through routing, which is the part most businesses skip and then regret. One rule set sits in front of every model and decides, per request, which model may see it. The decision is made on the data class of the payload, not on which tool the person happened to open.

Sensitive classes go to the private or local model. Everything else can use a hosted model where that is genuinely the better tool for the job. Redaction runs before the boundary, so a partly sensitive request can be split rather than blocked outright. Every decision is logged with the class, the destination and the outcome, which is what turns a policy into evidence. This is the same mechanism described on our AI model routing page, applied with privacy as the routing criterion rather than cost or speed. A boundary without a router is a policy nobody enforces.

How does a business start with private AI?

Start with a data map, not a server. List the systems holding sensitive records, mark which fields are genuinely restricted, and name the rule or the contract clause behind each mark. That exercise alone usually resolves half the argument, because most teams find the restricted set is narrower and sharper than the general anxiety around it suggested.

Then pick one workflow that is currently blocked by that restriction and build only that, with routing, redaction and logging in place from the first run rather than added later. Run it on real work for a few weeks with the people who will actually use it. What earns its place stays, and what does not gets cut before it becomes infrastructure nobody maintains. The business owns the models, the rules and the logs. We have built this way for 35+ companies over 3+ years, and the same discipline runs through our wider AI governance work.

Related capabilities. The same parts, your business.

Keep reading. Pages close to this one.

Tell us what cannot leave. We build around that.

Send one message describing which records the business cannot send to a public API and what work is currently stuck behind that. We reply with an honest read on what private AI can do for it, what it will cost in capability, and what it will take to run.