Skip to content

Home / AI Model Routing

AI Model Routing · South Africa

AI model routing sends each task to the model that suits it.

One model for every job is a commercial decision disguised as a technical one. AI model routing puts a layer between your systems and the models, so each task goes to the option that fits on capability, cost, latency and where the data is allowed to live, and keeps running when a provider does not. Built in Cape Town for South African companies, on the stack you already run.

Built around your workflowBased in South AfricaHuman oversight by design

Routing log · todayExample view
Karoo Logistics delivery note extraction at 07:12, small model, structured output validRouted
VDM Attorneys contract clause review at 09:38, long context model, data kept in regionRestricted route
Stargas Energies primary provider timed out at 11:05, retried, fallback model answeredFailover
Bayside Pools support reply drafted at 14:47, cheapest passing model, human approves before sendAwaiting sign

What is AI model routing?

AI model routing is the practice of sending each task in a system to the model that suits that task, judged on capability, cost, latency and where the data is allowed to go, instead of committing a whole business to one provider. AI model routing is a layer, not a product. It sits between the business logic and the models, reads the task, and picks.

In a live system that looks ordinary. Classification and tagging go to a small fast model. A contract summary goes to a stronger one. Anything carrying personal information goes to a model hosted where the business permits it, which is often the reason a private AI deployment exists in the first place. The application code does not change when the choice does. We build AI model routing for South African companies from Cape Town, and we have delivered systems on this pattern for 35+ companies over 3+ years.

Why does AI model routing matter to a business?

Because a single model choice hardens into a commercial position. Without routing, every task pays the price and the latency of the strongest model, including the tagging, extraction and formatting work a small model handles perfectly well. The bill grows with volume, not with value.

The exposure is the larger half. A price change, a rate limit, a deprecated model or a region that goes unavailable reaches every workflow at once when there is only one path out. Routing turns model choice into configuration rather than architecture, so a swap becomes a decision instead of a rebuild. It also separates the two kinds of work most companies run together: the regulated processing that must stay on approved infrastructure, and the ordinary volume that can use whatever performs best this quarter. That separation is what makes a shared context layer safe to use across teams, because the rules travel with the task rather than living in someone's head.

How does a routing layer decide which model to use?

On the properties of the task, not on preference. Each task type is registered with what it actually needs: reasoning depth, context size, structured output, tool calling, an acceptable response time, and a data classification. A router without that register is guesswork with extra steps.

Requests carry the classification through, so a task touching personal information is only ever offered models running on approved infrastructure, and the rest of the rules follow it. Within the remaining candidates the router takes the cheapest and fastest option that has passed evaluation for that task, and escalates only when the cheap path returns something that fails validation. Every decision is logged with the model, the reason, the latency and the token counts. That log is what turns cost from a monthly surprise into a number per task, and it feeds the same reporting we build for decision intelligence, where spend per workflow sits next to the outcome it produced.

What happens when a model provider fails?

It gets designed for, because providers do fail. They time out, they rate limit, they return malformed output, and they deprecate on their own schedule. A system with one path out treats each of those as an outage. A routed system treats them as a route change.

Every route carries an ordered fallback chain of models that have passed the same evaluation for that task. A timeout or an error retries once, then moves to the next model in the chain. Repeated failures trip a breaker that parks that provider for a cooling period so the system stops queueing behind a dead endpoint. Structured output is validated before it is accepted, and a failed schema check counts as a failure rather than being passed downstream. Work that cannot complete queues for retry or waits for a person, it does not disappear. This is the same discipline that runs through our execution infrastructure: retries, queues, breakers and an operator screen showing what re-routed and why.

How does AI model routing keep a system usable as models change?

Models change faster than business processes do. A workflow that reconciles invoices or triages support tickets is still the same workflow after a model release, but the system around it often is not, because the model was wired straight into the logic.

Routing keeps the system usable by making the model a replaceable part behind a stable interface. The prompts, the tools, the context and the approval steps outlive any single model release. Each task type keeps a small evaluation set drawn from the company's own work, real documents and real messages, so a new model is measured on cases the business recognises rather than on a public benchmark. New models are shadowed against live traffic first, compared on quality, latency and cost, and promoted only on the routes where they win. Rollback is a configuration change, not an incident. The result is a system that absorbs the market rather than being redesigned by it.

How does a company start with AI model routing?

Start with an inventory, not a platform decision. We list the tasks the system actually performs, what each one needs from a model, and which of them touch data that carries rules. That conversation costs nothing and usually takes under an hour.

Most companies find that a large share of their calls are routine work paying premium prices, and that the rules about sensitive data live in conversation rather than in code. We then put a routing layer in front of the existing calls without changing the business logic, register the task types, add the fallback chains, and turn on logging so cost and latency become visible per task instead of per invoice. Evaluation sets are built from the company's own cases before anything is promoted. The company owns the routing rules, the evaluation sets, the prompts and the provider keys. Nothing here is rented back.

Related capabilities. The same parts, your business.

Keep reading. Pages close to this one.

Tell us what you are locked into. We build the way out.

Send one message describing where your AI spend goes, which tasks cannot leave the country, and what breaks when a provider has a bad morning. We reply with an honest read on what AI model routing can fix and what it will take.