Skip to content

Home / Multilingual NLP

Multilingual NLP · South Africa

AI that understands English, SA languages and code-switching.

Customers do not speak one clean language. They mix Afrikaans, isiZulu, isiXhosa, Sesotho, Setswana and English, often inside a single sentence. Our multilingual NLP detects the language, retrieves evidence from your own documents, and replies fluently in the language the customer prefers, with safe human handover when confidence drops. A multilingual answer system grounded in your content, not in guesses. Built in Cape Town, POPIA-aware, on the channels your customers already use.

Built around your workflowBased in South AfricaHuman oversight by design

Multilingual inbox · todayExample view
Kaapstad Meubels Afrikaans and English in one message, segments tagged at 08:41Code-switch
Ndlovu Plumbing isiZulu delivery question matched to the English logistics policyRetrieved
Sandton Dental Studio refund dispute, confidence below threshold at 14:07Human handover
Bloem Agri Supplies Sesotho warranty query answered from the approved library at 19:12Resolved

What is multilingual NLP?

Multilingual NLP is language technology that reads a customer message in English, Afrikaans, isiZulu, isiXhosa, Sesotho or Setswana, works out what the person is actually asking for, finds the answer inside your own documents, and replies in the language the customer used. Multilingual NLP is not translation bolted onto a chatbot.

Detection, intent, retrieval and generation stay separate governed steps, and every step is measurable on its own. That structure is what lets multilingual NLP report accuracy per language and per intent instead of a single vague score. Customers ask for exact policy rules, exact turnaround times and exact warranty wording, so the reply has to match what your documents actually say. We build multilingual NLP for South African businesses from Cape Town, and we have delivered systems like this for 35+ companies over 3+ years.

How does multilingual NLP handle code-switching?

Multilingual NLP handles code-switching by tagging language at segment level rather than guessing one language for a whole message. A customer who opens in isiXhosa, drops an English product name in the middle, and closes with an Afrikaans sign-off produces three tagged segments, one intent, and one preferred reply language.

Single-language detection is where most chatbots break. The classifier picks the dominant language, misses the intent carried by the other one, and answers a question nobody asked. Intent cues differ per language too, so we evaluate the classifier per language instead of on an English-only test set. A preferred-language rule then keeps every reply consistent, whether the customer set that preference explicitly or the system inferred it from the opening message. Entities such as invoice numbers, branch names and product codes are extracted regardless of the surrounding language, so routing works the same in all of them.

How does multilingual NLP answer from our own documents?

Multilingual NLP answers from your documents using retrieval augmented generation. Your PDFs, FAQs, policies, price sheets and web pages are chunked and indexed as embeddings, the question is embedded in whatever language it arrived in, and cross-lingual retrieval pulls the matching chunks even when an Afrikaans query has to reach an English source document.

Top-k retrieval narrows the candidates, optional reranking sharpens precision, and the model writes the reply using only what was retrieved. When the documents do not contain the answer, the system says so. That fallback is deliberate, because hallucinations are expensive in support: a wrong policy quoted confidently in a customer language costs trust that takes months to rebuild. Retrieved snippets stay attached to the answer, so the team can see the evidence behind any reply and correct the source document rather than patching the prompt.

Which architecture does multilingual NLP use?

Multilingual NLP runs on one of three common patterns, chosen against your languages, latency needs and document complexity. Translate-then-retrieve moves queries and documents into a pivot language, which launches quickly on English-heavy content bases and carries translation edge cases as the trade-off.

Native multilingual keeps a single stack handling queries directly, with no translate step in the middle. That pattern suits genuinely multi-language corpora and demands strong per-language evaluation to stay honest. Retrieve-then-generate fetches the best evidence first and writes in the customer language afterwards, which is the strongest option for governed factual answers and the one that carries code-switching most cleanly. The pattern is a decision, not a default. We pick it with you, against the documents you actually hold and the languages your customers actually use, then evaluate it per language before anything reaches a live channel.

Is multilingual NLP POPIA compliant, and when does a human take over?

Multilingual NLP built by us is POPIA-aware from the first design session, because multilingual automation is only valuable if it is safe. Confidence thresholds decide when the assistant answers and when a person does. Sensitive categories such as refunds, disputes and legal questions escalate by rule rather than by judgement.

Critical policy wording comes from an approved answer library, so the sentences customers see in any language are sentences the business signed off. Consent is captured with source and time stamps, opt-in and opt-out flows run where they are relevant, and transcripts are deleted or anonymised on a retention schedule. Role-based access and action logs record who opened what. Handover is a feature, not a failure. A clean escalation with the conversation history attached beats a confident wrong answer in every language on the list.

How does a business start with multilingual NLP?

Starting with multilingual NLP is a pilot, not a platform migration. One channel, roughly three languages, and a smaller content set prove the pattern first: language detection, routing rules, and basic retrieval over the documents that already answer most incoming questions. That scope launches fast enough to validate the outcome before the budget conversation gets serious.

From there the build widens. WhatsApp, web chat and email run on the same logic with only the interface changing, more languages come on with glossary support for product and brand terms, and CRM or workflow integration connects the answers to the rest of the business. Scorecards travel with every stage, covering intent accuracy, retrieval quality, hallucination risk and language consistency. Language distribution, top intents and the questions the assistant could not answer become the input to the next improvement cycle.

Related capabilities. The same parts, your business.

Keep reading. Pages close to this one.

Tell us which languages you lose. We build what answers them.

Send one message describing where your support breaks down, whether that is Afrikaans queries going unanswered after hours, isiZulu messages routed to the wrong team, or policy questions that nobody has time to look up. We reply with an honest read on what multilingual NLP can fix, which architecture pattern fits your documents, and what the pilot would take.