Healthcare Technology

Why We're Not Putting AI on the Phone Lines (Yet): A Reality Check on African Healthcare Automation

A framework for evaluating AI voice agents in East African healthcare — self-hosted LLMs, cloud APIs, and vendor platforms, and why we split by channel instead.

maxwell.kimaiyoSep 21, 20268 min read

Every healthcare platform operating at scale in East Africa eventually gets the same pitch: an AI voice agent that can handle patient calls, transcribe conversations, and automate scheduling perfectly. The demos are always slick, and the accuracy statistics sound incredibly impressive.

We recently put these systems through a rigorous internal evaluation for our network in Kenya and Uganda. The process forced us to take a hard look at where the technology actually stands today. What we discovered is that the real challenge isn’t a simple debate of “AI versus humans”. It is about understanding exactly where each option breaks under real-world conditions and building a strategy that keeps our patients safe when it does.

Here is the exact framework we used, and where we actually landed.

The Four Options on the Table

When you are managing a hospital network and looking to scale patient support, you aren’t choosing between a chatbot and a call centre. There are actually four distinct paths you can take:

  1. In-House Self-Hosted LLM — Running an open-weight model (like Llama or Mistral) on physical, company-owned servers.
  2. Cloud-Hosted AI APIs — Offloading the conversational logic to a managed cloud engine (like Azure OpenAI) via web endpoints.
  3. Third-Party Managed Platform — Buying a packaged, end-to-end product from a communications vendor (like ConnexAI) to handle telephony and automation.
  4. Human Agents — Keeping your existing support team—the baseline that every alternative must be measured against.

Most vendor sales pitches skip straight to the third option and frame it as an all-or-nothing upgrade. That is a mistake. When you look under the hood, each option carries massive, hidden liabilities.

1. Self-Hosting Your Own Model: The Infrastructure Trap

The appeal of building your own setup sounds great on paper: absolute data residency, total control, and no recurring vendor fees. The reality is an incredibly expensive logistics headache.

To run an enterprise-grade text brain with acceptable speed, you cannot use standard web servers. You have to buy specialized AI hardware, specifically NVIDIA A100 or H100 GPU clusters. A single adequate server node easily runs into $30,000 to $100,000+ USD before you even factor in high Kenyan customs duties and shipping.

Once the hardware arrives, the operational risks shift directly into your server room. These chips draw massive amounts of power and generate extreme heat. In an environment like Nairobi, where grid reliability fluctuates, you have to invest heavily in industrial-grade, precision cooling systems and uninterruptible power backups (UPS) just to keep a phone line online. If a single specialized component burns out, you face massive import lead times and a total lack of local replacement support. Your entire support infrastructure could sit offline for weeks waiting for a spare part.

Even if you get the hardware running, you hit the fine-tuning mirage. Training an AI to handle complex clinical scheduling requires thousands of hours of perfectly clean, structured, and legally consented local conversation data. Most networks simply don’t have this volume of curated data yet. If you try to train the model on messy internal system notes instead, you get an AI that sounds completely confident while inventing fake doctor names, closed clinic hours, and scheduling rules that do not exist.

Finally, managing this infrastructure introduces a fragile staffing dependency. It requires advanced MLOps (Machine Learning Operations) skills that standard software developers do not have. If you hire a lone specialist to manage this cluster and they leave the company, you are left holding millions of shillings in complex, stranded hardware that nobody else on your team knows how to maintain or patch.

Where this option actually wins: Absolute data residency. If your legal team requires that patient data physically never leaves your building, this is the only path where compliance is structurally guaranteed rather than contractually negotiated.

2. Cloud-Hosted AI APIs: The Global Benchmark Blindspot

This path separates the conversational “brain” from the hardware headache. You plug into ready-made web endpoints and pay for usage instead of physical servers. Since most software engineering teams already know how to work with APIs, this is by far the easiest path for internal developers to build and test.

However, the risks don’t disappear; they just change shape:

  • Unpredictable Billing — Usage-based token models make budgeting incredibly volatile at scale. If the AI gets stuck in a “retry storm,” loops indefinitely with a confused patient, or processes long, malformed audio streams, your monthly cloud bill can spike drastically without warning.
  • The Accent Barrier — You are entirely at the mercy of global foundational models. These AI brains are trained and benchmarked almost exclusively on standard American and British English. Their accuracy on East African accents, Swahili medical terms, and conversational Sheng is completely unproven. Because you do not own the underlying code, you have zero power to tweak or tune the model when it consistently misunderstands local patients.
  • Network Fragility — On local mobile networks, packet loss and dropped connections are a daily reality—especially for patients calling from rural facility catchments. If a patient’s internet link stutters mid-call to an overseas API endpoint, the system frequently loses the entire conversation state. Without complex, custom-built state management software, the patient is forced to start their call completely over from the beginning.

3. The Third-Party Platform: The Budgeting Illusion

This is the all-in-one product pitch that dominates most enterprise sales meetings. Vendors promise a clean, predictable flat software fee—usually quoted around KES 2.2 million to KES 2.7 million per month.

In practice, this fixed fee is a budgeting illusion because it excludes the actual cost of operationalizing the platform:

  • Hidden Telecom Fees — The software licence does not include your carrier costs. Company X remains entirely responsible for paying the provider for the actual SIP trunks and per-minute voice traffic.
  • The Engineering Burden — Your internal development team doesn’t escape the workload. They still have to build, monitor, and maintain custom middleware connectors to keep the vendor’s closed API in sync whenever your internal database schemas change.
  • The Double Payroll — Because local accent accuracy is completely unvalidated, you cannot simply fire your support staff. You have to keep a human team active on standby to handle inevitable automation failures and escalations. Ultimately, you end up paying the massive vendor fee plus your existing human payroll.

There is also a severe legal liability trap. Vendors will happily display compliance certifications like SOC2, HIPAA, or ISO. However, these are Western regulatory frameworks that hold zero legal standing under the Kenya Data Protection Act (2019). Furthermore, the vendor’s enterprise contract will include strict liability limitation clauses. If a system bug leaks patient data or causes a cross-border breach, the Office of the Data Protection Commissioner (ODPC) will hold Company X 100% legally and reputationally accountable—not the software vendor.

Once your routing rules, conversation logic, and system workflows are fully embedded into a vendor’s proprietary ecosystem, walking away becomes incredibly difficult. This option carries the highest vendor lock-in of all four paths.

4. Keeping the Humans: Our Strongest Baseline

It is incredibly easy for technology teams to look at a traditional, human-staffed call center and dismiss it as the “boring, legacy” option. That is a major strategic error. Your human support team is a highly optimized, predictable, and risk-managed business asset.

Humans provide flawless, native comprehension of local accents, immediate contextual understanding of code-switching, and effortless navigation of casual Sheng or Swahili phrasing right out of the box. You don’t need a single byte of training data or a months-long engineering pipeline for them to understand a patient perfectly.

Financially, human support features a linear, predictable cost structure built entirely around standard salaries, making it the easiest model to budget and forecast.

Most importantly, human errors are logical and explainable. If a human agent is tired or suffers a lapse in communication, they might enter a typo or mishear a date. But if a human does not know the answer to a complex medical question, they will stop and say, “I’m a scheduler, let me connect you to a triage nurse.” They will never suffer from “hallucinations”—they will never confidently fabricate entirely false clinical data, closed clinic hours, or non-existent doctors the way a generative language model will.

The Real Limitation: The only genuine constraint of a human team is scale and availability. Providing 24/7 coverage requires clear hiring timelines and explicit night-shift premium wages, and you cannot instantly double your staff capacity overnight during a sudden traffic spike.

The Breakthrough Strategy: Split by Channel, Not by Technology

The turning point in our evaluation came when we realized we were asking the wrong question. We were looking for one technology to replace everything. Instead, we found a clear strategic boundary: voice and text are fundamentally different problems.

Phone calls combine every high-risk operational vulnerability at once. They suffer from accent ambiguity, mobile network fragility, and the highest regulatory stakes, because processing raw voice files means handling sensitive biometric data across borders.

Text-based messaging channels (like WhatsApp, SMS, and web portal queries) completely sidestep these issues. There is no accent to mishear, no background noise, and no audio to lose mid-call. The data footprint is incredibly clean, and modern language models are genuinely excellent at parsing written Swahili and English text.

This realization led us directly to a highly efficient, defensible Channel-Split Strategy:

                      ┌─── [ Patient Contacts Company X ] ───┐
                      │                                        │
                      ▼                                        ▼
             [ INCOMING VOICE CALL ]                   [ TEXT-BASED CHAT ]
                      │                                        │
                      ▼                                        ▼
             [ Human Support Agent ]                  [ Cloud AI Engine ]
             • Handles local accents natively         • Direct text parsing
             • Zero network fragility risk            • Multi-lingual translation
             • High clinical empathy                  • "Suggest, Never Modify" guardrails
  • Keep Phone Calls 100% Human-Driven — We preserve our human team where the compounding risks are highest and where human agents hold an uncontested linguistic and cultural advantage.
  • Automate Digital Chats via Cloud AI — We route written text queries through our existing Azure OpenAI integration layer. This avoids all local server hardware costs, bypasses the speech-to-text accent barrier, requires zero specialized hiring, and gives us an immediate path to a live pilot using infrastructure our developers already run every day.
  • Build a Reliable Escape Hatch — The AI’s job is to handle the routine, high-volume 80% of text queries (like clinic hours or basic appointment reminders). The moment a conversation becomes complex, clinical, or frustrating, the system must instantly and seamlessly route the live chat ticket to a human agent.

Pick one — your choice is public to other readers

Notes from the Arcnull workbench.

Engineering notes and release news, sent when there's something worth sending. No cadence, no sales sequence.

Arcnull, 2026