
Who we are
A family-built AI lab in Cape Town.
NPU Labs is a small, family-built lab. We are engineers first, and we build private AI the way we would want it built for our own business. Years of designing, building, deploying, and maintaining large enterprise systems in production gave us the engineering discipline that AI work runs on.
The people in the room
AI engineers
Agents, RAG, fine-tuning, evaluation. We ship AI that holds up in production, long past the demo.
AI architects
The shape of the whole system: models, retrieval, tools, guardrails, and how they hold together under load.
Senior developers
Hands on the keyboard, and on the systems long after launch. We build software we would maintain ourselves.
Software architects
Systems designed to last and to change. The difference between a prototype and a platform a business depends on.
Data engineers
Pipelines, evaluation, and the unglamorous plumbing that decides whether an AI system actually works.
We are a lab
Grounded in research. Sharpened in production.
NPU Labs began with years of building and running the large systems businesses depend on. We moved into AI the way a lab would: reading the research, testing it against reality, and keeping what survived production. What we learned became practice.
Grounded in research.
We build on published work and cite it. The papers behind our claims sit on the product pages, because a lab shows its sources.
Proven in production.
Years running real AI and enterprise systems, at real cost. We build for the failure modes we have seen and had to fix ourselves.
Written into practice.
A defined way of working, golden scenarios, and quality gates that turn hard-won lessons into a repeatable standard for the next build.
What we do
Four ways we work with a business.
Most clients start with one and grow into the others. The thread through all four is the same: the stack runs in your environment, and your team ends up able to run it.
Private AI infrastructure
NVIDIA DGX Spark nodes and AI workstations, supplied, configured, and supported in South Africa. Fixed-cost compute you own, sized to the workload.
See the infrastructureAI engineering services
Agents, retrieval, fine-tuning, evaluation, and system modernization. We design the stack, build it, and stay on it after it goes live.
See what we buildAI guardrails and governance
AI usage policy turned into controls that run in the request path, built to align with ISO/IEC 42001 and what POPIA and GDPR require.
See the guardrailsProducts we build and run
Software we use ourselves before we sell it: relationship intelligence, customer agents, and the way of working behind how we deliver.
See the productsWhat we focus on now
AI, system modernization, and private AI stacks.
We moved our focus to AI from the inside out. We have paid the cloud bills, operated the MLOps lifecycle, and learned what it really takes to get model inference correct, fast, and observable. That experience is why our AI work performs in production.
We have paid the cloud bills
We run AI agents and LLMs in production at roughly US$7,800 a month, so we know exactly where the money goes and how owning the stack changes the economics.
We have run MLOps in production
Prepare, fine-tune, evaluate, deploy, monitor, retrain. We run the full lifecycle and build for the day after the model goes live.
Correct, observable inference
Throughput, latency, hallucinations, drift. All of it on a dashboard, in front of the team who can act on it. Getting this right is most of the work.

Cape Town, South Africa
Built here, for businesses here.
South African businesses carry their own rules, and we build to fit them: compute you own, data that stays where the law expects it, and the people who built it in the same timezone as you.
Your data stays in the country
Private nodes run in your environment, so POPIA data residency is a property of the architecture itself.
Hardware supplied and supported here
Nodes are sourced, configured, burned in, and installed locally, so support is a call to the people who built it.
Engineers in your working day
The people who built the stack are reachable in your hours, and they stay on it after handover.
How we work with you
The same method, every engagement.
Every build is different. The path through them is the same one we have run on our own systems. Four steps, each with a person answerable for it.
Assess
We start with your systems, costs, and constraints, then map where AI genuinely earns its place and what to sequence first.
Build
Private compute and a lab on hardware you own, with the models and tooling that fit your data-residency rules and your budget.
Ship
Fine-tuning, evaluation, guardrails, and observability, built for the day after launch, with a named owner for every decision.
Hand over
A documented way of working so your team owns it, or a managed arrangement where we keep it healthy. Your call, either way.
Tell us where you are, and we will tell you straight whether we are the right people to build it.
Talk to usWant to talk to the people who would build it?
Tell us what you are trying to run. We will come back with what it takes, what it costs, and a straight answer on whether we are the right people to build it.