NPU LabsNPU LABS
Two NPU Labs engineers working together at a home desk in Cape Town, a private AI computer beside the laptop, Table Mountain through the window.

Who we are

A family-built AI lab in Cape Town.

NPU Labs is a small, family-built lab. We are engineers first, and we build private AI the way we would want it built for our own business. Years of designing, building, deploying, and maintaining large enterprise systems in production gave us the engineering discipline that AI work runs on.

The people in the room

AI engineers

Agents, RAG, fine-tuning, evaluation. We ship AI that holds up in production, long past the demo.

AI architects

The shape of the whole system: models, retrieval, tools, guardrails, and how they hold together under load.

Senior developers

Hands on the keyboard, and on the systems long after launch. We build software we would maintain ourselves.

Software architects

Systems designed to last and to change. The difference between a prototype and a platform a business depends on.

Data engineers

Pipelines, evaluation, and the unglamorous plumbing that decides whether an AI system actually works.

We are a lab

Grounded in research. Sharpened in production.

NPU Labs began with years of building and running the large systems businesses depend on. We moved into AI the way a lab would: reading the research, testing it against reality, and keeping what survived production. What we learned became practice.

01

Grounded in research.

We build on published work and cite it. The papers behind our claims sit on the product pages, because a lab shows its sources.

02

Proven in production.

Years running real AI and enterprise systems, at real cost. We build for the failure modes we have seen and had to fix ourselves.

03

Written into practice.

A defined way of working, golden scenarios, and quality gates that turn hard-won lessons into a repeatable standard for the next build.

What we focus on now

AI, system modernization, and private AI stacks.

We moved our focus to AI from the inside out. We have paid the cloud bills, operated the MLOps lifecycle, and learned what it really takes to get model inference correct, fast, and observable. That experience is why our AI work performs in production.

We have paid the cloud bills

We run AI agents and LLMs in production at roughly US$7,800 a month, so we know exactly where the money goes and how owning the stack changes the economics.

We have run MLOps in production

Prepare, fine-tune, evaluate, deploy, monitor, retrain. We run the full lifecycle and build for the day after the model goes live.

Correct, observable inference

Throughput, latency, hallucinations, drift. All of it on a dashboard, in front of the team who can act on it. Getting this right is most of the work.

NPU Labs engineers working together over a laptop in a Cape Town office, Table Mountain through the window.

Cape Town, South Africa

Built here, for businesses here.

South African businesses carry their own rules, and we build to fit them: compute you own, data that stays where the law expects it, and the people who built it in the same timezone as you.

Your data stays in the country

Private nodes run in your environment, so POPIA data residency is a property of the architecture itself.

Hardware supplied and supported here

Nodes are sourced, configured, burned in, and installed locally, so support is a call to the people who built it.

Engineers in your working day

The people who built the stack are reachable in your hours, and they stay on it after handover.

How we work with you

The same method, every engagement.

Every build is different. The path through them is the same one we have run on our own systems. Four steps, each with a person answerable for it.

01

Assess

We start with your systems, costs, and constraints, then map where AI genuinely earns its place and what to sequence first.

02

Build

Private compute and a lab on hardware you own, with the models and tooling that fit your data-residency rules and your budget.

03

Ship

Fine-tuning, evaluation, guardrails, and observability, built for the day after launch, with a named owner for every decision.

04

Hand over

A documented way of working so your team owns it, or a managed arrangement where we keep it healthy. Your call, either way.

Tell us where you are, and we will tell you straight whether we are the right people to build it.

Talk to us

Want to talk to the people who would build it?

Tell us what you are trying to run. We will come back with what it takes, what it costs, and a straight answer on whether we are the right people to build it.