
NVIDIA
DGX Spark
Premium build, NVIDIA-supported OS, the gold-standard option.
- Memory
- 128 GB unified
- Chip
- GB10 Grace Blackwell
- Storage
- 4 TB NVMe
- Draw
- ~240W
We provision, configure, and hand it over running.
Hardware we build on
The cluster architecture is the same across all four. Pick by build quality, support story, and budget. We provision, configure, and deploy on whichever you choose.

NVIDIA
Premium build, NVIDIA-supported OS, the gold-standard option.
We provision, configure, and hand it over running.
What runs on the node
NVIDIA
Ubuntu
Docker
Kubernetes
Python
PyTorch
Hugging Face
Ollama
Qdrant
PostgreSQL
LangChain
Grafana
Prometheus
Open source where it counts, so the lab stays portable and your team keeps the skills they already have.
Private AI Stack
How much private compute you need is set by the inference throughput you run, not a fixed product tier. A single node sits on a desk. Four nodes make a 512 GB cluster on one switch. Past that, clusters scale out for as much aggregate throughput as the workload demands.

Four nodes pooled into a single 512 GB cluster, sized for a department's day-to-day experimentation and inference.
512 GB
Unified memory
pooled across 4 nodes
~400B-class
Model size
pooled inference
4× 200G
Switched fabric
any-to-any
~960W
Total draw
standard wall circuit
Topology
Powered from a standard 10A wall circuit. 10GbE management per node over your existing network. No dedicated power or networking infrastructure required.
In the cluster
128 GB unified memory and up to 4 TB NVMe per node. Four nodes pooled into one cluster for distributed serving of a single large model.
Four 400G QSFP56-DD ports run at 200G to match each node's NIC. One per node, fully populated for a cluster. A second cluster adds its own switch and routes across.
One identical short-run passive-copper cable per node, switch to node NIC. Nothing exotic to source.
Indicative figures only. They include a safety margin and will move with supplier pricing, exchange rates, and import costs. Final pricing is confirmed on a written quote. As an all-in example, a single-node deployment lands around R0.4M to R1.2M depending on the node SKU you pick: ~R260k-R900k hardware plus the 13 to 20 day setup below. Final scope depends on your throughput targets, identity provider, and compliance environment.
Model strategy
A private AI platform runs a portfolio of model routing, RAG, guardrails, observability, and workload-specific endpoints. A cluster is four nodes; capacity scales from a single node to a multi-cluster fabric, with the model chosen to fit the workload.
1 node · 128 GB · up to ~200B
Embeddings, reranking, private chatbot, summarisation, policy Q&A, voice pre/post-processing.
4 nodes · 512 GB · 405B-class
Private enterprise assistant, developer guardrails, repo analysis, agentic coding, RAG, multimodal document understanding, model bake-offs.
8+ nodes · distributed inference
Frontier coding agents, long-horizon autonomous workflows, 1M-token reasoning, multi-agent orchestration, regulated high-value workloads.
A single node typically handles models up to ~200B, and a pooled four-node cluster reaches ~405B-class. Larger frontier models run via NVFP4 variants, sharding, and multi-cluster serving, all validated per workload before client production.
See the full model registryServing architecture
Every model sits behind the same gateway. Requests clear policy and guardrails, then route to the endpoint and node sized for the workload, with audit and observability on every path.
Workload endpoints
Node pool
Tier 1-2 workloads
Multi-cluster fabric
Frontier models
Professional services
The hardware arrives configured at the DGX OS level. From there, six configuration steps turn it into a governed, observable platform wired into your network. Roughly 13 to 20 days for a standard stack.
Prometheus and Grafana across the hardware, inference, and model-cost layers, tracking utilization, TTFT, latency, and per-team token spend, with alerts.
Built on DeepEval's cost and efficiency metrics, tracking spend and tokens per user and per task, with insights into which models complete the work economically and which burn budget.
An API gateway in front of every endpoint, per-team keys with rotation, LDAP/AD or SSO integration, and TLS everywhere.
A central registry of models, versions, and quantisations, with staged promotion to production, one-step rollback, and provenance for every deployed weight.
PII detection and redaction pre- and post-model, prompt-injection detection, output filtering, tool allow-lists, and red-team testing.
The lab wired into your corporate network over site-to-site or client VPN. Private endpoints only, firewall rules scoped per team, nothing exposed to the public internet.
Deployment
Hardware is purchased from a trusted vendor, configured and burned in at our lab, then installed and handed over on-site. Three phases, with a person answerable at every handoff.
All hardware ordered at once
Compute nodes, switch and cables.
Begins on hardware arrival
Six workstreams, ~13-20 days, run on-site once everything is racked.
Working, governed, observed
Total order-to-operational for a standard deployment on an existing network.
Common questions
It depends on the workload, so we size it before we quote. A single NVIDIA DGX Spark class node suits a team running inference, embeddings, and fine-tuning up to 70B. Larger stacks add nodes. The point of owning it is a fixed capital cost and an electricity bill, instead of a monthly cloud bill that grows with usage.
Four and a half to eight weeks. Hardware is purchased from a trusted vendor, configured and burned in at our lab, then installed and handed over on-site. Stack configuration on top of that is roughly 13 to 20 days for a standard setup.
Yes. Nodes are sourced, configured, burned in, and installed locally, and they run in your building or your data centre. Your data stays where POPIA expects it because the compute never leaves.
You have a support agreement with us and the people who built your stack are the ones who answer. Hardware is covered by the vendor warranty, and we carry the configuration, the runtime, and the models on top of it.
Yes. A workstation-class box is enough to prove a workload, and the stack we install on it is the same one that runs on a cluster. Teams often start with one unit and add nodes once the workload is known.
The stack lands configured and operational, with monitoring, model registry, guardrails, and authentication already wired in. We hand over a documented way of working, and we stay on a support arrangement for as long as it is useful.
Not the question you came with?
Ask us about your setupStart with a conversation. We will work through your workloads, the data-residency constraints, and the budget envelope, and come back with a concrete spec.