
NVIDIA
DGX Spark
Premium build, NVIDIA-supported OS, the gold-standard option.
- Memory
- 128 GB unified
- Chip
- GB10 Grace Blackwell
- Storage
- 4 TB NVMe
- Draw
- ~240W
Hardware we build on
The cluster architecture is the same across all four. Pick by build quality, support story, and budget. We provision, configure, and deploy on whichever you choose.

NVIDIA
Premium build, NVIDIA-supported OS, the gold-standard option.
Private AI Stack
How much private compute you need is set by the inference throughput you run, not a fixed product tier. A single node sits on a desk. Four nodes make a 512 GB cluster on one switch. Past that, clusters scale out for as much aggregate throughput as the workload demands.

Four nodes pooled into a single 512 GB cluster, sized for a department's day-to-day experimentation and inference.
512 GB
Unified memory
pooled across 4 nodes
~400B-class
Model size
pooled inference
4× 200G
Switched fabric
any-to-any
~960W
Total draw
standard wall circuit
Topology
Powered from a standard 10A wall circuit. 10GbE management per node over your existing network. No dedicated power or networking infrastructure required.
In the cluster
128 GB unified memory and up to 4 TB NVMe per node. Four nodes pooled into one cluster for distributed serving of a single large model.
Four 400G QSFP56-DD ports run at 200G to match each node's NIC. One per node, fully populated for a cluster. A second cluster adds its own switch and routes across.
One identical short-run passive-copper cable per node, switch to node NIC. Nothing exotic to source.
Indicative figures only. They include a safety margin and will move with supplier pricing, exchange rates, and import costs. Final pricing is confirmed on a written quote. As an all-in example, a single-node deployment lands around R0.4M to R1.2M depending on the node SKU you pick: ~R260k-R900k hardware plus the 13 to 20 day setup below. Final scope depends on your throughput targets, identity provider, and compliance environment.
Model strategy
A private AI platform runs a portfolio of model routing, RAG, guardrails, observability, and workload-specific endpoints. A cluster is four nodes; capacity scales from a single node to a multi-cluster fabric, with the model chosen to fit the workload.
1 node · 128 GB · up to ~200B
Embeddings, reranking, private chatbot, summarisation, policy Q&A, voice pre/post-processing.
4 nodes · 512 GB · 405B-class
Private enterprise assistant, developer guardrails, repo analysis, agentic coding, RAG, multimodal document understanding, model bake-offs.
8+ nodes · distributed inference
Frontier coding agents, long-horizon autonomous workflows, 1M-token reasoning, multi-agent orchestration, regulated high-value workloads.
A single node typically handles models up to ~200B, and a pooled four-node cluster reaches ~405B-class. Larger frontier models run via NVFP4 variants, sharding, and multi-cluster serving, all validated per workload before client production.
See the full model registryServing architecture
No model is reached directly. Requests enter through one gateway, clear policy and guardrails, then route to the endpoint and node sized for the workload, with audit and observability on every path.
Workload endpoints
Node pool
Tier 1-2 workloads
Multi-cluster fabric
Frontier models
Professional services
The hardware arrives configured at the DGX OS level. From there, six configuration steps turn it into a governed, observable platform wired into your network. Roughly 13 to 20 days for a standard stack.
Prometheus and Grafana across the hardware, inference, and model-cost layers, tracking utilization, TTFT, latency, and per-team token spend, with alerts.
Built on DeepEval's cost and efficiency metrics, tracking spend and tokens per user and per task, with insights into which models complete the work economically and which burn budget.
An API gateway in front of every endpoint, per-team keys with rotation, LDAP/AD or SSO integration, and TLS everywhere.
A central registry of models, versions, and quantisations, with staged promotion to production, one-step rollback, and provenance for every deployed weight.
PII detection and redaction pre- and post-model, prompt-injection detection, output filtering, tool allow-lists, and red-team testing.
The lab wired into your corporate network over site-to-site or client VPN. Private endpoints only, firewall rules scoped per team, nothing exposed to the public internet.
Deployment
Hardware is purchased from a trusted vendor.
All hardware ordered at once
Compute nodes, switch and cables.
Begins on hardware arrival
Six workstreams, ~13-20 days, run on-site once everything is racked.
Working, governed, observed
Total order-to-operational for a standard deployment on an existing network.
Start with a conversation. We will work through your workloads, the data-residency constraints, and the budget envelope, and come back with a concrete spec.