NPU LabsNPU LABS

Governance and monitoring

AI you can put in front of a regulator.

A shared quality operating model: production monitoring, human review, incident response, and regression learning, all built to align with ISO/IEC 42001, POPIA and GDPR.

Live Traffic

Real user interactions across web, mobile, and API.

24.6kreq/min
1.2kusers/min
98.7%success

Sampling Policy

Intelligent sampling for value, risk, and coverage.

10%overall sample rate
Risk-based boostOn
PII protectionOn
Sampled Traces1.2k

Async Evaluation

Score quality, safety, and policy compliance.

Quality Score0-100
Safety Score0-100
Policy Checks12
Evaluated1.2k

Alerting & Watchlists

Detect issues early and surface what matters.

Score Drift
Risk Spikes
Route Anomalies
Customer Complaints17

Analyst Review Queue

Prioritized queue for human review.

High Risk23
Score Drift31
New Anomaly14
Pending Reviews68
Findings from the review queue

Failure Clusters

Group similar issues to find recurring root causes.

Top Clusters

  • Incorrect Tool Choice42%
  • Missing Context21%
  • Policy Violation15%

Recurring Failure Pattern

Continuous Improvement Loop

MONITORINVESTIGATEIMPROVERE-TEST
Clusters become regression coverage

Regression Updates

Turn findings into tests that prevent regressions.

New Golden Cases24
Updated Test Sets6
Guardrails Added3
Create Golden Case

Prompt / Knowledge / Policy Improvements

Apply targeted fixes across prompts, knowledge & tools.

Tune Prompts
Fix Retrieval
Update Tool Rules
Changes Deployed

Release Revalidation

Re-run evaluations and gates to validate improvements.

Rerun Gates
Quality GatePass
Safety GatePass
Ready for Release
What the loop delivers

Outcome

  • Monitor

    Observe production health

  • Investigate

    Find issues that matter

  • Improve

    Fix root causes

  • Re-test

    Validate with confidence

  • Release

    Ship quality, safer AI

Forward Flow
Feedback Loop
Revalidation Loop

The operating model

One quality brain, six linked capabilities.

Test evidence, production quality signals, human review and regression learning all feed a single quality intelligence layer. DeepEval defines the evaluations; the observability estate stays vendor-neutral and portable.

01

Golden scenario library

Versioned journeys, expected outcomes, rules, evidence and review notes.

02

CI/CD quality gates

Release decisions based on quality, safety, tool, workflow and regression thresholds.

03

Production monitoring

Sampled traces, online evaluations, risk alerts, quality trends and cost signals.

Quality Intelligence Layer

One shared system for test evidence, production quality signals, human review and regression learning.

04

Human review queue

Low-confidence, risky, disputed or high-value interactions reviewed by accountable owners.

05

Regression memory

Confirmed failures are converted into permanent test coverage and release protection.

06

Executive reporting

Quality scorecards, risk movement, release readiness and improvement backlog.

DeepEval defines evaluations; observability remains vendor-neutral.

The production loop

Real traffic becomes release protection.

Production traces are sampled by risk, scored asynchronously, triaged, and reviewed by accountable humans. Confirmed issues turn into permanent regression coverage, so the system gets safer every week it runs.

What happens to a production trace

  1. 01

    Sample traces

    risk-based and random

  2. 02

    Online evaluation

    quality, safety, tools, RAG

  3. 03

    Risk triage

    severity and ownership

  4. 04

    Human review

    annotation and expected behaviour

  5. 05

    Regression update

    new permanent coverage

What we are watching for

Conversation risk
Poor tone, unresolved intent, weak refusal, repeated correction.
Grounding risk
Unsupported claim, stale context, missing source evidence.
Tool risk
Wrong tool, failed action, unsafe input, approval bypass.
Operational risk
Latency, errors, cost spike, queue pressure, transfer failures.

The signal stack

Every layer of the system is visible.

A layered metric model, from conversation quality down to platform cost, feeds separate dashboards for engineering, product, compliance and leadership. No single generic score hides what is really happening.

1
Deterministic Checks

Hard rules & guardrails

Schema validityMandatory disclosurePolicy blocks
2
LLM Judge Signals

Model-as-a-judge evaluations

RelevancyFaithfulnessClarityHelpfulness
3
Retrieval Signals

Quality of retrieved context

Context precisionContext recallCitation support
4
Tool & Workflow Signals

Actions & orchestration quality

Tool correctnessTool sequenceWorkflow completion
5
Experience Signals

Perceived experience quality

ToneLatency perceptionHandover quality
6
Business Outcome Signals

Customer & business impact

ResolutionDeflectionEscalation rateCustomer outcome
7
Operational Telemetry

System health & observability

LatencyError rateTrace completenessDrift alerts

Evidence

Trace
trace_9f3a...e8fe12m ago
Scores
Overall92.4
Quality 95.1Safety 91.0Grounding 92.7
Reviewer
Anita C.May 16, 2025 · 10:42 AM
Dataset
Eval Set · Support v2.31,248 samples
Version
Prod

Deployed May 15, 2025

Thresholds
Quality ≥ 90Safety ≥ 90Grounding ≥ 85

Assessed against our ISO/IEC 42001-aligned assurance framework

ISO/IEC 42001

AI management system

POPIA

South Africa

GDPR

EU and UK

NIST AI RMF

Risk framework

Common questions

Questions we get asked about AI governance.

What does AI governance mean in practice?

A shared operating model rather than a policy document: production monitoring, risk-based sampling of real traces, scored evaluation, human review by accountable people, incident response, and confirmed failures turned into permanent regression coverage. It is built to align with ISO/IEC 42001, POPIA and GDPR.

Is NPU Labs a certification body?

No, and we are careful about this. We are not accredited to certify anyone against ISO/IEC 42001. We build and assess AI systems to align with it, and we produce the evidence a real auditor or regulator would ask for. Certification, if you want it, comes from an accredited body.

How do you monitor an AI system in production?

Live traffic is sampled by value, risk and coverage rather than at a flat rate, with PII protection applied at capture. Sampled traces are scored asynchronously, grouped into failure clusters to find recurring root causes, and routed to human reviewers where the impact justifies it.

What data do you need access to?

Sampled production traces, and only what the evaluation needs. Sampling is policy-driven with PII protection on by default, and the observability estate stays vendor-neutral and portable, so the evidence remains yours and is not locked into a tool we chose.

How does human review fit in?

It is the point, not a fallback. High-impact and high-risk decisions route to a named human, reviewers are calibrated against each other so the bar does not drift, and every confirmed issue becomes a regression test. The system gets safer every week it runs because people keep deciding what safer means.

Which standards and regulations does this align with?

ISO/IEC 42001 for AI management, POPIA for South Africa, GDPR for the EU and UK, and the NIST AI Risk Management Framework. We treat these as the assurance framework we assess against, and the monitoring and review evidence is what demonstrates it.

Not the question you came with?

Ask us about your setup

Find out what your AI is actually doing in production.

Sampled trace evaluation, quality signals your team can act on, and a human review loop that turns confirmed failures into permanent coverage. Set up once, run every week.