Governance and monitoring
AI you can put in front of a regulator.
A shared quality operating model: production monitoring, human review, incident response, and regression learning, all built to align with ISO/IEC 42001, POPIA and GDPR.
Live Traffic
Real user interactions across web, mobile, and API.
Sampling Policy
Intelligent sampling for value, risk, and coverage.
Async Evaluation
Score quality, safety, and policy compliance.
Alerting & Watchlists
Detect issues early and surface what matters.
Analyst Review Queue
Prioritized queue for human review.
Failure Clusters
Group similar issues to find recurring root causes.
Top Clusters
- Incorrect Tool Choice42%
- Missing Context21%
- Policy Violation15%
Recurring Failure Pattern
Continuous Improvement Loop
Regression Updates
Turn findings into tests that prevent regressions.
Prompt / Knowledge / Policy Improvements
Apply targeted fixes across prompts, knowledge & tools.
Release Revalidation
Re-run evaluations and gates to validate improvements.
Outcome
Monitor
Observe production health
Investigate
Find issues that matter
Improve
Fix root causes
Re-test
Validate with confidence
Release
Ship quality, safer AI
Live Traffic
Real user interactions across web, mobile, and API.
Sampling Policy
Intelligent sampling for value, risk, and coverage.
Async Evaluation
Score quality, safety, and policy compliance.
Alerting & Watchlists
Detect issues early and surface what matters.
Analyst Review Queue
Prioritized queue for human review.
Failure Clusters
Group similar issues to find recurring root causes.
Top Clusters
- Incorrect Tool Choice42%
- Missing Context21%
- Policy Violation15%
Recurring Failure Pattern
Outcome
Monitor
Observe production health
Investigate
Find issues that matter
Improve
Fix root causes
Re-test
Validate with confidence
Release
Ship quality, safer AI
Release Revalidation
Re-run evaluations and gates to validate improvements.
Prompt / Knowledge / Policy Improvements
Apply targeted fixes across prompts, knowledge & tools.
Regression Updates
Turn findings into tests that prevent regressions.
The operating model
One quality brain, six linked capabilities.
Test evidence, production quality signals, human review and regression learning all feed a single quality intelligence layer. DeepEval defines the evaluations; the observability estate stays vendor-neutral and portable.
01
Golden scenario library
Versioned journeys, expected outcomes, rules, evidence and review notes.
02
CI/CD quality gates
Release decisions based on quality, safety, tool, workflow and regression thresholds.
03
Production monitoring
Sampled traces, online evaluations, risk alerts, quality trends and cost signals.
Quality Intelligence Layer
One shared system for test evidence, production quality signals, human review and regression learning.
04
Human review queue
Low-confidence, risky, disputed or high-value interactions reviewed by accountable owners.
05
Regression memory
Confirmed failures are converted into permanent test coverage and release protection.
06
Executive reporting
Quality scorecards, risk movement, release readiness and improvement backlog.
DeepEval defines evaluations; observability remains vendor-neutral.
The production loop
Real traffic becomes release protection.
Production traces are sampled by risk, scored asynchronously, triaged, and reviewed by accountable humans. Confirmed issues turn into permanent regression coverage, so the system gets safer every week it runs.
What happens to a production trace
- 01
Sample traces
risk-based and random
- 02
Online evaluation
quality, safety, tools, RAG
- 03
Risk triage
severity and ownership
- 04
Human review
annotation and expected behaviour
- 05
Regression update
new permanent coverage
What we are watching for
- Conversation risk
- Poor tone, unresolved intent, weak refusal, repeated correction.
- Grounding risk
- Unsupported claim, stale context, missing source evidence.
- Tool risk
- Wrong tool, failed action, unsafe input, approval bypass.
- Operational risk
- Latency, errors, cost spike, queue pressure, transfer failures.
The signal stack
Every layer of the system is visible.
A layered metric model, from conversation quality down to platform cost, feeds separate dashboards for engineering, product, compliance and leadership. No single generic score hides what is really happening.
Deterministic Checks
Hard rules & guardrails
LLM Judge Signals
Model-as-a-judge evaluations
Retrieval Signals
Quality of retrieved context
Tool & Workflow Signals
Actions & orchestration quality
Experience Signals
Perceived experience quality
Business Outcome Signals
Customer & business impact
Operational Telemetry
System health & observability
Evidence
Deployed May 15, 2025
Assessed against our ISO/IEC 42001-aligned assurance framework
ISO/IEC 42001
AI management system
POPIA
South Africa
GDPR
EU and UK
NIST AI RMF
Risk framework
Common questions
Questions we get asked about AI governance.
What does AI governance mean in practice?
A shared operating model rather than a policy document: production monitoring, risk-based sampling of real traces, scored evaluation, human review by accountable people, incident response, and confirmed failures turned into permanent regression coverage. It is built to align with ISO/IEC 42001, POPIA and GDPR.
Is NPU Labs a certification body?
No, and we are careful about this. We are not accredited to certify anyone against ISO/IEC 42001. We build and assess AI systems to align with it, and we produce the evidence a real auditor or regulator would ask for. Certification, if you want it, comes from an accredited body.
How do you monitor an AI system in production?
Live traffic is sampled by value, risk and coverage rather than at a flat rate, with PII protection applied at capture. Sampled traces are scored asynchronously, grouped into failure clusters to find recurring root causes, and routed to human reviewers where the impact justifies it.
What data do you need access to?
Sampled production traces, and only what the evaluation needs. Sampling is policy-driven with PII protection on by default, and the observability estate stays vendor-neutral and portable, so the evidence remains yours and is not locked into a tool we chose.
How does human review fit in?
It is the point, not a fallback. High-impact and high-risk decisions route to a named human, reviewers are calibrated against each other so the bar does not drift, and every confirmed issue becomes a regression test. The system gets safer every week it runs because people keep deciding what safer means.
Which standards and regulations does this align with?
ISO/IEC 42001 for AI management, POPIA for South Africa, GDPR for the EU and UK, and the NIST AI Risk Management Framework. We treat these as the assurance framework we assess against, and the monitoring and review evidence is what demonstrates it.
Not the question you came with?
Ask us about your setupFind out what your AI is actually doing in production.
Sampled trace evaluation, quality signals your team can act on, and a human review loop that turns confirmed failures into permanent coverage. Set up once, run every week.