NPU LabsNPU LABS

Testing and evaluation, as a service

Proving an AI system works. Then proving it again.

Anyone can demo an agent once. We build the repeatable quality system that catches regressions and turns every production failure into permanent test coverage.

Change Events

Prompt
Model
Knowledge Base
Retriever
Tool
Workflow
Policy

Golden Scenarios

TC-001Customer Refundv3.2

Verify refund policy answer

Expected: Policy summary

LLM Judge2
TC-002Plan Comparisonv2.1

Compare features across plans

Expected: Table w/ key diff

LLM Judge2
TC-003Data Extractionv1.4

Extract order details from email

Expected: JSON with fields

LLM Judge2
Add Scenario

Evaluation Engine

Built-in Metrics
G-Eval
DAGMetric

Release Gates

Pass≥ 85
Warn60 – 84
Block< 60

Gate Policy

Production Quality

Production Sampling

Trace 9f3a...c81e2m ago

Plan comparison request

Sampled
Trace 7b21...a9d07m ago

Refund policy question

Sampled
Trace 3c44...e2f612m ago

Upgrade eligibility check

Sampled
Sampling Rate10%

Human Review

Trace 9f3a...c81eNeeds Work

Answer missed key plan limitations.

Reviewer

ACAnita C.1h ago
Trace 7b21...a9d0Looks Good

Clear and accurate.

Reviewer

JMJames M.2h ago
Add Annotation
And back around
Feeds Golden Scenarios

Regression Library

RS-017Plan comparison – edge casesv1.2
RS-012Refund policy – exceptionsv1.3
RS-009Data extraction – receiptsv1.1
Save as Regression Scenario
Feeds Release Gates

Continuous Improvement

Triage recurring issues
Improve prompts & knowledge
Update policies & tools
Closes the loop on Human Review

Audit & Evidence

VersionsPrompt v1.8 · Model gpt-4.1
ScoresOverall 92.4 · Gate Pass
Trace ID9f3a8c81e4b24619
ReviewerACAnita C.
Reviewed AtMay 16, 2025 · 10:42 AM
Forward Flow
Feedback Loop
Improvement Loop

End to end: from a prompt change to an audited, released version.

Release quality gates

Every meaningful change passes evaluation before it ships.

A prompt, model, retrieval, tool, policy or workflow change enters the pipeline, runs the full evaluation suite, and is held, risk-accepted with an owner, or approved. Nothing reaches production on a hunch.

  1. 01

    Change intake

    Prompt, model, RAG, workflow, tool, policy or data update.

  2. 02

    Evaluation suite

    Golden scenarios, RAG checks, tool checks, safety checks and conversation tests.

  3. 03

    Gate review

    Blocker, warning and advisory results reviewed against thresholds.

  4. 04

    Controlled release

    Approved version moves through staging, pilot or production rollout.

  5. 05

    Production quality

    Real traces are sampled, scored, trended and routed for review.

  6. 06

    Learning loop

    Confirmed issues update datasets, thresholds and backlog priorities.

Release hold

Blocker failures, severe regressions, approval bypasses or hard policy violations.

Risk accepted

Non-blocking warnings proceed only with a named owner and target date.

Release approved

Quality, safety, workflow and monitoring criteria are ready for rollout.

What we test

Nine test groups, each with a job.

We test the system, not just the prompt: retrieval, model output, tool use, workflow routing, policy enforcement, voice stages, and the final response. Every group is reported on its own, because a single blended score hides the one thing you needed to see.

Testing Portfolio Overview

Unified view of quality, risk, and readiness across the AI system.

Functional Tasks

Validate core capabilities and task success.

Success paths
Happy path coverage
Output correctness
Exceptions & fallbacks

Tests

412

Pass Rate

92%

Retrieval & Grounding

Ensure answers are grounded and well-cited.

Citation quality
Grounding accuracy
Hallucination risk
Source diversity

Tests

358

Pass Rate

88%

Tool Use

Verify safe and effective tool execution.

Tool permissions
Input validation
Error handling
Result verification

Tests

296

Pass Rate

90%

Policy & Guardrails

Confirm policy adherence and safe behavior.

Content safety
Escalation rules
Refusal accuracy
PII protection

Tests

275

Pass Rate

95%

Multi-turn Conversation

Evaluate context handling across turns.

Context retention
Clarification handling
Intent tracking
Long conversation

Tests

322

Pass Rate

91%

Voice Experience

Validate spoken interactions end-to-end.

STT accuracy
TTS quality
Interruptions
Latency & timing

Tests

268

Pass Rate

89%

Edge Cases & Adversarial

Stress the system with hard and risky scenarios.

Adversarial prompts
Jailbreak attempts
Abuse & safety cases
Robustness & limits

Tests

310

Pass Rate

87%

Human Review

Human-in-the-loop validation for high-impact areas.

Approved decisions
Risky responses
Reviewer notes
Dispute outcomes

Reviews

128

Agreement

94%

Regression Suite

Detect regressions and prevent quality drift.

Saved failures
Snapshot comparisons
Drift detection
Baseline stability

Tests

1,024

Pass Rate

93%

Release Readiness Summary

Overall quality signals and actionable next steps.

Coverage

91%

2,865 / 3,150 tests

Across 9 domains

Risk Mix

  • High18%
  • Medium37%
  • Low45%

Critical Gaps

  • Hallucination risk7
  • Tool permission gaps5
  • Adversarial jailbreaks4

View all gaps

Recommended Action

  • Address hallucination riskHigh
  • Tighten tool permissionsHigh
  • Expand adversarial coverageMedium

Create action plan

And the suites that cover them

Nine test suites, each running at its own point in the lifecycle: some on every commit, some only at the release gate, some continuously against production traffic.

Foundation tests

Always on

Prompt contract, schema validation, rules, policy, tool contracts and deterministic checks.

Golden scenario tests

Release gate

Known journeys with expected outcomes, evidence requirements, approval paths and failure examples.

Conversation simulation

Regression

Multi-turn journeys, memory behaviour, interruptions, ambiguous inputs and persona drift.

RAG and knowledge tests

High value

Retrieval quality, source grounding, freshness, conflicting content and unsupported claims.

Agent workflow tests

Critical

Routing, tool sequence, tool result use, exception handling and escalation behaviour.

Safety and privacy tests

Blocker

PII exposure, unsafe requests, refusal quality, policy coverage and jailbreak resistance.

Voice interaction tests

Voice

STT quality, VAD and EOU timing, barge-in, silence, TTS quality and transfer hand-off.

Load and resilience tests

Operational

Latency, concurrency, failures, retries, back pressure, cost and queue behaviour.

Production replay tests

Learning loop

Real traces promoted to offline tests after review, annotation and root-cause tagging.

The reusable IP

The golden scenario library.

This is the centre of repeatable AI testing, and it is a product asset, not a temporary QA file. Every important behaviour, happy-path, edge case, adversarial, policy, tool failure, voice issue, escalation, and every known production regression, lives here as a versioned, reusable scenario.

When we work with you, this library is yours. It compounds: every failure we find becomes coverage that protects you forever.

Each scenario captures

  • Scenario name
  • Journey type
  • User input
  • Expected outcome
  • Required evidence
  • Prohibited behaviour
  • Required tools
  • Escalation rule
  • Evaluation metrics
  • Owner
  • Version metadata

Scoring and routing

Every release decision is explained.

Weighted score rails, decision zones, and routing outcomes, with a policy override that forces a block or a human handover on any critical safety or compliance failure.

Score Rails

(Dimensions)

Quality

88

Safety

72

Retrieval

91

Tool Use

64

Policy

58

Business Criticality

76
Scores update continuously

Overall Score

Live
76Overall Score

Decision Zones

Below 60

RED

High Risk

Requires human review

60 – 84

AMBER

Moderate Risk

Proceed with caution

85 – 100

GREEN

Low Risk

Safe to automate

Weighted factors · Calibrated to your risk profile

Quality

25%

Safety

25%

Retrieval

15%

Tool Use

15%

Policy

10%

Business Criticality

10%

Routing Outcomes

Auto Approve

Release automatically

Score ≥ 85

All rails ≥ 60

Deploy with Watchlist

Monitor closely in production

Score 70 – 84

No critical risks

Require Human Review

Queue for expert review

Score 60 – 69

Or any rail < 60

Block Release

Do not release

Score < 60

Or critical failure

Mandatory Human Handover

Escalate before any action

High-risk or

Policy override

Human Handover Triggers

Any trigger can escalate

High-risk intent

Detected in prompt

Policy uncertainty

Ambiguous or unclear

Low confidence

Model confidence low

Conflicting retrieval

Sources disagree

Failed tool action

Tool error / exception

Sensitive customer impact

PII, financial, legal, safety

Repeated low score

Consistently below threshold

Policy Override

(Always On)

Critical safety or compliance failures bypass score averages and force Block or Mandatory Handover.

Examples

Harm / Safety violationPII exposure riskLegal / Regulatory breachProhibited contentSecurity vulnerability
Healthy
Monitor
High Risk
Escalated
All scores are explained · Versioned · AuditableLast updated just now

What we do for you

We run the quality system, or we assess yours.

  • Stand up a golden scenario library for your agents, treated as a product asset, not a throwaway QA file.
  • Wire release quality gates into your delivery so unsafe changes cannot ship.
  • Run production sampling and evaluation so real failures become permanent test coverage.
  • Assess an existing AI system against our ISO/IEC 42001-aligned assurance framework.

Assessed against our ISO/IEC 42001-aligned assurance framework

ISO/IEC 42001

AI management system

POPIA

South Africa

GDPR

EU and UK

NIST AI RMF

Risk framework

Common questions

Questions we get asked about AI testing.

What does AI testing and evaluation actually cover?

We test the system, not just the prompt: retrieval quality, model output, tool use, workflow routing, policy enforcement, voice stages and the final response. Each group is scored and reported on its own, because a single blended score hides the one thing you needed to see.

How is this different from normal software QA?

An AI system is not deterministic, so a pass or fail assertion does not survive contact with it. We score behaviour against graded expectations using a golden scenario library, run every meaningful change through the same suite, and use weighted scoring with explicit decision zones instead of a binary result.

What is a golden scenario library?

A versioned, reusable set of the behaviours that matter to your business: happy path, edge case, adversarial, policy, tool failure, voice and escalation cases, plus every production regression we have ever confirmed. It is a product asset rather than a throwaway QA file, and when we work with you it is yours to keep.

Can you assess an AI system we already built?

Yes. An assessment is one of the two ways we work: we evaluate an existing system against real scenarios, safety and privacy checks and our ISO/IEC 42001-aligned assurance framework, then give you an honest report and a plan to make it release-ready. The alternative is that we run the quality system for you on an ongoing basis.

What happens when a release fails its quality gate?

One of three things, and all of them are recorded. The change is blocked, or it is risk-accepted with a named owner and a stated reason, or it is approved. Any critical safety or compliance failure triggers a policy override that forces a block or a human handover regardless of the score average.

Do you replace our engineers?

No. We build the quality system and the evidence trail around your team's work. Human review stays in the loop by design, particularly on high-impact and high-risk decisions, and accountable people sign off rather than a model doing it silently.

Not the question you came with?

Ask us about your AI system

Bring us your AI. We will tell you the truth about it.

An honest evaluation against real scenarios, real safety checks, and an ISO/IEC 42001-aligned framework. Then a plan to make it release-ready.