NPU LabsNPU LABS

Flash

Flash - Decision Engine

Flash is a small decision model for the bounded questions inside an AI system, such as which category a request belongs to, which route to take, or whether a person should take over. It returns a score and its confidence. Your software keeps the thresholds and decides the action. Flash is in its baseline and validation phase.

How Flash returns a decision

Suppose Flash estimates an 87% probability that a customer interaction needs escalation. Flash returns the score, and the rule that 87% means a transfer to a human belongs in software. Another deployment may use 80%, and a regulated workflow may require manual review whenever confidence falls inside a set range. Changing those rules does not require retraining the model.

Input

Customer message

I was charged twice for my subscription.

Flash reads it

Flash

Flash

  • Classify
  • Screen
  • Score

Returns scores, not actions

Returns three scores

Scores

Classification

Billing

0.97

Guardrail

Pass

0.99

Escalation

Handover

0.96
Checked against your thresholds

Policy

Your policy

Thresholds in your software

  • Billing ≥ 0.90Billing queue
  • Guardrail passesContinue
  • Handover ≥ 0.90Human agent
The policy picks the action

Action

Action

  • Route to billing
  • Hand to a human

What Flash answers

Flash is intended to answer small questions quickly. It is not intended to compete with a frontier model at general reasoning.

Questions it answers

  • Which category does this belong to?
  • Which route should be selected?
  • How strongly does this input match a condition?
  • How confident is the model?
  • Is the uncertainty high enough that another system should take over?

Where it applies

  • A support platform classifies intent before it routes a conversation
  • A voice agent decides whether a caller may need escalation
  • A retrieval pipeline decides whether a document is relevant enough for a more expensive reasoning stage
  • A development agent decides which tools or validation rules apply to a change
  • A guardrail estimates risk
  • A workflow engine selects the next path

In our view, large generative models belong where open-ended reasoning adds value, and small decision models belong at the repetitive boundaries around them. Deterministic software stays responsible for policy and enforcement.

The planned runtime

A client application sends a request to the Flash runtime and gets back a structured decision. ONNX Runtime is the primary deployment direction because it separates model execution from a single hardware target.

Client application

Sends a request and receives a structured decision

Request
Response

Flash runtime

  • Task routing
    • Select model
    • Load checkpoint
    • Apply configuration
  • Model execution
    • ONNX Runtime
    • Optimised inference
    • CPU and GPU support
  • Decision logic
    • Confidence scoring
    • Thresholds and rules
    • Fallback and escalation

    Changed without retraining

  • Result handling
    • Structured output
    • Metadata
    • Logging
Loads models
Metrics and logs

Checkpoints and models

Base and fine-tuned models

  • Laya reference
  • Domain models
  • Quantised variants

Observability

Metrics, logs, and evaluation

  • Latency
  • Accuracy
  • Calibration
  • Tracing
The planned Flash runtime, simplified.

Design decisions

Use Laya as a baseline

Reproduce known behaviour before introducing NPU Labs-specific changes. Laya is the reference implementation, not the final Flash architecture.

Laya upstream repository

Keep Flash to bounded decisions

Classification, scoring, routing, and screening are the target workloads.

Make CPU deployment a requirement

Small decisions should not need dedicated accelerator infrastructure.

Deploy through ONNX Runtime

The model should be portable across hardware classes.

ONNX Runtime execution providers

Keep policy outside the model

Thresholds and consequences stay deterministic and auditable.

Measure uncertainty explicitly

Confidence and abstention are part of the product's behaviour, not optional diagnostics.

Talk to us about Flash

If your system makes the same bounded decisions many times, such as routing, screening, or escalation, talk to us about where a small decision model could fit.