Flash
Flash - Decision Engine
Flash is a small decision model for the bounded questions inside an AI system, such as which category a request belongs to, which route to take, or whether a person should take over. It returns a score and its confidence. Your software keeps the thresholds and decides the action. Flash is in its baseline and validation phase.
How Flash returns a decision
Suppose Flash estimates an 87% probability that a customer interaction needs escalation. Flash returns the score, and the rule that 87% means a transfer to a human belongs in software. Another deployment may use 80%, and a regulated workflow may require manual review whenever confidence falls inside a set range. Changing those rules does not require retraining the model.
Input
Customer message
I was charged twice for my subscription.
Flash
Flash
- Classify
- Screen
- Score
Returns scores, not actions
Scores
Classification
Billing
Guardrail
Pass
Escalation
Handover
Policy
Your policy
Thresholds in your software
- Billing ≥ 0.90Billing queue
- Guardrail passesContinue
- Handover ≥ 0.90Human agent
Action
Action
- Route to billing
- Hand to a human
Input
Flash
Scores
Policy
Action
Customer message
I was charged twice for my subscription.
Flash
- Classify
- Screen
- Score
Returns scores, not actions
Classification
Billing
Guardrail
Pass
Escalation
Handover
Your policy
Thresholds in your software
- Billing ≥ 0.90Billing queue
- Guardrail passesContinue
- Handover ≥ 0.90Human agent
Action
- Route to billing
- Hand to a human
What Flash answers
Flash is intended to answer small questions quickly. It is not intended to compete with a frontier model at general reasoning.
Questions it answers
- Which category does this belong to?
- Which route should be selected?
- How strongly does this input match a condition?
- How confident is the model?
- Is the uncertainty high enough that another system should take over?
Where it applies
- A support platform classifies intent before it routes a conversation
- A voice agent decides whether a caller may need escalation
- A retrieval pipeline decides whether a document is relevant enough for a more expensive reasoning stage
- A development agent decides which tools or validation rules apply to a change
- A guardrail estimates risk
- A workflow engine selects the next path
In our view, large generative models belong where open-ended reasoning adds value, and small decision models belong at the repetitive boundaries around them. Deterministic software stays responsible for policy and enforcement.
The planned runtime
A client application sends a request to the Flash runtime and gets back a structured decision. ONNX Runtime is the primary deployment direction because it separates model execution from a single hardware target.
Client application
Sends a request and receives a structured decision
Flash runtime
Task routing
- Select model
- Load checkpoint
- Apply configuration
Model execution
- ONNX Runtime
- Optimised inference
- CPU and GPU support
Decision logic
- Confidence scoring
- Thresholds and rules
- Fallback and escalation
Changed without retraining
Result handling
- Structured output
- Metadata
- Logging
Checkpoints and models
Base and fine-tuned models
- Laya reference
- Domain models
- Quantised variants
Observability
Metrics, logs, and evaluation
- Latency
- Accuracy
- Calibration
- Tracing
Design decisions
Use Laya as a baseline
Reproduce known behaviour before introducing NPU Labs-specific changes. Laya is the reference implementation, not the final Flash architecture.
Laya upstream repositoryKeep Flash to bounded decisions
Classification, scoring, routing, and screening are the target workloads.
Make CPU deployment a requirement
Small decisions should not need dedicated accelerator infrastructure.
Deploy through ONNX Runtime
The model should be portable across hardware classes.
ONNX Runtime execution providersKeep policy outside the model
Thresholds and consequences stay deterministic and auditable.
Measure uncertainty explicitly
Confidence and abstention are part of the product's behaviour, not optional diagnostics.
Talk to us about Flash
If your system makes the same bounded decisions many times, such as routing, screening, or escalation, talk to us about where a small decision model could fit.