A surprising amount of work inside an AI system is not really a language-generation problem.
Which workflow should handle this request? Is this message likely to require escalation? Which tool should run next? Does this input fall into a known risk category? Should a document continue through the pipeline or be rejected?
These are decisions.
Large language models can make them. The more useful question is whether they should.
Not every AI decision needs a large language model.
Our view is simple: use generative models for ambiguity, and smaller decision models when the answer space is already known.
Flash is being built around fast classification, scoring, and routing with measurable confidence and portable inference.
The goal is not a smaller chatbot. It is a smaller decision surface.
The challenge
Modern AI systems increasingly route every intelligent decision through a generative model.
That is understandable. Large language models are flexible, easy to integrate, and capable of solving many different tasks through one interface.
But flexibility has a cost.
If the required output is one of five categories, a probability, a score, a route, or a yes-or-no decision, generation introduces capabilities the application may not need.
The model still has to process the request as a generative task. The application then has to validate the answer, constrain its format, and decide what to do with it.
One unnecessary model call is rarely important.
An agentic system can make dozens of these decisions during a single interaction. At that point the overhead becomes part of the architecture.
Our view
We think generative models should be used where generation and open-ended reasoning are actually required.
They are useful when a system has to interpret ambiguous language, combine context, reason across incomplete information, or produce a nuanced response.
When the answer space is already bounded, a different type of model can be a better fit.
Flash is our attempt to build specifically for that boundary.
It is not intended to compete with a frontier model at general reasoning.
It is intended to answer much smaller questions quickly:
- Which category does this belong to?
- Which route should be selected?
- How strongly does this input match a condition?
- How confident is the model?
- Is the uncertainty high enough that another system should take over?
A decision model should provide evidence to the application.
The application should remain responsible for the consequence.
The constraints
- Latency — these decisions often happen inside a larger workflow, so each call needs to add as little delay as possible.
- Calibration — a confidence score is useful only if the surrounding system can interpret it reliably.
- Predictability — bounded decisions should produce bounded outputs rather than prose that needs to be interpreted again.
- Portability — a routing or classification model should not automatically require a dedicated GPU service.
- Operational control — thresholds, escalation rules, and business policy should remain outside the model.
- Evaluation — a model that is accurate on average can still be dangerous if it is confidently wrong in the cases that matter.
What we changed
The first important change was conceptual.
Flash started as a small-model project.
We narrowed that definition to a small decision-model project.
That gives the work a much clearer engineering boundary.
We are using Laya as the reference implementation and initial baseline rather than starting without a point of comparison. Flash itself remains an NPU Labs implementation.
The intention is to reproduce a known baseline first, understand its behaviour, establish the evaluation contract, and then make architectural changes deliberately.
Architecture and approach
The boundary we want is straightforward:
Client application
Sends a request and receives a structured decision
Flash runtime
Task routing
- Select model
- Load checkpoint
- Apply configuration
Model execution
- ONNX Runtime
- Optimised inference
- CPU and GPU support
Decision logic
- Confidence scoring
- Thresholds and rules
- Fallback and escalation
Changed without retraining
Result handling
- Structured output
- Metadata
- Logging
Checkpoints and models
Base and fine-tuned models
- Laya reference
- Domain models
- Quantised variants
Observability
Metrics, logs, and evaluation
- Latency
- Accuracy
- Calibration
- Tracing
Flash handles interpretation.
The application handles enforcement.
Suppose Flash estimates an 87% probability that a customer interaction requires escalation.
Flash should return the score.
It should not own the rule that says 87% means the interaction must be transferred to a human.
That rule belongs in software.
A different deployment may use an 80% threshold. Another may require 95%. A regulated workflow may require manual review whenever confidence falls inside a particular range.
Changing those rules should not require retraining the model.
ONNX Runtime is our primary deployment direction because it separates model execution from a single hardware target.
Key decisions
- Use Laya as a baseline, not as the final Flash architecture — reproduce known behaviour before introducing NPU Labs-specific changes.
- Keep Flash focused on bounded decisions — classification, scoring, routing, and screening are the target workloads.
- Make CPU deployment a first-class requirement — small decisions should not automatically require dedicated accelerator infrastructure.
- Use ONNX Runtime as the deployment abstraction — the model should be portable across hardware classes.
- Keep operational policy outside the model — thresholds and consequences remain deterministic and auditable.
- Measure uncertainty explicitly — confidence and abstention are part of the product behaviour, not optional diagnostics.
What we tested
Flash is still in the baseline and validation phase.
That means our immediate focus is not publishing a headline performance number.
It is defining the conditions under which a performance number would actually mean something.
Evidence
| Test or evaluation | Result | Practical effect |
|---|---|---|
| Reference selection | Laya selected as the initial behavioural reference | Gives Flash a reproducible starting point |
| Workload definition | Flash narrowed to bounded decision tasks | Prevents the project drifting into another general-purpose assistant |
| Runtime direction | ONNX Runtime selected for inference | Creates a portable deployment path across CPU and accelerator targets |
| Policy boundary | Model output separated from operational thresholds | Business rules can change without retraining the model |
| Evaluation scope | Accuracy, macro F1, calibration, abstention, latency, throughput, and memory identified as core measures | Flash can be judged against production behaviour rather than one metric |
We deliberately do not yet claim that Flash is faster or more accurate than its reference implementation.
Those claims should come after the same model, dataset, runtime, and hardware configuration can be reproduced and measured consistently.
What we learned
The architecture becomes much cleaner once the model is not responsible for the final operational decision.
Flash decides what the input appears to represent.
Software decides what the organisation does about it.
The first problem is probabilistic.
The second is policy.
Mixing those responsibilities makes both harder to reason about.
The lesson
Use models to interpret ambiguity. Use software to enforce consequences.
A model is useful because it can recognise patterns in uncertain inputs.
Software is useful because it can apply an exact rule every time.
A production AI system should use each where it is strongest.
What this demonstrates
This work demonstrates our ability to:
- design models around specific production responsibilities — Flash starts with the decision boundary rather than a model-size target;
- separate probabilistic judgement from deterministic policy — operational behaviour remains explicit and auditable;
- build evaluation into the architecture — calibration, abstention, latency, and resource usage are part of the definition of success;
- design for portable inference — the deployment model is intended to span ordinary CPUs, private compute, and accelerated environments;
- use upstream work as evidence rather than dependency — Laya provides a baseline while Flash remains free to evolve around NPU Labs requirements.
Where this applies
There are small decisions throughout almost every production AI system.
A support platform needs to classify intent before routing a conversation.
A voice agent needs to determine whether a caller may require escalation.
A retrieval pipeline needs to decide whether a document is relevant enough to send into a more expensive reasoning stage.
A development agent needs to decide which tools or validation rules apply to a change.
A guardrail needs to estimate risk.
A workflow engine needs to select the next path.
None of those tasks necessarily needs a frontier model.
A useful architecture can therefore contain both.
Large generative models handle the places where open-ended reasoning adds value.
Small decision models handle the repetitive boundaries around them.
Deterministic software remains responsible for policy and enforcement.
That is the role we see for Flash.
Not as a replacement for large language models.
As a way to stop using them where they are unnecessary.
What comes next
- Reproduce the selected Laya baseline.
- Establish compatibility and regression tests before making architectural changes.
- Define the first Flash datasets around classification and bounded decision workloads.
- Measure accuracy, macro F1, calibration, and abstention behaviour.
- Establish p50, p95, and p99 inference latency rather than relying on an average.
- Measure throughput, memory usage, model loading, and tokenisation separately from inference.
- Export and benchmark the first ONNX Runtime path.
- Compare AMD64 and ARM64 CPU deployment.
- Introduce NPU Labs-specific architecture and training changes only after the baseline is reproducible.
Related reading
- A private AI, small enough to run on a Raspberry Pi. — how specialised models can move inference from central infrastructure to ordinary hardware.
- Two Claude models were pulled worldwide. Your architecture should survive that. — why controlling more of the inference path can reduce operational dependence on external providers.
References
- Laya upstream repository — reference implementation for the initial Flash baseline.
- ONNX Runtime execution providers — deployment abstraction for CPU, GPU, and specialised accelerator targets.