NPU LabsNPU LABS
All articles
small-models5 October 2026

A tiny language model as a decision engine.

Input

Customer message

I was charged twice for my subscription.

Flash reads it

Flash

Flash

  • Classify
  • Screen
  • Score

Returns scores, not actions

Returns three scores

Scores

Classification

Billing

0.97

Guardrail

Pass

0.99

Escalation

Handover

0.96
Checked against your thresholds

Policy

Your policy

Thresholds in your software

  • Billing ≥ 0.90Billing queue
  • Guardrail passesContinue
  • Handover ≥ 0.90Human agent
The policy picks the action

Action

Action

  • Route to billing
  • Hand to a human

A surprising amount of work inside an AI system is not really a language-generation problem.

Which workflow should handle this request? Is this message likely to require escalation? Which tool should run next? Does this input fall into a known risk category? Should a document continue through the pipeline or be rejected?

These are decisions.

Large language models can make them. The more useful question is whether they should.

Not every AI decision needs a large language model.

Our view is simple: use generative models for ambiguity, and smaller decision models when the answer space is already known.

Flash is being built around fast classification, scoring, and routing with measurable confidence and portable inference.

The goal is not a smaller chatbot. It is a smaller decision surface.

The challenge

Modern AI systems increasingly route every intelligent decision through a generative model.

That is understandable. Large language models are flexible, easy to integrate, and capable of solving many different tasks through one interface.

But flexibility has a cost.

If the required output is one of five categories, a probability, a score, a route, or a yes-or-no decision, generation introduces capabilities the application may not need.

The model still has to process the request as a generative task. The application then has to validate the answer, constrain its format, and decide what to do with it.

One unnecessary model call is rarely important.

An agentic system can make dozens of these decisions during a single interaction. At that point the overhead becomes part of the architecture.

Our view

We think generative models should be used where generation and open-ended reasoning are actually required.

They are useful when a system has to interpret ambiguous language, combine context, reason across incomplete information, or produce a nuanced response.

When the answer space is already bounded, a different type of model can be a better fit.

Flash is our attempt to build specifically for that boundary.

It is not intended to compete with a frontier model at general reasoning.

It is intended to answer much smaller questions quickly:

  • Which category does this belong to?
  • Which route should be selected?
  • How strongly does this input match a condition?
  • How confident is the model?
  • Is the uncertainty high enough that another system should take over?

A decision model should provide evidence to the application.

The application should remain responsible for the consequence.

The constraints

  • Latency — these decisions often happen inside a larger workflow, so each call needs to add as little delay as possible.
  • Calibration — a confidence score is useful only if the surrounding system can interpret it reliably.
  • Predictability — bounded decisions should produce bounded outputs rather than prose that needs to be interpreted again.
  • Portability — a routing or classification model should not automatically require a dedicated GPU service.
  • Operational control — thresholds, escalation rules, and business policy should remain outside the model.
  • Evaluation — a model that is accurate on average can still be dangerous if it is confidently wrong in the cases that matter.

What we changed

The first important change was conceptual.

Flash started as a small-model project.

We narrowed that definition to a small decision-model project.

That gives the work a much clearer engineering boundary.

We are using Laya as the reference implementation and initial baseline rather than starting without a point of comparison. Flash itself remains an NPU Labs implementation.

The intention is to reproduce a known baseline first, understand its behaviour, establish the evaluation contract, and then make architectural changes deliberately.

Architecture and approach

The boundary we want is straightforward:

Client application

Sends a request and receives a structured decision

Request
Response

Flash runtime

  • Task routing
    • Select model
    • Load checkpoint
    • Apply configuration
  • Model execution
    • ONNX Runtime
    • Optimised inference
    • CPU and GPU support
  • Decision logic
    • Confidence scoring
    • Thresholds and rules
    • Fallback and escalation

    Changed without retraining

  • Result handling
    • Structured output
    • Metadata
    • Logging
Loads models
Metrics and logs

Checkpoints and models

Base and fine-tuned models

  • Laya reference
  • Domain models
  • Quantised variants

Observability

Metrics, logs, and evaluation

  • Latency
  • Accuracy
  • Calibration
  • Tracing
The planned Flash runtime, simplified. Thresholds and rules sit in configuration, outside the model.

Flash handles interpretation.

The application handles enforcement.

Suppose Flash estimates an 87% probability that a customer interaction requires escalation.

Flash should return the score.

It should not own the rule that says 87% means the interaction must be transferred to a human.

That rule belongs in software.

A different deployment may use an 80% threshold. Another may require 95%. A regulated workflow may require manual review whenever confidence falls inside a particular range.

Changing those rules should not require retraining the model.

ONNX Runtime is our primary deployment direction because it separates model execution from a single hardware target.

Key decisions

  • Use Laya as a baseline, not as the final Flash architecture — reproduce known behaviour before introducing NPU Labs-specific changes.
  • Keep Flash focused on bounded decisions — classification, scoring, routing, and screening are the target workloads.
  • Make CPU deployment a first-class requirement — small decisions should not automatically require dedicated accelerator infrastructure.
  • Use ONNX Runtime as the deployment abstraction — the model should be portable across hardware classes.
  • Keep operational policy outside the model — thresholds and consequences remain deterministic and auditable.
  • Measure uncertainty explicitly — confidence and abstention are part of the product behaviour, not optional diagnostics.

What we tested

Flash is still in the baseline and validation phase.

That means our immediate focus is not publishing a headline performance number.

It is defining the conditions under which a performance number would actually mean something.

Evidence

Test or evaluationResultPractical effect
Reference selectionLaya selected as the initial behavioural referenceGives Flash a reproducible starting point
Workload definitionFlash narrowed to bounded decision tasksPrevents the project drifting into another general-purpose assistant
Runtime directionONNX Runtime selected for inferenceCreates a portable deployment path across CPU and accelerator targets
Policy boundaryModel output separated from operational thresholdsBusiness rules can change without retraining the model
Evaluation scopeAccuracy, macro F1, calibration, abstention, latency, throughput, and memory identified as core measuresFlash can be judged against production behaviour rather than one metric

We deliberately do not yet claim that Flash is faster or more accurate than its reference implementation.

Those claims should come after the same model, dataset, runtime, and hardware configuration can be reproduced and measured consistently.

What we learned

The architecture becomes much cleaner once the model is not responsible for the final operational decision.

Flash decides what the input appears to represent.

Software decides what the organisation does about it.

The first problem is probabilistic.

The second is policy.

Mixing those responsibilities makes both harder to reason about.

The lesson

Use models to interpret ambiguity. Use software to enforce consequences.

A model is useful because it can recognise patterns in uncertain inputs.

Software is useful because it can apply an exact rule every time.

A production AI system should use each where it is strongest.

What this demonstrates

This work demonstrates our ability to:

  • design models around specific production responsibilities — Flash starts with the decision boundary rather than a model-size target;
  • separate probabilistic judgement from deterministic policy — operational behaviour remains explicit and auditable;
  • build evaluation into the architecture — calibration, abstention, latency, and resource usage are part of the definition of success;
  • design for portable inference — the deployment model is intended to span ordinary CPUs, private compute, and accelerated environments;
  • use upstream work as evidence rather than dependency — Laya provides a baseline while Flash remains free to evolve around NPU Labs requirements.

Where this applies

There are small decisions throughout almost every production AI system.

A support platform needs to classify intent before routing a conversation.

A voice agent needs to determine whether a caller may require escalation.

A retrieval pipeline needs to decide whether a document is relevant enough to send into a more expensive reasoning stage.

A development agent needs to decide which tools or validation rules apply to a change.

A guardrail needs to estimate risk.

A workflow engine needs to select the next path.

None of those tasks necessarily needs a frontier model.

A useful architecture can therefore contain both.

Large generative models handle the places where open-ended reasoning adds value.

Small decision models handle the repetitive boundaries around them.

Deterministic software remains responsible for policy and enforcement.

That is the role we see for Flash.

Not as a replacement for large language models.

As a way to stop using them where they are unnecessary.

What comes next

  • Reproduce the selected Laya baseline.
  • Establish compatibility and regression tests before making architectural changes.
  • Define the first Flash datasets around classification and bounded decision workloads.
  • Measure accuracy, macro F1, calibration, and abstention behaviour.
  • Establish p50, p95, and p99 inference latency rather than relying on an average.
  • Measure throughput, memory usage, model loading, and tokenisation separately from inference.
  • Export and benchmark the first ONNX Runtime path.
  • Compare AMD64 and ARM64 CPU deployment.
  • Introduce NPU Labs-specific architecture and training changes only after the baseline is reproducible.

References

Talk to us

Want to talk about what this changes for your team?

We will scope the cluster, the runtime, and what it would take to stand the stack up.