NPU LabsNPU LABS

Model registry

The models we support.

The open-weight models NPU Labs tracks, evaluates, and deploys on private nodes across reasoning, coding, multimodal, retrieval, and voice. A portfolio we keep current, not a benchmark claim: final selection is always proven on your workload.

Size is not an exclusionNVFP4 variants preferredLicensing checked before productionValidated per workload

Core models

Reasoning, text, code, and vision.

Grouped by what the model is for. Coverage aligned with the current open-source frontier. NVFP4 variants preferred where available. Status is not a benchmark claim. It reflects what we run in production today.

Reasoning

SmolLM3-3B

Hugging Face

Size3B dense

Context128K (YaRN)

LicenseApache 2.0

Specialized

Runs1 node (small)

Phi-4-reasoning-plus

Microsoft

Size14B dense

Context32K

LicenseMIT

Preferred

Runs1 node (small)

QwQ-32B

Alibaba

Size32B dense

Context131K

LicenseApache 2.0

Preferred

Runs1 node

GLM-4.5-Air

Z.ai

Size106B · 12B active

Context128K

LicenseMIT

Production default

Runs1 node

Nemotron-H-47B-Reasoning

NVIDIA

Size47B hybrid Mamba-Transformer

Context128K

LicenseNVIDIA Open Model

Preferred

Runs1 node

DeepSeek-R1-Distill-Llama-70B

DeepSeek

Size70B dense

Context128K

LicenseMIT

Production default

Runs1 node

GLM-4.5

Z.ai

Size355B · 32B active

Context128K

LicenseMIT

Preferred

Runs4-node cluster

DeepSeek-R1

DeepSeek

Size671B · 37B active

Context128K

LicenseMIT

Preferred

RunsMulti-cluster

NVIDIA Nemotron 3 Ultra

NVIDIA

Size550B · 55B active

Context128K

LicenseOpenMDW-1.1

Premium frontier

RunsMulti-cluster

General text

Llama 3.2 3B Instruct

Meta

Size3B dense

Context128K

LicenseLlama 3.2 community

Production default

Runs1 node (small)

Phi-4-mini-instruct

Microsoft

Size3.8B dense

Context128K

LicenseMIT

Production default

Runs1 node (small)

Qwen3 4B

Alibaba

Size4B dense

Context128K

LicenseApache 2.0

Production default

Runs1 node (small)

Gemma 3 4B

Google

Size4B dense

Context128K

LicenseGemma terms

Specialized

Runs1 node (small)

OLMoE-1B-7B Instruct

Allen AI

Size7B · 1B active

Context4K

LicenseApache 2.0

Specialized

Runs1 node (small)

Nemotron-H-8B-Instruct

NVIDIA

Size8B hybrid Mamba-Transformer

Context8K

LicenseNVIDIA Open Model

Specialized

Runs1 node (small)

Falcon 3 10B Instruct

TII

Size10B dense

Context32K

LicenseFalcon LLM License

Specialized

Runs1 node (small)

Gemma 3 12B

Google

Size12B dense

Context128K

LicenseGemma terms

Production default

Runs1 node (small)

Phi-4

Microsoft

Size14B dense

Context16K

LicenseMIT

Production default

Runs1 node (small)

Qwen3 14B

Alibaba

Size14B dense

Context128K

LicenseApache 2.0

Production default

Runs1 node (small)

DeepSeek-V2-Lite Chat

DeepSeek

Size16B · 2.4B active

Context32K

LicenseDeepSeek License

Specialized

Runs1 node (small)

Mistral Small 3.2 24B Instruct

Mistral

Size24B dense

Context128K

LicenseApache 2.0

Preferred

Runs1 node (small)

Gemma 4 31B

Google

Size31B dense

Context256K

LicenseGemma terms

Production default

Runs1 node

OLMo 2 32B Instruct

Allen AI

Size32B dense

Context4K

LicenseApache 2.0

Specialized

Runs1 node

Granite 4.1 Small

IBM

Size32B · 9B active (Mamba-2 hybrid)

Context128K

LicenseApache 2.0

Production default

Runs1 node

Qwen3 32B

Alibaba

Size32B dense

Context128K

LicenseApache 2.0

Preferred

Runs1 node

Qwen3.6-35B-A3B

Alibaba

Size35B · 3B active

Context262K

LicenseApache 2.0

Production default

Runs1 node

Llama 3.3 70B Instruct

Meta

Size70B dense

Context128K

LicenseLlama 3.3 Community

Production default

Runs1 node

Llama 4 Scout

Meta

Size109B · 17B active

Context10M

LicenseLlama 4 community

Preferred

Runs1 node

NVIDIA Nemotron 3 Super

NVIDIA

Size120B · 12B active

Context~1M

LicenseNVIDIA Open

Production default

Runs1 node

Hunyuan-Large

Tencent

Size389B · 52B active

Context256K

LicenseTencent Hunyuan Community

Experimental

Runs4-node cluster

Jamba Large 1.6

AI21

Size398B · 94B active (SSM-Transformer)

Context256K

LicenseJamba Open Model License

Preferred

Runs4-node cluster

Llama 4 Maverick

Meta

Size402B · 17B active

Context1M

LicenseLlama 4 community

Preferred

RunsMulti-cluster

Mistral Large 3

Mistral

Size675B · 41B active

Context256K

LicenseApache 2.0

Premium frontier

RunsMulti-cluster

GLM-5.2

Z.ai

Size753B · 40B active

Context1M

LicenseMIT

Premium frontier

RunsMulti-cluster

Kimi K2.6

Moonshot AI

Size1T · 32B active

Context128K

LicenseModified MIT

Premium frontier

RunsMulti-cluster

DeepSeek V4 Pro

DeepSeek

Size1.6T · 49B active

Context1M

LicenseMIT

Premium frontier

RunsMulti-cluster

Code

Mamba-Codestral 7B v0.1

Mistral

Size7B Mamba-2

Context256K

LicenseApache 2.0

Specialized

Runs1 node (small)

Ornith-1.0 9B

DeepReinforce AI

Size9B dense

Context256K

LicenseMIT

Preferred

Runs1 node (small)

Ornith-1.0 35B

DeepReinforce AI

Size35B dense

Context256K

LicenseMIT

Production default

Runs1 node

CodeGemma 7B

Google

Size7B dense

Context8K

LicenseGemma terms

Support model

Runs1 node (small)

StarCoder 2 15B

BigCode

Size15B dense

Context16K

LicenseBigCode OpenRAIL

Support model

Runs1 node

Devstral Small 2

Mistral

Size24B dense

Context256K

LicenseApache 2.0

Production default

Runs1 node (small)

Qwen2.5-Coder 32B Instruct

Alibaba

Size32B dense

Context128K

LicenseApache 2.0

Preferred

Runs1 node

Qwen3-Coder-Next

Alibaba

Size80B · 3B active

Context256K

LicenseApache 2.0

Production default

Runs1 node

Devstral 2

Mistral

Size123B dense

Context256K

LicenseModified MIT

Preferred

Runs4-node cluster

DeepSeek-V3 Coder

DeepSeek

Size236B · 21B active

Context128K

LicenseMIT

Preferred

RunsMulti-cluster

Ornith-1.0 397B

DeepReinforce AI

Size397B dense

Context256K

LicenseMIT

Preferred

RunsMulti-cluster

Qwen 3 Coder 480B

Alibaba

Size480B

Context256K

LicenseApache 2.0

Preferred

RunsMulti-cluster

Vision and multimodal

Phi-4-multimodal-instruct

Microsoft

Size5.6B (text + vision + speech)

Context128K

LicenseMIT

Preferred

Runs1 node (small)

Pixtral 12B

Mistral

Size12B dense

Context128K

LicenseApache 2.0

Production default

Runs1 node (small)

Gemma 3 27B

Google

Size27B dense + 400M SigLIP

Context128K

LicenseGemma

Production default

Runs1 node

InternVL3.5-38B

OpenGVLab (Shanghai AI Lab)

Size38B dense

Context32K

LicenseMIT

Production default

Runs1 node

Qwen2.5-VL 72B

Alibaba

Size72B vision

Context128K

LicenseQwen license

Support model

Runs1 node

Molmo 72B

Allen AI

Size72B dense

Context4K

LicenseApache 2.0

Specialized

Runs1 node

InternVL3-78B

OpenGVLab (Shanghai AI Lab)

Size78B dense

Context32K

LicenseMIT

Preferred

Runs1 node

Llama 3.2 Vision 90B

Meta

Size90B vision

Context128K

LicenseLlama 3.2 community

Preferred

Runs1 node

Pixtral Large

Mistral

Size124B vision

Context128K

LicenseMistral Research

Preferred

Runs1 node

Parameters, context, and licensing are recorded from model cards and verified before deployment. The list updates as the open-source frontier moves.

Support models

Retrieval, speech, voice.

The models behind RAG and the voice pipeline. Transcription stays separate from the reasoning model; TTS licensing is confirmed before production.

Embeddings and retrieval

Qwen3 Embedding 8B

Alibaba

Support model

General embeddings / RAG

Apache 2.0

BGE-M3

BAAI

Support model

Hybrid retrieval (dense + sparse + multi-vector)

MIT

bge-reranker-v2

BAAI

Support model

Reranking after retrieval

MIT

Llama-Embed-Nemotron-8B

NVIDIA

Preferred

General embeddings (8B Llama backbone, 8K context)

NVIDIA Open Model

NV-Embed-v2

NVIDIA

Specialized

MTEB-leading 7B Mistral-backbone embedder

CC-BY-NC 4.0

Stella en 1.5B v5

NovaSearch

Support model

Lightweight retrieval, strong English coverage

MIT

Speech-to-text

Voxtral Small 24B

Mistral

Support model

Transcription and audio understanding

Apache 2.0

Voxtral Mini 4B Realtime

Mistral

Support model

Real-time transcription

Apache 2.0

Kimi-Audio 7B Instruct

Moonshot AI

Preferred

Universal audio foundation model (STT + understanding)

Apache 2.0

NVIDIA Parakeet EN 0.6B

NVIDIA

Support model

English transcription

NVIDIA Open

Deepgram Nova-3

Deepgram

Support model

Production STT (hosted, multilingual)

Commercial

Text-to-speech

Kokoro 82M

hexgrad

Support model

Lightweight TTS

Apache 2.0

Sesame CSM-1B

Sesame AI Labs

Preferred

Conversational speech model (1B backbone + 100M decoder)

Apache 2.0

ElevenLabs

ElevenLabs

Support model

Production TTS, branded voice clones

Commercial

Fish Audio S2 Pro

Fish Audio

Restricted

High-quality TTS

Verify with vendor

Voxtral 4B TTS

Mistral

Restricted

TTS (evaluation)

CC BY-NC 4.0

Step-Audio-EditX

StepFun

Watchlist

Audio editing / expressive TTS

Verify with vendor

How to read it

Registry status.

Every model carries a status. A release date is recorded separately from deployment status, so a model can be released but still need node validation before it leads a client deployment.

Production default

Safe to lead with for common private enterprise workloads after normal validation.

Preferred

Available in the platform and used where the workload matches.

Premium frontier

Strong model positioned for high-value workloads; validated on our nodes.

Support model

Part of the platform for embeddings, reranking, speech-to-text, or text-to-speech.

Watchlist

Promising, but weights, license, runtime, or stability still need confirmation.

Restricted commercial use

Not sold into client production unless commercial terms are confirmed.

Which model fits your workload?

We benchmark candidates on your data before anything goes to production, measuring retrieval quality, tool use, latency, and cost. Tell us the workload and we'll scope the model set and the node it runs on.