Model registry
The models we support.
The open-weight models NPU Labs tracks, evaluates, and deploys on private nodes across reasoning, coding, multimodal, retrieval, and voice. A portfolio we keep current, not a benchmark claim: final selection is always proven on your workload.
Core models
Reasoning, text, code, and vision.
Grouped by what the model is for. Coverage aligned with the current open-source frontier. NVFP4 variants preferred where available. Status is not a benchmark claim. It reflects what we run in production today.
Reasoning
SmolLM3-3B
Hugging Face
Size3B dense
Context128K (YaRN)
LicenseApache 2.0
Runs1 node (small)
Phi-4-reasoning-plus
Microsoft
Size14B dense
Context32K
LicenseMIT
Runs1 node (small)
QwQ-32B
Alibaba
Size32B dense
Context131K
LicenseApache 2.0
Runs1 node
GLM-4.5-Air
Z.ai
Size106B · 12B active
Context128K
LicenseMIT
Runs1 node
Nemotron-H-47B-Reasoning
NVIDIA
Size47B hybrid Mamba-Transformer
Context128K
LicenseNVIDIA Open Model
Runs1 node
DeepSeek-R1-Distill-Llama-70B
DeepSeek
Size70B dense
Context128K
LicenseMIT
Runs1 node
GLM-4.5
Z.ai
Size355B · 32B active
Context128K
LicenseMIT
Runs4-node cluster
DeepSeek-R1
DeepSeek
Size671B · 37B active
Context128K
LicenseMIT
RunsMulti-cluster
NVIDIA Nemotron 3 Ultra
NVIDIA
Size550B · 55B active
Context128K
LicenseOpenMDW-1.1
RunsMulti-cluster
General text
Llama 3.2 3B Instruct
Meta
Size3B dense
Context128K
LicenseLlama 3.2 community
Runs1 node (small)
Phi-4-mini-instruct
Microsoft
Size3.8B dense
Context128K
LicenseMIT
Runs1 node (small)
Qwen3 4B
Alibaba
Size4B dense
Context128K
LicenseApache 2.0
Runs1 node (small)
Gemma 3 4B
Size4B dense
Context128K
LicenseGemma terms
Runs1 node (small)
OLMoE-1B-7B Instruct
Allen AI
Size7B · 1B active
Context4K
LicenseApache 2.0
Runs1 node (small)
Nemotron-H-8B-Instruct
NVIDIA
Size8B hybrid Mamba-Transformer
Context8K
LicenseNVIDIA Open Model
Runs1 node (small)
Falcon 3 10B Instruct
TII
Size10B dense
Context32K
LicenseFalcon LLM License
Runs1 node (small)
Gemma 3 12B
Size12B dense
Context128K
LicenseGemma terms
Runs1 node (small)
Phi-4
Microsoft
Size14B dense
Context16K
LicenseMIT
Runs1 node (small)
Qwen3 14B
Alibaba
Size14B dense
Context128K
LicenseApache 2.0
Runs1 node (small)
DeepSeek-V2-Lite Chat
DeepSeek
Size16B · 2.4B active
Context32K
LicenseDeepSeek License
Runs1 node (small)
Mistral Small 3.2 24B Instruct
Mistral
Size24B dense
Context128K
LicenseApache 2.0
Runs1 node (small)
Gemma 4 31B
Size31B dense
Context256K
LicenseGemma terms
Runs1 node
OLMo 2 32B Instruct
Allen AI
Size32B dense
Context4K
LicenseApache 2.0
Runs1 node
Granite 4.1 Small
IBM
Size32B · 9B active (Mamba-2 hybrid)
Context128K
LicenseApache 2.0
Runs1 node
Qwen3 32B
Alibaba
Size32B dense
Context128K
LicenseApache 2.0
Runs1 node
Qwen3.6-35B-A3B
Alibaba
Size35B · 3B active
Context262K
LicenseApache 2.0
Runs1 node
Llama 3.3 70B Instruct
Meta
Size70B dense
Context128K
LicenseLlama 3.3 Community
Runs1 node
Llama 4 Scout
Meta
Size109B · 17B active
Context10M
LicenseLlama 4 community
Runs1 node
NVIDIA Nemotron 3 Super
NVIDIA
Size120B · 12B active
Context~1M
LicenseNVIDIA Open
Runs1 node
Hunyuan-Large
Tencent
Size389B · 52B active
Context256K
LicenseTencent Hunyuan Community
Runs4-node cluster
Jamba Large 1.6
AI21
Size398B · 94B active (SSM-Transformer)
Context256K
LicenseJamba Open Model License
Runs4-node cluster
Llama 4 Maverick
Meta
Size402B · 17B active
Context1M
LicenseLlama 4 community
RunsMulti-cluster
Mistral Large 3
Mistral
Size675B · 41B active
Context256K
LicenseApache 2.0
RunsMulti-cluster
GLM-5.2
Z.ai
Size753B · 40B active
Context1M
LicenseMIT
RunsMulti-cluster
Kimi K2.6
Moonshot AI
Size1T · 32B active
Context128K
LicenseModified MIT
RunsMulti-cluster
DeepSeek V4 Pro
DeepSeek
Size1.6T · 49B active
Context1M
LicenseMIT
RunsMulti-cluster
Code
Mamba-Codestral 7B v0.1
Mistral
Size7B Mamba-2
Context256K
LicenseApache 2.0
Runs1 node (small)
Ornith-1.0 9B
DeepReinforce AI
Size9B dense
Context256K
LicenseMIT
Runs1 node (small)
Ornith-1.0 35B
DeepReinforce AI
Size35B dense
Context256K
LicenseMIT
Runs1 node
CodeGemma 7B
Size7B dense
Context8K
LicenseGemma terms
Runs1 node (small)
StarCoder 2 15B
BigCode
Size15B dense
Context16K
LicenseBigCode OpenRAIL
Runs1 node
Devstral Small 2
Mistral
Size24B dense
Context256K
LicenseApache 2.0
Runs1 node (small)
Qwen2.5-Coder 32B Instruct
Alibaba
Size32B dense
Context128K
LicenseApache 2.0
Runs1 node
Qwen3-Coder-Next
Alibaba
Size80B · 3B active
Context256K
LicenseApache 2.0
Runs1 node
Devstral 2
Mistral
Size123B dense
Context256K
LicenseModified MIT
Runs4-node cluster
DeepSeek-V3 Coder
DeepSeek
Size236B · 21B active
Context128K
LicenseMIT
RunsMulti-cluster
Ornith-1.0 397B
DeepReinforce AI
Size397B dense
Context256K
LicenseMIT
RunsMulti-cluster
Qwen 3 Coder 480B
Alibaba
Size480B
Context256K
LicenseApache 2.0
RunsMulti-cluster
Vision and multimodal
Phi-4-multimodal-instruct
Microsoft
Size5.6B (text + vision + speech)
Context128K
LicenseMIT
Runs1 node (small)
Pixtral 12B
Mistral
Size12B dense
Context128K
LicenseApache 2.0
Runs1 node (small)
Gemma 3 27B
Size27B dense + 400M SigLIP
Context128K
LicenseGemma
Runs1 node
InternVL3.5-38B
OpenGVLab (Shanghai AI Lab)
Size38B dense
Context32K
LicenseMIT
Runs1 node
Qwen2.5-VL 72B
Alibaba
Size72B vision
Context128K
LicenseQwen license
Runs1 node
Molmo 72B
Allen AI
Size72B dense
Context4K
LicenseApache 2.0
Runs1 node
InternVL3-78B
OpenGVLab (Shanghai AI Lab)
Size78B dense
Context32K
LicenseMIT
Runs1 node
Llama 3.2 Vision 90B
Meta
Size90B vision
Context128K
LicenseLlama 3.2 community
Runs1 node
Pixtral Large
Mistral
Size124B vision
Context128K
LicenseMistral Research
Runs1 node
Parameters, context, and licensing are recorded from model cards and verified before deployment. The list updates as the open-source frontier moves.
Support models
Retrieval, speech, voice.
The models behind RAG and the voice pipeline. Transcription stays separate from the reasoning model; TTS licensing is confirmed before production.
Embeddings and retrieval
Qwen3 Embedding 8B
Alibaba
General embeddings / RAG
Apache 2.0
BGE-M3
BAAI
Hybrid retrieval (dense + sparse + multi-vector)
MIT
bge-reranker-v2
BAAI
Reranking after retrieval
MIT
Llama-Embed-Nemotron-8B
NVIDIA
General embeddings (8B Llama backbone, 8K context)
NVIDIA Open Model
NV-Embed-v2
NVIDIA
MTEB-leading 7B Mistral-backbone embedder
CC-BY-NC 4.0
Stella en 1.5B v5
NovaSearch
Lightweight retrieval, strong English coverage
MIT
Speech-to-text
Voxtral Small 24B
Mistral
Transcription and audio understanding
Apache 2.0
Voxtral Mini 4B Realtime
Mistral
Real-time transcription
Apache 2.0
Kimi-Audio 7B Instruct
Moonshot AI
Universal audio foundation model (STT + understanding)
Apache 2.0
NVIDIA Parakeet EN 0.6B
NVIDIA
English transcription
NVIDIA Open
Deepgram Nova-3
Deepgram
Production STT (hosted, multilingual)
Commercial
Text-to-speech
Kokoro 82M
hexgrad
Lightweight TTS
Apache 2.0
Sesame CSM-1B
Sesame AI Labs
Conversational speech model (1B backbone + 100M decoder)
Apache 2.0
ElevenLabs
ElevenLabs
Production TTS, branded voice clones
Commercial
Fish Audio S2 Pro
Fish Audio
High-quality TTS
Verify with vendor
Voxtral 4B TTS
Mistral
TTS (evaluation)
CC BY-NC 4.0
Step-Audio-EditX
StepFun
Audio editing / expressive TTS
Verify with vendor
How to read it
Registry status.
Every model carries a status. A release date is recorded separately from deployment status, so a model can be released but still need node validation before it leads a client deployment.
Safe to lead with for common private enterprise workloads after normal validation.
Available in the platform and used where the workload matches.
Strong model positioned for high-value workloads; validated on our nodes.
Part of the platform for embeddings, reranking, speech-to-text, or text-to-speech.
Promising, but weights, license, runtime, or stability still need confirmation.
Not sold into client production unless commercial terms are confirmed.
Which model fits your workload?
We benchmark candidates on your data before anything goes to production, measuring retrieval quality, tool use, latency, and cost. Tell us the workload and we'll scope the model set and the node it runs on.