Cookbook: serve Qwen3.8-2.4T-A95B-NVFP4 with vLLM¶
Scope: a deployment reference for nvidia/Qwen3.8-2.4T-A95B-NVFP4, NVIDIA's mixed-precision quantization of Alibaba's Qwen3.8-2.4T-A95B: what is actually quantized and what is not, the tensor-payload census derived from the checkpoint's own config and checked against its safetensors index, the per-GPU memory the card's own --tensor-parallel-size 8 implies, the engine version floors the card omits, and the four places the card contradicts its own artifact. The NVFP4 format itself (E2M1 codes, two-level scaling, why Blackwell is required) is NVFP4 and is not restated here; the method catalog is quantization for inference; the total-versus-active parameter economics are MoE sparse scaling; the two-cache problem a hybrid-attention stack creates for prefix caching is serving hybrid recurrent and full-attention models.
Reference template. The launch commands and the manifest are not executed here, and no inference was run: the checkpoint is 1.44 TB and Blackwell-only. What was executed is the Python block below, which derives the footprint from published values and asserts it against the published index; its output is pasted verbatim. Every config value, benchmark number, and quoted sentence was read from the model repository itself on 2026-08-30, not from a summary. The engine version floors are source-presence floors: they are the earliest release tags carrying the model-specific code, established by inspecting each tag, and neither serving path was executed. Pin the model revision and an exact engine build before production.
What it is¶
nvidia/Qwen3.8-2.4T-A95B-NVFP4 is a post-training quantization of Qwen/Qwen3.8-2.4T-A95B, published under the NVIDIA Open Model License and ungated (the Hugging Face API reports gated: false, and the raw files fetch without a token). The base model's own license is the separate Qwen3.8-Max License, which NVIDIA's repository does not reproduce.
The architecture, read from config.json, is a hybrid-attention sparse MoE: 92 layers at hidden size 8192, with layer_types alternating three linear_attention layers to one full_attention layer (full_attention_interval: 4), giving 69 gated-DeltaNet layers and 23 gated-attention layers. Every layer carries 512 experts with num_experts_per_tok: 10 plus one shared expert, at moe_intermediate_size: 2048. Full attention uses 64 query heads and 4 KV heads at head_dim: 256 with partial_rotary_factor: 0.25. Vocabulary is 248,320, and max_position_embeddings is 262,144. There is one MTP (multi-token prediction) block, mtp_num_hidden_layers: 1, carrying its own full set of 512 experts.1
What makes it a mixed-precision checkpoint is the quantization config, whose quant_algo is literally MIXED_PRECISION with two groups:
| Group | Precision | Targets | Modules |
|---|---|---|---|
group_1 |
NVFP4, W4A4, group_size: 16 |
model.layers.N.mlp.experts |
92 |
group_0 |
FP8 E4M3, W8A8, per-tensor | linear_attn.{in_proj_qkv, in_proj_z, in_proj_a, in_proj_b, out_proj} and self_attn.{q,k,v,o}_proj |
437 |
group_0, declared only |
declared FP8, stored BF16 with no scale tensors | linear_attn.conv1d |
69 |
| (unquantized) | BF16 | embeddings, lm_head, shared experts, routers, layernorms |
remainder |
ignore |
BF16 | ["mtp*", "mtp.layers.0*"] |
the whole MTP block |
kv_cache_scheme is {"dynamic": false, "num_bits": 8, "type": "float"}, so the KV cache is FP8, not NVFP4. The routed experts are 96.92% of the parameters and take 92.33% of the bytes; everything else stays at 8 or 16 bits. That is the standard frontier-MoE shape NVFP4 describes, applied at 2.4T.
Why use it¶
- It is the difference between a 4.9 TB checkpoint and a 1.44 TB one. The BF16 original is 2,446,182,725,504 parameters at two bytes each; this build is 1,444,420,107,432 bytes of tensor payload, a 3.387x reduction. That moves the model from roughly 20 GPUs to 8.
- Accuracy holds on the card's own comparison. Unusually for a quantized release, the card publishes a
Baseline(BF16)row, so the comparison is against the exact model it quantized rather than against nothing.2 NVFP4 matches or exceeds BF16 on five of six benchmarks, with HLE the single regression at 40.55 against 41.43. - Weights and activations both drop to four bits on the experts, so this is a W4A4 build that uses Blackwell's tensor cores rather than an FP4 storage format dequantized in software.
- The expensive tensors were left alone. Embeddings,
lm_head, the shared expert, and the routers stay BF16. Quantizing a router is how a sparse model loses its expert selection, and this checkpoint does not.
When to use it (and when not)¶
Use it when:
- The hardware is Blackwell and there are eight GPUs of 288 GB class available. The card names its test hardware as "NVIDIA GB200, NVIDIA B300".
- The serving stack can be pinned to vLLM 0.27.0 or newer, or SGLang 0.5.17 or newer. The card states no version at all. These are the earliest release tags that contain the model-specific source, established here by inspecting each tag; neither launch path was executed, so they are source-presence floors rather than validated minimum serving versions.3
- 262,144 tokens of context is enough. See the context caveat below before planning for more.
Avoid it when:
- The GPUs are 192 GB parts. The lower bound derived below leaves at most 2.0 GB per GPU for KV cache and activations on a nominally 192 GB part read decimally, and 15.0 GB read as 192 GiB, against 90.3 GB on a 288 GB part. Those are ceilings on the budget, not estimates, because the weight figure is itself a floor. Measure the driver-reported usable memory and the engine's own reported weight footprint on the actual part rather than trusting the label.
- A 1M-token route is the requirement. The card advertises "Context length up to 1 million tokens" without qualification, but this checkpoint's
max_position_embeddingsis 262,144 and itsrope_typeis"default"with no scaling factor. The base model card is the precise one: "262,144 natively and extensible up to 1,010,000 tokens." Both of NVIDIA's own launch commands set 262,144. The 1M figure is not reachable from this checkpoint without an explicit RoPE-scaling override that neither card supplies. - Thinking mode must be disabled. The base model card states the model "requires thinking mode for all interactions" and that "thinking cannot be disabled", and the base chat template enforces it with
{{- raise_exception('Disabling thinking is not supported.') }}. NVIDIA's repackagedchat_template.jinjaremoves that guard and adds image and video content handling that the base template does not have, even though the base card calls the model text-only. Settingenable_thinking=falseagainst this build will therefore not raise; it will silently produce whatever an untested path produces. - Only one or two GPUs are available. This is not a model that degrades gracefully to a small box; for that end of the range see Colibri or prima.cpp.
Architecture¶
flowchart TB
TOK["Input tokens<br/>vocab 248,320"] --> STACK
subgraph STACK["92 layers, pattern repeated 23 times"]
GDN["3x gated DeltaNet<br/>(linear attention, FP8 projections)"] --> ATT["1x gated attention<br/>64 Q / 4 KV heads, head_dim 256, FP8"]
end
ATT --> MOE["MoE block, every layer<br/>512 experts, top-10 routed + 1 shared"]
MOE --> RT["Router (BF16)<br/>selects 10 of 512"]
RT --> EXP["Routed experts<br/>NVFP4 W4A4, block 16"]
RT --> SH["Shared expert (BF16)"]
MOE --> KV["KV cache: FP8 E4M3"]
EXP --> OUT["lm_head (BF16)"]
SH --> OUT
MTP["MTP block, 1 layer + 512 experts<br/>BF16, in the 'ignore' list"] -.->|"dropped at load by vLLM"| OUT
Two structural facts drive the deployment. First, the model is hybrid: 69 of 92 layers are linear-attention (gated DeltaNet) and only 23 are full attention, so the KV cache is far smaller than a 92-layer dense-attention model of this size would need, but the cache manager must hold two state shapes at once. That is the problem hybrid recurrent and full-attention serving covers, and it is why prefix caching on this family needs care. Second, the MTP block is dead weight on the default path: vLLM's Qwen3_5MoeForCausalLM maps orig_to_new_prefix={"model.language_model.": "model.", "mtp.": None}, which drops every mtp.* tensor at load. That 52.756 GB of tensor payload is a download, storage, and load-time cost, not a VRAM cost, unless MTP speculative decoding is enabled through the separate Qwen3_5MoeMTP path.
How to use it¶
1. Size the replica¶
| Item | Value |
|---|---|
| Model | nvidia/Qwen3.8-2.4T-A95B-NVFP4 |
| Base | Qwen/Qwen3.8-2.4T-A95B (BF16, Qwen3.8-Max License) |
| Params | 2,446,182,725,504 total, 95B active (10 routed + 1 shared expert per token) |
| Layers | 92 (69 gated DeltaNet, 23 gated attention) |
| Precision | NVFP4 W4A4 experts, FP8 W8A8 attention, FP8 KV cache, BF16 remainder |
| Tensor payload | 1,444,420,107,432 bytes (1444.4 GB) across 200 safetensors shards; the files are larger by their headers |
| Context | 262,144 per config.json (the card's "1 million" is the base model's extended ceiling, not this checkpoint's) |
| License | NVIDIA Open Model License, ungated |
| Quantizer | config.json records producer: {"name": "modelopt", "version": "0.42.0"}; the card's prose says v0.46.0 |
| Starting hardware | 8x Blackwell, 288 GB class (B300). Card's test hardware: GB200 and B300 |
| vLLM | 0.27.0 or newer (not stated on the card) |
| SGLang | 0.5.17 or newer (not stated on the card) |
2. vLLM server, exactly as the card publishes it¶
vllm serve nvidia/Qwen3.8-2.4T-A95B-NVFP4 \
--port 8000 \
--tensor-parallel-size 8 \
--max-model-len 262144 \
--reasoning-parser qwen3
This is the card's command verbatim. Two things it does not say. It sets no --gpu-memory-utilization, so vLLM's default of 0.92 applies (the value at both v0.27.0 and v0.28.0) and the weights alone take at least 174.63 GB per GPU before any KV cache (see below). And it enables no expert parallelism: with 512 experts per layer, --enable-expert-parallel is worth measuring against pure TP for this shape (expert parallelism).
3. SGLang server, as the card publishes it¶
python -m sglang.launch_server \
--model-path nvidia/Qwen3.8-2.4T-A95B-NVFP4 \
--port 8000 \
--tp-size 8 \
--context-length 262144 \
--reasoning-parser qwen3
Prefer the vLLM path unless SGLang is already the standard. python/sglang/srt/models/qwen3_5_text.py first appears at tag v0.5.17 (it 404s at v0.5.16), and the checkpoint declares transformers_version: "5.13.0" while SGLang pins transformers below that, so a stock install cannot satisfy the version that wrote the file.
4. Smoke test¶
curl -s http://<service-url>:8000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "nvidia/Qwen3.8-2.4T-A95B-NVFP4",
"messages": [{"role": "user", "content": "Prove that the sum of two odd integers is even."}],
"temperature": 1.0,
"top_p": 0.95,
"max_tokens": 2048,
"top_k": 20
}'
Pass criteria:
- The response separates reasoning from the answer, which is what
--reasoning-parser qwen3is for. If reasoning content arrives inline incontent, the parser is not attached. GET /v1/modelsreports the served context length as 262144, not 1000000.- Sampling matches the card's own evaluation settings,
temperature=1.0, top_p=0.95, top_k=20, before any comparison against its published scores. - Startup did not silently fall back: confirm the log reports the ModelOpt mixed-precision path and not a dequantize-to-BF16 route, which would not fit.
How to develop with it¶
The one number that decides whether this model runs on the hardware available is the checkpoint's tensor payload, and it can be derived from config.json without downloading 1.44 TB. Doing so also exposes where the card's own compression claim goes wrong. This block reconstructs the checkpoint's byte census from the config identity plus the published dtype counts, asserts it against the total_size recorded in the safetensors index, and then prices the card's own --tensor-parallel-size 8:
# qwen38_nvfp4_footprint.py - derive this checkpoint's tensor payload from
# config.json alone, check it against the published safetensors index, and bound
# the per-GPU weight memory the card's own TP=8 launch implies. Inputs are
# published values cited on this page; the arithmetic is this page's. Nothing
# here was run on a GPU, and no serving path was executed.
import numpy as np
# --- config.json (nvidia/Qwen3.8-2.4T-A95B-NVFP4) --------------------------
L, E, D, I = 92, 512, 8192, 2048 # layers, experts, hidden, moe_intermediate
GROUP = 16 # quantization_config group_size
# --- published censuses ----------------------------------------------------
U8_STORAGE = 1_185_410_973_696 # HF safetensors census, packed FP4 bytes
BF16_PARAMS = 35_470_849_920 # HF safetensors census
FP8_PARAMS = 39_889_928_192 # HF safetensors census (attention weights)
TOTAL_SIZE = 1_444_420_107_432 # index metadata.total_size (tensor payload)
BASE_PARAMS = 2_446_182_725_504 # Qwen/Qwen3.8-2.4T-A95B, all BF16
MTP_BYTES = 52_756_055_040 # summed from the shard headers, all BF16
# 1) The expert tensors are a closed-form identity from config.json: every layer
# holds E experts of three D x I projections. Two FP4 codes pack per byte.
experts = E * L * 3 * I * D
assert experts == 2_370_821_947_392
assert experts == U8_STORAGE * 2, "packed U8 bytes must be exactly half the FP4 codes"
# 2) NVFP4 costs 4 data bits plus one FP8 block scale per GROUP values.
bits_per_expert_param = 4 + 8 / GROUP
assert bits_per_expert_param == 4.5
expert_bytes = experts * bits_per_expert_param / 8
assert expert_bytes == U8_STORAGE + experts / GROUP
# 3) Census the whole checkpoint and compare with the published index total.
census = expert_bytes + FP8_PARAMS * 1 + BF16_PARAMS * 2
residual = TOTAL_SIZE - census
assert 0 < residual < 2e6 and residual / TOTAL_SIZE < 2e-6
# 4) Experts dominate parameters more than they dominate bytes: the 3.1% of
# parameters left in FP8/BF16 cost 8x to 16x more each.
total_params = experts + FP8_PARAMS + BF16_PARAMS
assert total_params == BASE_PARAMS, "no parameter added or lost by quantizing"
p_share, b_share = experts / total_params, expert_bytes / TOTAL_SIZE
assert p_share > b_share
# 5) Payload compression against the BF16 original. This is a storage ratio, not
# a live-VRAM ratio, and it cannot reach 4x: only the transformer-block linear
# operators were converted, and NVFP4 carries block scales on top of 4 bits.
ratio = (BASE_PARAMS * 2) / TOTAL_SIZE
assert 3.3 < ratio < 3.5
assert not (2.4 < ratio < 2.6), "the card's 'approximately 2.5x' is not the measured ratio"
assert ratio < 4.0
eff_bits = TOTAL_SIZE * 8 / total_params
assert eff_bits > bits_per_expert_param, "mixed precision costs more than pure NVFP4"
# 6) Adversarial: four bits times the exact base parameter count still understates
# the payload, because it ignores block scales and the FP8/BF16 remainder.
four_bit_floor = BASE_PARAMS * 4 / 8
assert four_bit_floor < TOTAL_SIZE
floor_err = (TOTAL_SIZE - four_bit_floor) / TOTAL_SIZE
# 7) Per-GPU weight memory at the card's --tensor-parallel-size 8. Dividing the
# payload by TP is a LOWER BOUND, not an estimate: TP shards the experts,
# attention, embeddings and shared expert, but REPLICATES the MoE router and
# the layernorms on every rank, so each replicated tensor costs (TP-1)/TP more
# than the quotient already charges. vLLM drops every mtp.* tensor on the
# Qwen3_5MoeForCausalLM path, so subtract the measured MTP payload.
TP = 8
router_total = L * E * D * 2 # BF16, replicated per rank
router_correction = router_total * (TP - 1) / TP # what the quotient omits
assert abs(router_total - 771_751_936) == 0
quotient_all = TOTAL_SIZE / TP
quotient_nomtp = (TOTAL_SIZE - MTP_BYTES) / TP
lower_all = quotient_all + router_correction
lower_nomtp = quotient_nomtp + router_correction
assert lower_nomtp > quotient_nomtp, "replication can only raise the requirement"
assert MTP_BYTES > E * 3 * I * D * 2, "MTP is more than its routed experts alone"
# vLLM's default --gpu-memory-utilization is 0.92 at v0.27.0 and v0.28.0.
UTIL = 0.92
def left(nominal, weights):
"""Upper bound on the KV + activation budget per GPU: it charges only the
weight lower bound, and no allocator, CUDA-graph or activation overhead."""
return nominal * UTIL - weights
parts = [("192 GB (decimal)", 192e9), ("192 GiB (binary)", 192 * 2**30),
("288 GB (decimal)", 288e9)]
budget = {n: left(c, lower_nomtp) for n, c in parts}
assert 0 < budget["192 GB (decimal)"] < 3e9, "192 GB decimal: no usable KV budget even at the bound"
assert 10e9 < budget["192 GiB (binary)"] < 20e9
assert budget["288 GB (decimal)"] > 85e9
assert left(192e9, lower_all) < 0, "with MTP resident, a 192 GB decimal part cannot load"
G = 1e9
print(f"experts {experts:>19,} params ({100*p_share:.2f}% of the model)")
print(f"expert bytes {int(expert_bytes):>19,} bytes ({100*b_share:.2f}% of the payload)")
print(f" = packed FP4 {U8_STORAGE:>19,} + FP8 block scales {int(experts/GROUP):,}")
print(f"census total {int(census):>19,} bytes")
print(f"published index total {TOTAL_SIZE:>19,} bytes (residual {int(residual):,}, {residual/TOTAL_SIZE:.2e})")
print()
print(f"BF16 payload {BASE_PARAMS*2/G:>10.1f} GB ({BASE_PARAMS:,} params)")
print(f"NVFP4 payload {TOTAL_SIZE/G:>10.1f} GB compression {ratio:.3f}x (card says ~2.5x)")
print(f"effective bits/param {eff_bits:>10.3f} (NVFP4 alone is {bits_per_expert_param})")
print(f"4 bits x base params {four_bit_floor/G:>10.1f} GB understates the payload by {100*floor_err:.1f}%")
print()
print(f"TP={TP} per-rank weight LOWER BOUND (quotient + {router_correction/1e6:.0f} MB replicated router):")
print(f" MTP resident {lower_all/G:>7.2f} GB (quotient alone {quotient_all/G:.2f} GB)")
print(f" MTP dropped {lower_nomtp/G:>7.2f} GB (quotient alone {quotient_nomtp/G:.2f} GB)")
print(f"KV + activation budget left at --gpu-memory-utilization {UTIL}, MTP dropped, at the bound:")
for n, _ in parts:
print(f" {n:18s} {budget[n]/G:>8.2f} GB")
print(f" with MTP resident, 192 GB (decimal) is short by {-left(192e9, lower_all)/G:.2f} GB")
Executed output (.venv/bin/python, Python 3.12.3, numpy 2.5.1, pasted verbatim):
experts 2,370,821,947,392 params (96.92% of the model)
expert bytes 1,333,587,345,408 bytes (92.33% of the payload)
= packed FP4 1,185,410,973,696 + FP8 block scales 148,176,371,712
census total 1,444,418,973,440 bytes
published index total 1,444,420,107,432 bytes (residual 1,133,992, 7.85e-07)
BF16 payload 4892.4 GB (2,446,182,725,504 params)
NVFP4 payload 1444.4 GB compression 3.387x (card says ~2.5x)
effective bits/param 4.724 (NVFP4 alone is 4.5)
4 bits x base params 1223.1 GB understates the payload by 15.3%
TP=8 per-rank weight LOWER BOUND (quotient + 675 MB replicated router):
MTP resident 181.23 GB (quotient alone 180.55 GB)
MTP dropped 174.63 GB (quotient alone 173.96 GB)
KV + activation budget left at --gpu-memory-utilization 0.92, MTP dropped, at the bound:
192 GB (decimal) 2.01 GB
192 GiB (binary) 15.03 GB
288 GB (decimal) 90.33 GB
with MTP resident, 192 GB (decimal) is short by 4.59 GB
The readings that matter for capacity planning. The census reproduces the published index total to seven parts in ten million, so the mixed-precision layout is fully accounted for and nothing large is hiding. The effective rate is 4.724 bits per parameter, not 4, because NVFP4 itself costs 4.5 once the FP8 block scale per 16 values is counted, and the 3.08% of parameters left in FP8 and BF16 cost 8 and 16 bits each. That is why four bits times the exact base parameter count, 1223.1 GB, understates the real payload by 15.3%: budgeting from it puts a 1.22 TB number in a capacity plan for a 1.44 TB payload. Rounding the parameter count to a literal 2.4T understates it further.
The compression claim on the card understates its own result while phrasing it in a way that invites overstating it. The card says the quantization "reduces the number of bits per parameter from 16 to 4, reducing the disk size and GPU memory requirements by approximately 2.5x". The measured payload ratio is 3.387x, so 2.5x is too low. The "16 to 4" phrasing reads as 4x, which the checkpoint cannot reach, and the card itself says why one sentence earlier: "Only the weights and activations of the linear operators within transformer blocks are quantized." Everything else stays at 8 or 16 bits, and NVFP4 carries a block scale on top of its four bits. Plan against 3.387x, and note that it is a storage ratio: live VRAM adds the KV cache, activations, and allocator overhead.
The per-GPU figure is a lower bound, not an estimate, and the distinction matters here. Dividing the payload by the TP degree is right for the tensors TP actually shards, which is the experts (92.33% of the bytes), the attention projections, the embeddings and the shared expert. But the MoE router and the layernorms are replicated on every rank, so the quotient charges each of them once where the machine pays for them eight times. The router is the largest of these at 92 layers by 512 experts by 8192 in BF16, 771.75 MB in total, of which the quotient already counts 96.47 MB; the remaining 675.28 MB is a correction the quotient omits. Adding it gives 174.63 GB per rank with the MTP tensors dropped and 181.23 GB with them resident. Allocator padding, CUDA graphs, activations, and non-Torch allocations all sit above that, so the engine's reported weight footprint will be higher again.
What that leaves to serve with is the number that matters, and because the weight figure is a floor, the budget figures are ceilings. At the default --gpu-memory-utilization 0.92 with MTP dropped, a 288 GB part leaves at most 90.33 GB per GPU for KV cache and activations. A part labelled 192 GB leaves at most 2.01 GB read decimally and 15.03 GB read as 192 GiB. So a 192 GB part is not a viable target: on the decimal reading of its own label there is no usable KV budget even before the omitted overheads, and keeping MTP resident puts it 4.59 GB short of loading at all. Since the card names GB200 among its test hardware and states no VRAM figure anywhere, measure the driver-reported usable memory and the engine's own weight report on the actual part before committing capacity.
How to run it in production¶
# Reference template, not executed here. Pin the model revision and an exact
# vLLM image (0.27.0 or newer) before promoting this.
apiVersion: apps/v1
kind: Deployment
metadata:
name: vllm-qwen38-nvfp4
namespace: serving
spec:
replicas: 1
selector:
matchLabels: { app: vllm-qwen38-nvfp4 }
template:
metadata:
labels: { app: vllm-qwen38-nvfp4 }
spec:
nodeSelector:
accelerator.nvidia.com/class: b300-8gpu
containers:
- name: vllm
image: vllm/vllm-openai:<pinned-tag-0.27.0-or-newer>
args:
- --model=nvidia/Qwen3.8-2.4T-A95B-NVFP4
- --served-model-name=qwen38-nvfp4
- --tensor-parallel-size=8
- --max-model-len=262144
- --reasoning-parser=qwen3
ports:
- { containerPort: 8000, name: http }
resources:
limits: { nvidia.com/gpu: 8 }
requests: { nvidia.com/gpu: 8 }
readinessProbe:
httpGet: { path: /health, port: 8000 }
initialDelaySeconds: 1800
failureThreshold: 60
volumeMounts:
- { name: dshm, mountPath: /dev/shm }
- { name: hf-cache, mountPath: /root/.cache/huggingface }
volumes:
- name: dshm
emptyDir: { medium: Memory, sizeLimit: 64Gi }
- name: hf-cache
persistentVolumeClaim: { claimName: hf-cache-2tb }
---
apiVersion: v1
kind: Service
metadata:
name: vllm-qwen38-nvfp4
namespace: serving
spec:
selector: { app: vllm-qwen38-nvfp4 }
ports:
- { name: http, port: 8000, targetPort: 8000 }
- Give the readiness probe a very long grace period. 200 shards and 1.44 TB is not a two-minute load; the 1800-second
initialDelaySecondsabove is a starting point to measure, not a recommendation. - Provision at least 2 TB of persistent cache and never let a pod re-download the checkpoint on reschedule. Weight loading at this size is its own discipline (engine weight loading).
- Cap
--max-model-lento the route's real need. 262,144 is a ceiling. The KV cache is FP8 and only 23 of 92 layers are full attention, which helps, but the remaining budget after 174 GB of weights per GPU is what decides concurrency (KV cache management). - Measure expert parallelism against pure TP. 512 experts per layer at top-10 is a routing shape where the all-to-all cost and the grouped-GEMM efficiency both matter (expert parallelism, MoE expert backends).
- Gate on your own evals. The card's numbers come from a single unstated harness with no seeds or intervals; treat them as a claim about NVIDIA's setup (eval gate).
How to maintain it¶
- Pin the model revision. The repository ships no LICENSE file of its own and the card's text has already proven inconsistent with its artifact in four places; a silent re-upload would be hard to notice.
- Track the engine floor, not just the engine.
Qwen3_5MoeForCausalLMis absent from vLLM's registry atv0.26.0and present fromv0.27.0; SGLang'sqwen3_5_text.pyis absent atv0.5.16and present atv0.5.17. Both are source-presence floors, not versions anyone has been observed to serve this checkpoint on. Enablement for this family was still landing upstream around the card's publication, so prefer a recent pinned build and re-run the smoke test after any move. - Re-check the chat template on every revision bump. NVIDIA's template differs from the base model's in two behaviourally significant ways (the removed thinking guard, the added image and video branches). If a future revision restores the guard, clients that currently pass
enable_thinking=falsewithout error will start raising. - Watch the
conv1dmodules. All 69linear_attn.conv1dentries are declared{"quant_algo": "FP8"}inquantized_layers, but the shipped tensors areBF16of shape[20480, 1, 4]with noinput_scaleand noweight_scale, while every genuinely FP8 module has both. The declaration and the artifact disagree; which of a stale target list, a skipped export, or something else produced that is not recoverable from the files. It is inert on vLLM, which buildsconv1dwith no quantization config attached, but a loader that trusts the declaration and looks for scales that do not exist will fail. Verify on any engine other than vLLM before promoting. - Do not chase the card's version string. The card says Model Optimizer v0.46.0; the checkpoint's own
config.jsonrecordsproducer: {"name": "modelopt", "version": "0.42.0"}. Reproduction attempts should start from the recorded producer version.
Failure modes¶
- No usable KV budget on a 192 GB part. At least 174.63 GB of weights per GPU against a 0.92 utilization budget leaves at most 2.01 GB on the decimal reading of the label, before allocator and activation overhead. The symptom is either a failure during weight loading or a server that starts and then admits almost no concurrency. Move to 288 GB parts, or increase tensor parallelism beyond 8.
- Requests rejected above 262,144 tokens after capacity was planned for the card's advertised 1M. The checkpoint's
max_position_embeddingsis 262,144 with no RoPE scaling configured. enable_thinking=falsesilently accepted. NVIDIA's template dropped the base model's guard. The base card says thinking cannot be disabled, so this path is untested rather than supported.- Multimodal input silently accepted. NVIDIA's template has image and video branches; the base card says the model is text-only and "Multimodal inputs are not supported".
- Reasoning text leaking into
content.--reasoning-parser qwen3is missing or the engine build predates its registration. - Capacity planned from "2.4T at 4 bits". Taken literally that estimate is 1200 GB against a real 1444 GB, a 16.9% shortfall; even four bits times the exact 2,446,182,725,504 parameter count gives 1223.1 GB and a 15.3% shortfall, because both ignore the FP8 block scales and the FP8/BF16 remainder.
- Benchmarks that will not reconcile with the card. The card's results table has an HLE column, but HLE is absent from its own "Datasets" list, and the τ²-Bench Telecom benchmark described at length in its Properties prose has no row in the table at all. No harness, seed count, or interval is given for any of them.
References¶
- NVIDIA Qwen3.8-2.4T-A95B-NVFP4 model card, config, and safetensors index (the primary source for every quantization, size, and benchmark figure here): https://huggingface.co/nvidia/Qwen3.8-2.4T-A95B-NVFP4
- Qwen3.8-2.4T-A95B base model card (parameter counts, the 262,144-native context statement, the thinking-mode and text-only statements): https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B
- NVIDIA Open Model Agreement (the governing terms the card names): https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-open-model-agreement/
- NVIDIA TensorRT Model Optimizer (the
modeloptproducer recorded inconfig.json): https://github.com/NVIDIA/Model-Optimizer - vLLM
Qwen3_5MoeForCausalLMimplementation, including themtp.weight-drop mapping: https://github.com/vllm-project/vllm/blob/main/vllm/model_executor/models/qwen3_5.py - vLLM model registry (used to establish the 0.27.0 source-presence floor): https://github.com/vllm-project/vllm/blob/main/vllm/model_executor/models/registry.py
- SGLang (used to establish the 0.5.17 source-presence floor via
python/sglang/srt/models/qwen3_5_text.py): https://github.com/sgl-project/sglang
Related: NVFP4, Quantization for inference, MoE sparse scaling, Expert parallelism, MoE expert backends, Hybrid recurrent and full-attention serving, Blackwell platform, Engine weight loading, KV cache management, Generic vLLM recipe, Open-weight serving, Qwen3-235B-A22B, Kimi K3 multi-node, LLM evaluation harness
-
All architecture values quoted from
https://huggingface.co/nvidia/Qwen3.8-2.4T-A95B-NVFP4/raw/main/config.json, retrieved 2026-08-30:architectures: ["Qwen3_5MoeForCausalLM"],model_type: "qwen3_5_moe_text",num_hidden_layers: 92,hidden_size: 8192,num_experts: 512,num_experts_per_tok: 10,moe_intermediate_size: 2048,shared_expert_intermediate_size: 2048,num_attention_heads: 64,num_key_value_heads: 4,head_dim: 256,vocab_size: 248320,max_position_embeddings: 262144,full_attention_interval: 4,mtp_num_hidden_layers: 1,dtype: "bfloat16",transformers_version: "5.13.0", andrope_parameters: {"mrope_interleaved": true, "mrope_section": [11, 11, 10], "partial_rotary_factor": 0.25, "rope_theta": 10000000, "rope_type": "default"}. Note the field isdtype, nottorch_dtype, andrope_thetasits insiderope_parametersrather than at top level. The 69/23 split is a count over the 92-elementlayer_typesarray. The quantization block reportsquant_algo: "MIXED_PRECISION",quant_method: "modelopt",producer: {"name": "modelopt", "version": "0.42.0"},ignore: ["mtp*", "mtp.layers.0*"],kv_cache_scheme: {"dynamic": false, "num_bits": 8, "type": "float"}, and twoconfig_groups:group_0atnum_bits: 8, type: "float"over 506 targets andgroup_1atnum_bits: 4, type: "float", group_size: 16over 92mlp.expertstargets. A shard-header read confirms the layout: expert.weightisU8of shape[2048, 4096](two FP4 codes per byte) and.weight_scaleisF8_E4M3of shape[2048, 512], which is 8192/16 block scales per row. The byte total 1,444,420,107,432 ismetadata.total_sizefrommodel.safetensors.index.json, which is the tensor payload and excludes the 200 files' own headers; the parameter censuses are the Hugging Face API'ssafetensors.parametersfor each repository, which does not include the NVFP4 block-scale tensors in its dtype counts. Two figures were measured here rather than published: the 69linear_attn.conv1d.weighttensors areBF16of shape[20480, 1, 4]and the index contains noconv1dtensor other than.weight, so the{"quant_algo": "FP8"}declaration inquantized_layersis not reflected in the artifact; and the totalmtp.*payload is 52,756,055,040 bytes, all BF16, summed over the 1,553 MTP tensors across the eight shards that hold them, which is more than the 51,539,607,552 bytes of its routed experts alone. ↩ -
Card results table, verbatim, both rows: Precision / GPQA Diamond / HLE / SciCode / AA-LCR / IFBench / Terminal Bench 2.1;
Baseline(BF16)92.55 / 41.43 / 54.44 / 71.5 / 79.93 / 76.03;NVFP492.58 / 40.55 / 56.21 / 71.63 / 81.73 / 76.4. Its footnote: "Baseline: Qwen3.8-2.4T-A95B. Benchmarked with temperature=1.0, top_p=0.95, top_k=20. GPQA Diamond, SciCode, AA-LCR and IFBench used max_new_tokens=65,536; HLE used max_new_tokens=131,072; Terminal Bench 2.1 used max_new_tokens=262,144." The card's own "Datasets" line lists "GPQA Diamond, SciCode, IFBench, AA-LCR, Terminal Bench 2.1", omitting HLE, and its Properties paragraph describes τ²-Bench Telecom, which appears in neither the list nor the table. No eval harness, seed count, or confidence interval is stated. ↩ -
Established by checking the architecture's presence at each release tag rather than from any documentation.
Qwen3_5MoeForCausalLMreturns zero matches invllm/model_executor/models/registry.pyat tagv0.26.0and one match atv0.27.0andv0.28.0.python/sglang/srt/models/qwen3_5_text.pyreturns HTTP 404 at tagv0.5.16and HTTP 200 atv0.5.17andv0.5.18. Themtp.drop isorig_to_new_prefix={"model.language_model.": "model.", "mtp.": None}at line 322 ofvllm/model_executor/models/qwen3_5.pyonmain; a separateQwen3_5MoeMTPentry exists in the registry for the speculative-decoding path. The chat-template comparison is between the two repositories'chat_template.jinjafiles: the base contains{{- raise_exception('Disabling thinking is not supported.') }}and no image or video branch; NVIDIA's contains neither the guard nor its absence of media handling, testingitem.type == 'image'anditem.type == 'video'. vLLM publishes no recipe for this model family as of 2026-08-30; theQwen/recipe directory contains aQwen3.5.mdcovering a different checkpoint. ↩