Runbook: fabric qualification gate before a platform change¶
Scope: running the fabric ladder as a gate rather than as an investigation. What one qualification run has to record, how a baseline is keyed so two topologies are never compared, the parsing traps that let a run report success having measured nothing, and the decision rule that separates a node fault from a fleet-wide change without flapping.
Run this before a driver, CUDA, NCCL, NIC firmware, switch software or cabling change reaches production, when commissioning new nodes, and on a rotating sample of the fleet. If you are already in an incident, you want the regression ladder instead.
Commands are reference templates on real APIs and were not run against a fabric during authoring.
nccl-testsflag availability was verified by walking the upstream release tags; DCGM suite membership and option parsing were read from NVIDIA's published reference and the DCGM source; perftest units were read fromperftest_parameters.h. The Python block was executed; its output is pasted verbatim.
The ladder in the regression runbook answers "which layer diverged". A gate answers a narrower question under time pressure: may this change proceed. That difference is why a gate needs things an investigation does not, namely a machine-readable record, a baseline keyed to a topology rather than a number, and a rule that is stable enough to run unattended. The rung procedures themselves are not repeated here; neither are the DCGM run levels, the standing per-pair matrix, the node-level scheduler plumbing, or the acceptance framing.
The failure this runbook is built around is not a slow fabric. It is a gate that says pass because it never measured anything.
Trigger¶
- A driver, CUDA, NCCL, GPU firmware, NIC firmware or switch software change is staged for the fleet.
- New nodes are being commissioned, or capacity is being added (capacity add).
- Cabling or topology changed, including a rail moved between leaves.
- A scheduled re-qualification of a rotating fleet sample is due.
- A distributed workload regressed and the fleet needs re-baselining after the fix.
Pre-checks¶
- The nodes under test are drained. The level-3 DCGM plugins are explicit that they need idle GPUs: the
diagnosticplugin "should run on GPUs that are not serving production workloads", andtargeted_stressandtargeted_power"should run on idle GPUs".1 Level 4 escalates that to an idle system. This is the opposite of the pre-job dispatch gate, which stays at level 1 precisely because the node is busy (GPU health gating). - A baseline exists for this exact topology class. If it does not, this run is establishing one, not gating against one. Say which it is.
- The tool versions are pinned and recorded, because two of the gate's inputs are version-dependent in ways that fail silently. See the parsing traps below.
DCGM_NCCL_TESTS_BIN_PATHis set, in the host engine's environment, if you intend the DCGMnccl_testsplugin to run. NVIDIA requires it to be set beforenv-hostengineor thenvidia-dcgmservice starts, because the plugin reads it in the host-engine process; exporting it in the operator's shell beforedcgmi diaghas no effect. The two failure shapes differ and only one of them is visible: with the variable unset the plugin registers no test at all, sonccl_testsis simply absent from the report, while a variable pointing at a missing or non-regular path produces a genuine skip. Other bad paths do not skip at all: unsafe ownership or permissions, a launch failure, a non-zero exit or missing expected output are reported as a failure, which is a third outcome the gate has to keep distinct from both.4
Flow¶
flowchart TB
A["Change staged"] --> B["Drain sample nodes"]
B --> C["Run suite: DCGM, nccl-tests, perftest"]
C --> D["Parse into records"]
D --> E{"Did every step<br/>actually measure?"}
E -->|"no"| F["no-measurement: fix the harness"]
E -->|"yes"| G{"Same topology class<br/>as the baseline?"}
G -->|"no"| H["incomparable: get the right baseline"]
G -->|"yes"| I{"Cohort moved too?"}
I -->|"yes"| J["fleet-regression: blame the change"]
I -->|"no"| K["pass / warn / block on this node"]
Procedure¶
1. Pin the profile¶
Record the same profile the ladder's rung 0 records, and treat it as the baseline key rather than as metadata. A baseline is a sentence, not a number: "8 x H100 SXM, two nodes, one rail" and "8 x H100 SXM, eight nodes, four rails" are different baselines and never versions of each other.
2. Run the suite on drained nodes¶
-j is a boolean switch, not a path. --fail-early is long-only and has no short form, and --check-interval is rejected unless --fail-early is also set. The level aliases that parse are quick, short, medium, long and xlong; extended appears in legacy help text but is not accepted by the parser, so -r extended is read as a test name and fails.2
Then the collective and RDMA rungs, using the same invocations that produced the baseline (regression ladder, rungs 2 to 5) and the error-counter procedure around them.
3. Parse into records, and treat every parse failure as a failure¶
Four traps turn a green gate into a meaningless one. Each is silent.
An unknown nccl-tests flag exits 0 without benchmarking. The per-iteration timing flag -I first appears in release v2.19.2, and -K in v2.19.6; both are absent from every earlier tag. Given an option it does not recognise, getopt_long returns '?', so glibc writes invalid option -- 'I' to stderr, the binary then prints invalid option '?' to stdout, and it executes return 0 after the usage block.3 A harness keyed on exit status records a successful qualification run containing no measurement. Gate on a parsed number, never on an exit code, and probe the binary first:
#wrong reads N/A rather than 0 when checking is off. Correctness reporting follows the check-iteration count, so -c 0 prints N/A in that column. Parsing N/A as zero claims correctness was verified when it was never checked.
A skipped DCGM test is not a passed one. Read per-test status out of the JSON rather than the overall verdict, which does not distinguish them.
perftest reports MiB/sec by default, not MB/sec. On current releases the header reads BW peak[MiB/sec] and BW average[MiB/sec], and --report_gbits switches it to BW peak[Gb/sec]; builds at or before the 24.04 release printed MB/sec for the identical mebibyte value, so read the header rather than assuming it. perftest's own help gives the conversion as Factor = 10^9/(2^20*8) = 119.2.6 Treating a MiB/sec figure as MB/sec and multiplying by 8 divides by 125 instead of 119.2, which reads about 4.6% low: 12,000 MiB/s is 100.7 Gb/s and the mistake reports 96.0. That is inside the range a gate is trying to resolve, and it errs toward failing healthy hardware.
For DCGM specifically, the exit code carries information worth keeping: 226 means the diagnostic ran and reported an error, 205 means it reported a condition that requires isolation, 217 means another diagnostic was already running, and 204 means the NVVS binary was not found.5 205 is unambiguously a node verdict and the strongest of them, since DCGM is asking for the GPU to be isolated. 217 and 204 are harness problems and must not be recorded as node failures. 226 is mixed: it covers a genuine test failure but DCGM also returns it for harness conditions such as a failed child wait or a malformed response, so parse the per-test results out of the JSON before attributing it to hardware rather than cordoning on the exit code alone. If dcgmi itself is killed by a signal the shell reports 128 plus the signal number, which describes the process rather than any diagnostic result.
4. Decide¶
The rule is in the executed block below. It is deliberately conservative in one direction: anything it cannot interpret blocks rather than passes.
5. Land the verdict¶
A verdict that does not change scheduling is a report, not a gate. Every verdict the rule can return needs a landing place:
| Verdict | Where it lands |
|---|---|
corrupt |
Stop the rollout. Wrong answers outrank every performance question. |
block |
Cordon the node and route to the regression ladder. |
fleet-regression |
Escalate to the change owner, not the node's owner. Do not cordon the fleet. |
warn |
Record and re-measure. Repeated warnings escalate rather than accumulate silently. |
diag-failed |
A DCGM plugin reported a failure. Route to the node's owner, not the harness queue. |
no-measurement, diag-incomplete, no-baseline, unstable, implausible, invalid-input |
Hold in inventory and fix the harness. These are not clean nodes; they are absent measurements. |
bimodal |
The median looks fine while a quarter of the run is well under baseline. Re-measure and look for a per-rank or per-rail split before grading. |
insufficient-samples |
Re-run with more samples. The question was not answered, so neither passing nor blocking is supported. |
correctness-unverified |
Re-run with checking on. Never release on it. |
incomparable |
Get the right baseline, or establish one and say so. |
The decision rule¶
import numbers
# gate.py -- validated: decide a verdict for one qualification run against a baseline,
# and refuse to pass anything the run did not actually measure.
# numpy only.
import numpy as np
WARN, BLOCK = 0.05, 0.10 # matches the fabric-regression runbook's policy table
# Chosen conventions, not derived quantities. Both bound "the run measured something
# other than what the baseline measured" rather than "the fabric is slow".
FAST, SLOW = 1.5, 0.25
# The gate asserts on a NAMED SET, because a test that never registered contributes no
# status at all and an absent status is not a failed one.
REQUIRED_DIAG = ("pcie", "memory", "nccl_tests")
def _drop(samples, baseline_median):
"""Fractional shortfall of this run's median against the baseline median."""
return 1.0 - float(np.median(samples)) / baseline_median
def _noise_bar(cv, n):
"""Two standard errors of the baseline's own run-to-run spread.
cv/sqrt(n) is the standard error of the MEAN, used here for the median's, which is
wider by about 25% on normal data. The substitution makes the bar mildly permissive.
It is the requirement to repeat, not the constant, that carries the argument.
"""
return 2.0 * cv / np.sqrt(n)
def _significant(drop, bar, cv, n):
"""A shortfall counts only if it clears both the policy bar and the measurement noise."""
return drop > bar and drop > _noise_bar(cv, n)
# Verdicts that are not statements about performance. None of them may be admitted, and
# none may be rewritten by hysteresis, which only smooths pass / warn / block.
HARD = ("corrupt", "correctness-unverified", "invalid-input", "no-measurement",
"no-baseline", "diag-incomplete", "diag-failed", "incomparable", "unstable",
"implausible", "insufficient-samples")
DIAG_STATES = ("pass", "skip", "fail")
def _finite(*vals):
"""NaN and infinity compare false against every threshold, so a rule that does not
reject them fails OPEN: every rejection branch is skipped and the verdict is pass.
numbers.Real rather than (int, float) so numpy scalars are accepted consistently:
np.float64 subclasses float but np.int64 and np.float32 do not, and an accept rule
resting on that inheritance accepts or rejects by accident.
"""
for v in vals:
if isinstance(v, bool) or not isinstance(v, numbers.Real):
return False
if v != v or v in (float("inf"), float("-inf")):
return False
return True
def _unmeasured(record, baseline):
"""Every way a run can fail to be a measurement. Checked before any arithmetic."""
# Corruption is baseline-independent and outranks everything, so it really is first.
# Ordering it after the sample and baseline checks lets a rollout-stopping result be
# downgraded to a harness repair.
wrong = record.get("wrong", "missing")
if wrong == "missing" or (wrong is not None and not _finite(wrong)):
return "invalid-input"
if wrong is None:
return "correctness-unverified" # the #wrong column read N/A
if wrong < 0:
return "invalid-input" # a negative count of wrong elements is not a count
if wrong > 0:
return "corrupt"
if not record.get("samples"):
return "no-measurement" # exit status 0 is not a measurement
if not _finite(*record["samples"]) or any(s <= 0 for s in record["samples"]):
return "invalid-input" # a non-positive bandwidth is not a measurement
if (not _finite(baseline.get("median"), baseline.get("cv"))
or baseline["median"] <= 0 or baseline["cv"] <= 0):
return "no-baseline" # a non-positive spread cannot describe run-to-run noise # absent, zero or non-finite
diag = record.get("diag") or {}
if any(s not in DIAG_STATES for s in diag.values()):
return "invalid-input" # an unrecognised status is not a pass
# Absent, skipped and failed are three different things, and the page says so. A rule
# that collapses them reports a diagnostic that FAILED as merely missing, which routes
# a hardware finding to the harness queue.
if any(t not in diag for t in REQUIRED_DIAG) or any(s == "skip" for s in diag.values()):
return "diag-incomplete"
if any(s == "fail" for s in diag.values()):
return "diag-failed"
if record.get("class") != baseline.get("class"):
return "incomparable" # a baseline is a topology, not a number
spread = float(np.std(record["samples"]) / np.mean(record["samples"]))
if len(record["samples"]) > 1 and spread > 3.0 * baseline["cv"]:
return "unstable" # the run disagrees with itself
return None
def gate(record, baseline, cohort_drops=()):
"""Verdict for one candidate measurement. Anything not clearly measured blocks."""
unmeasured = _unmeasured(record, baseline)
if unmeasured:
return unmeasured
ratio = float(np.median(record["samples"])) / baseline["median"]
if ratio >= FAST or ratio <= SLOW:
return "implausible"
drop, n = _drop(record["samples"], baseline["median"]), len(record["samples"])
# A median describes the middle, not the population. A run that is half at baseline and
# half well under it has a clean median and a spread that can stay inside the
# instability bar, so the slow fraction is checked separately, and only where the
# aggregate already looks acceptable. A uniformly slow run is not bimodal, it is slow,
# and belongs in the ordinary warn/block path below.
slow = float(np.mean(np.asarray(record["samples"]) < baseline["median"] * (1 - BLOCK)))
if not _significant(drop, WARN, baseline["cv"], n):
if slow >= 0.25:
return "bimodal"
# Failing to detect a regression is not evidence of equivalence. If the run could
# not have resolved the policy bar at this sample count, say so instead of passing.
if _noise_bar(baseline["cv"], n) > WARN and drop > 0:
return "insufficient-samples"
return "pass"
# A fleet event excuses a node only if the node is NO WORSE than the fleet, within
# noise. The test is deliberately one-sided: a node better than the cohort is still
# explained by the fleet event. Comparing against the cohort median with no bound at
# all is what would excuse an arbitrarily bad node.
# Cohort data is validated exactly as strictly as candidate data. Unvalidated, five
# infinities or five booleans satisfy the majority test and excuse an arbitrarily bad
# node, which makes the alibi easier to obtain than the accusation.
cohort = [d for d in cohort_drops if _finite(d) and -1.0 <= d <= 1.0]
if len(cohort) == len(cohort_drops) and len(cohort) >= 5 and np.mean(np.asarray(cohort) > WARN) >= 0.6:
if drop - float(np.median(cohort)) <= _noise_bar(baseline["cv"], n):
return "fleet-regression"
if _significant(drop, BLOCK, baseline["cv"], n):
return "block"
return "warn"
def with_hysteresis(verdict, raw_history, blocked=False, clear_after=2, escalate_after=3): # noqa: C901
"""Damp both edges, from RAW gate verdicts plus an explicit blocked flag.
Taking history as raw verdicts and blocked-ness separately is what makes this
well-defined. Feeding it back its own output cannot work: store effective verdicts and
a block never clears, because each attempted pass is recorded as another block; store
raw ones with no flag and a block raised by accumulated warnings clears on one pass.
Returns (effective verdict, still blocked).
"""
# bool is a subclass of int and True >= 1, so an unexcluded bool passes as the count 1.
if any(isinstance(v, bool) or not isinstance(v, int) or v < 1
for v in (clear_after, escalate_after)):
raise ValueError("thresholds are counts of runs, minimum 1")
# Hysteresis smooths a noisy performance signal. It has no business rewriting a verdict
# that is not about performance: a corrupt or unmeasured run must stay non-admissible,
# and must not be relabelled "block" either, which would erase the reason.
if verdict in HARD:
return verdict, True
recent = [v for v in raw_history if v not in HARD] + [verdict]
if blocked:
cleared = len(recent) >= clear_after and all(v == "pass" for v in recent[-clear_after:])
return ("pass", False) if cleared else ("block", True)
if verdict == "block":
return "block", True
if len(recent) >= escalate_after and all(v == "warn" for v in recent[-escalate_after:]):
return "block", True
return verdict, False
BASE = {"class": "8xH100-SXM/2node/1rail", "median": 240.0, "cv": 0.04}
def rec(**kw):
r = {"class": BASE["class"], "samples": [240.0] * 9, "wrong": 0,
"diag": {t: "pass" for t in REQUIRED_DIAG}}
r.update(kw)
return r
# (1) A run at baseline passes.
assert gate(rec(), BASE) == "pass"
# (2) Adversarial, and the reason this gate exists: nccl-tests answers an unknown option
# by printing usage and returning 0. A gate keyed on exit status records a pass having
# measured nothing. No samples is never a pass.
assert gate(rec(samples=[]), BASE) == "no-measurement"
# (3) Adversarial: an ABSENT diagnostic block must not read as a pass. DCGM does not
# report a skip when DCGM_NCCL_TESTS_BIN_PATH is unset; the plugin registers no test at
# all, so its status is missing rather than failed, and `any()` over an empty dict is
# False. The gate therefore asserts on a named set, not on the statuses it happens to find.
bare = rec()
del bare["diag"]
assert gate(bare, BASE) == "diag-incomplete"
assert gate(rec(diag={"pcie": "pass", "memory": "pass"}), BASE) == "diag-incomplete"
# (4) A test that ran and was skipped, which is what a WRONG path produces, is also not a pass.
assert gate(rec(diag={"pcie": "pass", "memory": "pass", "nccl_tests": "skip"}), BASE) == "diag-incomplete"
# (5) With -c 0 the #wrong column prints N/A. Reading absent as zero claims correctness
# was verified when it was never checked.
assert gate(rec(wrong=None), BASE) == "correctness-unverified"
# (6) Adversarial ordering: data corruption needs no baseline to be meaningful, so it must
# not be masked by a mismatched topology. Checking the class first would report this run as
# merely incomparable and discard the corruption signal.
assert gate(rec(wrong=3, **{"class": "8xH100-SXM/8node/4rail"}), BASE) == "corrupt"
# (7) A baseline is a topology, not a number. An 8-node result against a 2-node baseline
# compares two systems, and the arithmetic works perfectly on both.
assert gate(rec(**{"class": "8xH100-SXM/8node/4rail"}), BASE) == "incomparable"
# (8) A missing or zero baseline divides by nothing. Fail closed rather than raise.
assert gate(rec(), {"class": BASE["class"], "median": 0.0, "cv": 0.04}) == "no-baseline"
# (9) Adversarial: a run that disagrees with ITSELF has not measured a stable quantity,
# even though its median sits exactly on baseline and the sample count looks reassuring.
assert gate(rec(samples=[100., 100., 240., 380., 380.]), BASE) == "unstable"
# (10) Too good is a defect, not a win, and so is absurdly bad: both mean the run measured
# something other than what the baseline measured. 1.5 and 0.25 are chosen conventions.
assert gate(rec(samples=[380.0] * 9), BASE) == "implausible"
assert gate(rec(samples=[1.0] * 9), BASE) == "implausible" # a parser fault, not a fabric fault
# (11) Adversarial statistics: a 7% shortfall from ONE sample, on a fleet whose own
# run-to-run cv is 4%, is inside the noise: one sample admits nothing under 8%. The policy
# bar alone would raise a warning and send someone to investigate a fabric that is fine.
# But failing to DETECT a regression is not evidence of equivalence, so the honest verdict
# is that this run could not resolve the question, not that the node passed.
one = rec(samples=[223.2]) # 7% under 240
assert _drop(one["samples"], BASE["median"]) > WARN # the bar says warn
assert gate(one, BASE) == "insufficient-samples" # and one sample cannot settle it
# (12) The same 7% measured nine times is a finding: nine samples narrow the noise bar from
# 8% to 2.7%. Repetition, not a lower threshold, turns a suspicion into a verdict.
assert gate(rec(samples=[223.2] * 9), BASE) == "warn"
# (12b) The noise bar is a strict inequality, so a shortfall landing exactly on two
# standard errors reads as pass. Do not build a demonstration on that boundary: a value
# nominally equal to 2 SE can compute either side of it depending on how it was produced.
assert _significant(0.08000001, WARN, 0.04, 1) and not _significant(0.08, WARN, 0.04, 1)
# (13) Adversarial attribution: a 12% drop the whole cohort shares is a fleet event, a
# driver or NCCL rollout, not this node. Blaming the node sends a healthy card to RMA.
fleet = (0.12, 0.115, 0.125, 0.118, 0.122)
assert gate(rec(samples=[211.2] * 9), BASE, cohort_drops=fleet) == "fleet-regression"
# (14) Adversarial: the fleet event must not become an alibi for an arbitrarily bad node.
# Comparing only against the cohort MEDIAN excuses a node at half of baseline because the
# fleet moved 6%. The node has to be indistinguishable from the fleet, not merely joined to it.
mild = (0.06, 0.055, 0.051, 0.06, 0.058)
assert gate(rec(samples=[120.0] * 9), BASE, cohort_drops=mild) == "block"
# (15) Adversarial: two faulty nodes in a three-member cohort are a majority of its median,
# so a small cohort lets the second faulty node alibi the first. Requiring five members and
# a clear majority past the bar is what stops that, not a count of three.
assert gate(rec(samples=[211.2] * 9), BASE, cohort_drops=(0.12, 0.11, 0.0)) == "block"
# (16) The identical measurement against a healthy cohort blocks the node.
assert gate(rec(samples=[211.2] * 9), BASE, cohort_drops=(0.0, 0.01, 0.0, 0.005, 0.0)) == "block"
# (17) Hysteresis damps both edges. One good run does not clear a block, and warnings
# accumulate into one. Without the first, a node near the bar rejoins the fleet every
# other run; without the second, a node warns forever and is never acted on.
assert with_hysteresis("pass", ["warn"], blocked=True) == ("block", True)
assert with_hysteresis("pass", ["pass"], blocked=True) == ("pass", False)
assert with_hysteresis("warn", ["warn", "warn"], blocked=False) == ("block", True)
# (17b) Adversarial: the history has to be RAW verdicts plus an explicit blocked flag.
# Feeding the function its own output breaks both ways, and the naive single-argument
# version silently picked one of these failures depending on what the caller stored.
raw = ["warn", "warn", "warn"] # third warn escalates to block
v, blocked = with_hysteresis(raw[-1], raw[:-1], blocked=False)
assert (v, blocked) == ("block", True)
# One good run must NOT clear that block, which is what a raw-history-only rule would do.
assert with_hysteresis("pass", raw, blocked=blocked) == ("block", True)
# Two consecutive good runs clear it.
assert with_hysteresis("pass", raw + ["pass"], blocked=blocked) == ("pass", False)
# (17c) Both thresholds are counts of runs and must be positive integers. The guard raises
# rather than asserts, so `python -O` cannot strip it.
for bad in [0, -1, 2.5, True]:
try:
with_hysteresis("pass", [], blocked=True, clear_after=bad)
raise SystemExit(f"threshold guard did not fire for {bad!r}")
except ValueError:
pass
# (17d) Adversarial: hysteresis smooths a performance signal and must not touch a verdict
# that is not one. Returning (verdict, False) for a corrupt run hands a caller that trusts
# the boolean an admissible result; rewriting it to "block" erases why it was refused.
assert with_hysteresis("corrupt", [], blocked=False) == ("corrupt", True)
assert with_hysteresis("invalid-input", [], blocked=False) == ("invalid-input", True)
assert with_hysteresis("corrupt", [], blocked=True) == ("corrupt", True)
# A hard verdict in the history does not count toward clearing a block either.
assert with_hysteresis("pass", ["corrupt"], blocked=True) == ("block", True)
# (17e) A diagnostic that FAILED is not a diagnostic that is missing. Collapsing them
# routes a hardware finding to the harness queue.
assert gate(rec(diag={"pcie": "pass", "memory": "pass", "nccl_tests": "fail"}), BASE) == "diag-failed"
assert gate(rec(diag={"pcie": "pass", "memory": "pass", "nccl_tests": "bogus"}), BASE) == "invalid-input"
# (17f) Adversarial: cohort data must be validated as strictly as candidate data, or the
# alibi is easier to obtain than the accusation. Five infinities satisfy the majority test.
assert gate(rec(samples=[120.0] * 9), BASE, cohort_drops=(float("inf"),) * 5) == "block"
assert gate(rec(samples=[120.0] * 9), BASE, cohort_drops=(True,) * 5) == "block"
# (17g) Adversarial: a median describes the middle, not the population. Half a run at
# baseline and half 15% under it has a clean median and a spread inside the instability
# bar, so the aggregate says pass while a quarter of the observations are slow.
split = [240.0] * 501 + [204.0] * 499
assert gate(rec(samples=split), BASE) == "bimodal"
# A uniformly slow run is not bimodal; it is slow, and takes the ordinary path.
assert gate(rec(samples=[211.2] * 9), BASE, cohort_drops=(0.0, 0.01, 0.0, 0.005, 0.0)) == "block"
# (17h) A negative count of wrong elements is not a count.
assert gate(rec(wrong=-1), BASE) == "invalid-input"
# (18) Adversarial: NaN compares false against every threshold, so a rule that does not
# reject it skips every rejection branch and returns pass. Failing OPEN is the worst
# outcome for a gate whose stated promise is that anything unmeasured blocks.
NAN = float("nan")
assert gate(rec(samples=[NAN] * 9), BASE) == "invalid-input"
assert gate(rec(), {"class": BASE["class"], "median": NAN, "cv": 0.04}) == "no-baseline"
assert gate(rec(), {"class": BASE["class"], "median": 240.0, "cv": NAN}) == "no-baseline"
assert gate(rec(wrong=NAN), BASE) == "invalid-input"
# (19) Adversarial ordering: corruption outranks every other rejection, so it must survive
# a run that ALSO has no samples and no usable baseline. Checking samples first downgrades
# a rollout-stopping result to a harness repair.
assert gate(rec(wrong=7, samples=[]), BASE) == "corrupt"
assert gate(rec(wrong=3), {"class": BASE["class"], "median": 0.0, "cv": 0.04}) == "corrupt"
# (20) The implausibility bounds are inclusive, so a run exactly on them is rejected.
assert gate(rec(samples=[360.0] * 9), BASE) == "implausible" # ratio exactly 1.5
assert gate(rec(samples=[60.0] * 9), BASE) == "implausible" # ratio exactly 0.25
# (21) Adversarial: non-positive inputs. An all-zero sample set divides by zero in the
# spread test and yields NaN, which then compares false everywhere; a non-positive
# baseline cv makes the same test trivially true. Both happened to fail closed further
# down, which is luck rather than a rule, so both are rejected up front.
assert gate(rec(samples=[0.0] * 9), BASE) == "invalid-input"
assert gate(rec(samples=[-240.0] * 9), BASE) == "invalid-input"
assert gate(rec(), {"class": BASE["class"], "median": 240.0, "cv": 0.0}) == "no-baseline"
assert gate(rec(), {"class": BASE["class"], "median": 240.0, "cv": -0.04}) == "no-baseline"
# (22) The "missing" sentinel for an absent #wrong key cannot be spoofed by a record that
# literally carries that string, because a string is not a finite number either way.
assert gate(rec(wrong="missing"), BASE) == "invalid-input"
assert gate(rec(wrong=True), BASE) == "invalid-input"
# (23) insufficient-samples must not swallow a real block: a single sample far enough
# under baseline still clears both bars and blocks.
assert gate(rec(samples=[120.0]), BASE) == "block"
ROWS = [
("at baseline", rec(), ()),
("exit 0, nothing parsed", rec(samples=[]), ()),
("diag block absent", bare, ()),
("nccl_tests skipped", rec(diag={"pcie": "pass", "memory": "pass",
"nccl_tests": "skip"}), ()),
("#wrong read N/A", rec(wrong=None), ()),
("corrupt + wrong baseline", rec(wrong=3, **{"class": "other"}), ()),
("run disagrees with itself", rec(samples=[100., 100., 240., 380., 380.]), ()),
("0.4% of baseline", rec(samples=[1.0] * 9), ()),
("7% down, 1 sample", rec(samples=[223.2]), ()),
("NaN in samples", rec(samples=[float("nan")] * 9), ()),
("corrupt with no samples", rec(wrong=7, samples=[]), ()),
("7% down, 9 samples", rec(samples=[223.2] * 9), ()),
("12% down, cohort also down", rec(samples=[211.2] * 9), fleet),
("50% down, cohort mildly down", rec(samples=[120.0] * 9), mild),
("12% down, cohort healthy", rec(samples=[211.2] * 9), (0.0, 0.01, 0.0, 0.005, 0.0)),
]
print(f"{'case':32s} {'n':>2s} verdict")
for label, r, coh in ROWS:
print(f"{label:32s} {len(r.get('samples', [])):>2d} {gate(r, BASE, coh)}")
print()
print("all assertions passed")
Executed output:
case n verdict
at baseline 9 pass
exit 0, nothing parsed 0 no-measurement
diag block absent 9 diag-incomplete
nccl_tests skipped 9 diag-incomplete
#wrong read N/A 9 correctness-unverified
corrupt + wrong baseline 9 corrupt
run disagrees with itself 5 unstable
0.4% of baseline 9 implausible
7% down, 1 sample 1 insufficient-samples
NaN in samples 9 invalid-input
corrupt with no samples 0 corrupt
7% down, 9 samples 9 warn
12% down, cohort also down 9 fleet-regression
50% down, cohort mildly down 9 block
12% down, cohort healthy 9 block
all assertions passed
Seven of those rows are why the rule is code rather than a threshold in a wiki page.
exit 0, nothing parsed is the failure mode that motivates the whole gate. Every arithmetic check downstream is correct, and the run measured nothing. This is not hypothetical: it is what an older nccl-tests binary does when handed -I 1.
#wrong read N/A, diag block absent, nccl_tests skipped and NaN in samples are the same defect in four shapes. In each the absence of a result is representable as something that looks like a pass, and the gate has to distinguish "checked and fine" from "not checked". The diag block absent row is the sharpest: any() over an empty collection is False, so a diagnostic section that never ran passes every status check written as a filter over the statuses present. That is why the rule asserts on a named set of required tests instead. NaN in samples is the same hole one level down: IEEE comparisons against NaN are all false, so an unvalidated rule skips every rejection branch and reports a pass. A gate whose promise is that anything unmeasured blocks has to fail closed on input it cannot interpret, and corrupt with no samples shows the ordering that follows from it, since corruption outranks the missing samples that would otherwise have answered first.
The two 7% down rows are the same shortfall measured once and measured nine times, and they get different verdicts. On a fleet whose own run-to-run coefficient of variation is 4%, one sample carries a 4% standard error, so the noise bar sits at 8% and a single run admits nothing below it. Nine samples narrow that bar to 2.7% and the same 7% becomes a finding. The policy bar alone would have warned on the single sample and sent someone to investigate a healthy fabric, which is why the gate specifies a sample count rather than only a percentage. Note the one-sample verdict is insufficient-samples, not pass: failing to detect a regression is not evidence that there is none, and a gate that reports the two identically converts a measurement it could not make into a clean bill of health. Two honest limits on this: the noise bar is a strict inequality, so a shortfall landing exactly on two standard errors reads as pass, and a value nominally equal to 2 SE can compute either side of that point depending on how it was derived, which is why the assertions avoid demonstrating on it; and cv/sqrt(n) is the standard error of the mean standing in for the median's, which is about 25% wider on normal data, so the bar is mildly permissive.
The 12% down rows are identical measurements on the node, differing only in what the rest of the cohort did. When the whole cohort moved together the node is not the story and the change is, and blaming the node sends a healthy card to RMA while leaving the actual regression deployed. Two guards keep that from becoming an excuse. The node must be no worse than the cohort within noise, not merely accompanied by it. The test is deliberately one-sided, since a node better than its cohort is still explained by the fleet event: the 50% down row is joined to a fleet that moved 6%, and grading it against the cohort median alone would excuse a node running at half of baseline. And the cohort must be large enough that a majority is not two bad nodes: with three members, two faults are the median, so the second faulty node alibis the first, which is why the rule wants five members with a clear majority past the bar.
Verification¶
- The gate reproduces its own verdict on a re-run of the same profile against the same baseline.
- Every record contains a parsed number, and the count of records equals the count of runs attempted. A missing record is a harness failure, not a silent pass.
- A deliberately broken run is caught. Point the harness at a binary without
-Isupport, or unsetDCGM_NCCL_TESTS_BIN_PATH, and confirm the gate returnsno-measurementordiag-incompleterather thanpass. A gate that has never failed has not been tested. - The baseline used is the one for this topology class, confirmed by reading the key rather than trusting the lookup.
blockactually cordons. Confirm the scheduler state changed, not just that a line was logged.
Rollback¶
The gate blocks changes; rolling it back means letting a change through, so treat an override as a decision with a name attached.
- An override needs a recorded reason and an expiry. A permanently overridden gate is a gate that has been deleted slowly.
- If the gate itself is wrong, fix the parser or the baseline and re-run, rather than lowering the bar. A threshold lowered to make a fleet pass tells you nothing on the next change.
- If a baseline turns out to be stale, re-establish it deliberately on known-good hardware and record when and on what it was taken. Re-baselining on top of an undiagnosed regression bakes the regression into every future comparison.
- Cordoned nodes are released only after the rung that failed is back at baseline (regression ladder).
Related runbooks¶
- NCCL / fabric performance regression: the ladder this gate automates, and where you go when the gate fails.
- Fabric link errors and congestion counters: the counter procedure that runs alongside each rung.
- Add GPU capacity: per-node admission for new hardware.
- NCCL socket fallback: the transport check a gate should include.
- operational runbooks: the runbook index.
References¶
- NVIDIA nccl-tests, argument parsing and the unknown-option path: https://github.com/NVIDIA/nccl-tests/blob/master/src/common.cu
- NVIDIA nccl-tests releases (the tags at which
-Iand-Kfirst appear): https://github.com/NVIDIA/nccl-tests/tags - NVIDIA DCGM diagnostics, suite levels and plugin membership: https://docs.nvidia.com/datacenter/dcgm/latest/user-guide/dcgm-diagnostics.html
- NVIDIA
dcgmi diagcommand reference (level aliases,-j,--fail-early,--check-interval,-p, exit codes): https://docs.nvidia.com/datacenter/dcgm/latest/reference/command-line-reference/dcgmi/dcgmi-diag.html - NVIDIA DCGM source, suite-to-plugin mapping: https://github.com/NVIDIA/DCGM
- linux-rdma perftest, result format and unit handling: https://github.com/linux-rdma/perftest/blob/master/src/perftest_parameters.h
- NVIDIA NCCL troubleshooting, performance and tuning: https://docs.nvidia.com/deeplearning/nccl/user-guide/docs/troubleshooting/performance_and_tuning.html
Related: Fabric Performance Regression · Fabric Link Errors · Diagnostics and Validation · GPU Health Gating · Commissioning and Acceptance · Continuous Fabric Benchmarking · Add GPU Capacity · Operational Runbooks · Glossary
-
NVIDIA DCGM plugin references, per-plugin prerequisites. Diagnostics: "The plugin requires CUDA support and should run on GPUs that are not serving production workloads because it applies sustained compute, memory, and power load." (https://docs.nvidia.com/datacenter/dcgm/latest/reference/diagnostics/plugins/diagnostic.html). Targeted Stress: "should run on idle GPUs so that competing compute work does not distort the achieved-performance result." (https://docs.nvidia.com/datacenter/dcgm/latest/reference/diagnostics/plugins/targeted-stress.html). Targeted Power: "Run it on idle GPUs: competing workloads and an enforced power limit below the configured target can prevent a meaningful target-power result." (https://docs.nvidia.com/datacenter/dcgm/latest/reference/diagnostics/plugins/targeted-power.html). DCGM states no level-based drain rule; the requirement is per plugin, and the level-3 set is where these plugins first appear. ↩
-
NVIDIA
dcgmi diagreference: "The names quick and short, medium, long, and xlong are aliases for levels 1 through 4, respectively." Level 1 runs the built-in software deployment checks; level 2 addsmemoryandpcie; level 3 addsdiagnostic,memory_bandwidth,targeted_stress,targeted_power,nvbandwidthandnccl_tests, plus the EUD plugins when DCGM runs as root; level 4 addsmemtestandpulse_test. NVIDIA's user-guide table and itsdcgmi diagreference disagree over whethermemory_bandwidthruns at level 2 or level 3; the reference and the DCGM source both place it at level 3 and above. https://docs.nvidia.com/datacenter/dcgm/latest/reference/command-line-reference/dcgmi/dcgmi-diag.html ↩ -
Verified by walking the upstream tags: the
getopt_longoption string insrc/common.cufirst containsI:atv2.19.2andK:atv2.19.6, and neither appears atv2.19.1or earlier. Atv2.19.1the parser's fallthrough readscase 'h': default: if (c != 'h') printf("invalid option '%c'\n", c);followed by the usage block andreturn 0, so an unrecognised option terminates the process with status 0 and no benchmark. Notecthere isgetopt_long's return value for an unknown option, which is'?'rather than the offending letter, andopterris never cleared, so glibc emits its owninvalid option -- 'I'on stderr first. Reproduced locally by compiling the same optstring and default arm. https://github.com/NVIDIA/nccl-tests/blob/master/src/common.cu ↩ -
DCGM source,
nvvs/plugin_src/nccl_tests/NcclTestsWrapper.cpp:info->numTests = 0; if (SupportNcclTestsPlugin()) { ... }, whereSupportNcclTestsPlugin()returns whether theDCGM_NCCL_TESTS_BIN_PATHenvironment variable exists. With the variable unset no test is registered, so the suite omitsnccl_testsentirely rather than reporting it skipped; theSetResult(testName, NVVS_RESULT_SKIP)paths inNcclTestsPlugin.cppare reached when the variable is set but the path is missing or unusable. https://github.com/NVIDIA/DCGM ↩ -
NVIDIA
dcgmi diagreference, exit codes:226 DCGM_ST_NVVS_ERROR("The diagnostic ran but reported an error"),205 DCGM_ST_NVVS_ISOLATE_ERROR("The diagnostic reported a condition that requires isolation"),217 DCGM_ST_DIAG_ALREADY_RUNNING,204 DCGM_ST_NVVS_BINARY_NOT_FOUND, and203 DCGM_ST_NVVS_KILLEDfor a diagnostic process terminated by a signal (https://docs.nvidia.com/datacenter/dcgm/latest/reference/command-line-reference/dcgmi/dcgmi-diag.html). The 128-plus-signal convention concernsdcgmiitself and is documented separately: "If dcgmi is terminated by a signal, POSIX-style shells conventionally report 128 plus the signal number, such as 130 for SIGINT and 143 for SIGTERM. Those values describe process termination rather than a dcgmReturn_t result." https://docs.nvidia.com/datacenter/dcgm/latest/reference/command-line-reference/dcgmi/index.html ↩ -
linux-rdma perftest,
src/perftest_parameters.h:#define RESULT_FMT " #bytes #iterations BW peak[MiB/sec] BW average[MiB/sec] MsgRate[Mpps]"and#define RESULT_FMT_G " #bytes #iterations BW peak[Gb/sec] BW average[Gb/sec] MsgRate[Mpps]". The--report_gbitshelp text, which lives insrc/perftest_parameters.crather than the header, reads "Report Max/Average BW of test in Gbit/sec (instead of MiB/sec) (SYMMETRIC)" and "Note: MiB=2^20 byte, while Gb=10^9 bits. Use these formulas for conversion: Factor=10^9/(2^20*8)=119.2; MiB=Gb_result * factor; Gb=MiB_result / factor". Builds at or before the 24.04 release labelled the same valueMB/sec; the value was always mebibytes and only the label was corrected. https://github.com/linux-rdma/perftest/blob/master/src/perftest_parameters.h and https://github.com/linux-rdma/perftest/blob/master/src/perftest_parameters.c ↩