JevAdvBench: robustness benchmark for typed-decision models¶
Scope: JevAdvBench (arXiv 2609.31142, September 2026), a benchmark and black-box attack suite for models trained with reinforcement learning for calibrated decisions (RLCD), which answer a typed question with a probability, a choice, or a score that software acts on without a person reading it. Covers the measurement design, the reported results on one model version, the limits of those results, and what they imply for builders. It is descriptive only: attack templates are not reproduced. It complements prompt-injection defense and LLM red-teaming with promptfoo, which target text-generating systems.
Status: read from the arXiv abstract page, the project page, and the repository README (2026-09-30). The paper PDF was not read in full, and the repository's offline analysis script was not executed here, so no reported number below is reproduced by this KB. The code block on this page is an original toy that illustrates the metric, not the benchmark data.
Overview¶
An RLCD model such as Jev returns a well-formed typed answer even when manipulated, so text-oriented attack benchmarks (which score generated or executed content) do not apply. Three difficulties shape the design:
- identical requests can return different answers;
- 669 of 812 labels (82.4%) are the model's own five-run consensus, and 143 are human-reviewed;
- the hosted API preprocesses requests out of view, so an edit may never reach the model.
The benchmark answers these with three choices: score each attacked decision against the model's own clean decision (no external labels), subtract the change caused by an identical re-run (the noise floor, 1.0% with interval [0.3, 1.9] per the project page), and confirm delivery through billed input tokens.
flowchart LR
Q["Typed question + state"] --> C["Clean run"]
Q --> R["Identical re-run"]
Q --> A["One single-edit variant"]
C --> F["Flip rate vs clean"]
R --> N["Noise floor"]
A --> F
A --> T["Billed tokens: edit delivered?"]
F --> X["Excess over noise floor"]
N --> X
Core knowledge¶
Scale (repository README and abstract): 66 scenarios, 812 questions (314 Noul probability, 337 Choice, 161 Score), 9,744 variants (12 per question), 11,368 recorded responses (10,556 evaluation requests and 812 re-runs). Domains include customer-support triage, information extraction, finance command parsing, moderation, and claims and hiring decisions.
Attack families, each a single fixed-template edit with no gradients, search, or repeated queries: question-text edits (spacing, paraphrase, unrelated sentences), state edits (unrelated note, observer opinion, opinion by analogy), injected commands (override, authority impersonation, fake validation note), and structural edits (extra or escaped fields, opinion in an extra field).
Reported results on jev-1.13.0 (claims from the authors):
| Finding | Reported value |
|---|---|
| Rewording vs re-run baseline | within 1.2 percentage points |
| Fields outside the schema | never reach the model |
| One unverified opinion appended to the state | flips 12.1% of decisions |
| Strongest injected command (authority impersonation) | flips 10.1%, statistically tied with the opinion |
| Confident answers pushed below the 0.8 human-review gate | 38% |
| Flips landing on the attacker-named target | 88.8% (project page) |
The authors conclude that the state must be treated as untrusted, argued input, that user text should not reach instruction fields, and that human-review capacity should be sized for attack-induced escalation.
Metric illustration (toy, not the benchmark)¶
The block below is original code executed for this page. It shows why the excess over an identical re-run is the quantity to report: a nonzero floor exists even with no attack.
import random
random.seed(0)
N = 812
def decide(p_true: float, noise: float) -> str:
return "yes" if p_true + random.gauss(0, noise) > 0.5 else "no"
def flip_rate(clean: list[str], other: list[str]) -> float:
assert len(clean) == len(other) > 0
return sum(a != b for a, b in zip(clean, other)) / len(clean)
p = [random.random() for _ in range(N)]
clean = [decide(x, 0.03) for x in p]
rerun = [decide(x, 0.03) for x in p]
shifted = [decide(min(1, max(0, x + 0.05)), 0.03) for x in p]
floor = flip_rate(clean, rerun)
attacked = flip_rate(clean, shifted)
print(f"re-run noise floor: {floor:.3f}")
print(f"attacked flip rate: {attacked:.3f}")
print(f"excess over floor: {attacked - floor:.3f}")
assert floor > 0
assert attacked > floor
assert flip_rate(clean, clean) == 0.0
try:
flip_rate([], [])
except AssertionError:
print("empty input rejected")
Executed output:
Don't-miss checklist¶
- Compare any flip rate with an identical re-run of the same model; raw flip rates overstate attack effect.
- Verify the edit reached the model (token or request accounting) before counting a null result as robustness.
- Keep user-writable text out of instruction and criteria fields.
- Budget reviewer capacity for the confidence gate: the reported escalation is a workload increase, not only an accuracy loss.
- Data is CC BY-NC 4.0; code is MIT; scenarios derived from the vendor's examples keep their original terms.
Failure modes¶
- Fixed templates with no adaptive search: the authors state reported rates are lower bounds.
- One model version (
jev-1.13.0, hosted, non-deterministic): results do not transfer to other RLCD systems or later versions. - Label quality: 82.4% of labels derive from the model itself; a single annotator checked 40 variants and 25 preserved the answer, so variant validity is weakly evidenced.
- Mismatch to note: the project page lists the authors' affiliations loosely ("and others"); the arXiv listing is the authoritative author record (Hu, Zhang, Liu, Zeng, Zeng, Wang, Wang, Zhang per the repository BibTeX).
Open questions and validation¶
- Whether adaptive attackers exceed the fixed-template rates, and by how much.
- Whether the 12.1% opinion-flip rate holds on other model versions.
- The repository ships an offline script that claims 256 of 256 reported values match the paper; this KB has not run it. Re-deriving the headline numbers from the released responses is the validation step.
References¶
- JevAdvBench paper, arXiv 2609.31142: https://arxiv.org/abs/2609.31142
- Project page: https://jevadvbench.github.io/JevAdvBench/
- Repository: https://github.com/JevAdvBench/JevAdvBench
- Dataset: https://huggingface.co/datasets/Hangtao/JevAdvBench
Related: Prompt-injection defense · LLM red-teaming with promptfoo · Agent security threat model · LLM benchmarks · Evaluation integrity