APPA: recoverable information-flow control for LLM agents¶
Scope: APPA (Agentic Permissions Policy Algebra, arXiv 2607.24625v2, Archestra AI) and its open implementation OpenAPPA, a reference monitor that applies information-flow control (IFC) to an LLM agent's tool calls and adds recovery paths when a call is blocked. Covers the label lattice, the two-gate monitor, the repair taxonomy, call-scoped authority, gradual typing, branch confinement with attest-schema, a runnable numpy/stdlib reconstruction of the core rules, the reported results with their inconsistencies, and what the repository does and does not contain. Complements the heuristic defences in prompt injection defense and the allow/deny layer in agent policy engine; it does not replace either. This page is the paper-level view (algebra, proofs, results audit, a runnable reconstruction); the engine, its TOML policy, and deployment are on OpenAPPA, and the hosting platform is on Archestra.
The paper is a 2026 preprint from the tool's authors, and the benchmark they built for it (Bench-Corp) ships in the same repository. Figures below are quoted from the paper or recomputed from its tables; none were reproduced on a live agent. The Python block is this page's own reduced reconstruction of Sections 3 to 6, not the OpenAPPA engine.
What it is¶
APPA sits between an agent and its tools (for example at a Model Context Protocol gateway or a CLI dispatcher) and decides, before every tool call, whether the data the agent has read so far may flow to that tool. The model is treated as untrusted for policy: it can propose any call, but it cannot forge authorisation. The trusted computing base (TCB) is the monitor, the declared tool contracts, the registered authority resolvers, and the registered sanitizers, schema validators, and casts.
Core objects, as the paper defines them:
- Label. A pair
L = (R, t): a setRof authorised readers and a trust leveltfrom a totally ordered chain (for example suspicious < trusted). Combining two labels is the component-wise meet(R ∩ R', min(t, t')), so a trajectory's label can only stay the same or get more restrictive. Proposition 2 states that a trajectory over the finite lattice performs only finitely many strict descents. - Tool contract. Each tool declares the label contribution
dits output will add, the effect tokens it commits on success, and preconditions: a trust floor, a recipient cover, history predicatesprior(k)andno_prior(k)over a shared append-only effect log, and authority gates. - Prospective label. For a proposed call
τunder current labelL, the monitor computesc_τ(L) = L ∧ d_τand checks release requirements against that, not againstL. - Partial label. An unannotated tool contributes an
Unknownsource instead of a guessed lattice position. A registered cast resolves it later, bounded by a declared ceiling (may_cast).
The paper's contribution over earlier agent IFC (Fides, CaMeL) is the recovery semantics: a blocked call returns structured remedies (satisfy a prerequisite, obtain a call-scoped ruling, apply a registered transformation, or run the read in a disposable child branch) rather than an opaque refusal.
Why use it¶
- Monotone taint strands agents. Once an agent reads an untrusted web page, its label descends to suspicious and every later write that needs a trusted context is blocked for the rest of the session. The usual workarounds (retry prompts, application glue, blanket human approval) either fail or cause approval fatigue.
- Reactive checks miss composite tools. A single tool that reads a confidential ledger and emails it externally passes a check made against the pre-call label. Proposition 4 states that such a contract exists, so the check has to use the post-call label.
- Deterministic enforcement. The engine decides from the event log alone. Classifiers and PII detectors are probabilistic; the flow check is not, and it returns the same decision on the same log.
- Reported effect on the authors' benchmarks. Zero observed attacks in 1,320 guarded episodes at 64.2% to 91% utility, against Fides configurations that recorded 5.0% to 35% attack success in the chaos profile or on Bench-Corp. See the results section for the caveats.
When to use it (and when not)¶
Use it when an agent holds private context, calls tools with external side effects, and ingests untrusted content in the same session, and when tool calls can be mediated at a gateway or harness the operator controls.
Do not treat it as sufficient when:
- tools cannot be intercepted (no gateway, no harness hook);
- the contracts are wrong or missing for the sensitive sinks: the paper assumes accurate contracts and trusted sanitizers, and warns that misconfigured contracts create false assurance;
- covert channels matter: timing and other side channels are out of scope;
- the concern is exactly-once effects against the real world:
no_prioronly sees outcomes reported to the monitor, so an effect that succeeded externally but closed as indeterminate is absent from the log unless the tool boundary adds idempotency keys or reconciliation; - a child branch must be rolled back: confinement isolates model context and labels, not external side effects a child already committed.
Architecture¶
flowchart LR
M["LLM proposes call"] --> G1["Gate 1: prospective check<br/>c(L) = L meet d, history E, rendered call"]
G1 -->|"all requirements hold"| T["Tool executes (outside TCB)"]
G1 -->|"gap"| R["Remedies: prerequisite, ruling,<br/>sanitizer or cast, child branch"]
R -->|"ruling bound to call hash"| T
T --> G2["Gate 2: verify realized return<br/>unannotated becomes Unknown"]
G2 -->|"fold L = L meet d"| C["Agent context"]
C --> M
C -.->|"fork"| B["Child branch: inherits L,<br/>absorbs taint locally"]
B -->|"attest-schema exit"| C
Three deployment tiers, without model fine-tuning: a protocol gateway tier (Gate 1 and Gate 2 at MCP or CLI dispatch), an inference-harness tier (branch management, which needs something that can bound model context), and an authority tier (single-use rulings bound to the rendered call hash).
How to use it¶
Policy model¶
Policy is declarative. In OpenAPPA it is TOML, loaded by the appa CLI; the repository README documents two pre-merge checks that do not execute any tool:
The first checks that the configuration loads; the second replays scripted tool calls against expected decisions. These are quoted from the repository README at commit 6710c7f; they were not run for this page.
The repair taxonomy (Theorem 5)¶
When a call is blocked by gap g, the paper proves which state changes can clear it:
| Gap | What clears it |
|---|---|
| Trust floor or recipient cover (release-side, monotone in the label) | No sequence of admitted reads. Only a call-scoped ruling, a registered transformation, or a cast. |
prior(k) unmet |
A committed effect that establishes k, or a ruling under a registered history waiver. |
no_prior(k) violated |
No effect (the log only grows). Only a ruling under a registered history waiver. |
| Unaccepted narrowing | In-session acceptance, or running the call in a child branch so the parent never descends. |
Bounded search over the finite recovery graph finds a clearing path whenever one exists in the registered abstraction and returns empty otherwise. Completeness is model-relative: a widened ACL or an expired credential is outside the graph.
Call-scoped authority¶
An authority (deterministic rules, a human, or a model evaluator) rules on the exact rendered call, inside a pre-declared mandate (trust-floor ceiling, recipient coverage, named history waivers). Theorem 6 states that an inadequate release dispatches only by atomically consuming rulings over its exact rendered form, and that rulings never change a trajectory label or the effect log. Rulings therefore do not carry over to later calls, and the paper states the result is an authorised-and-audited release guarantee, not noninterference.
Branch confinement¶
fork gives a child the pre-branch transcript and the parent's current label (L0c := Lp; starting from a top label would launder already-ingested data). The child absorbs taint locally and exits by one of three routes: abandonment (parent unchanged), a labelled return (an ordinary narrowing check on the parent), or a sanitized or schema-attested return. The parent fixes a return_schema at fork time; the dialect admits booleans, bounded integers, fixed-precision decimals, closed enums and formats, and bounded arrays and objects, and rejects free text, unbounded numbers, and open collections. The paper is explicit that this bounds shape and bandwidth, not meaning: a malicious child can still pick a false allowed value, and promoting the payload's integrity is a TCB decision.
How to develop with it¶
The repository is a Rust workspace (crates include appa-engine, appa-policy, appa-eventlog, appa-runtime, and adapters for Claude Code, Amp, and kagent) plus a bench/ directory that holds Bench-Corp. The engine's label model in appa-engine/src/label.rs is richer than the paper's (reader identities, group references, chain audiences), so the paper's two-component lattice is a simplification of what ships.
The block below reconstructs the paper's rules so they can be tested in isolation: the lattice laws on random labels, the prospective-versus-reactive difference, the repair taxonomy, single-use and hash-bound rulings with mandates, Gate 2 folding, gradual casts with ceilings, and branch isolation with an attest-schema check. Assertions cover the negative cases (an over-ceiling cast, a ruling reused or replayed on changed arguments, a waiver outside the mandate, free text through the schema gate).
"""APPA core, reduced: label lattice, two-gate monitor, casts, branches, rulings.
Reference reconstruction of the paper's Sections 3-6, not the OpenAPPA engine."""
from __future__ import annotations
import hashlib
import json
from dataclasses import dataclass, field
import numpy as np
SUSPICIOUS, TRUSTED = 0, 1
@dataclass(frozen=True)
class Label:
readers: frozenset[str]
trust: int
def meet(self, o: Label) -> Label:
return Label(self.readers & o.readers, min(self.trust, o.trust))
def __le__(self, o: Label) -> bool:
return self.readers <= o.readers and self.trust <= o.trust
@dataclass(frozen=True)
class Contract:
name: str
delta: Label | None # declared contribution; None = unannotated
effects: frozenset[str] = frozenset()
min_trust: int = SUSPICIOUS
recipients: frozenset[str] = frozenset()
prior: frozenset[str] = frozenset()
no_prior: frozenset[str] = frozenset()
def call_hash(name: str, args: dict) -> str:
return hashlib.sha256(json.dumps([name, args], sort_keys=True).encode()).hexdigest()
@dataclass
class Monitor:
label: Label
contracts: dict[str, Contract]
casts: dict[str, tuple[Label, Label]] = field(default_factory=dict) # src -> (result, ceiling)
mandate: frozenset[str] = frozenset() # gap kinds the authority may rule on
effects: set[str] = field(default_factory=set) # shared E
unknown: set[str] = field(default_factory=set) # partial-label sources
rulings: set[str] = field(default_factory=set) # unused call hashes
def prospective(self, c: Contract) -> Label:
return self.label if c.delta is None else self.label.meet(c.delta)
def gaps(self, c: Contract) -> set[str]:
post, g = self.prospective(c), set()
if post.trust < c.min_trust:
g.add("trust")
if not c.recipients <= post.readers:
g.add("audience")
g |= {f"prior:{k}" for k in c.prior - self.effects}
g |= {f"no_prior:{k}" for k in c.no_prior & self.effects}
if post != self.label:
g.add("narrowing")
return g
def rule(self, c: Contract, args: dict) -> bool:
"""Authority issues a single-use ruling bound to the rendered call."""
kinds = {g.split(":")[0] for g in self.gaps(c)} - {"narrowing"}
if not kinds <= self.mandate:
return False
self.rulings.add(call_hash(c.name, args))
return True
def dispatch(self, name: str, args: dict, accept: bool = False,
realized: Label | None = None) -> str:
c = self.contracts[name]
g = self.gaps(c) - ({"narrowing"} if accept else set())
h = call_hash(name, args)
if g and h not in self.rulings:
return "BLOCK " + ",".join(sorted(g))
self.rulings.discard(h) # single use
self.effects |= c.effects # committed on success
d = c.delta if realized is None else (realized if c.delta is None else realized.meet(c.delta))
if d is None:
self.unknown.add(name) # Gate 2: unannotated -> Unknown
else:
self.label = self.label.meet(d)
return "OK"
def resolve(self, src: str) -> bool:
if src not in self.casts:
return False # fail closed
lam, ceiling = self.casts[src]
if not lam <= ceiling:
return False
self.label = self.label.meet(lam)
self.unknown.discard(src)
return True
def fork(self) -> Monitor:
return Monitor(self.label, self.contracts, self.casts, self.mandate,
self.effects, set(self.unknown), set()) # E shared, rulings dropped
def attest_schema(value: object, schema: dict) -> bool:
"""Closed-shape check: bool, bounded int, closed enum. Rejects free text."""
if schema["type"] == "bool":
return isinstance(value, bool)
if schema["type"] == "int":
return isinstance(value, int) and not isinstance(value, bool) and schema["lo"] <= value <= schema["hi"]
if schema["type"] == "enum":
return value in schema["values"]
return False
PUB, INT = frozenset({"public", "internal"}), frozenset({"internal"})
TOP = Label(PUB, TRUSTED)
C = {
"web_fetch": Contract("web_fetch", Label(PUB, SUSPICIOUS)),
"read_ledger": Contract("read_ledger", Label(INT, TRUSTED), frozenset({"ledger_read"})),
"send_email": Contract("send_email", None, recipients=frozenset({"public"})),
"share_legal_packet": Contract("share_legal_packet", Label(INT, TRUSTED),
recipients=frozenset({"public"})),
"write_db": Contract("write_db", Label(PUB, TRUSTED), min_trust=TRUSTED),
"refund": Contract("refund", None, prior=frozenset({"manager_ok"})),
"archive": Contract("archive", None, frozenset({"archived"}), no_prior=frozenset({"archived"})),
}
new = lambda **kw: Monitor(TOP, C, **kw)
# 1. Lattice laws on random labels (numpy-generated), plus monotone descent.
rng = np.random.default_rng(0)
U = ["a", "b", "c", "d", "e"]
def rand_label() -> Label:
return Label(frozenset(np.array(U)[rng.random(5) < 0.5].tolist()), int(rng.integers(0, 2)))
for _ in range(2000):
x, y, z = rand_label(), rand_label(), rand_label()
assert x.meet(y) == y.meet(x) and x.meet(y).meet(z) == x.meet(y.meet(z)) and x.meet(x) == x
assert x.meet(y) <= x and x.meet(y) <= y
assert (x.meet(y) == x) == (x <= y)
print("1 lattice laws ok on 2000 random triples")
# 2. Prop 4: composite read-and-send. Reactive check (pre-call label) passes, prospective blocks.
m = new()
assert m.label.readers >= C["share_legal_packet"].recipients # reactive view: allowed
print("2 prospective:", m.dispatch("share_legal_packet", {"to": "press"}))
# 3. Thm 5(i): no admitted read clears a trust-floor gap; only a ruling does.
m = new()
assert m.dispatch("web_fetch", {}, accept=True) == "OK"
for _ in range(3):
m.dispatch("read_ledger", {}, accept=True)
assert "trust" in m.gaps(C["write_db"])
print("3 after reads:", m.dispatch("write_db", {"row": 1}))
m.mandate = frozenset({"trust"})
assert m.rule(C["write_db"], {"row": 1})
print("3 after ruling:", m.dispatch("write_db", {"row": 1}))
print("3 ruling is single use:", m.dispatch("write_db", {"row": 1}))
assert m.rule(C["write_db"], {"row": 1})
print("3 ruling bound to call hash:", m.dispatch("write_db", {"row": 2}))
m.mandate = frozenset({"trust"})
print("3 ruling beyond mandate (audience gap) refused:", not new(mandate=m.mandate).rule(C["share_legal_packet"], {}))
# 3b. Gate 2: the realized return is folded, so an over-claiming contract cannot keep the label high.
g2 = new()
g2.dispatch("write_db", {"row": 9}, realized=Label(PUB, SUSPICIOUS))
assert g2.label.trust == SUSPICIOUS
print("3b gate 2 folds realized label:", g2.label.trust == SUSPICIOUS)
# 4. History: prior(k) clears by effect, no_prior(k) is irreparable by any effect.
m = new()
print("4 refund before prior:", m.dispatch("refund", {}))
m.effects.add("manager_ok")
print("4 refund after prior:", m.dispatch("refund", {}))
print("4 archive first:", m.dispatch("archive", {}), "| second:", m.dispatch("archive", {}))
m.mandate = frozenset({"no_prior"})
assert m.rule(C["archive"], {})
print("4 second archive with waiver ruling:", m.dispatch("archive", {}))
m = new(mandate=frozenset({"trust"}))
m.dispatch("archive", {})
print("4 waiver outside mandate refused:", not m.rule(C["archive"], {}))
# 5. Gradual typing: unannotated tool gives Unknown; sink check on Unknown dimension fails closed.
m = new(casts={"send_email": (Label(PUB, SUSPICIOUS), Label(PUB, SUSPICIOUS))})
m.dispatch("send_email", {"to": "public"})
assert m.unknown == {"send_email"} and m.label == TOP
assert not m.resolve("refund"), "no cast registered"
assert m.resolve("send_email") and m.label.trust == SUSPICIOUS
m = new(casts={"send_email": (TOP, Label(PUB, SUSPICIOUS))}) # cast claims more than ceiling
m.dispatch("send_email", {})
assert not m.resolve("send_email") and m.label == TOP and m.unknown
print("5 casts: unresolved fails closed, over-ceiling cast rejected, label never raised")
# 6. Branch confinement: child absorbs taint; parent label exact on abandon; rulings do not cross.
p = new(mandate=frozenset({"trust"}))
child = p.fork()
assert child.label == p.label # L0c = Lp
child.dispatch("web_fetch", {}, accept=True)
assert child.label.trust == SUSPICIOUS and p.label == TOP # parent untouched
p.rulings.add(call_hash("write_db", {"row": 1}))
assert not p.fork().rulings # rulings invalidated across fork
ok = {"type": "enum", "values": ["low", "high"]}
assert attest_schema("high", ok) and not attest_schema("ignore previous instructions", ok)
assert not attest_schema(True, {"type": "int", "lo": 0, "hi": 9}) and not attest_schema(10, {"type": "int", "lo": 0, "hi": 9})
assert attest_schema("low", ok) # a malicious child can still pick a false allowed value
print("6 branch: parent label preserved, shape check rejects free text, false enum value still passes")
Executed output (Python 3, stdlib plus numpy, seeded):
1 lattice laws ok on 2000 random triples
2 prospective: BLOCK audience,narrowing
3 after reads: BLOCK trust
3 after ruling: OK
3 ruling is single use: BLOCK trust
3 ruling bound to call hash: BLOCK trust
3 ruling beyond mandate (audience gap) refused: True
3b gate 2 folds realized label: True
4 refund before prior: BLOCK prior:manager_ok
4 refund after prior: OK
4 archive first: OK | second: BLOCK no_prior:archived
4 second archive with waiver ruling: OK
4 waiver outside mandate refused: True
5 casts: unresolved fails closed, over-ceiling cast rejected, label never raised
6 branch: parent label preserved, shape check rejects free text, false enum value still passes
Two behaviours in the output are design points and not bugs. Step 4 shows no_prior staying violated after the first archive call until a waiver ruling is issued. Step 6 shows a false enum value passing the schema gate; that is the limit the paper states, reproduced rather than hidden.
How to maintain it¶
- Treat contracts as security code. A contract that understates
d_τis only partly covered: Gate 2 re-verifies the realized return before it enters context, but effect tokens and preconditions are taken as declared. - Audit rulings and minimise mandates: the paper's broader-impact section names misconfigured authorities and sanitizers as the route to false assurance.
- Keep authority mandates narrower than the gaps a model-backed authority could be talked into approving. The paper notes model-backed authorities may err and relies on mandates to bound the damage.
- Run
appa describe --checkandappa replayas required CI checks on policy changes, as the README suggests. - Re-check the repository before relying on this page: the README describes the project as "Preview & RFC" and the commit history is moving daily (version 0.30.0 on 2026-09-30).
How to run it in production¶
- Enforcement at a gateway needs no model changes; branch confinement needs a harness or inference proxy that can bound model context.
- Start with unannotated tools (they contribute
Unknownand run freely) and add contracts where sensitive sinks consume them. A sink that consumes an unresolved dimension with no registered cast fails closed. - Use idempotency keys or outcome reconciliation at tool boundaries wherever
no_prioris meant to give exactly-once behaviour. - Combine with sandboxing for what IFC does not cover (process-level side effects) and with risk-tiered approval for the human rung of the authority tier.
- OpenAPPA can run in process or as a sidecar; the README also lists an Archestra LLM-proxy integration for several coding agents. Those integrations were not evaluated here.
Reported results and inconsistencies¶
Setup as stated in the paper: AgentThreatBench (24 Inspect tasks: 10 Memory Poisoning, 6 Autonomy Hijacking, 8 Data Exfiltration) and Bench-Corp (20 multi-step enterprise scenarios, built by the authors), three models (named in the paper as GPT-5.6 Luna, DeepSeek V4 Flash, Gemini 3.7 Flash), five arms, five repetitions, two prompt profiles. Comparators are two Microsoft Agent Framework Fides mappings implemented by the authors (F-M, F-N), not vendor-tuned policies.
| Benchmark | APPA utility / attack success | Comparator (Fides-middleware) |
|---|---|---|
| AgentThreatBench | 64.2% to 76.7% / 0 of 720 | 55.8% to 74.2% utility; 0% to 30.8% attack success |
| Bench-Corp | 87% to 91% / 0 of 600 | 36% to 45% utility; 26% to 35% attack success |
Ablation on Bench-Corp with one model (paper Table 3): removing remedy plans leaves 35.0% utility, removing child branching leaves 56.5%, both at 0% attack success, against 88.0% for the full system. The paper notes the two deficits interact and are not additive.
Checks made for this page against the paper's own tables:
- Episode count. The abstract's 6,600 episodes equals 3,600 AgentThreatBench plus 3,000 Bench-Corp. The body also reports 300 separately counted recipient-control episodes, so the total including them is 6,900. The 1,320 guarded figure does reconcile (720 plus 600).
- Utility range. 64.2% to 91% spans two benchmarks with different baselines; 64.2% to 76.7% is AgentThreatBench and 87% to 91% is Bench-Corp. A single "64 to 91%" is not one distribution.
- Memory subset. The text says APPA trails Fides-middleware on DeepSeek standard by 10 episodes, entirely in Memory (71 of 150 against 150 of 150). Appendix Table 5 lists APPA DeepSeek standard Memory as 17 of 50 and Table 4 gives 77 of 120 total against 87; the 10-episode gap reconciles, but 71 of 150 does not match any cell this page could locate, and the denominators (150 versus 50 per cell) are not explained.
- Attack-free completions. The text cites 35.8% to 45.0% for Fides-middleware under chaos. Table 1 utility minus attack success gives 35.8%, 44.2%, and 33.4%. Because an episode can be both a success and an attack, attack-free completions are at least utility minus attack rate, so the quoted range is not contradicted, but the lower bound alone does not reproduce it.
- Zero is finite. Zero observed attacks in 1,320 episodes is an observation over 20 or 24 task designs repeated five times. The paper states repetitions are not independent tasks and that zero observations do not prove attacks impossible.
- Benchmark coupling. The paper states the Fides policies "do not encode" Bench-Corp's exact-recipient, prerequisite-ordering, exactly-once, external-authority, or fine-grained trust constraints. Bench-Corp was built to exercise APPA's features; a large gap there is not evidence about Fides tuned for the task. AgentThreatBench, an external suite, shows the smaller gap.
- Repository versus paper. The README table compares against "Claude Auto mode" (90% completion, 10% attacks). The paper's text, as extracted for this page, contains no Claude Code auto-mode arm, so that row is unsupported by the cited paper version.
Failure modes¶
- Wrong or optimistic contracts: Gate 1 trusts the declared
d_τ; only the realized-return fold at Gate 2 catches a tool that returns more sensitive data than declared. - A sanitizer or cast that over-promises: casts are bounded by
may_castceilings, but the classifier behind a cast is inside the TCB and unvetted classifiers can become laundering sinks if ceilings are loose. - Allowed false values: an
attest-schemaexit cannot stop a compromised child from choosing the wrong enum member or integer. Promote payload integrity only when the value is independently checkable. - Reported versus real outcomes:
no_priorrelies on outcomes the monitor saw; an indeterminate close leaves the log blind to an effect that happened. - Child side effects: abandoning a branch does not undo effects it committed; they stay in the shared effect log.
- Fork-crossing authority: rulings and remedy plans are invalidated at the fork; exporting a child-derived value through a later parent call needs a fresh ruling if its label fails the sink.
Open questions and validation¶
- No independent reproduction exists at the time of writing; all numbers trace to the authors' harness.
- Only three models and one policy per defended arm were evaluated; policy ergonomics and approval fatigue are listed as future work.
- The engine's richer label model (groups, chain audiences) means the paper's proofs cover the simplified lattice, and the page has not checked that the shipped engine satisfies them.
- Not checked here: building or running OpenAPPA, the Bench-Corp scorers, and the website claims.
References¶
- Kravchenko, Liventsev, Konstantinov, Iskhakov, Kukuy, "APPA: Recoverable Information-Flow Control for Real-World LLM Agents", arXiv 2607.24625 (v2, 26 Aug 2026): https://arxiv.org/abs/2607.24625
- OpenAPPA repository (MIT, Rust workspace; read at commit 6710c7f): https://github.com/archestra-ai/OpenAPPA
- Project site: https://openappa.com
- Costa et al., "Securing AI Agents with Information-Flow Control" (Fides): https://arxiv.org/abs/2505.23643
- Debenedetti et al., "Defeating prompt injections by design" (CaMeL): https://arxiv.org/abs/2503.18813
- Kolluri et al., "Optimizing agent planning for security and autonomy" (Prudentia): https://arxiv.org/abs/2602.11416
- Siddiqui et al., "Permissive information-flow analysis for large language models": https://arxiv.org/abs/2410.03055
- Willison, "The Dual LLM pattern for building AI assistants that can resist prompt injection": https://simonwillison.net/2023/Apr/25/dual-llm-pattern/
- AgentThreatBench (Inspect Evals): https://ukgovernmentbeis.github.io/inspect_evals/evals/agent_threat_bench/index.html
- Model Context Protocol: https://modelcontextprotocol.io
Related: OpenAPPA · Archestra agent platform · Prompt injection defense · Agent policy engine · Risk-tiered approval · Agent identity and access · Agent sandboxing and isolation · Agent intent verification · Agent action execution boundary · Agent security threat model