Skip to content
Markdown

OpenAPPA: deterministic information-flow guardrails for agents

Scope: OpenAPPA, the open-source engine built on APPA (Agentic Permissions Policy Algebra, arXiv 2607.24625) that sits between an agent and its tools and decides, before each call, whether the data the session has read may go to the call's destination. Covers the security label (audience, trust), tool contracts, remedy plans (sanitizers, authorities, subagent reads), policy testing with appa replay, deployment shapes, and an audit of the published benchmark claims. The hosting platform that ships it at the LLM-proxy layer is on Archestra; the generic policy-gate pattern is policy, guardrails and governance.

flowchart LR
  A["Agent loop"] -->|"tool call"| E["APPA engine (outside the loop)"]
  L["Event log: what the session read and did"] --> E
  P["appa.toml: tool contracts"] --> E
  E -->|"allow"| T["Tool runs"]
  E -->|"block + remedy plans"| A
  T -->|"result: label narrows"| L

What it is

  • A decision engine, not a model. The Rust crate appa-engine is documented as a pure function of the event log: it "performs no IO, reads no clock, and never mutates a store", so the same log yields the same decision.
  • Each session (a "trajectory") carries one security label, audience x trust. Reading a private record narrows the audience; reading text written by an outsider (web page, external commenter) lowers trust. The label only becomes more restrictive. Two further concepts are tracked: effects (what the agent already did, accumulated) and attention (a per-action approval requirement that an approval clears for that action only).
  • Policy is declarative TOML (appa.toml, version = 2). Each tool has a contract with delta (how its result changes the label), requires (what the session must satisfy), and effects (what is recorded after success). The first explicit contract whose argument selectors match is used; a * contract routes unmatched tools through an annotator; a schema error on a matched contract refuses the call.
  • When a call is blocked, the engine returns remedy plans instead of a bare refusal: a sanitizer (mask or redact so data may reach a wider audience), an authority (a person, approval service, or LLM evaluator that approves one action without loosening the session), or a subagent read (an untrusted read in a disposable child branch that returns only what the policy allows).
  • Licence MIT. Repository archestra-ai/OpenAPPA; the README labels it a "preview and an RFC" whose "config and wire surfaces may break without shims". The paper is accepted to the NeurIPS 2026 Workshop on Agents in the Wild. Development is sponsored by Archestra; the site says it is vendor-agnostic.

Why use it

  • Approval fatigue and classifier judges are the usual alternatives. The project's argument is that a second-model judge cannot track data flow across calls and is itself injectable, whereas a label algebra is not influenced by prompt text: the injected instruction may be followed by the model, but the exfiltrating call is still refused.
  • Blacklists (block curl, block rm -rf /) have no coverage measure and are bypassed by rewriting the command. A declarative contract set can be checked in CI (appa describe --check) and exercised with scripted traces (appa replay).
  • Refusals come with a legal way forward, which is the project's explanation for utility staying high (see the benchmark audit below).

When to use it (and when not)

Use it when an agent holds private context and also has outbound tools (email, issue trackers, shell, git push), which is the combination where prompt injection turns into data exfiltration. It fits coding agents, MCP gateways, and LLM proxies.

Do not rely on it for:

  • Anything outside the tool boundary. The guarantee is about calls mediated by the engine; a tool left unannotated under a permissive wildcard, or an integration that does not route a tool through it, is unprotected.
  • Correctness of labels. The engine trusts delta declarations and annotator answers. A contract that labels a private source as public defeats it; the algebra cannot detect a wrong declaration.
  • Integrity or hallucination of content that stays inside the allowed audience. It controls where data flows, not whether the data is true.
  • Stability. It is a preview; pin a commit.

Architecture

The engine is a pure core; an outer runtime owns state. The repository workspace splits into appa-engine (decision core), appa-policy (TOML to compiled policy), appa-runtime and appa-eventlog (state and log), and adapters for Claude Code, kagent, and Amp. The audience dimension is symbolic: a built-in chain self within internal within public, group references such as @finance, and literal readers (for example, the channel membership for a Slack channel ID pulled from the tool arguments). Trust is a finite, configurable rank chain (the three-trust-ranks example defines a custom one). The fold is a restrictive meet: minimum trust, intersected audience.

Subagent handling distinguishes three cases in the docs: resuming a session (same root, no fork), spawning a subagent (child bound to an approved spawn in the same family, answer crosses a checked return path), and a root fork (new family that inherits the source label and effect history frozen at opening, via the embedding call Runtime::open_root_fork, not automatically in the native Claude Code hooks).

How to use it

Claude Code path from the README (piping an installer to a shell is the project's documented route; read it first):

curl -fsSL https://openappa.com/install.sh | sh &&
  ~/.local/bin/appa plugin install claude-code
clappa          # protected session
/appa-guide     # policy setup skill

The README calls the Claude Code integration "a playground for the model, not the product". Other shapes: embed the runtime in an agent from another language, connect via hooks, or use an LLM proxy (Archestra).

A minimal contract set, taken from examples/tests/secret-stays-inside/appa.toml in the repository:

[policy]
version = 2

[[policy.tool]]
name = "mcp/files/read(path:/hr/*)"
delta = { audience = ["[email protected]"] }

[[policy.tool]]
name = "mcp/files/read"
delta = {}

[[policy.tool]]
name = "mcp/mail/send"
parameters = { type = "object", properties = { to = { type = "string" } }, required = ["to"] }
requires = { audience = { contains = ["$to"] } }
delta = {}

Tool names match exactly or as *; globs belong in argument selectors, not names (per the Archestra integration notes).

How to develop with it

Tests are trace files listing tool calls in order, each with expect allow | deny | authority | sanitizer. Real run against the repository at commit 6710c7f (built with cargo; no real tools execute, allowed calls return empty results):

$ appa replay -v --config examples/tests/secret-stays-inside/appa.toml examples/tests/secret-stays-inside
ok    ...secret-stays-inside.appa:2 mcp/mail/send allow
ok    ...secret-stays-inside.appa:7 mcp/files/read allow (after accepting the narrowing)
ok    ...secret-stays-inside.appa:12 mcp/mail/send deny
ok    ...secret-stays-inside.appa:17 mcp/mail/send allow
1 file: 1 ok, 0 failed, 0 could not run

The same command on untrusted-web-content gave bash allow, fetch allow (after accepting the narrowing), bash deny, files/write allow. A negative check: the trace copied with every expect deny changed to expect allow printed FAIL, reported 0 ok, 1 failed, and exited 1, so a wrong expectation breaks CI as intended. appa describe --check on the example printed that 3 declared tools have "no host observation to compare against", i.e. it cannot validate tool coverage outside a protected session.

Reference model of the label rule (stdlib Python; re-implements the documented meet and the two traces above, does not run the Rust engine):

from __future__ import annotations
from dataclasses import dataclass
from itertools import permutations

PUBLIC = None  # None = unrestricted audience


@dataclass(frozen=True)
class Label:
    trust: int = 1
    audience: frozenset[str] | None = PUBLIC

    def meet(self, other: "Label") -> "Label":
        if self.audience is PUBLIC:
            aud = other.audience
        elif other.audience is PUBLIC:
            aud = self.audience
        else:
            aud = self.audience & other.audience
        return Label(min(self.trust, other.trust), aud)


def allows(label: Label, need_trust: int, recipient: str | None) -> bool:
    if label.trust < need_trust:
        return False
    if recipient is None or label.audience is PUBLIC:
        return True
    return recipient in label.audience


def run(calls: list[tuple[str, str]]) -> list[bool]:
    label, out = Label(), []
    for tool, arg in calls:
        if tool == "read" and arg.startswith("/hr/"):
            label = label.meet(Label(1, frozenset({"[email protected]"})))
            out.append(True)
        elif tool == "read":
            out.append(True)
        elif tool == "fetch":
            label = label.meet(Label(0, PUBLIC))
            out.append(True)
        elif tool == "send":
            out.append(allows(label, 1, arg))
        elif tool == "bash":
            out.append(allows(label, 1, None))
        else:
            out.append(False)  # unknown tool: refused
    return out


trace = [("send", "[email protected]"), ("read", "/hr/salaries.csv"),
         ("send", "[email protected]"), ("send", "[email protected]")]
assert run(trace) == [True, True, False, True]
print("secret-stays-inside:", run(trace))

web = [("bash", ""), ("fetch", "http://x"), ("bash", ""), ("read", "/tmp/a")]
assert run(web) == [True, True, False, True]
print("untrusted-web:", run(web))

reads = [Label(1, frozenset({"a", "b"})), Label(1, frozenset({"b", "c"})), Label(0, PUBLIC)]
results = set()
for p in permutations(reads):
    l = Label()
    for r in p:
        nl = l.meet(r)
        assert nl.trust <= l.trust
        assert l.audience is PUBLIC or (nl.audience is not PUBLIC and nl.audience <= l.audience)
        l = nl
    results.add(l)
assert len(results) == 1
print("order-independent final label:", next(iter(results)))

l = Label().meet(Label(1, frozenset({"hr"}))).meet(Label(1, PUBLIC))
assert l.audience == frozenset({"hr"})          # public read does not widen
l = Label().meet(Label(1, frozenset({"a"}))).meet(Label(1, frozenset({"b"})))
assert not allows(l, 1, "a") and not allows(l, 1, "b")  # empty intersection
assert run([("curl", "")]) == [False]           # unknown tool refused
print("edge cases ok")

Executed output:

secret-stays-inside: [True, True, False, True]
untrusted-web: [True, True, False, True]
order-independent final label: Label(trust=0, audience=frozenset({'b'}))
edge cases ok

How to maintain it

  • Keep appa.toml in a repository, require appa describe --config appa.toml --check and appa replay in CI (the project documents a GitHub Actions workflow on its validation page), and add a trace for every incident or confusing block.
  • Re-run traces when a tool server changes its argument names; selectors and parameters schemas are matched literally, and a schema error on a matched contract refuses the call.
  • Sanitizer, authority, and annotator services are externals with timeouts and body limits (timeout_ms, max_body_bytes in the examples). Treat them as part of the trusted base: a sanitizer that fails to mask defeats the to = public permit it backs.
  • Pin the release or commit; the project states breaking changes will not ship with shims.

How to run it in production

  • Observability and reporting: the docs list an observability section and appa yell for reporting confusing blocks; this page did not exercise them.
  • Decide the wildcard policy deliberately. A catch-all annotator that returns empty changes leaves unknown tools open; a call no contract covers is refused, and a missing annotation is an operational refusal rather than a policy denial.
  • Approvals: human-in-the-loop authority is built in (builtin = "hitl"); an approval applies to one action and does not lift the session label.
  • Sessions need stable identity. In-process use gets it from the integration; proxy use needs a session header (see Archestra).

Benchmark claims audit

Claim (source) Status
"100% resistant to data exfiltration" (home page) Stronger than the evidence. Paper abstract: "zero observed attacks across 1,320 guarded episodes". Zero observed is not a proof of 100% resistance; the paper separately claims formal invariants, which this page did not check.
89% task completion, 0% attacks (home and README table) Reproduced from the README text only. Not re-run.
"37% to 90%" completion lift (home page) Not reconciled with the 89% table figure or with the paper's "64.2-91% utility" range; the conditions differ and were not compared.
FIDES lets 31% (site) or "28-35%" (README) through The site and README quote different forms of one result; not re-derived.
Auto mode 90% completion, 10% attacks The benchmark's auto arms use a scenario-specific auto-ifc variant and a bundled Claude Code runtime pinned in bench/corp/README.md; results depend on that setup and the chosen model.
6,600 episodes (paper v2) Stated in the abstract; not re-run.

The benchmark is built and published by the same group that builds the engine, on a mock environment (corp-systems: hr, finance, task tracker, public forum, vendor, email) plus the UK AISI AgentThreatBench. Independent replication was not found during this review.

Failure modes

  • Mislabelled source: a delta that under-restricts a private tool leaks silently.
  • Over-tight policy: every read narrows the session permanently, so long sessions drift into needing approvals; subagent reads exist to contain this, at the cost of summarisation loss.
  • Unmediated path: a tool or connector that bypasses the integration point is invisible to the engine.
  • Approval rubber-stamping: an authority that auto-approves, or an LLM evaluator used as authority, reintroduces a probabilistic decision.
  • Preview churn: config or wire breaks between commits.

References

  • OpenAPPA site: https://www.openappa.com/ , how it works: https://www.openappa.com/how-it-works , evaluation: https://www.openappa.com/evaluation , validation: https://www.openappa.com/validation
  • Source repository (MIT): https://github.com/archestra-ai/OpenAPPA (read at commit 6710c7f)
  • Paper: APPA, Recoverable Information-Flow Control for Real-World LLM Agents, https://arxiv.org/abs/2607.24625
  • NeurIPS 2026 Agents in the Wild workshop: https://agentwild-workshop.github.io/neurips2026/
  • AgentThreatBench (inspect_evals): https://github.com/UKGovernmentBEIS/inspect_evals

Related: APPA paper page · Archestra agent platform · Policy, guardrails and governance · Prompt-injection defense · Agent threat model · Sandboxing and isolation · Action execution boundary