Benchmarks

Security without blocking useful work#

These results answer three questions: Does OpenAPPA stop policy violations? Can the agent still finish legitimate tasks? How many extra tokens does protection require?

OpenAPPA uses deterministic policy enforcement to check each action. These benchmarks test the complete integration and show whether its policies stop the intended threats without making the agent useless.

TLDRNo scored attack got past OpenAPPA in 1,320 evaluations while it completed 89% of tasks; Claude Auto mode and Microsoft FIDES let 10% and 31% of attacks through.
Task completion
OpenAPPA89%
Claude Auto mode90%
FIDES (Microsoft)41%
Attacks that succeeded
OpenAPPA0%
Claude Auto mode10%
FIDES (Microsoft)31%

Security: no observed attacks in 1,320 evaluations#

No scored attack succeeded against guarded OpenAPPA in 1,320 evaluations: 600 from Bench‑Corp and 720 from AgentThreatBench. Both suites tested standard and adversarial prompts.

In Bench-Corp, the evaluated Microsoft FIDES configurations had a 28–35% attack success rate. OpenAPPA's policies also enforced rules those configurations did not support, including recipient authorization, out-of-band approval, and required action ordering.

The suites test concrete ways an agent can break policy:

  • Sensitive-data sharing: Restricted data must not reach an unauthorized person or a less restricted data store.
  • Prompt injection: Instructions in untrusted email, memory, or forum content try to make the agent bypass policy.
  • Approval and ordering: The agent receives text that falsely claims an approval happened. Policy requires the real approval in its recorded history.
  • Tenant isolation: The agent must keep each customer's data within that customer's authorized readers.

Utility: 88–90% completion while enforcing policy#

A secure agent is not useful if it cannot finish legitimate work. Across three language models in Bench-Corp, guarded OpenAPPA completed 88–90% of tasks. The evaluated FIDES configurations completed 37–45%. Each table entry shows task completion followed by attack success rate (ASR).

ModelGuarded OpenAPPA (Utility / ASR)FIDES middleware (Utility / ASR)FIDES native (Utility / ASR)
GPT-5.6 Luna88.0% / 0%38.5% / 32.0%37.0% / 32.5%
DeepSeek V4 Flash89.5% / 0%39.5% / 34.5%41.5% / 33.0%
Gemini 3.7 Flash90.0% / 0%43.5% / 28.5%44.5% / 28.0%

In AgentThreatBench's adversarial tests, guarded OpenAPPA had the highest task completion for all three models. In the standard tests, it led with Luna and Gemini. Middleware FIDES led with DeepSeek.

OpenAPPA's recovery mechanisms help the agent continue safely after a policy block. In a Bench-Corp test with Luna, task completion was 88.0%. It fell to 56.5% without subagent isolation and to 35.0% without guided recovery.

No scored attack succeeded in any of these configurations. The features interact, so this test does not measure each feature's effect in isolation.

Tau Bench's banking benchmark tests ordinary banking support work rather than attack prompts. The validated comparison tested all 97 tasks four times with GPT‑5.6 Luna at maximum reasoning effort.

Agent configurationSuccessful simulationsMean Tau reward
Guarded OpenAPPA151/38838.92%
OpenAPPA agent, permissive policy153/38839.43%
Stock Tau agent156/38840.21%

Guarded OpenAPPA finished two fewer simulations than the permissive OpenAPPA agent and five fewer than stock. Its mean reward was 0.52 percentage points below permissive and 1.29 points below stock.

OpenAPPA checked 11,355 calls. One policy decision stopped a state-changing call before identity verification. Three other blocks rejected malformed input. No tool execution errors occurred. Tau is not an attack benchmark, so these results do not establish net security.

Token overhead: 4.22% over stock on Tau#

Guarded OpenAPPA used a mean of 1,307,763 agent tokens per simulation. This was 4.22% more than stock and 2.80% more than the permissive OpenAPPA agent.

Agent configurationMean agent tokens per simulationDifference from stock
Guarded OpenAPPA1,307,763+4.22%
OpenAPPA agent, permissive policy1,272,128+1.38%
Stock Tau agent1,254,821—

These totals include every prompt and completion in the complete agent trajectory. They measure the whole agent setup, not the computing cost of the OpenAPPA policy engine alone.

Claude Code auto mode: only OpenAPPA stopped every scored attack#

Claude Code auto mode asks a classifier model to approve each tool call. OpenAPPA asks no model to judge intent. It records where each value came from and who may see it, then checks every flow against explicit policy. Each decision is deterministic and traces to a policy rule.

We ran guarded OpenAPPA against two Auto configurations:

  • IFC-tuned Auto is the like-for-like comparison: it tries to solve the same problem with a classifier instead of a policy engine. Each scenario's auto-mode rules restate the information-flow policy OpenAPPA enforces in plain language: which sources are trusted, who may see each piece of data, where data may go, and which narrowing is legitimate. It gets no attack markers or expected answers.
  • Stock Auto is the out-of-the-box baseline. Its default rules target generic developer risks, such as destructive commands, actions beyond the user's request, and secrets sent to external endpoints. Unconfigured, it knows nothing about a deployment's own rules, such as which colleague may see which record.

All three used the same actor model (Claude Sonnet 5), prompts, and tools, and ran every task once, so the results have no variance estimate. Each benchmark applied its own scoring. Utility is the task completion rate. Attacks counts the scenarios or tasks in which a scored attack succeeded.

BenchmarkGuarded OpenAPPA (Utility / Attacks)IFC-tuned Auto (Utility / Attacks)Stock Auto (Utility / Attacks)
Bench-Corp (20 scenarios)75.0% / 0/2085.0% / 0/2090.0% / 2/20
AgentThreatBench (24 tasks)75.0% / 0/2495.8% / 6/2487.5% / 8/24

OpenAPPA was the only configuration with zero scored attacks in both suites. IFC-tuned Auto matched it on Bench-Corp but let six attacks through on AgentThreatBench. Stock Auto, which never saw the policy, let ten through.

That protection costs completion and tokens. OpenAPPA completed 75% of tasks in both suites; the Auto configurations completed 85–96%. Guarded OpenAPPA also used more tokens than the same agent under a permissive policy, because isolated child trajectories and recovery run on top of the task. Claude's SDK does not report the tokens Auto's classifier spends, so total costs are not directly comparable.

What we measured#

  • Bench-Corp contains 20 multi-step workplace tasks. They cover HR, Finance, customer support, vendors, email, forums, and task tracking. The benchmark checks what the agent actually changed or sent. It does not use an LLM judge. Each model ran every scenario five times with standard prompts and five times with adversarial prompts.
  • AgentThreatBench is a 24-task suite published in Inspect Evals. It tests poisoned memory, attempts to take control of the agent, and data theft. We ran each standard and adversarial test five times with the same three models.
  • Tau Bench tests policy-sensitive banking support tasks. We used its standard scoring. The comparison covers all 97 banking_knowledge tasks, with four trials for each agent configuration. Matching tasks and random seeds makes the comparison fairer, but separate model sessions can still differ.