Archestra released OpenAPPA, an open-source security engine that achieved zero successful attacks across two major security benchmarks designed to test AI agent vulnerabilities. The tool addresses a critical weakness in agentic systems: data leaks caused by prompt injection or model hallucination. Unlike existing approaches that rely on probabilistic models to judge each tool call, OpenAPPA operates outside the agent's prompt and execution loop, using deterministic rules to track data flow and enforce security policies.

Testing on Bench-Corp, which includes 20 multi-step enterprise workflows, and the OWASP AgentThreatBench showed OpenAPPA maintained an 89% task completion rate while blocking every attempted breach. By comparison, Claude Code's auto mode allowed 10% of attacks to succeed with a 90% completion rate, while Microsoft FIDES permitted 31% of attacks and completed only 41% of tasks. The benchmarks test explicit policy violations including sensitive data sharing, prompt injection, approval bypasses, and tenant isolation failures. AgentThreatBench, which operationalizes the OWASP Top 10 for Agentic Applications (2026), uses a dual-metric scoring system that evaluates both utility and security—OpenAPPA earned a perfect security score. When the research team disabled recovery strategies in ablation experiments, task completion fell to 35%, suggesting these mechanisms are essential to maintaining operational usefulness under strict security constraints.

The documentation explains why current industry solutions fail: second models that judge tool calls can't track data flow across multiple calls, and because these classifiers are themselves vulnerable to prompt injection, systems hide tool outputs from them entirely. Even the best probabilistic approaches cap out at 99.3% accuracy, which translates to thousands of breaches when agents make millions of calls. The team writes that rule sets end up "either so tight they break the agent or so intricate nobody can audit what they permit." Agents have proven adept at circumventing simple restrictions—a blocked rm -rf command may be replaced with equivalent Python code—while overly restrictive policies that prevent such workarounds also block legitimate tasks. The GitHub repository frames the challenge this way: an agent that permits unauthorized data flows is unsafe, while an agent that refuses valid work is useless.

OpenAPPA resolves this tension through what Archestra calls an Agentic Permissions Policy Algebra (APPA), detailed in a paper by Arseny Kravchenko, Vadim Liventsev, Innokentii Konstantinov, Ildar Iskhakov, and Matvey Kukuy. The engine runs outside the agent's loop, defeating any attempts by the underlying language model to inspect, negotiate with, or bypass policy rules. Security policies are configured in a single appa.toml file that defines data sources, audiences, trust levels, and authorities. The system jointly labels and monitors both audience—the authorized set of consumers—and trust, the degree of data verification. Labels compose using lattice algebra and can only become more restrictive: reading restricted records narrows the audience, while reading unvetted external web pages lowers trust. Each tool contract specifies what permissions it requires to run, what restrictions it applies when returning data, and what audit trail it leaves. When an agent attempts an illegal action, the engine halts dispatch and provides structured recovery pathways: sanitizers strip personally identifiable information to expand the permitted audience, authorities route requests to human operators for approval, and disposable child branches isolate untrusted data reads in transient subagent branches that return only sanitized outputs.

The formal algebra and recovery guarantees are published on arXiv, and OpenAPPA is currently available as a preview with documentation containing additional technical details and empirical results. The research team encourages readers to examine both the paper and documentation for a deeper understanding of the approach. Organizations deploying agentic systems now face a stark choice between solutions that leak data or block legitimate work—but deterministic policy enforcement may finally offer a path forward. The architecture's success hinges on whether enterprises can translate their implicit security intuitions into explicit lattice rules, a translation challenge that probabilistic shortcuts were designed to avoid but ultimately couldn't solve.