Open, privacy-bounded assurance for AI agents: containment provenance, identity passports, authorization twins, OTel evidence, CI gates, and OSCAL.
-
Updated
Aug 29, 2026 - Python
Open, privacy-bounded assurance for AI agents: containment provenance, identity passports, authorization twins, OTel evidence, CI gates, and OSCAL.
Action-graded severity scoring (L0-L6) for tool-using AI agents, computed from red-team execution traces.
A local-first research scaffold for evaluating models, agent harnesses, and complete agent products on realistic, stateful tasks.
Formal runtime verification for tool-using LLM agents: MFOTL/MonPoly replayed offline on AgentDojo, STAC and R-Judge. Paper, MFOTL specifications, experiments, and the benchmark-readiness audit (CPSIoTSec 2026).
Measure how well an agent tool-call authorization layer separates legitimate actions from injected ones, on AgentDojo ground truth. No agent runs, no model calls, seconds to run. Ships the controls that make a block rate meaningful.
LangChain-native AgentDojo benchmark: utility + ASR evaluation across banking, slack, travel, and workspace suites.
Benchmarking schema-valid false tool observations and defense baselines for tool-using LLM agents.
Evaluating provenance-gated tool calls as a prompt injection defense on AgentDojo. Reproducible runs, per-case analysis, and published results.
Security audit of LLM-based multi-agent systems with indirect prompt-injection PoCs and mitigations.
Reproducible evaluation of deterministic sequence rules for AI-agent tool-call traces.
Personal research project — solo, unaffiliated. Inspect AI evaluation framework for LLM agent security: ASR, benign utility, and Transparency Rate across prompt injection, tool poisoning, and psych attacks.
Add a description, image, and links to the agentdojo topic page so that developers can more easily learn about it.
To associate your repository with the agentdojo topic, visit your repo's landing page and select "manage topics."