Open-source, 100% reproducible AI Agent Runtime Security Benchmark & Sandbox Environment (RFC-010 Draft Protocol).
-
Updated
Aug 28, 2026 - HTML
Open-source, 100% reproducible AI Agent Runtime Security Benchmark & Sandbox Environment (RFC-010 Draft Protocol).
Privacy-first security and mission-assurance lab for AI agents: TraceProof, authorization twins, MCP/OTel evidence, CI gates, and OSCAL exports.
Adversarial security benchmark for agent authorization: does a compromised agent's policy-violating proposal become an unauthorized external effect? 73 trials, nine families, an independent oracle, per-mechanism ablation, confidence intervals. 0 unauthorized effects in 61 attack trials (95% CI [0.0%, 5.9%]). Reproduction is partial.
Open deterministic security tests for unsafe multi-agent handoffs and authority escalation.
Open-source benchmark for adversarial evidence attacks on LLM-based cybersecurity auditors, targeting ACM AsiaCCS 2027.
Local-first workbench to run, inspect, compare, report, and gate OpenAI Codex Security scans.
Deterministic security benchmark for tool-using AI agents
Vendor-neutral benchmark measuring how MCP security proxies/gateways DEFEND against 22+ attack vectors — crosswalked to NIST AI RMF & OWASP LLM/Agentic Top 10. CI-gated, reproducible, DOI-cited. Submit your tool to the leaderboard.
Internal PyPI SCA precision and recall benchmark corpus
Production-grade microservices security benchmark featuring OWASP Top 10 logic exploits, automated remediation, custom Semgrep SAST rules, and CI/CD DevSecOps gates.
FreightSkillBench is a reproducible benchmark for evaluating document-to-transaction integrity, prompt-injection risk, and security controls in AI-enabled shipping and logistics workflows.
Internal Go Modules SCA precision and recall benchmark corpus
Smart contract vulnerability detection and benchmarking framework for Solidity, including SDB and enhanced DeFi security test suites.
GitHub action for Maester
Open AI-for-security validation benchmark: non-LLM scorer + a SOTA-validation loop. Labeled positive corpus withheld pending coordinated disclosure.
The core repository for the Maester module with helper cmdlets that will be called from the Pester tests.
Automated adversarial security testing for AI agents. Deploys an LLM-powered attacker against tool-using systems, validates violations via deterministic oracles, and produces reproducible vulnerability reports with causal attack graphs.
Reproducible benchmark for smart-contract security tools, measuring precision, recall, and false positives against executable PoCs and versioned ground truth.
Product-security LLM benchmark harness for realistic AppSec, supply-chain, and LLM application security evaluations.
ReplayBench-IoT: reproducible IoT replay-defense benchmark with Monte Carlo sweeps, CI, static demo, and hardware-validation artifacts.
Add a description, image, and links to the security-benchmark topic page so that developers can more easily learn about it.
To associate your repository with the security-benchmark topic, visit your repo's landing page and select "manage topics."