Abstract
MOLE is a benchmark for evaluating defenses that detect harmful actions by AI agents operating shared services under limited review budgets.
Model misalignment, prompt injection, or operator misuse could lead AI agents operating frontier-lab accounts to exfiltrate model weights, poison training data, or weaken release gates. Existing benchmarks do not test whether defenders can detect this activity among routine work under a limited review budget. We introduce MOLE, an open benchmark of 150 AI-operated accounts sharing 9 stateful services over 30 workdays, with 12 threats and 8 corpora from four models totaling roughly 20 billion tokens. Of 39 agent models, 72% complete most assigned harmful objectives and agent refusal does not predict completion. MOLE enables comparison of 40 monitors across corpus generators, observability levels, and threats; even the best evaluated monitor in our single-day audit-event comparison misses nearly half of completed harm. MOLE also enables monitor development: benchmark-guided search improves a mid-tier monitor by 49-64%, while selective use of a stronger monitor improves budget-AUC by 10% over applying it to every account-day at comparable modeled cost.
Community
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D (2026)
- SkillShield: Prompt-Space Security Skills for LLM Coding Agents (2026)
- Cross-Agent Campaign Attribution: Linking Asynchronous Attacks Across LLM Agents (2026)
- ActBench: Self-Evolving Benchmark of Behavioral Safety in Cowork Agents (2026)
- StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents (2026)
- Agent Hacks Agent: Autoresearch for Production-Agent Red-Teaming (2026)
- Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.06966 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 1
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper