Reliable Weak-to-Strong Monitoring of LLM Agents
Neil Kale , Chen Bo Calvin Zhang , Kevin Zhu , Ankit Aich , Paula Rodriguez , Scale Red Team , Christina Q. Knight , Zifan Wang
- 🏛 Institutions
- Scale AI , Carnegie Mellon University , Massachusetts Institute of Technology
- 📅 Date
- August 26, 2025
- 📑 Publisher
- ICLR 2025
- 💻 Env
- 🔑 Keywords
TLDR
Stress-tests LLM agent monitoring systems for detecting covert misbehavior using a monitor red-teaming (MRT) workflow varying agent/monitor awareness and adversarial evasion strategies, evaluated on SHADE-Arena for tool-calling agents and CUA-SHADE-Arena for computer-use agents.
Related papers (9)
- SIR: Self-improving Red-teaming for Compute Use AgentsAugust 31, 2026 · arXiv
- Genesis: Evolving Attack Strategies for LLM Web Agent Red-TeamingOctober 21, 2025 · ICME 2026
- AgentHijack: Visual Patch Attacks on Multimodal Computer-Use AgentsSeptember 6, 2026 · arXiv
- WiP: Characterizing and Defending Against Mobile-Agent-Driven MFA AutomationSeptember 2, 2026 · arXiv
- Beyond the Verdict: Evidence-Aligned Evaluation of Visual Prompt-Injection GuardrailsSeptember 2, 2026 · arXiv
- WebMCP-Phalanx: Enforcing and Characterizing Trust Boundaries for Browser-Integrated LLM AgentsAugust 25, 2026 · arXiv
- Are Android GUI Agents Robust Against Runtime Anomalies? AnTrap: Evaluating Agents in Dynamic Adversarial EnvironmentsAugust 25, 2026 · arXiv
- ADeptS-Bench: Measuring the Trustworthiness of Computer Use Agents Across DevicesAugust 25, 2026 · arXiv
- MobileWorldSafety: Benchmarking GUI Agent Safety Against Environmental Injection Attacks in Android AppsAugust 18, 2026 · arXiv