Is Monitoring Enough? Strategic Agent Selection For Stealthy Attack in Multi-Agent Discussions
Abstract
Multi-agent discussions have been widely adopted, motivat-ing growing efforts to develop attacks that expose their vulnerabilities.In this work, we study a practical yet largely unexplored attack sce-nario, the discussion-monitored scenario, where anomaly detectors con-tinuously monitor inter-agent communications and block detected ad-versarial messages. Although existing attacks are effective without dis-cussion monitoring, we show that they exhibit detectable patterns andlargely fail under such monitoring constraints. But does this imply thatmonitoring alone is sufficient to secure multi-agent discussions? To an-swer this question, we develop a novel attack method explicitly tailored tothe discussion-monitored scenario. Extensive experiments demonstratethat effective attacks remain possible even under continuous monitoring,indicating that monitoring alone does not eliminate adversarial risks.