AI 编程 › AI 资讯 › 正文

MADBench: Benchmarking the Security of Multi-Agent Debate

arxiv · arXiv · 2026-09-30 15:12 · 评分 88

提出MADBench,系统评估多智能体辩论在多种对抗攻击下的安全性。辩论虽能相互纠错,也可能传播对抗错误并误导群体答案,该基准考察辩论机制能否缓解攻击,为多智能体安全提供评测框架。

原文:arXiv | 返回 AI 资讯列表