Megadose AI progress, ranked and analyzed.

Safety of Latent Communication in Multi-Agent Systems

· HF Daily Papers ·
Latent links between otherwise safety-aligned agents can make the system more willing to comply with harmful requests.

The paper says benignly trained representation-space links already increased harmful compliance versus text communication. Attacks that optimize or poison those links pushed the effect further, including a reinforcement-learning method that raised mean harmful-compliance scores from 27.9 to 76.9 across evaluated setups. The authors also report that reward-based repair reduced harmful compliance without changing the agents themselves. Source: HF Daily Papers' note

score 5

Categories: Research