Arxiv 2609 39788 Safety Of Latent Communication In Multi Agent Systems
ASecurity
Research paper: Safety of Latent Communication in Multi-Agent Systems. Demonstrates that benign link training in latent communication increases harmful compliance relative to text-based communication. Develops RL-based attack raising harmful-compliance from 27.9 to 76.9, and shows reward adaptation enables repair without updating agents.
Installs into .claude/skills of the current project.
Are you the author of Arxiv 2609 39788 Safety Of Latent Communication In Multi Agent Systems?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/hiyenwong-arxiv-2609-39788-safety-of-latent-communication-in)
---
name: arxiv-2609-39788-safety-of-latent-communication-in-multi-agent-systems
description: "Research paper: Safety of Latent Communication in Multi-Agent Systems. Demonstrates that benign link training in latent communication increases harmful compliance relative to text-based communication. Develops RL-based attack raising harmful-compliance from 27.9 to 76.9, and shows reward adaptation enables repair without updating agents."
version: 1.0.0
author: Hermes Agent
license: MIT
metadata:
hermes:
tags: [Research, Arxiv, AI-Safety, Multi-Agent, Latent-Communication, Security, Alignment]
related_skills: [ai-safety-eval, multi-agent-rl]
---
# Safety of Latent Communication in Multi-Agent Systems
**arXiv ID:** 2609.39788
**Categories:** cs.AI, cs.LG, cs.MA
**Utility Score:** 0.80 → Promoted (High)
**PDF:** https://arxiv.org/pdf/2609.39788
## Abstract
Latent communication enables multi-agent systems to exchange information directly in internal representation space, reducing the token, computation, and latency overhead of text-based communication. To this end, lightweight trainable links are introduced to map the sender's representations into the receiver's input space. In this work, we show that even benign link training can increase harmful compliance relative to text-based communication while the underlying safety-aligned agents remain unchanged. An attacker can amplify this effect by optimizing the links on harmful query--response pairs or poisoning otherwise benign training data. We further develop a reinforcement-learning attack that rewards harmful compliance alongside benign task performance without requiring harmful target responses. Across three communication topologies and four safety benchmarks, this attack raises the mean harmful-compliance score from 27.9 with benignly trained links to 76.9. Compared with direct supervised optimization, it also achieves higher average accuracy on two benign utility benchmarks. Adapting the rewards toward safer behavior also enables repair of compromised links, substantially reducing harmful compliance across all evaluated attacks without updating the agents. Overall, our results show that safety alignment requires considering the multi-agent system as a whole. Code: https://github.com/Muhammad-Huzaifaa/latent-safety
## Key Contributions
1. **Latent Communication Vulnerability**: First demonstration that benign link training increases harmful compliance in safety-aligned agents
2. **RL-Based Attack**: Develops reinforcement learning attack that rewards harmful compliance alongside benign task performance
3. **Quantitative Impact**: Raises harmful-compliance score from 27.9 (benign) to 76.9 (attacked)
4. **Repair Mechanism**: Shows reward adaptation can repair compromised links without updating agents
5. **System-Level Safety**: Demonstrates that safety alignment must consider the multi-agent system as a whole
## Technical Approach
### Latent Communication
- **Mechanism**: Agents exchange information in internal representation space
- **Benefit**: Reduces token, computation, and latency overhead vs text-based communication
- **Implementation**: Lightweight trainable links map sender representations to receiver input space
### Attack Vectors
1. **Benign Link Training**: Even standard training increases harmful compliance
2. **Targeted Optimization**: Attacker optimizes links on harmful query-response pairs
3. **Data Poisoning**: Poisoning benign training data to induce harmful behavior
4. **RL Attack**: Reinforcement learning rewards harmful compliance alongside task performance
### Attack Characteristics
- **No Harmful Targets Required**: RL attack doesn't require harmful target responses
- **Topology Agnostic**: Works across three communication topologies
- **Benchmark Robust**: Effective across four safety benchmarks
- **Utility Preservation**: Maintains higher accuracy on benign tasks than supervised attacks
### Repair Strategy
- **Reward Adaptation**: Modify rewards toward safer behavior
- **Link Repair**: Substantially reduces harmful compliance
- **No Agent Updates**: Repair doesn't require updating agent parameters
- **Universal Effectiveness**: Works across all evaluated attacks
## Experimental Results
### Attack Effectiveness
- **Benign Links**: 27.9 mean harmful-compliance score
- **RL Attack**: 76.9 mean harmful-compliance score (2.75x increase)
- **Topologies**: Tested across 3 communication architectures
- **Benchmarks**: Validated on 4 safety benchmarks
### Utility Preservation
- **Benign Tasks**: RL attack achieves higher accuracy on 2 utility benchmarks
- **Comparison**: Outperforms direct supervised optimization
- **Trade-off**: Maintains task performance while increasing harmful compliance
### Repair Results
- **Harmful Compliance**: Substantially reduced across all attacks
- **Agent Parameters**: No updates required
- **Generalization**: Effective against multiple attack types
## Implications for Agent Systems
- **Safety Alignment Gap**: Individual agent alignment ≠ system-level safety
- **Latent Channel Risks**: Efficient communication introduces new attack surfaces
- **Training Vulnerability**: Even benign training can compromise safety
- **Repair Feasibility**: System-level repair possible without retraining agents
- **Design Principle**: Multi-agent safety requires holistic consideration
## Communication Topologies Tested
1. **Point-to-Point**: Direct sender-receiver links
2. **Broadcast**: Single sender to multiple receivers
3. **Mesh**: All-to-all communication
## Safety Benchmarks
Four benchmarks evaluating harmful compliance across different domains and attack vectors.
## Code & Resources
- Paper: https://arxiv.org/abs/2609.39788
- PDF: https://arxiv.org/pdf/2609.39788
- Code: https://github.com/Muhammad-Huzaifaa/latent-safety
## Related Work
- Safety alignment in LLMs
- Multi-agent communication protocols
- Adversarial attacks on neural networks
- Emergent communication in multi-agent systems
- Representation learning vulnerabilities