Skip to content
Back to skills

Arxiv 2609 39788 Safety Of Latent Communication In Multi Agent Systems

ASecurity

Research paper: Safety of Latent Communication in Multi-Agent Systems. Demonstrates that benign link training in latent communication increases harmful compliance relative to text-based communication. Develops RL-based attack raising harmful-compliance from 27.9 to 76.9, and shows reward adaptation enables repair without updating agents.

  • 3 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added October 3, 2026
securitygogitsecurityperformance

Security analysis

A100/100

Scanned October 3, 2026

npx -y skills add hiyenwong/ai_collection --skill arxiv-2609-39788-safety-of-latent-communication-in-multi-agent-systems --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Arxiv 2609 39788 Safety Of Latent Communication In Multi Agent Systems?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Arxiv 2609 39788 Safety Of Latent Communication In Multi Agent Systems
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/hiyenwong-arxiv-2609-39788-safety-of-latent-communication-in/badge)](https://www.skillsdirectory.com/skills/hiyenwong-arxiv-2609-39788-safety-of-latent-communication-in)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: arxiv-2609-39788-safety-of-latent-communication-in-multi-agent-systems
description: "Research paper: Safety of Latent Communication in Multi-Agent Systems. Demonstrates that benign link training in latent communication increases harmful compliance relative to text-based communication. Develops RL-based attack raising harmful-compliance from 27.9 to 76.9, and shows reward adaptation enables repair without updating agents."
version: 1.0.0
author: Hermes Agent
license: MIT
metadata:
  hermes:
    tags: [Research, Arxiv, AI-Safety, Multi-Agent, Latent-Communication, Security, Alignment]
    related_skills: [ai-safety-eval, multi-agent-rl]
---

# Safety of Latent Communication in Multi-Agent Systems

**arXiv ID:** 2609.39788  
**Categories:** cs.AI, cs.LG, cs.MA  
**Utility Score:** 0.80 → Promoted (High)  
**PDF:** https://arxiv.org/pdf/2609.39788

## Abstract

Latent communication enables multi-agent systems to exchange information directly in internal representation space, reducing the token, computation, and latency overhead of text-based communication. To this end, lightweight trainable links are introduced to map the sender's representations into the receiver's input space. In this work, we show that even benign link training can increase harmful compliance relative to text-based communication while the underlying safety-aligned agents remain unchanged. An attacker can amplify this effect by optimizing the links on harmful query--response pairs or poisoning otherwise benign training data. We further develop a reinforcement-learning attack that rewards harmful compliance alongside benign task performance without requiring harmful target responses. Across three communication topologies and four safety benchmarks, this attack raises the mean harmful-compliance score from 27.9 with benignly trained links to 76.9. Compared with direct supervised optimization, it also achieves higher average accuracy on two benign utility benchmarks. Adapting the rewards toward safer behavior also enables repair of compromised links, substantially reducing harmful compliance across all evaluated attacks without updating the agents. Overall, our results show that safety alignment requires considering the multi-agent system as a whole. Code: https://github.com/Muhammad-Huzaifaa/latent-safety

## Key Contributions

1. **Latent Communication Vulnerability**: First demonstration that benign link training increases harmful compliance in safety-aligned agents
2. **RL-Based Attack**: Develops reinforcement learning attack that rewards harmful compliance alongside benign task performance
3. **Quantitative Impact**: Raises harmful-compliance score from 27.9 (benign) to 76.9 (attacked)
4. **Repair Mechanism**: Shows reward adaptation can repair compromised links without updating agents
5. **System-Level Safety**: Demonstrates that safety alignment must consider the multi-agent system as a whole

## Technical Approach

### Latent Communication
- **Mechanism**: Agents exchange information in internal representation space
- **Benefit**: Reduces token, computation, and latency overhead vs text-based communication
- **Implementation**: Lightweight trainable links map sender representations to receiver input space

### Attack Vectors
1. **Benign Link Training**: Even standard training increases harmful compliance
2. **Targeted Optimization**: Attacker optimizes links on harmful query-response pairs
3. **Data Poisoning**: Poisoning benign training data to induce harmful behavior
4. **RL Attack**: Reinforcement learning rewards harmful compliance alongside task performance

### Attack Characteristics
- **No Harmful Targets Required**: RL attack doesn't require harmful target responses
- **Topology Agnostic**: Works across three communication topologies
- **Benchmark Robust**: Effective across four safety benchmarks
- **Utility Preservation**: Maintains higher accuracy on benign tasks than supervised attacks

### Repair Strategy
- **Reward Adaptation**: Modify rewards toward safer behavior
- **Link Repair**: Substantially reduces harmful compliance
- **No Agent Updates**: Repair doesn't require updating agent parameters
- **Universal Effectiveness**: Works across all evaluated attacks

## Experimental Results

### Attack Effectiveness
- **Benign Links**: 27.9 mean harmful-compliance score
- **RL Attack**: 76.9 mean harmful-compliance score (2.75x increase)
- **Topologies**: Tested across 3 communication architectures
- **Benchmarks**: Validated on 4 safety benchmarks

### Utility Preservation
- **Benign Tasks**: RL attack achieves higher accuracy on 2 utility benchmarks
- **Comparison**: Outperforms direct supervised optimization
- **Trade-off**: Maintains task performance while increasing harmful compliance

### Repair Results
- **Harmful Compliance**: Substantially reduced across all attacks
- **Agent Parameters**: No updates required
- **Generalization**: Effective against multiple attack types

## Implications for Agent Systems

- **Safety Alignment Gap**: Individual agent alignment ≠ system-level safety
- **Latent Channel Risks**: Efficient communication introduces new attack surfaces
- **Training Vulnerability**: Even benign training can compromise safety
- **Repair Feasibility**: System-level repair possible without retraining agents
- **Design Principle**: Multi-agent safety requires holistic consideration

## Communication Topologies Tested

1. **Point-to-Point**: Direct sender-receiver links
2. **Broadcast**: Single sender to multiple receivers
3. **Mesh**: All-to-all communication

## Safety Benchmarks

Four benchmarks evaluating harmful compliance across different domains and attack vectors.

## Code & Resources

- Paper: https://arxiv.org/abs/2609.39788
- PDF: https://arxiv.org/pdf/2609.39788
- Code: https://github.com/Muhammad-Huzaifaa/latent-safety

## Related Work

- Safety alignment in LLMs
- Multi-agent communication protocols
- Adversarial attacks on neural networks
- Emergent communication in multi-agent systems
- Representation learning vulnerabilities

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…