HySTAR - fix the value-decomposition basis (overlapping sparse hypergraph scaffold) while learning adaptive spatiotemporal representations, eliminating structural target drift in cooperative multi-agent RL. Use for credit assignment under partial observability, stable high-order coalition value decomposition, or when dynamic grouping keeps reshuffling your learning targets.
Scanned 10/3/2026
npx -y skills add hiyenwong/ai_collection --skill anchored-hypergraph-credit-assignment --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Anchored Hypergraph Credit Assignment?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/hiyenwong-anchored-hypergraph-credit-assignment)More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.
---
name: anchored-hypergraph-credit-assignment
description: HySTAR - fix the value-decomposition basis (overlapping sparse hypergraph scaffold) while learning adaptive spatiotemporal representations, eliminating structural target drift in cooperative multi-agent RL. Use for credit assignment under partial observability, stable high-order coalition value decomposition, or when dynamic grouping keeps reshuffling your learning targets.
category: ai_collection
trigger_words: multi-agent RL, credit assignment, structural target drift, hypergraph, MAPPO, value decomposition, cooperative MARL, coalition value, SMAC, Google Research Football, agent death, grouping topology
---
# HySTAR: Anchored Hypergraphs for Stable Credit Assignment
**Source**: arXiv:2609.31531v1 (2026-09-25) — Luo, Zhang, Kuang, Yuan, Zeng et al. (8 authors), cs.LG.
## Problem: Structural Target Drift
Cooperative MARL under partial observability with shared rewards must assign team outcomes to individual agents **and high-order coalitions**. Two failure modes:
- **MAPPO-style critics**: compress joint behavior into one global value → no fine-grained credit.
- **Dynamic grouping critics**: adaptively reconstruct the interaction topology at every step → the *mapping from agents/coalitions to value components keeps changing* → the value-decomposition learning target is a moving object. The paper names this **structural target drift**: representation learning and target assignment are entangled, so neither converges cleanly.
## Core Mechanism: Anchor the Basis, Adapt the Features
HySTAR separates the two roles:
1. **Anchored decomposition scaffold** (fixed): an **overlapping sparse hypergraph** with uniform coverage — the *basis* mapping agents↔coalitions to value components never changes during training.
2. **Adaptive representation** (learned): a **spatiotemporal encoder** represents physical and task-dependent interactions, feeding *features* into the fixed scaffold.
3. **Agent-specific advantages**: combine **temporal relevance** (TD advantages) with **structural relevance** (hypergraph-derived coalition membership) to construct per-agent advantage estimates.
The scaffold being fixed kills structural target drift; the encoder being adaptive preserves expressiveness.
## Recipe
1. Start from MAPPO (shared reward, centralized critic).
2. Define a hypergraph over agents: hyperedges = candidate coalitions (overlapping allowed); ensure uniform coverage of all agents. Freeze this topology for the whole run.
3. Learn a spatiotemporal encoder (e.g., attention over agent observations + positions) that produces interaction-aware embeddings.
4. Value head: decompose global value over hypergraph nodes/edges using the encoder features — the *decomposition structure is fixed, only feature-dependent weights are learned*.
5. Advantage per agent = f(temporal advantage, structural advantage from its hyperedge memberships).
6. Train actors with per-agent advantages (PPO-style clipping as in MAPPO).
## Validation Results
- **SMAC (hardest settings)**: +16.7% relative gain over MAPPO, +15.6% over HYGMA (prior hypergraph method with dynamic topology).
- **Google Research Football**: ranks first on all 6 scenarios.
- **Traffic Junction**: convergence epochs reduced by up to 40.2% vs. MAGIC.
- **MPE**: highest episode rewards.
- Controlled experiments (topology variants, agent-death, neighborhood sizes, parameter sweeps) all support **anchoring the scaffold while adapting representations**.
## Reusable Design Principle
**When a learning target depends on a structure that the learner itself adapts, split the structure into (fixed scaffold) + (adaptive features).** Fixed scaffolds stabilize optimization targets; adaptive features carry the expressiveness. The failure pattern — dynamic grouping → target drift → unstable credit assignment — recurs in: continual learning (task-boundary drift), neural architecture search (supernet reparameterization), and any self-referential learning system.
**Activation**: credit assignment, MARL, hypergraph value decomposition, structural target drift, MAPPO variants, coalition credit, anchored scaffold.
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!