Skip to content
Back to skills

Apst Association Profile Frozen Bci

ASecurity

Use for cross-session BCI decoding with frozen weights. APST.

  • 3 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added October 3, 2026
businessgogit

Works with

  • cursor
  • cli

Security analysis

A100/100

Scanned October 3, 2026

npx -y skills add hiyenwong/ai_collection --skill apst-association-profile-frozen-bci --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Apst Association Profile Frozen Bci?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Apst Association Profile Frozen Bci
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/hiyenwong-apst-association-profile-frozen-bci/badge)](https://www.skillsdirectory.com/skills/hiyenwong-apst-association-profile-frozen-bci)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: apst-association-profile-frozen-bci
description: "Use for cross-session BCI decoding with frozen weights. APST."
category: ai_collection
trigger_words: [cross-session intracortical decoding, frozen-weight BCI adaptation, association profile, FALCON benchmark, permutation-invariant set attention, few-shot calibration, unit drift, directional tuning profile, SVD-compressed profile, sliding-window causal transformer, FiLM conditioning, streaming neural decoding]
---

# APST: Association Profile-Conditioned Set-Temporal Transformer for Cross-Session Intracortical Motor Decoding (Frozen Weights)

**Source**: arXiv:2609.39080 — Zhang, Mo, Wen, Liang, Yang, Zeng, Wang, Wang (HKU + SUSTech, 30 Sep 2026)
**Code**: https://github.com/HsinyuanZhang/APST

## Core Problem

Intracortical BCI decoders degrade across recording sessions because electrode-tissue motion **changes the unit set**: units are lost, gained, and *drift* (a persisting unit can keep its firing statistics while changing its behavioral association — the Rokni et al. 2007 "unstable representations" problem). Existing fixes split into two camps, each with a flaw:

- **Gradient-based adaptation** (NDT2 fine-tuning, RNN-FT): needs backprop at deployment, risks overfitting a few calibration trials, and (on FALCON H1) costs 2.29× normalized latency.
- **Gradient-free, unsupervised adaptation** (SPINT, latent alignment): uses *unlabeled* activity — cannot resolve drift where firing statistics are preserved but the tuning changed.

**Key insight**: clinical BCIs already run periodic cued-movement calibration with a handful of labeled trials. Those labels are sufficient to estimate, **in closed form**, a compact per-unit "functional identity" that re-anchors the decoder — no gradient step needed.

## Method Architecture (three reusable patterns)

### Pattern 1 — Closed-form association profile (the 4-D interface)

Per unit u, from M labeled calibration trials (M as low as 4–8):

**Generic weighted ridge form** (Eq. 1):
```
a_raw[u] = argmin_v Σ_t ω[t]·(r[t,u] − φ[t]ᵀv)² + η‖v‖²
```
Behavior enters via design φ or weights ω. Two task-specific instantiations, both collapsed to the **same 4-D profile interface**:

1. **Directional tuning** (2-D cursor/finger tasks, FALCON M2 + DANDI688): first-harmonic cosine tuning (Georgopoulos 1982). φ = [1, cosθ, sinθ], uniform weights. Profile = [a, d, ρ=√(a²+d²), b] (cos/sin coefficients, modulation depth, baseline), z-scored by source-session moments (μ_dir, σ_dir).
2. **SVD-compressed association** (high-dim tasks, FALCON M1 16-ch EMG, H1 7-D velocity): behavior-weighted association matrix A[u,k] with softplus-split weights (K=16/14 states), normalized responses (Poisson-floor rate scale), then project onto **top-4 right singular vectors V₄ of pooled uncentered source associations** + source standardization. The source SVD basis is the cross-session bridge: target profiles live in the same 4-D space as source ones.

Cost: **< 15k operations** (< 0.1% of total calibration MACs). One calibration pass, zero target gradients.

### Pattern 2 — Profile-conditioned unit identity (FiLM at zero-init)

Calibration activity → trial-averaged signature h[u]; profile p[u] modulates it:
```
[γ,β] = W_o·ReLU(W_i·p + b_i) + b_o     (W_o, b_o ZERO-initialized — adaLNZero)
h̃[u] = (1+γ)⊙h[u] + β
e[u] = MLP([h̃[u]; p[u]])                  ← unit identity, computed once, cached
```
Zero-init means training starts from concatenation alone; the FiLM path is a pure residual add-on. **The cached identity costs nothing at streaming time.**

### Pattern 3 — Permutation-invariant set-temporal decoder

Two conditioning sites, dual use:
- **Site 1** (calibration): identity e[u] as above.
- **Site 2** (streaming token): live activity → 5-bin causal conv → add projected identity, concat profile → token z[u,t].

Unit tokens aggregate via **8 learned slot queries + masked multi-head set-attention** (Set Transformer / Perceiver style) — invariant to any unit count/order, padding mask handles variable N_s. Population token (256-D) feeds a **4-layer pre-LN causal transformer** with per-layer ALiBi-style relative-time bias and **sliding-window KV cache bounded by receptive field R = 5 + Σ(w_ℓ−1)** — streaming state is O(R), not O(session length). Inference-only EMA (α=1/3, resets at session boundary) smooths output. **Whole-unit dropout during source training** teaches the set encoder to survive unit loss.

## Headline Results

| Benchmark | APST (frozen) | Activity-only | RNN-FT (gradient) | SPINT (unlabeled) |
|---|---|---|---|---|
| DANDI688 Sub-C (held-out) | **0.78** | 0.40 | 0.77 (FT) | — |
| DANDI688 Sub-M (held-out) | **0.81** | 0.58 | 0.70 | — |
| FALCON M1 (private held-out) | 0.65 | 0.62 | 0.59 (NDT2) | 0.66 |
| FALCON M2 | **0.42** | 0.24 | 0.43 (NDT2) | 0.26 |
| FALCON H1 | **0.44** | 0.24 | 0.52 (NDT2) | 0.29 |

- **Sample efficiency**: 8 APST calibration trials beat 16-trial RNN fine-tuning on Sub-C (0.75 vs 0.74); on Sub-M APST leads by 0.11–0.38 at *every* budget 4–32.
- **Compute asymmetry**: APST calibration = 3.83–21.75M MACs vs RNN fine-tune = 383–1195 *billion* forward MACs (~5 orders of magnitude).
- **Held-in vs held-out gap (HI−HO)**: APST (0.12/0.22/0.16) — gains don't trade away within-session accuracy; smaller gaps than SPINT on M2/H1.
- **Drift robustness**: on the most distant session (Sub-C 12-01, 138 days post-source), APST retains 0.58 vs RNN-FT 0.48; on Sub-M sessions 8–11 days out, APST beats both baselines on all three.
- **Latency**: 0.09–0.15 normalized (real-time capable), vs 2.29 for NDT2 Multi on H1.

## Critical Ablations (why it works)

1. **Profile shuffle scores BELOW activity-only** (0.00/0.03 on Sub-C/Sub-M): profiles act as *unit-specific identities*, not session-level summaries — shuffling keeps the session profile set constant but destroys per-unit binding.
2. **Static identity table fails** (−0.32 to 0.38): units re-sorted per session can't be tracked by one frozen table — recalibration is genuinely needed.
3. **Token-only ≈ dual conditioning** (Δ ≤ 0.01); identity-only site drops H1 to 0.40 vs 0.48 — the streaming token conditioning carries most benefit; site-1 kept because it's free (cached).

## Reusable Patterns & Transfer Guide

1. **Closed-form identity estimation beats fine-tuning on few shots**: whenever a domain gives a few labeled probes + a stable low-dim basis (here: motor manifold SVD / cosine tuning), project per-sensor associations into it analytically instead of backprop-ing. Applicable to any drifting-sensor decoding (neuropixels across days, wearable IMU re-placement, EEG cap shifts).
2. **Two-level conditioning — cache the static part**: split conditioning into calibration-time (identity, cached forever) vs streaming-time (live token concat). Zero-init (adaLNZero/FiLM) makes the added path a safe residual.
3. **Permutation-invariance is the correct inductive bias for unit-set drift**: set-attention over unit tokens with slot queries handles any N; whole-unit dropout during source training is essential regularization.
4. **Shared low-dim interface across heterogeneous tasks**: directional tuning (cos/sin) and SVD-projection both land in the SAME 4-D space — one decoder conditioned on a unified "functional fingerprint" interface. Generalizes the idea: design per-task estimators, share the conditioning interface.
5. **Bounded streaming state**: sliding-window causal attention with per-layer KV cache + ALiBi bias keeps decoding O(R) — for clinical streaming decoders this matters as much as R².

## Limitations (stated honestly)

- Profile estimators and SVD bases are **still designed per task** (directional vs SVD variants); unified learned estimators are future work.
- Division of labor between the two conditioning sites not fully characterized.
- Uses target labels (unlike SPINT) — requires the clinical recalibration workflow to exist.
- FALCON M1: SPINT slightly ahead (0.66 vs 0.65) — unlabeled methods can win where tuning profiles are weak (EMG association).

## Key References

- FALCON benchmark: Karpowicz et al., NeurIPS 2024
- SPINT (main competitor): Le et al., NeurIPS 2025 — permutation-invariant unit embeddings from unlabeled activity
- NDT2 Multi: Ye et al., NeurIPS 2023 — labeled few-shot fine-tuning baseline
- Rokni et al., Neuron 2007 — unstable neural representations (the drift problem origin)
- Gallego et al., Nat Neurosci 2020 — motor manifold stability (DANDI688 source)
- Set Transformer: Lee et al., ICML 2019; Perceiver: Jaegle et al., ICML 2021

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…