Attention backend benchmarking for tabular foundation models - row/column alternating attention has asymmetric shape profiles (long row sequences, short column sequences, strided memory layout) so optimal backend (FlashAttention-2/3/4, cuDNN, vLLM, SageAttention) differs per attention type and hardware. Use when optimizing TabPFN/Mitra/ConTextTab-style inference, choosing attention kernels for 2D tabular layouts, or benchmarking attention on A100/H100/B200.
Installs into .claude/skills of the current project.
Are you the author of Tabular Attention Benchmark?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/hiyenwong-tabular-attention-benchmark)
---
name: tabular-attention-benchmark
description: Attention backend benchmarking for tabular foundation models - row/column alternating attention has asymmetric shape profiles (long row sequences, short column sequences, strided memory layout) so optimal backend (FlashAttention-2/3/4, cuDNN, vLLM, SageAttention) differs per attention type and hardware. Use when optimizing TabPFN/Mitra/ConTextTab-style inference, choosing attention kernels for 2D tabular layouts, or benchmarking attention on A100/H100/B200.
category: ai_collection
trigger_words: tabular foundation model, TabPFN, attention benchmark, FlashAttention, cuDNN, SageAttention, vLLM, row attention, column attention, 2D attention, tabular in-context learning, GPU kernel selection, A100 H100 B200, strided memory layout
---
# Benchmarking Attention Backends for Tabular Foundation Models
**Source**: arXiv:2609.31306v1 (2026-09-25) — Schambach, Biehl, Thelin (SAP), cs.LG. Code: github.com/SAP-samples/tabular-attention-benchmark
## Problem: 2D Attention ≠ 1D Attention
Tabular in-context learners (TabPFN, Mitra, ConTextTab) alternate **row attention** (over samples — long sequences) and **column attention** (over features — short sequences) on 2D grids of latent embeddings. This breaks the assumptions of 1D-tuned attention kernels:
- **Row attention**: long sequences (thousands of rows)
- **Column attention**: much shorter sequences (tens-hundreds of features)
- **Memory layout**: tabular data is strided → producing contiguous tensors for kernel calls is costly
- **Hidden dims**: small compared to modern LLMs (kernel heuristics tuned for d_model≥1024 mispredict here)
Efficient-attention research is almost entirely 1D; the 2D tabular setting was unexplored.
## Benchmark Design (reusable)
Reproducible harness measuring forward + backward throughput on **realistic tabular shapes** across:
- **Backends**: Torch SDPA (efficient + cuDNN), FlashAttention-2/3/4, vLLM (inference-only), SageAttention (inference-only)
- **Hardware**: A100, H100, B200 (three GPU generations)
- **Axes**: row vs. column attention separately; head dimension; sequence length
## Key Findings
1. **No single winner** — the optimal backend choice differs between column and row attention, varies across hardware generations and model specifics.
2. **FlashAttention generation-matched builds perform best overall** — but are at times **outperformed by cuDNN for column attention** at longer sequences, with **cross-over points depending on head dimension**.
3. **SageAttention wins for row attention at very large sequence lengths** (>16k rows).
4. Practical guidance: benchmark *your* attention type × head-dim × hardware combination; do not assume the 1D-LLM default transfers.
## Reusable Methodology: Shape-Asymmetric Kernel Benchmarking
When attention has **structurally heterogeneous profiles** (row vs. column, spatial vs. channel, token vs. feature), benchmark each profile **separately** before choosing kernels:
1. Enumerate backends per attention type (not globally).
2. Sweep the actual deployment shapes (sequence lengths, head dims) — cross-over points are shape-dependent, not global.
3. Account for memory-layout costs (strided→contiguous copies) in the measurement, not just kernel FLOPs.
4. Pin results to hardware generation — FlashAttention builds are generation-specific.
5. Keep an inference-only track (vLLM/SageAttention) separate from training-capable backends.
## When This Matters
- Deploying TabPFN-style in-context tabular prediction at scale
- Any architecture with alternating attention over different axes (vision transformers on non-standard grids, multimodal token mixing)
- Choosing between serving frameworks for tabular foundation models
**Activation**: tabular attention backend, FlashAttention vs cuDNN, SageAttention long sequences, 2D attention kernels, TabPFN optimization.