Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Cpu Cache Opt

ASecurity

Use when diagnosing cache misses with perf, fixing false sharing, choosing AoS or SoA layout, or adding software prefetch. Not for cache theory: use memory-hierarchy-and-caches.

54 stars
0 votes
0 copies
1 views
Added 9/20/2026
ai-agentsc++bashnodeperformance

Security Analysis

A100/100

Pro scans all 3 files and shows the line behind each finding

Scanned 9/20/2026

$npx -y skills add OutlineDriven/outline-driven-development --skill cpu-cache-opt --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Cpu Cache Opt?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Cpu Cache Opt
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/outlinedriven-cpu-cache-opt-outline-driven-development/badge)](https://www.skillsdirectory.com/skills/outlinedriven-cpu-cache-opt-outline-driven-development)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: cpu-cache-opt
description: 'Use when diagnosing cache misses with perf, fixing false sharing, choosing AoS or SoA layout, or adding software prefetch. Not for cache theory: use memory-hierarchy-and-caches.'
---

# CPU cache optimization

Cache performance is a layout property before it is a code property. Measure with hardware counters, move the data so the working set fits, and re-measure. Every change must show up in the counters, or it reverts.

## Contract

| Field | Bound contract |
|---|---|
| Trigger | The task diagnoses cache misses, detects or fixes false sharing, restructures data layout, evaluates AoS versus SoA, or decides on software prefetch. |
| Authority | Read-only. The skill runs performance tools that read the program and prints guidance; source and build changes land through the normal coding path. No remote mutation. |
| Side effect | None beyond tool output files that the profiling tools write to their own scratch locations. |
| Done | A measured before and after for the proposed layout change exists, with the cache counters quoted, or the diagnosis names the miss class and the evidence for it. |

## Inputs

- The program and a reproducible workload: required. A benchmark that runs the hot loop long enough to fill counters.
- The build: required, with symbols for attribution and without changing optimization between measurements.
- The target machine: required. Counters and latencies are properties of the specific microarchitecture.

## Procedure

1. Measure before changing anything. Take the generic counters first, then the level-specific ones. Done when: a baseline of counts and miss rates for the real workload is recorded.

```bash
perf stat -e cache-references,cache-misses,cycles,instructions ./bench
perf stat -e L1-dcache-loads,L1-dcache-load-misses,LLC-loads,LLC-load-misses ./bench
```

Interpret relatively. An L1 miss rate that is acceptable for a pointer-chasing graph walk is severe for a streaming numeric kernel; judge each rate against the same program on the same machine, before against after, and against the memory-bound share of the run.

2. Confirm false sharing before padding it. Threads writing to distinct variables that share one line invalidate each other constantly. The Intel HITM events count the modified-line hits that mark it. Done when: the counter either shows the traffic or rules the hypothesis out.

```bash
perf stat -e mem_load_l3_hit_retired.xsnp_hitm,machine_clears.memory_ordering ./bench
```

Event availability differs across vendors and generations; check `perf list` on the target machine and treat a missing event as unknown, not as zero.

3. Apply the layout fixes in order of cost, cheapest first. Done when: each applied fix has a re-measurement from step 1.

- Split hot from cold fields so the hot struct stays one line:

```c
struct record_hot { int id; int value; };            // touched every iteration
struct record_cold { char name[128]; char desc[256]; };
```

- Pad or align per-thread data to its own line:

```c
struct alignas(64) padded_counter {
    int value;
    // the padding keeps the next counter on another line
};
```

In C++17 prefer `std::hardware_destructive_interference_size` over the literal 64, and read the real line size at runtime with `sysconf(_SC_LEVEL1_DCACHE_LINESIZE)`. Sixty-four bytes is the common line on x86-64 and ARM server cores; some consumer and Apple cores use 128.

- Convert array-of-structs to struct-of-arrays when the loop touches few fields. Reading `x[i]` from an SoA stream loads only `x`, and the access vectorizes; the same loop over an AoS drags every field through every line.

4. Remove the pointer chasing that no prefetcher predicts. Linked structures miss once per node. Pool-allocate nodes, replace links with indices into an array, or sort the traversal order to match memory order. Done when: the hot loop's accesses are sequential or the chasing is provably off the critical path.

5. Add software prefetch only where the pattern defeats the hardware prefetcher, such as linked lists and irregular graphs. The hint is a hint; measure or remove it. Done when: a counter or timing improvement proves the hint earns its place.

```c
// issue the hint far enough ahead to cover latency,
// and early enough that the line is not evicted first
for (Node *n = head; n; n = n->next) {
    if (n->next) __builtin_prefetch(n->next, 0, 1);
    process(n);
}
```

`_mm_prefetch` on x86 takes a locality hint (`_MM_HINT_T0` for L1, `T1` for L2, `T2` for L3, `NTA` for non-temporal). Prefetch distance is a tunable per machine, not a constant.

6. Block the loop when the working set exceeds the cache. Choose the block size so one block of the working arrays fits the data cache, and tune the size on the target machine. Done when: the blocked version beats the naive one in step 1's counters.

```c
// process cache-sized tiles instead of whole rows
#define BLOCK 64   // tune on the target machine
for (int i = 0; i < N; i += BLOCK)
for (int k = 0; k < N; k += BLOCK)
for (int j = 0; j < N; j += BLOCK)
    for (int ii = i; ii < i + BLOCK && ii < N; ii++)
    for (int kk = k; kk < k + BLOCK && kk < N; kk++)
    for (int jj = j; jj < j + BLOCK && jj < N; jj++)
        C[ii*N+jj] += A[ii*N+kk] * B[kk*N+jj];
```

7. Verify allocation alignment when the transformation depends on it. `aligned_alloc(64, size)` and `posix_memalign` give line-aligned buffers; check struct layout with `pahole -C MyStruct ./prog`, and let `-Wpadded` report compiler-side padding. Done when: the layout the code assumes is the layout `pahole` prints.

## Failure and recovery

| Failure class | Behavior |
|---|---|
| Counters read zero | The event is unavailable or virtualization hides it. Run `perf list` on the target and pick supported events. |
| No change after a layout fix | The loop was not miss-bound. Re-profile for the real bottleneck before the next transform. |
| SoA made it slower | The loop needs whole records, so SoA splits them across lines. Keep AoS and split only the hot fields. |
| Prefetch hurt | The hint came too early and evicted useful lines, or the hardware prefetcher already covered it. Remove the hint. |
| Two threads still slow | The sharing moved. Re-run the HITM counters from step 2 on the new layout. |

## Output

A measurement report: baseline counters, each transformation applied, and the counters after it, with the winning layout named. Counter names, Cachegrind usage, and layout tooling are in `references/cache-counters.md`.

Attribution

OutlineDrivenOutlineDriven
View sourceSee grades on GitHubMore from OutlineDriven →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Terse caveman voice: answer first, fluff gone, every technical fact kept. Use for /caveman, "caveman mode", "talk like caveman", "be brief", "less tokens". Stays on until "stop caveman" or "normal mode".

1100021 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

698431 votes

Writing Skills

Create and manage Claude Code skills in HASH repository following Anthropic best practices. Use when creating new skills, modifying skill-rules.json, understanding trigger patterns, working with hooks, debugging skill activation, or implementing progressive disclosure. Covers skill structure, YAML frontmatter, trigger types (keywords, intent patterns), UserPromptSubmit hook, and the 500-line rule. Includes validation and debugging with SKILL_DEBUG. Examples include rust-error-stack, cargo-dep...

3931 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3421 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Amp, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Grok Build, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

741 votes
View all in ai-agents →