
Claude Skills by kekzl
github.com/kekzlUse when adding support for a new model architecture to imp, porting a model family, or debugging a model that loads but produces wrong output - "add support for <model>", "new arch", loader detection, chat template, tokenizer parity, RoPE variant, "outputs garbage", "prompt-blind", "digits scrambled", "NaN logits", "describes a different picture", "does it fit in VRAM". Do NOT use for kernel performance (sm120-cuda-expert) or quant-format questions (quant-formats).
Use when benchmarking, profiling, or A/B-testing CUDA kernels or end-to-end perf in the imp inference engine on RTX 5090 (sm_120), including refreshing tests/perf_baseline.json or publishing numbers to docs/BENCHMARKS.md and the README. Triggers on "benchmark kernel", "profile cuda", "ncu", "nsys", "kernel timing", "kernel sum", "occupancy", "bandwidth bound", "compute bound", "roofline", "perf baseline", "is this regression real", "decode dropped", "aggregate throughput", "two-image A/B", "p...
Use when building imp, running its test suite, checking CI status, or debugging build/test failures - "make build", "run the tests", "test-gpu", "verify-fast", GTEST_FILTER, Docker/CUDA toolchain, dependency bumps, "CI is red/blocked", "which gate failed", stale objects / segfault after a header edit, hook edits, docker-entrypoint env vars, determinism or perplexity checks. Do NOT use for benchmarking/profiling (benchmark-cuda) or output-quality batteries (check-degeneration).
Use when verifying that a model in the imp inference engine produces coherent output without repetition loops, token-stuck states, or state corruption across turns or streams. Triggers on "degenerates", "check degeneration", "repetition loop", "own own own", "stuck token", "empty content", "multi-turn regression", "does it still work", "NIAH", and after enabling CUDA graphs / changing forward pass / MoE routing / KV cache or KV dtype / GDN state or scan / sparse attention / PDL / speculation ...
Use when a question is about *structure* rather than text - who calls or launches a symbol, where it is defined, what a change would reach, whether something is dead, how a request gets from the API to a kernel. Triggers on "who calls", "who launches this kernel", "where is X defined", "what breaks if I change", "is this still used", "is this dead", "trace the path from X to Y", "blast radius", "what depends on this header". Do NOT use for free-text search (`rg` is better and cheaper) or to o...
Use when auditing the imp codebase for structural debt, dead code, god-objects/files, duplication, flag sprawl, or deciding whether a cleanup is worth shipping - "structure audit", "tech debt", "is this still used", "dead code", "refactor for clarity", "should we split this file", "remove this flag", "file size gate red", "alloc sites gate red". Do NOT use for build/test mechanics (building-and-testing), perf/kernel work (benchmark-cuda / sm120-cuda-expert), or output quality (check-degenerat...
Use when writing, moving or auditing any .md in imp - deciding which file a paragraph belongs in, adding a doc, fixing a stale claim, or when docs_lint.py fails in CI. Covers the four reader layers (L0 README / L1 operators / L2 kernel devs / L3 agents), the HTML-comment metadata header, [PROV:] provenance, the single-source-of-truth map, generated perf blocks, which numbers may appear in prose, plan-doc closure. Triggers on "which doc does this go in", "docs lint failed", "add a doc", "this ...
Use when keeping imp's docs and config examples coherent after a change - updating docs/internals/ARCHITECTURE.md / README / docs/GOAL.md / docs/MODELS.md / docs/roadmap.md / imp.conf.example / CHANGELOG / MISSION_JOURNAL, or "is this doc stale", "document this change", "the example config is out of date", "the README says X but the code does Y", "add a roadmap ledger row". Do NOT use for layer/frontmatter/lint questions (docs-layers), structural code audits (codebase-audit), measuring perf o...
Use when asking whether something in imp is actually finished - hunting stubs, placeholders, request fields that are parsed and then ignored, code paths or kernels that never run, gate-based features that silently no-op, tests that assert nothing. Triggers on "is this implemented", "unfinished", "stub", "placeholder", "dead path", "does this flag do anything", "accepted but ignored", "kernel never launched", "test asserts nothing", "neutral A/B on a gated feature". Do NOT use for structural d...
Use when imp outputs differ between two paths that should agree - chunk sizes, prefill vs decode, prefix-cache resend, two engines (imp vs llama.cpp vs HF), two builds or kernels - "output changes with chunk size", "first token differs", "logprobs moved", "which layer diverges", "who is right, imp or llama.cpp", "near-tie". Do NOT use for repetition loops or garbage output (check-degeneration), a new arch that loads wrong (add-model-arch), or perf (benchmark-cuda).
Use when working on imp's quantization formats, loaders, or dequant paths - GGUF Q4_0...Q8_0/Q*_K/IQ4/MXFP4, NVFP4 two-level scaling (Modelopt vs compressed-tensors layouts), FP8 E4M3, StorageTier, decode cache, KV-cache dtypes (auto/fp16/fp8/int8/int4/nvfp4/mxfp4), "which quant should I use", "which KV dtype", scale-factor layout, dequant kernel wiring, imp-quantize / AWQ, judging quant quality (PPL). Do NOT use for writing/optimizing GEMM/GEMV kernels (sm120-cuda-expert) or measuring quant ...
Use when working on or testing imp-server and its HTTP APIs - OpenAI/Anthropic endpoints, /v1/chat/completions, /v1/messages, /v1/responses, SSE streaming, tool calling, json_schema/regex/GBNF constrained decoding, thinking/reasoning_content, cache_control/prefix cache, model loading and swapping, request priority, X-Request-Id, OTLP tracing / traceparent / Jaeger, /metrics, container env (IMP_SET), "server returns 404/garbage", API compliance. Do NOT use for CLI-only inference (imp-cli), ker...
Use when opening, merging, or releasing a PR for imp - branching off main, `gh pr create`, enabling auto-merge, writing the PR body or a CHANGELOG entry, cutting a tagged release (version bump + CHANGELOG + tag + GitHub release). Symptoms - "PR stuck BLOCKED", "my last commit didn't land on main", "which check is required", "how do I cut a release", "auto-merge merged too early", "check-release failed", "STALE.md blocks git pull". Do NOT use for build/test mechanics (building-and-testing) or ...
Use when writing, reviewing, or optimizing CUDA kernels targeting sm_120a (RTX 5090 / Consumer Blackwell, GB202) in the imp inference engine. Triggers on CUDA/PTX kernel code, shared-memory layout, bank conflicts, tensor-core MMA (mxf4nvf4, tf32, 3xTF32/3xFP16), GEMV/GEMM/attention/quantization/GDN-scan kernels, decode tok/s under expected, kernel emitting HMMA instead of mxf4nvf4, occupancy, register pressure, __launch_bounds__, spills, PDL. Pair with `benchmark-cuda` for measurement and `ch...