Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsCommunityBlog
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Add Model Arch

ASecurity

Use when adding support for a new model architecture to imp, porting a model family, or debugging a model that loads but produces wrong output - "add support for <model>", "new arch", loader detection, chat template, tokenizer parity, RoPE variant, "outputs garbage", "prompt-blind", "digits scrambled", "NaN logits", "describes a different picture", "does it fit in VRAM". Do NOT use for kernel performance (sm120-cuda-expert) or quant-format questions (quant-formats).

43 stars
0 votes
0 copies
0 views
Added 9/28/2026
ai-agentsrustgodebugginggitapiperformance

Works with

cliapi

Security Analysis

A100/100

Scanned 9/28/2026

Install to Claude Code

$npx -y skills add kekzl/imp --skill add-model-arch --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Add Model Arch?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Add Model Arch
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/kekzl-add-model-arch/badge)](https://www.skillsdirectory.com/skills/kekzl-add-model-arch)

More formats (shields.io, HTML) on the badges page.

Files
SKILL.md
---
name: add-model-arch
description: Use when adding support for a new model architecture to imp, porting a model family, or debugging a model that loads but produces wrong output - "add support for <model>", "new arch", loader detection, chat template, tokenizer parity, RoPE variant, "outputs garbage", "prompt-blind", "digits scrambled", "NaN logits", "describes a different picture", "does it fit in VRAM". Do NOT use for kernel performance (sm120-cuda-expert) or quant-format questions (quant-formats).
---

# Adding a Model Architecture - imp

## First: is it a new arch at all?

Diff `config.json` against a supported sibling before estimating. Qwen3.8 shipped with zero enum members (loads as `QWEN35`); the work was tokenizer parity, template goldens, KV dtype default and MTP head (#1750). HF reference values come from `curl`, never from memory.

## Integration checklist (gpt-oss #572 is the reference PR)

| Step | Where | Notes |
|---|---|---|
| 1. Enum + registry | `ModelArch` in `src/model/model_arch.h`; `parse_model_arch` (GGUF `general.architecture` in `gguf_loader.cpp`, HF `architectures`/`model_type` in `hf_config_loader.cpp`), `model_arch_name`, `apply_arch_defaults`, sampling defaults in `src/model/model.cpp` | then `ModelProfile` (`src/model/model_profile.h/.cpp`, SSoT since #622/#623; `AttnVariant { STANDARD, GEMMA4_SWA, GPTOSS_SWA, NOPE, MLA }`). No new `cfg.arch == X` in hot paths. Add the arch to the KV-dtype lists `kv_nvfp4_default_safe`, `kv_fp8_hint_default_safe`, `kv_fp8_no_hint_default_safe` (evidence per family in `model.cpp`): a missing entry silently gets FP16 KV, which on a GDN hybrid gates context |
| 2. Loader | `src/model/tensor_kind_matcher.cpp`, `weight_map.cpp`; SafeTensors `safetensors_loader.cpp`; NVFP4 via `llm_compressor_loader.cpp` (compressed-tensors) or Modelopt (`hf_quant_config.json`) | the two NVFP4 layouts have RECIPROCAL tensor scales (quant-formats) |
| 3. Arch config | `model_config.h` + `apply_arch_defaults` | RoPE pair layout (NeoX vs GPT-J), YaRN/`rope_freq_scale`, SWA layer pattern, NoPE, sinks, softcap, norm placement, MoE router |
| 4. Chat template | family in `src/model/chat_template_families.cpp` (`ChatTemplateFamily` in `chat_template.h`), rendering in `chat_template.cpp` (+ `jinja.cpp`) | a new family needs a golden pin (`make chat-goldens`, `tests/refs/chat_template_goldens.h`, nine families since #1721, three more Jinja gaps fixed in #1701; Jinja fails SILENTLY). `reasoning_effort` must reach `ChatTemplate::apply*`/`render_jinja` + the server snapshot: identical prompt-token counts across efforts = not threaded (#1750: 67/67 before, 41/11/53 after) |
| 5. Tokenizer parity | template `tests/test_tokenizer_qwen38.cpp` (32/32 encode+decode vs HF), `tests/test_qwen38_chat_template.cpp`; harness `tools/tokenizer_parity/run.sh` (from #2159, branch `fix/tokenizer-hf-parity` until merged) | BERT-family GGUFs use SPM, not WordPiece. An HF reference built from a GGUF tokenizer can split special tokens (`<think>` -> 6 text ids instead of 151667): compare token ids before logits |
| 6. Kernels | only for genuinely new ops; check `src/exec/` + `src/compute/` first | new RoPE variants go into `src/compute/rope_yarn.cuh` (shared with MTP heads, #913) |
| 7. Verify | loads -> greedy coherent (check-degeneration) -> `imp-cli --perplexity` vs HF (within ~10-20%, often closer: gpt-oss 4.68 vs bf16 4.607, #663) -> decode/prefill sanity (benchmark-cuda) | PPL with `runtime.deterministic=true`, `speculative.mtp_k=0`, `ppl_corpus_45k.txt` (`tools/analysis/make_ppl_corpus.sh`) |
| 8. Docs | `docs/MODELS.md` row (+ `docs/BENCHMARKS.md` if hero); known gaps to `docs/LIMITATIONS.md`; perf baseline entry if gated | |

A new checkpoint is UNTRUSTED INPUT: SafeTensors/`tokenizer.json` parsers are hardened (#1660, #1694); fuzz targets under `fuzz/`; no parsing shortcuts around the bounds checks.

## Wrong-output fingerprints

| Symptom | Root-cause class | Case |
|---|---|---|
| Fluent but ignores the prompt ("prompt-blind") | RoPE pair layout: HF SafeTensors need `rope_neox=true`, GGUF pre-permutes Q/K | SafeTensors Llama/Mistral, #503 |
| Words fine, digits scrambled | position encoding (NoPE layer as RoPE or vice versa) | Nemotron-H `rope_attn_disabled`, #518 |
| Argmax always token 0 | NaN logits upstream (residual overflow, bad scale) | gpt-oss FP16 residual |
| Coherent to ~1k ctx, then garbage | YaRN/`rope_freq_scale` inverted or fused-rope path without YaRN | gpt-oss 1024x error, #572 |
| Wrong only with chunked prefill at long ctx | continuation-chunk path | #553 |
| Wrong language / valid-but-wrong tokens | weight upload / dequant layout (MoE: `weight_upload.cu` expert promotion first) | Qwen3.6-35B NVFP4, #925 |
| Garbage from token 0 (`!!!`) | silent VRAM-alloc failure in a decode fallback | MXFP4 GDN hybrids, #935 |
| Multimodal: describes a DIFFERENT picture | M-RoPE per-token (t,h,w) layout, `src/model/mrope_positions.cpp` | Qwen3-VL |
| Vision fluent but generic | tower loaded partly or embeddings never reach the sequence: `tools/analysis/vision_sight_check.py` | |
| Drift only at very long positions | YaRN float trap: `__sinf/__cosf` on an unreduced argument; long-ctx tests run `ext_factor=0` and cannot see it | #1704 |
| Coherent-ish, drifts vs HF | MLA/YaRN `rope_mscale` on the wrong base (1.261x); pin the transformers oracle (4.44.2 for MLA) | #880 |
| Draft accept collapses, output correct | MTP head RoPE differs from the target (NeoX vs YaRN) | #913 |
| CLI fine, server broken | not arch: server-api | |
| Which block diverges | `diagnostics.dump_hidden_dir` + `tools/analysis/layer_diff.py` (vs llama.cpp `llama-eval-callback`) or `layer_ab_diff.py` (two imp runs); `dump_gdn_state_dir`, `dump_logits_dir` | 0 non-finite GDN states over a 46579-token prefill on Qwen3.8 |

## Known traps

- `rope_neox`: GGUF converters pre-permute Q/K; HF SafeTensors do not.
- `swa_layers` was Gemma-only hardcoded once; verify per-layer attention type on any interleaved-SWA arch.
- The fused rope+KV-write kernel must apply the same YaRN scaling as the standalone path.
- Arch control tokens (Harmony channels) must not land on the banned list; the spec-verify argmax applies the mask since #1796.
- Gemma-4: per-layer `rope_freqs` for non-SWA layers, `n_rot=hd`.
- GDN/hybrid state: BF16 storage + FP32 arithmetic is the default (`gdn.state_bf16`, #1776/#1778); FP16 state NaNs at depth; the old "must be FP32" was a layout bug.
- Multimodal checkpoints wrap the LM under `model.language_model.*` (Qwen3.5-VL, #647): strip in the loader or every tensor "is missing".
- Encoder/embedding archs are supported (nomic-bert #867, cosine 0.999 vs HF).
- VRAM arithmetic: card 32 607 MiB; CUDA primary context ~1680 MiB; library reserve is a per-model MEASURED value cached in `src/memory/library_reserve_cache.h` (#1119; `vram.library_reserve_cache`, override `vram.library_reserve_mb`); the `kMeasuredLibraryReserveBytes` ~3900 MiB constant in `src/memory/plan.h` is the first-run fallback (measured 0 MiB on Qwen3-4B-IQ4_NL, 7460 on Qwen3-8B-Q8_0). An MTP head adds ~0.79 GiB when `speculative.mtp_k` auto takes it. Peaks per config: `docs/internals/MEMORY.md`.
- Getting a checkpoint onto the NVFP4 path: `scripts/stage-model.sh` (download + `imp-quantize` in one command; imp fetches nothing itself).

## After it works

`make verify-fast`; a `DegenerationTest`-compatible probe prompt; think-channel archs through `tools/analysis/degen_suite.py` (think-leak); vision through `vision_sight_check.py` and `make test-vision`.

Attribution

kekzlkekzl
View sourceMore from kekzl →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Ultra-compressed communication mode that cuts output tokens while keeping technical accuracy. Levels: lite, full, ultra and the wenyan variants. Use for /caveman, "caveman mode", "talk like caveman", "be brief" or "less tokens".

1074701 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

695601 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3351 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

691 votes

math-skill

A comprehensive mathematical reasoning skill for AI assistants — handles arithmetic to research-level problems with rigorous step-by-step reasoning, systematic verification, and transparent uncertainty handling

381 votes
View all in ai-agents →