Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Quantization

ASecurity

Quantize models for efficient inference — precision formats, calibration, accuracy validation, and hardware considerations. Use when models are too big, slow, or expensive to serve at full precision.

2 stars
0 votes
0 copies
0 views
Added 9/29/2026
ai-agentsgo

Security Analysis

A100/100

Scanned 9/29/2026

$npx -y skills add aicodedecode/awesome-muse-skills --skill quantization --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Quantization?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Quantization
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/aicodedecode-quantization/badge)](https://www.skillsdirectory.com/skills/aicodedecode-quantization)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: quantization
description: Quantize models for efficient inference — precision formats, calibration, accuracy validation, and hardware considerations. Use when models are too big, slow, or expensive to serve at full precision.
category: ai-research
---

# Model Quantization

Quantization shrinks models by representing weights (and activations) in fewer bits — 16-bit to 
8-bit to 4-bit — cutting memory, speeding inference, and lowering cost. Modern methods preserve 
quality remarkably well, but the details matter.

## Overview

The trade: fewer bits per weight means smaller, faster models with some precision loss. 
Post-training quantization (PTQ) converts an existing model without retraining — fast and usually 
sufficient. Quantization-aware training (QAT) trains with quantization simulated — better 
quality, more effort. Weight-only quantization shrinks memory (the common bottleneck); 
weight+activation quantization also speeds compute. Calibration data tunes the quantization ranges 
— use representative data.

## When to use

- Serving large models on limited GPUs or CPUs.
- Reducing inference cost and latency in production.
- Deploying to edge devices or laptops.
- Anytime memory bandwidth (not compute) is the bottleneck — which is most LLM serving.

## Core concepts

- **Precision formats**: FP16/BF16 (near-lossless baseline), INT8 (2× smaller, minimal loss with 
care), INT4 (4× smaller, needs good methods), mixed precision (sensitive layers kept higher).
- **PTQ vs QAT**: post-training is fast and works well down to 8-bit and often 4-bit for weights; 
QAT recovers quality at aggressive precisions but requires training.
- **Weight-only vs weight+activation**: quantizing weights cuts memory (usually the binding 
constraint); quantizing activations too speeds compute but is more quality-sensitive.
- **Calibration**: a small representative dataset sets quantization ranges. Bad calibration data 
→ bad quantization. Match your deployment distribution.
- **Granularity**: per-tensor, per-channel, or per-group quantization. Finer granularity preserves 
quality at low bit-widths; group-wise is the modern sweet spot.
- **Quality validation**: perplexity plus downstream task evals. Some capabilities (long-context 
reasoning, rare tokens) degrade before average metrics show it.

## Practical workflow

1. Profile the bottleneck: memory capacity, memory bandwidth, or compute? Quantize accordingly.
2. Start with weight-only 8-bit or 4-bit PTQ — the best effort-to-gain ratio.
3. Calibrate on representative data from your actual workload, not generic text.
4. Validate quality: perplexity + task evals + targeted tests (long context, math, rare knowledge).
5. If quality drops unacceptably: try finer granularity, keep sensitive layers at higher precision, 
or step up one bit-width.
6. Benchmark end-to-end: tokens/sec, latency, and memory on your hardware — theoretical gains 
don't always materialize.

```text
Quantization decision path:
Memory-bound serving? → weight-only INT4/INT8 PTQ
Need max quality?     → INT8, or mixed precision
Quality dropped?      → finer granularity → higher bits for sensitive layers → QAT
Always: calibrate on YOUR data, validate on YOUR tasks
```

## Common pitfalls

- **Generic calibration data**: calibrating on web text for a code assistant. Match the deployment 
distribution.
- **Average-metric validation**: perplexity looks fine while long-context reasoning silently 
degrades. Test capabilities, not just averages.
- **Ignoring the hardware**: a format your GPU doesn't accelerate may run slower than FP16. Match 
format to hardware support.
- **Over-quantizing**: pushing to 4-bit or below everywhere for marginal gains while quality 
bleeds. Step down gradually.
- **No baseline comparison**: quantizing without measuring the FP16 original on your evals. You 
need the delta.
- **Activation quantization by default**: it's the quality-sensitive one. Start weight-only; add 
activation quantization deliberately.

Attribution

aicodedecodeaicodedecode
View sourceSee grades on GitHubMore from aicodedecode →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Terse caveman voice: answer first, fluff gone, every technical fact kept. Use for /caveman, "caveman mode", "talk like caveman", "be brief", "less tokens". Stays on until "stop caveman" or "normal mode".

1100021 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

698621 votes

Writing Skills

Create and manage Claude Code skills in HASH repository following Anthropic best practices. Use when creating new skills, modifying skill-rules.json, understanding trigger patterns, working with hooks, debugging skill activation, or implementing progressive disclosure. Covers skill structure, YAML frontmatter, trigger types (keywords, intent patterns), UserPromptSubmit hook, and the 500-line rule. Includes validation and debugging with SKILL_DEBUG. Examples include rust-error-stack, cargo-dep...

3931 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3421 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Amp, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Grok Build, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

741 votes
View all in ai-agents →