Skip to content
Back to skills

Llamacpp Guide

ASecurity

Run and optimize models with llama.cpp — GGUF quantization, CPU/GPU offloading, server mode, and performance tuning. Use when you need maximum local inference efficiency or cross-platform deployment.

  • 2 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 29, 2026
ai-agentsgoc++apiperformance

Works with

  • cli
  • api

Security analysis

A100/100

Scanned September 29, 2026

npx -y skills add aicodedecode/awesome-muse-skills --skill llamacpp-guide --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Llamacpp Guide?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Llamacpp Guide
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/aicodedecode-llamacpp-guide/badge)](https://www.skillsdirectory.com/skills/aicodedecode-llamacpp-guide)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: llamacpp-guide
description: Run and optimize models with llama.cpp — GGUF quantization, CPU/GPU offloading, server mode, and performance tuning. Use when you need maximum local inference efficiency or cross-platform deployment.
category: ai-research
---

# llama.cpp Guide

llama.cpp is the portable inference engine: C++-based, runs everywhere from phones to servers, with 
the best CPU inference in the business and broad GGUF model support. It's the power tool behind 
most serious local LLM setups.

## Overview

The mental model: models in GGUF format (quantized variants from 2-bit to 8-bit), an engine that 
splits work between CPU and GPU layers, and two interfaces — a CLI for direct use and a server 
mode exposing an OpenAI-compatible API. Performance comes from tuning thread counts, batch sizes, 
and GPU offloading to your specific hardware. It's lower-level than packaged runners, which means 
more control and more knobs.

## When to use

- Maximum local performance: squeezing the most tokens/sec from your hardware.
- CPU-only inference where other engines struggle.
- Cross-platform deployment: same engine on Mac, Linux, Windows, even mobile.
- Fine-grained control: custom quantization, speculative decoding, grammar-constrained output.

## Core concepts

- **GGUF format**: the model container — architecture, weights, and metadata in one file, with 
quantization variants (Q4_K_M is the popular sweet spot). Choose the largest quant that fits your 
RAM.
- **GPU offloading**: layers assigned to GPU (`-ngl`) vs. CPU. Full offload when VRAM fits; partial 
offload still beats CPU-only. More offloaded layers = faster.
- **Threading and batching**: CPU thread count matched to your cores (physical, not hyperthreaded, 
usually); batch size tuned for prompt processing vs. generation.
- **Server mode**: HTTP server with OpenAI-compatible endpoints — the integration point for apps 
and agents. Supports concurrent requests with continuous batching.
- **Context management**: context size set at load (`-c`); KV cache quantization to stretch long 
contexts into limited RAM.
- **Advanced features**: speculative decoding (draft model proposes, main model verifies — big 
speedups), grammar-constrained generation (JSON mode via grammars), multimodal support.

## Practical workflow

1. Pick the GGUF quant: largest that fits comfortably in RAM (leave headroom for context/KV cache).
2. Benchmark your hardware: try full GPU offload first; fall back to partial; tune thread count 
(start at physical core count).
3. Set context size for your use case; enable KV cache quantization if context is the constraint.
4. Run server mode for app integration; test concurrent requests and streaming.
5. Tune batch parameters for your workload: prompt-heavy (bigger batches) vs. generation-heavy 
(smaller).
6. Consider speculative decoding with a small draft model for interactive latency wins.

```text
Tuning order (biggest wins first):
1. GPU offload: max layers that fit VRAM
2. Quant choice: biggest that fits RAM
3. Threads: physical core count (CPU-bound parts)
4. Speculative decoding: small draft model
5. Batch/context knobs for your workload mix
```

## Common pitfalls

- **Wrong thread count**: too many threads on hyperthreaded cores can slow things down. Start with 
physical cores.
- **Quant too aggressive**: Q2 fits but quality suffers. Q4_K_M is the default for a reason — 
deviate deliberately.
- **Ignoring KV cache**: long contexts eat RAM beyond the model weights. Budget for it or quantize 
the cache.
- **No GPU offload configured**: running full CPU when a GPU sits idle. Check offload actually 
engaged.
- **Server defaults for production**: default concurrency and timeouts are demo settings. Tune for 
your load.
- **Model-architecture mismatch**: not every GGUF file works with every build. Match the engine 
version to the model format.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…