Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Llama Cpp

BSecurity

llama.cpp runs large language and vision models locally from GGUF files on CPUs, Apple Silicon and GPUs, and serves them over an OpenAI-compatible HTTP API. Use when a user asks to run a model locally with llama.cpp, start llama-server or "llama serve", download a GGUF model from Hugging Face, host a local OpenAI-compatible endpoint, offload layers to the GPU, get embeddings or schema-constrained JSON from a local model, quantize a model to Q4_K_M, or benchmark tokens per second.

155 stars
0 votes
0 copies
0 views
Added 10/4/2026
ai-agentspythonrustgoc++shellbashdockergitapi

Works with

terminalcliapimcp

Security Analysis

B80/100
mediumUses curl or wget to download content
criticalExfiltrates credentials via HTTP — exact pattern from Snyk ToxicSkills study
mediumInstalls packages at runtime which could introduce malicious dependencies

Pro scans all 2 files and shows the line behind each finding

Scanned 10/4/2026

$npx -y skills add TerminalSkills/skills --skill llama-cpp --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Llama Cpp?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Llama Cpp
[![Security: B — Skills Directory](https://www.skillsdirectory.com/api/skills/terminalskills-llama-cpp/badge)](https://www.skillsdirectory.com/skills/terminalskills-llama-cpp)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: llama-cpp
description: >-
  llama.cpp runs large language and vision models locally from GGUF files on
  CPUs, Apple Silicon and GPUs, and serves them over an OpenAI-compatible HTTP
  API. Use when a user asks to run a model locally with llama.cpp, start
  llama-server or "llama serve", download a GGUF model from Hugging Face, host
  a local OpenAI-compatible endpoint, offload layers to the GPU, get embeddings
  or schema-constrained JSON from a local model, quantize a model to Q4_K_M, or
  benchmark tokens per second.
license: Apache-2.0
compatibility: "macOS, Linux, Windows or FreeBSD on x86_64 or arm64. Source builds need CMake and a C++17 compiler. GPU offload needs a Metal, CUDA, ROCm or Vulkan device."
metadata:
  author: terminal-skills
  version: "1.0.0"
  category: data-ai
  tags: ["llama-cpp", "gguf", "local-llm", "inference-server", "openai-compatible"]
  repository: https://github.com/ggml-org/llama.cpp
---
# llama.cpp — Local LLM inference from GGUF files

## Overview

llama.cpp is a C/C++ inference engine for models stored in the GGUF format. Current builds ship one `llama` binary with subcommands: `llama serve` (HTTP server with OpenAI- and Anthropic-compatible routes plus a web UI), `llama cli` (terminal chat) and `llama download`. Docker images and older installs expose the same programs as `llama-server` and `llama-cli`; the flags are identical, so every `llama serve ...` line below also works as `llama-server ...`.

## Instructions

### Installation

```bash
# Package managers
brew install llama.cpp                   # macOS and Linux
winget install llama.cpp                 # Windows
conda install -c conda-forge llama.cpp   # CUDA, Vulkan and Metal builds

# Official installer from llama.app: probes CUDA, ROCm, Vulkan, Metal or CPU
# features and copies a matching prebuilt binary to ~/.local/bin/llama.
# Download it, read it, then run it — never pipe it straight into a shell.
curl -LsSf https://llama.app/install.sh -o llama-install.sh
sh llama-install.sh

llama version
llama help all        # also lists quantize, bench, perplexity, completion
```

Build from source when no prebuilt binary matches the machine:

```bash
git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp
cmake -B build -DGGML_CUDA=ON        # drop the flag for CPU; Vulkan: -DGGML_VULKAN=ON
cmake --build build --config Release -j 8
ls build/bin/                        # llama, llama-server, llama-cli, llama-quantize, llama-bench
```

Metal is enabled by default on macOS. Docker images are published as `ghcr.io/ggml-org/llama.cpp:server`, `:light` and `:full`, each with `-cuda`, `-rocm`, `-vulkan` and `-intel` variants.

### Get a model

```bash
# Download into the local cache and print the file path
llama download -hf bartowski/Qwen2.5-7B-Instruct-GGUF:Q4_K_M

# Pick an exact file instead of a quantization tag
llama download -hf ggml-org/Qwen3.5-0.8B-GGUF -hff Qwen3.5-0.8B-Q8_0.gguf

llama cli --cache-list
```

`-hf user/repo[:quant]` defaults to `Q4_K_M` and falls back to the first file in the repository. Gated repositories read the token from `HF_TOKEN`. `LLAMA_CACHE` moves the cache directory, and `--offline` forbids network access.

### Run a prompt from the terminal

```bash
# Interactive chat
llama cli -hf bartowski/Qwen2.5-7B-Instruct-GGUF:Q4_K_M

# One answer, then exit: the form to use from scripts and agents
llama cli -hf bartowski/Qwen2.5-7B-Instruct-GGUF:Q4_K_M \
  -sys "You are a release-notes writer." \
  -p "Summarize: fixed login timeout, added CSV export, removed legacy v1 API." \
  -st -n 200 --temp 0.2 --no-display-prompt

# Vision model: the multimodal projector is downloaded together with the model
llama cli -hf ggml-org/gemma-3-4b-it-GGUF --image receipt-2026-09-14.jpg \
  -p "List the merchant, date and total." -st
```

### Serve an OpenAI-compatible API

```bash
export LLAMA_API_KEY="$(openssl rand -hex 24)"    # read by the server, same as --api-key

llama serve -hf bartowski/Qwen2.5-7B-Instruct-GGUF:Q4_K_M \
  --alias qwen2.5-7b-instruct \
  -c 16384 -ngl all -np 4 \
  --host 127.0.0.1 --port 8080 --metrics
```

| Flag | Meaning |
|---|---|
| `-m FILE` / `-hf REPO` | Local GGUF file, or Hugging Face repository to fetch |
| `--alias NAME` | Model id returned by `/v1/models`; default is the file path |
| `-c N` | Context size in tokens; `0` takes the value stored in the model |
| `-ngl N` | Layers kept in VRAM: a number, `auto` (default) or `all` |
| `-np N` | Parallel request slots; `-1` (default) selects automatically |
| `--host`, `--port` | Defaults are `127.0.0.1` and `8080` |
| `--api-key KEY` | Require `Authorization: Bearer KEY`; also `--api-key-file` |
| `--metrics` | Enable the Prometheus endpoint `/metrics` |
| `--sleep-idle-seconds N` | Unload the model after N idle seconds, reload on demand |

Every flag has an environment equivalent such as `LLAMA_ARG_CTX_SIZE` or `LLAMA_ARG_N_GPU_LAYERS`; a command-line flag overrides the variable. Wait for readiness before sending work:

```bash
curl -s http://127.0.0.1:8080/health
# {"status":"ok"}   (HTTP 503 with "Loading model" until the weights are loaded)
curl -s http://127.0.0.1:8080/v1/models -H "Authorization: Bearer $LLAMA_API_KEY"
```

### Call the server

```bash
curl -s http://127.0.0.1:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $LLAMA_API_KEY" \
  -d '{
    "model": "qwen2.5-7b-instruct",
    "messages": [
      {"role": "system", "content": "Answer in one sentence."},
      {"role": "user", "content": "What does HTTP status 429 mean?"}
    ],
    "temperature": 0.2,
    "max_tokens": 120
  }'
```

```python
import os
from openai import OpenAI

client = OpenAI(
    base_url="http://127.0.0.1:8080/v1",
    api_key=os.environ.get("LLAMA_API_KEY", "sk-no-key-required"),
)

# Schema-constrained output: the sampler can only emit JSON matching the schema
reply = client.chat.completions.create(
    model="qwen2.5-7b-instruct",
    messages=[{"role": "user", "content": "Invoice INV-2041 from Nordwind GmbH, due 2026-10-15, total 1,284.50 EUR"}],
    response_format={
        "type": "json_schema",
        "schema": {
            "type": "object",
            "properties": {
                "invoice_number": {"type": "string"},
                "vendor": {"type": "string"},
                "due_date": {"type": "string"},
                "total": {"type": "number"},
            },
            "required": ["invoice_number", "vendor", "due_date", "total"],
        },
    },
)
print(reply.choices[0].message.content)
```

Other routes on the same port: `/v1/completions`, `/v1/responses`, `/v1/embeddings`, the Anthropic-compatible `/v1/messages`, plus native `/completion`, `/tokenize`, `/infill`, `/props` and `/slots`. Function calling uses the standard `tools` array and depends on the model's chat template; pass `"parallel_tool_calls": true` to allow several calls per turn.

### Embeddings

```bash
llama serve --embd-gemma-default --port 8081    # downloads EmbeddingGemma on first start

curl -s http://127.0.0.1:8081/v1/embeddings \
  -H "Content-Type: application/json" \
  -d '{"input": ["refund policy", "shipping times"], "encoding_format": "float"}'
```

Any dedicated embedding GGUF works with `--embeddings`; add `--pooling mean` if the model file defines no pooling type.

### Serve several models (router mode)

```bash
# No -m and no -hf: the server becomes a router over every model in the cache
llama serve --models-max 2 --port 8080

curl -s http://127.0.0.1:8080/models          # ids and load status of each model
curl -s -X POST http://127.0.0.1:8080/models/load \
  -H "Content-Type: application/json" \
  -d '{"model": "bartowski/Qwen2.5-7B-Instruct-GGUF:Q4_K_M"}'
```

Requests are routed by the `"model"` field of the JSON body, and a model is loaded on first use unless `--no-models-autoload` is set. `--models-dir DIR` routes over a directory of GGUF files instead of the cache, and `--models-preset FILE.ini` sets per-model flags. Take the ids from `GET /models` rather than guessing them.

### Convert, quantize and benchmark

```bash
# From a source checkout
python3 -m pip install -r requirements.txt
python convert_hf_to_gguf.py --outfile gemma-4-E2B-it-bf16.gguf --outtype bf16 --remote google/gemma-4-E2B-it
./build/bin/llama-quantize gemma-4-E2B-it-bf16.gguf gemma-4-E2B-it-Q4_K_M.gguf Q4_K_M

# Compare CPU-only against full GPU offload, 128 generated tokens
./build/bin/llama-bench -m gemma-4-E2B-it-Q4_K_M.gguf -p 512 -n 128 -ngl 0,99 -o md
```

## Examples

### Example 1: Local endpoint for an existing OpenAI client

**Request:** "Serve Qwen 2.5 7B on this workstation so our scripts can call it like the OpenAI API."

```bash
export LLAMA_API_KEY="$(openssl rand -hex 24)"
llama serve -hf bartowski/Qwen2.5-7B-Instruct-GGUF:Q4_K_M \
  --alias qwen2.5-7b-instruct -c 8192 -ngl all --port 8080 &

until curl -sf http://127.0.0.1:8080/health > /dev/null; do sleep 2; done
curl -s http://127.0.0.1:8080/v1/chat/completions \
  -H "Content-Type: application/json" -H "Authorization: Bearer $LLAMA_API_KEY" \
  -d '{"model":"qwen2.5-7b-instruct","messages":[{"role":"user","content":"Reply with the word ready."}]}'
```

**Result:** a standard chat completion object with llama.cpp's extra `timings` block:

```json
{
  "choices": [{"index": 0, "finish_reason": "stop",
               "message": {"role": "assistant", "content": "ready"}}],
  "model": "qwen2.5-7b-instruct",
  "usage": {"prompt_tokens": 35, "completion_tokens": 2, "total_tokens": 37},
  "timings": {"prompt_per_second": 412.6, "predicted_per_second": 58.3}
}
```

Existing clients only need `base_url="http://127.0.0.1:8080/v1"` and the key.

### Example 2: One-shot structured extraction without a server

**Request:** "Pull the ticket id, product and severity out of this support email as JSON, fully offline."

```bash
llama cli -m "$HOME/models/Qwen2.5-7B-Instruct-Q4_K_M.gguf" --offline \
  -f support-email-48213.txt -st --temp 0 --no-display-prompt \
  -j '{"type":"object","properties":{"ticket_id":{"type":"string"},"product":{"type":"string"},"severity":{"type":"string","enum":["low","medium","high"]}},"required":["ticket_id","product","severity"]}'
```

**Result:**

```json
{"ticket_id": "48213", "product": "Billing Portal", "severity": "high"}
```

## Guidelines

- **Memory:** a model needs roughly its file size plus the context cache. When loading fails or swaps, lower `-c`, choose a smaller quantization (`Q4_K_M` instead of `Q8_0`) or reduce `-ngl`. `--fit` is on by default and shrinks unset parameters to fit device memory.
- **Context and slots:** all `-np` slots draw from the context budget set by `-c`. Raise `-c` together with `-np`, or long prompts stop fitting.
- **Model id:** without `--alias`, the id is the path passed to `-m`. Set an alias so clients are not tied to a file location.
- **Network exposure:** the server binds to `127.0.0.1`. Before using `--host 0.0.0.0`, set `--api-key` and put a TLS reverse proxy in front. `/health` stays public even with a key.
- **Tool flags are dangerous:** `--tools`, `--agent` and `--mcp-servers-config` let the model read and write files and run shell commands with the server's privileges. Never enable them on a server that untrusted clients can reach.
- **Chat template:** a model without a recognized template falls back to ChatML, which degrades answers and tool calls. Supply the right one with `--chat-template-file`.
- **Downloads:** `-hf` fetches multi-gigabyte files. Check free disk space first, and use `--offline` in air-gapped runs so a missing file fails fast instead of attempting a download.
- **Docker:** inside a container the server must listen on `--host 0.0.0.0`, and the model directory is mounted with `-v "$HOME/models:/models"`. The `-cuda` images need `--gpus all` and the NVIDIA container toolkit on the host.
- **Requantizing:** quantize from a 16-bit or 32-bit GGUF. Quantizing an already quantized file with `--allow-requantize` loses quality.
- **Updating:** `llama update` works only for binaries installed by the llama.app installer. Package-manager installs update through the package manager.
- **When not to use:** for a managed model library with automatic pulls and a daemon, Ollama is simpler. For high-throughput multi-GPU serving of unquantized weights, use vLLM. llama.cpp only loads GGUF, so safetensors checkpoints must be converted first.

Attribution

TerminalSkillsTerminalSkills
View sourceSee grades on GitHubMore from TerminalSkills →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Terse caveman voice: answer first, fluff gone, every technical fact kept. Use for /caveman, "caveman mode", "talk like caveman", "be brief", "less tokens". Stays on until "stop caveman" or "normal mode".

1100021 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

698431 votes

Writing Skills

Create and manage Claude Code skills in HASH repository following Anthropic best practices. Use when creating new skills, modifying skill-rules.json, understanding trigger patterns, working with hooks, debugging skill activation, or implementing progressive disclosure. Covers skill structure, YAML frontmatter, trigger types (keywords, intent patterns), UserPromptSubmit hook, and the 500-line rule. Includes validation and debugging with SKILL_DEBUG. Examples include rust-error-stack, cargo-dep...

3931 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3421 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Amp, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Grok Build, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

741 votes
View all in ai-agents →