Skip to content
Back to skills

Ollama Pro

ASecurity

Ollama guidance — running local LLMs, model selection, Modelfiles, API usage, and RAG/agent integration.

  • 2 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 29, 2026
ai-agentsgobashdockergitapi

Works with

  • cli
  • api

Security analysis

A96/100
  • mediumUses curl or wget to download content

Pro shows the line behind each finding and how to fix it

Scanned September 29, 2026

npx -y skills add aicodedecode/awesome-muse-skills --skill ollama-pro --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Ollama Pro?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Ollama Pro
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/aicodedecode-ollama-pro/badge)](https://www.skillsdirectory.com/skills/aicodedecode-ollama-pro)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: ollama-pro
description: Ollama guidance — running local LLMs, model selection, Modelfiles, API usage, and RAG/agent integration.
category: development
---

## Overview

Ollama is the easiest way to run open LLMs locally: `ollama run llama3.1` and you have a model on your machine — no cloud, no API keys, no data leaving your network. It manages model downloads, quantization, GPU offloading, and exposes an OpenAI-compatible API that plugs into existing tooling.

Local models trade raw capability for privacy, cost (after hardware), latency control, and offline operation. This skill covers running Ollama well: model selection, Modelfiles, the API, and integrating local models into RAG and agent workflows.

## When to use

- Running LLMs locally (privacy, offline, cost).
- Choosing local models (size vs quality vs hardware).
- Customizing models with Modelfiles.
- Using the Ollama API (including OpenAI-compatible endpoints).
- Building RAG/agents on local models.
- Deciding local vs cloud models.

## Core concepts

- **The model library.** `ollama pull llama3.1:8b` — versioned models with size tags (`8b`, `70b`), quantization variants (`q4_0`, `q8_0`), and variants (instruct, vision, code, embedding). Tags encode the size/quantization tradeoff.
- **Quantization.** Q4_K_M default — ~4 bits per weight, big quality retention; Q8 for more quality, smaller quants for tight VRAM. Quantization is why 70B models run on consumer hardware.
- **Hardware fit.** Rule of thumb: ~0.5GB VRAM per billion params at Q4 (plus context overhead). 8B ≈ fits 8GB GPUs; 70B needs ~40GB+ or CPU offload (slow). Match model to hardware, not ambition.
- **Modelfiles.** `FROM`, `PARAMETER` (temperature, num_ctx), `SYSTEM`, `TEMPLATE`, `ADAPTER` — declarative model customization; versioned in git like Dockerfiles. Custom system prompts and defaults without code.
- **The API.** `POST /api/generate` (completion), `/api/chat` (messages), `/api/embed` (embeddings) — plus OpenAI-compatible endpoints (`/v1/chat/completions`) that drop into existing clients. `ollama ps` shows loaded models.
- **Context window.** `num_ctx` parameter — larger contexts need more memory (KV cache scales with context); set per use case, not maximally.
- **GPU offloading.** `num_gpu` layers offloaded; partial offload splits CPU/GPU (slower but fits bigger models). Ollama auto-detects; override when tuning.
- **Keep-alive.** Models unload after inactivity (`keep_alive`) — preload latency-sensitive models (`ollama run` once, or `keep_alive: -1`); balance VRAM across concurrently-needed models.
- **Embeddings.** `nomic-embed-text`, `mxbai-embed-large` — local embeddings for RAG; quality close enough for most retrieval, zero API cost, private.
- **Vision models.** `llama3.2-vision`, `llava` — image understanding locally; quality behind frontier APIs but sufficient for many tasks (document QA, classification).
- **Tool use.** Supported in recent models/APIs (`tools` parameter) — local agents are viable for well-scoped tools; weaker instruction-following than frontier models means simpler tools and more validation.
- **RAG integration.** Ollama (LLM + embeddings) + your vector store (Chroma, Qdrant, pgvector) = fully local RAG. LangChain/LlamaIndex both have Ollama integrations — swap the model layer, keep the pipeline.
- **Evals for local models.** Smaller models fail differently (instruction-following, JSON validity, reasoning depth) — eval your actual tasks on the actual model; don't assume cloud-model prompts transfer.
- **Updates.** `ollama pull` refreshes models; pin versions for reproducibility (`:8b` vs `:latest` semantics); test after updates — model updates change behavior.

## Practical workflow

1. **Size to hardware.** Check VRAM/RAM, pick the largest model that fits comfortably at Q4 with your context needs:
   ```bash
   ollama pull llama3.1:8b        # ~5GB, solid general use
   ollama pull qwen2.5-coder:32b  # coding tasks, needs ~20GB
   ollama pull nomic-embed-text   # local embeddings
   ```
2. **Customize with Modelfiles.** System prompts, parameters, context size — versioned:
   ```
   FROM llama3.1:8b
   PARAMETER temperature 0
   PARAMETER num_ctx 8192
   SYSTEM """
   You are a code reviewer. Be terse. Flag only real issues.
   """
   ```
   ```bash
   ollama create reviewer -f Modelfile
   ```
3. **Use the API.** Chat, generate, embeddings — or the OpenAI-compatible endpoint for existing clients:
   ```bash
   curl http://localhost:11434/api/chat -d '{
     "model": "llama3.1:8b",
     "messages": [{"role": "user", "content": "Explain KV caching briefly."}],
     "stream": false
   }'
   ```
4. **Wire into RAG.** Local embeddings + vector store + Ollama generation — the private RAG stack; evaluate retrieval and generation on your corpus.
5. **Tune serving.** `keep_alive` for hot models, `num_ctx` per use case, concurrent request limits; monitor tokens/sec per model.
6. **Evaluate honestly.** Run your task evals on the local model; identify where it underperforms cloud models (usually: complex reasoning, strict JSON, long instructions) and design around it (simpler prompts, validation, fallbacks).
7. **Hybrid architectures.** Local for private/bulk/simple, cloud for hard reasoning — route by task sensitivity and difficulty, not ideology.
8. **Keep updated deliberately.** Pin versions in production; test model updates against evals before rolling.

## Common pitfalls

- **Model bigger than hardware** — swapping to death; size to actual VRAM/RAM.
- **Default context too small/large** — truncated inputs or wasted memory; `num_ctx` per use case.
- **Cloud prompts on local models** — complex prompts failing on 8B models; simplify, validate outputs.
- **No output validation** — weaker JSON adherence; schema-validate everything.
- **Ignoring keep-alive** — cold-start latency on every request; preload hot models.
- **Unpinned `:latest`** — silent model updates changing behavior; pin versions.
- **Local-only dogma** — struggling with tasks needing frontier reasoning; hybrid routing.
- **Embedding mismatch** — different embedding model than the index was built with; keep them paired.
- **No evals** — assuming parity with cloud models; measure on your tasks.
- **VRAM contention** — multiple large models fighting; plan concurrent model sets.
- **CPU offload surprises** — 70B "running" at 2 tok/s; know your throughput requirements.
- **Secrets in Modelfiles** — API keys baked into shared files; keep secrets out.
- **Skipping quantization awareness** — tiny quants for quality-critical tasks; match quant to need.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…