Skip to content
Back to skills

Vllm Operator

ASecurity

Use when installing, running or tuning vLLM — small or consumer GPUs, DGX Spark, TTFT and throughput under concurrency, KV/prefix caching, quantization, OOM, tool calling.

  • 2 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added October 4, 2026
ai-agentspythonrustgoapibackendsecurityperformance

Works with

  • cli
  • api

Security analysis

A100/100

Pro scans all 20 files and shows the line behind each finding

Scanned October 4, 2026

npx -y skills add Getty/skills --skill vllm-operator --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Vllm Operator?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Vllm Operator
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/getty-vllm-operator/badge)](https://www.skillsdirectory.com/skills/getty-vllm-operator)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: vllm-operator
description: "Use when installing, running or tuning vLLM — small or consumer GPUs, DGX Spark, TTFT and throughput under concurrency, KV/prefix caching, quantization, OOM, tool calling."
metadata:
  version: "1.0.1"
  language: en
  researched: "2026-09-21"
  translated: "2026-09-30"
  scope: "1–4 GPUs; consumer hardware, small workstations, DGX Spark, small rented instances"
---

# vLLM Operator — Small Hardware, Reliable Serving

## Mission

Optimize **correct, useful answers within a latency budget per euro**,
not just nominal tokens/s. This skill is an operations and diagnostic tool,
not a model-training course or a data-center architecture kit.

Load references on demand. Do not load every file into every context.
Explain decisions in English; preserve API and CLI identifiers unchanged.
For installed skills, resolve script and reference paths relative to this file,
not an arbitrary project working directory. Write measurements to a separate,
approved experiment directory, not into the skill installation.

## Non-negotiable rules

1. **Version first.** Read [Versioning](references/03-installation-versioning.md).
   Check the local CLI, model revision, GPU/architecture, driver, and image digest.
   Research dated 2026-09-21; observed release v0.29.0. `stable` is a moving target.
   Do not mix old V0 guidance with Engine V1 or Model Runner V2.
2. **Quality is a gate.** A faster model, different quantization,
   different parser, or shorter context is a product change, not a free
   infrastructure upgrade.
3. **State measurement boundaries.** Client TTFT, first-answer latency, queue time,
   engine TTFT, TPOT, ITL, and SSE chunk spacing are different quantities.
4. **Distinguish caches.** APC ≠ response cache ≠ KV offload ≠ weight offload
   ≠ compile cache. A hit does not automatically accelerate decode.
5. **Operational safety.** No unprotected public serving, no prompts or secrets
   in default logs, and no arbitrary remote-code or administrative capabilities.
   Derive salts and cache namespaces from a trusted identity.
6. **Small steps.** One hypothesis, a fixed workload, and a rollback per experiment.
   Driver changes, rented instances, network exposure, power limits, and
   persistent-data changes require explicit approval. Scripts do not rent servers.
7. **No fabricated evidence.** Distinguish `documented`, `locally checked`,
   `measured`, `hypothesis`, and `not verified`.

## Collect inputs

Record hardware/OS; GPU count, VRAM or unified memory, PCIe/P2P; CPU/RAM/NVMe;
model ID and exact revision; task/modality; quantization/KV dtype;
context distribution; output/reasoning lengths; arrival rate/bursts/active sessions;
cache reuse and tenant boundaries; TTFT/answer/TPOT/E2E SLOs; cost ceiling.

Record missing data as assumptions. With little information, plan for **one GPU,
a suitable small model, a short bounded context, and no offload**.
Do not ask again for hardware details already provided.

Artifact: [Workload contract](templates/workload-contract.md).
Audit: `python scripts/audit_env.py --out ./audit` (local, read-only;
review output files before sharing).

## Workflow

### A. Does the requested combination work?

[Capability matrix](references/02-capability-map.md) →
[Hardware](references/16-consumer-hardware.md) or
[Spark](references/17-dgx-spark.md) →
[Installation](references/03-installation-versioning.md).

Check model architecture, task/runner, dtype, quantization format, kernel,
tokenizer, template, parser, and API. A successful server start does not prove
correct tools, embeddings, or long-context behavior. Run a basic prompt and
representative tasks before optimizing.

### B. Does the model fit in memory under the intended load?

Read [Memory sizing](references/04-memory-capacity.md), then run
`python scripts/capacity.py --help`.
Weights + KV + activations + CUDA Graphs + runtime/multimodal memory + reserve.
Do not apply the MHA/GQA calculator to MLA, Mamba, or hybrid models.
Compare the calculation with the actual KV capacity reported at startup.

### C. Establish a baseline

Read [Deployment](references/22-deployment-security.md) and
[Experiment recipes](references/24-recipes.md). Use a local or private endpoint;
bound output; initially retain `dtype=auto` and a working attention backend.
Record observations before and after warmup.

Client probe: `python scripts/stream_probe.py --help`.
APC probe: use the same file twice; missing `cached_tokens` means unknown, not zero.
Keep resolution/quality tests separate from load tests.

### D. Measure under load

Read [Measurement methodology](references/09-benchmarking.md),
[Metrics](references/10-metrics.md), and [Scheduler](references/08-scheduling-high-load.md).

`python scripts/bench_sweep.py --help` generates a plan by default.
Execution requires `--execute`. It uses the official `vllm bench serve`
CLI and checks flags against the locally installed help output.

Measure cold/warm caches, different prompt/output lengths, realistic arrival
rates, bursts, and saturation. Report errors, timeouts, cancellations, and
SLO goodput alongside p50/p95/p99. Do not present a p99 from a handful of requests
as reliable evidence.

### E. Select a bottleneck; do not tune flags at random

| Observation | First reference | Next controlled experiment |
|---|---|---|
| Long queue, GPU/KV full | [High load](references/08-scheduling-high-load.md) | Bound admission/output; evaluate another replica |
| Long uncached prefills | [Caching](references/05-cache-mechanics.md) | Prefix layout; chunk budget; context |
| Missing cache hits | [Prompt layout](references/06-cache-layout-isolation.md) | Compare tokens/templates/salts/routing |
| Insufficient KV / preemptions | [Memory](references/04-memory-capacity.md) | Reduce active sequences; evaluate KV quantization |
| Slow decode | [Kernels](references/14-kernels-compilation.md) | Isolate bandwidth, batch, quantization, speculation |
| GPU frequently idle, CPU busy | [Tracing](references/11-tracing.md) | Tokenization, parsers, IPC, SDK/proxy |
| Good TTFT, late answer | [API/reasoning](references/19-api-agents.md) | Reasoning/output budget; measure visible answer onset |
| TP slower than expected | [Multi-GPU](references/15-multi-gpu.md) | Topology/P2P; compare independent replicas |
| Offload makes performance worse | [KV offload](references/07-external-kv-offload.md) | Compare transfer time with avoided prefill |
| Slow only behind the proxy | [Deployment](references/22-deployment-security.md) | SSE buffering, timeouts, cancellation propagation |

### F. Deliver the result

Provide a concise operational decision with assumptions, a compatible
configuration, measurements, quality tests, cost calculations, known limits,
and rollback. Use [Experiment](templates/experiment-record.md) and
[Handover](templates/handover.md). Without a GPU test, explicitly state
`not validated on target hardware`; never present script unit tests as GPU benchmarks.

## Reference router

| Topic | File |
|---|---|
| Architecture, prefill/decode, PagedAttention | [01](references/01-engine-fundamentals.md) |
| Capabilities, compatible combinations, and limits | [02](references/02-capability-map.md) |
| Linux/WSL/containers, versions, and upgrades | [03](references/03-installation-versioning.md) |
| Weights, KV, reserves, GQA/MLA/hybrid | [04](references/04-memory-capacity.md) |
| All cache types and their effects | [05](references/05-cache-mechanics.md) |
| Agent prompts, stable prefixes, salts, isolation | [06](references/06-cache-layout-isolation.md) |
| Native KV tiers, LMCache, NVMe, transfer | [07](references/07-external-kv-offload.md) |
| Chunked prefill, queues, overload, and fairness | [08](references/08-scheduling-high-load.md) |
| TTFT/ITL/TPOT/E2E, load tests, and goodput | [09](references/09-benchmarking.md) |
| Prometheus, dashboards, GPU telemetry | [10](references/10-metrics.md) |
| OpenTelemetry and distributed request traces | [11](references/11-tracing.md) |
| PyTorch/Nsight, overhead, and safe diagnostics | [12](references/12-profiling.md) |
| Weight/KV quantization and quality gates | [13](references/13-quantization.md) |
| Attention backends, graphs, compilation, speculation | [14](references/14-kernels-compilation.md) |
| 2–4 GPUs, TP/PP/DP/EP/CP, P2P, and replicas | [15](references/15-multi-gpu.md) |
| 8–48 GB consumer/small-workstation classes | [16](references/16-consumer-hardware.md) |
| DGX Spark/GB10 and two small systems | [17](references/17-dgx-spark.md) |
| Vast.ai/Runpod, small rented GPUs, and costs | [18](references/18-rented-gpus-costs.md) |
| APIs, templates, tools, JSON, reasoning, agents | [19](references/19-api-agents.md) |
| Vision, audio, video, embeddings, and reranking | [20](references/20-multimodal-pooling.md) |
| LoRA, multiple models, sleep, RL boundaries | [21](references/21-lora-lifecycle.md) |
| Operations, proxy, auth, cancellation, privacy | [22](references/22-deployment-security.md) |
| Symptom → cause → test → rollback | [23](references/23-troubleshooting.md) |
| Concrete startup and tuning experiments | [24](references/24-recipes.md) |
| Specialized features and deliberate boundaries | [25](references/25-advanced-boundaries.md) |
| Flag register with consequences | [26](references/26-flag-register.md) |
| Glossary and formulas | [27](references/27-glossary.md) |
| Primary sources and updates | [28](references/28-sources.md) |

## Supporting material

`scripts/` contains an audit tool, capacity and cost calculators, a streaming
probe, a benchmark sweep, a metrics inventory, and package validation.
`configs/` contains deliberately bounded example configurations.
`templates/` structures decisions and evidence. [README](README.md) explains
usage and tests; [VALIDATION](VALIDATION.md) records checks actually performed.

Files in this skill

  • MANIFEST.json9.5 KB
  • README.md5.7 KB
  • SKILL.md9.9 KB
  • VALIDATION.md6.2 KB
  • configs/README.md1.7 KB
  • configs/cache-request.json4.7 KB
  • configs/client-request.json296 B
  • configs/kv-offload-native.json141 B
  • configs/nginx-vllm.conf1.1 KB
  • configs/otel-collector.yaml761 B
  • configs/prometheus-rules.yml1.6 KB
  • configs/prometheus.yml629 B
  • references/01-engine-fundamentals.md3.6 KB
  • references/02-capability-map.md4.2 KB
  • references/03-installation-versioning.md3.6 KB
  • references/04-memory-capacity.md4.2 KB
  • references/05-cache-mechanics.md3.8 KB
  • references/06-cache-layout-isolation.md3.5 KB
  • references/07-external-kv-offload.md3.8 KB
  • references/08-scheduling-high-load.md5.6 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…