Skip to content
Back to skills

Engine Sglang

CSecurity

Serve a Hugging Face model with SGLang on a Linux machine with an NVIDIA or AMD GPU, configured from the model's SGLang cookbook page — or, when it has none, from the model's own files — and join it to the person's fleet. Load before installing, configuring, starting or stopping SGLang.

  • 1,141 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 30, 2026
ai-agentspythongoshelldockergitapi

Works with

  • api

Security analysis

C67/100
  • criticalAccesses sensitive system or user directories
  • mediumInstalls packages at runtime which could introduce malicious dependencies

Pro shows the line behind each finding and how to fix it

Scanned September 30, 2026

npx -y skills add autonomous-ai/openharness --skill engine-sglang --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Engine Sglang?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Engine Sglang
[![Security: C — Skills Directory](https://www.skillsdirectory.com/api/skills/autonomous-ai-engine-sglang/badge)](https://www.skillsdirectory.com/skills/autonomous-ai-engine-sglang)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: engine-sglang
description: "Serve a Hugging Face model with SGLang on a Linux machine with an NVIDIA or AMD GPU, configured from the model's SGLang cookbook page — or, when it has none, from the model's own files — and join it to the person's fleet. Load before installing, configuring, starting or stopping SGLang."
---

# SGLang (Linux + NVIDIA CUDA or AMD ROCm)

Official sources, read 2026-09-29. Agent index: [docs.sglang.io/llms.txt](https://docs.sglang.io/llms.txt)
(every page is available as `.md`). [Cookbook](https://docs.sglang.io/cookbook/intro) ·
[Models](https://www.sglang.io/models) · [Quickstart](https://docs.sglang.io/docs/get-started/quickstart.md) ·
[Server arguments](https://docs.sglang.io/docs/advanced_features/server_arguments.md) ·
[Tool parser](https://docs.sglang.io/docs/advanced_features/tool_parser.md) ·
[Separate reasoning](https://docs.sglang.io/docs/advanced_features/separate_reasoning.md) ·
[Native API](https://docs.sglang.io/docs/basic_usage/native_api.md) ·
[Transformers fallback](https://docs.sglang.io/docs/supported-models/transformers_fallback.md) ·
source: [github.com/sgl-project/sglang](https://github.com/sgl-project/sglang)
Tested: not runnable on macOS (below); no GPU run yet. Tags: `[doc]` official source, `[run]` seen on a machine, `[?]` unverified.

## When to use it

- Linux with an NVIDIA GPU of compute capability 8.0 or newer (A10, A100, L4, L40S, H100), Python ≥ 3.10 [doc];
  AMD GPUs, Intel Xeon CPUs and TPUs have their own guides [doc]. Hugging Face models, many requests at once,
  several GPUs.
- **Never on macOS.** The current release ships Linux wheels only and requires CUDA builds of its kernels;
  on a Mac the installer falls back to an old release whose server does not import [run].
- A GGUF file for one person: Grid's own engine serves it with nothing to install.

## 1. The cookbook page — `"$GRID_FLEET" recipe sglang ORG/NAME`

Finds the model family's page through `llms.txt` (the most specific family name the model name starts
with) and prints: `source` (the page), `installation` (its version requirement — some families need a
release or the main branch — and install commands), `serveCommands` (the page's own launch commands per
GPU type), `toolCallParsers`, `reasoningParsers`, and `configurationTips`. Read the page itself for
anything else. Exit code 3 means no page: go to step 2.

## 2. No cookbook page — read the model: `"$GRID_FLEET" model-facts ORG/NAME`

1. `support.sglang.listed: true` means SGLang's model registry has the architecture; `null` means this
   check cannot tell. SGLang can run most decoder models through `--model-impl transformers` [doc] — a test,
   not a promise.
2. `modelCard.serveCommands` and parsers: the model authors' own SGLang command. Use it when present.
3. Otherwise `--tool-call-parser auto` and `--reasoning-parser auto`: SGLang detects both from the model's
   chat template [doc]. Then the tool-call acceptance check decides; if it fails, say so and stop — do not
   cycle through parser names.
4. `contextLength` must be ≥ 65536.

## Install and version

- Slow step, ask first. `uv pip install --prerelease=allow sglang` [doc]; a family whose cookbook page asks
  for the main branch: `uv pip install --prerelease=allow 'git+https://github.com/sgl-project/sglang.git#subdirectory=python'` [doc].
  `OSError: CUDA_HOME environment variable is not set` → `export CUDA_HOME=/usr/local/cuda-<version>` [doc].
- Docker: `lmsysorg/sglang:latest` (NVIDIA), `lmsysorg/sglang-rocm:<tag>` (AMD, tag per GPU generation on
  the cookbook page) [doc], run with `--gpus all --shm-size 32g --ipc=host -p 127.0.0.1:P:30000
  -v ~/.cache/huggingface:/root/.cache/huggingface` [doc].

## Common configs (the cookbook's flags win)

Each adds the parsers from step 1 or 2, `--host 127.0.0.1 --port P` and `--context-length` of at least 65536
(the default is the model's own maximum [doc]).

| Case | Add |
|---|---|
| One GPU, one agent | `--max-running-requests 4` |
| One GPU, many people | leave `--max-running-requests` to SGLang |
| Several GPUs, one machine | `--tp-size <GPU count>` [doc] |
| Out of memory | lower `--mem-fraction-static` (weights plus KV pool share) [doc]; `--kv-cache-dtype fp8_e4m3` [doc] |

## Start, ready, join

    "$GRID_FLEET" serve sglang-P --env HF_HUB_OFFLINE=1 -- ~/.grid/envs/sglang/bin/sglang serve \
      --model-path <model id or snapshot dir> --served-model-name <id> --host 127.0.0.1 --port P <flags>

(`python3 -m sglang.launch_server` takes the same arguments [doc].) Defaults are 127.0.0.1 and port 30000 [doc].
Install into `~/.grid/envs/sglang` so `fleet models` finds it; `fleet serve` keeps it alive after your shell
returns (log `run/sglang-P.log`).

- Verify: `"$GRID_FLEET" verify --at http://127.0.0.1:P/v1 --model <id> --kind sglang` (bounded, narrated).
- Ready: the log says `The server is fired up and ready to roll!` [doc]; `GET /health` → 200;
  `GET /health_generate` generates one token [doc]; `GET /v1/models` lists `<id>`; one bounded answer; one tool call.
- Join: `"$GRID_FLEET" run -- join GRID --at http://127.0.0.1:P/v1 -m <id> --advertise-as ALIAS`. Grid's detector has
  no SGLang probe and would mislabel it on 8000 or 8080 [run], so always `--at`, always with `/v1`.
- Thinking off per request: `"chat_template_kwargs": {"enable_thinking": false}` where the cookbook page shows it [doc].

## Stop

`"$GRID_FLEET" stop sglang-P` (or `docker stop` your container); confirm the port is free.

## Known failures → what to do

| Sign | Do |
|---|---|
| out of memory while serving | lower `--mem-fraction-static` [doc] |
| no `tool_calls` | the cookbook's `--tool-call-parser`, else `auto` [doc]; still none → report |
| a CUDA crash, a multi-GPU hang, a production incident | read SGLang's own maintainer skills: `.agents/skills/debug-cuda-crash`, `debug-distributed-hang`, `sglang-prod-incident-triage` (`SKILL.md` in each) at `https://raw.githubusercontent.com/sgl-project/sglang/main/` |
| 404 on `/models` | the `--at` URL lacks `/v1` |

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…