
Claude Skills by shen-shanshan
github.com/shen-shanshanGenerate three local bash scripts (docker run / server start / AgentX client replay) to run an InferenceX srt-slurm recipe YAML benchmark manually on a local GPU server, without Slurm or srt-slurm. Given a recipe like dsv4/vllm/mi355x-fp4-mtp/agentic.yaml and an override variant (e.g. override_tp8_c56 or override_dep8_c64), optionally capture one complete torch profiler prefill or decode step via vllm serve --profiler-config, with configurable host paths for model weights, the InferenceX chec...
Compare vLLM serving benchmark outputs before and after a code change. Parses plain-text vLLM benchmark output (the "Serving Benchmark Result" block), computes per-metric percentage changes with improvement/regression markers, generates a Markdown report with a full metrics table and a key-changes summary, and saves it to ./outputs/. Use when the user pastes or provides vLLM benchmark output and asks to compare, summarize, or analyze performance differences between two runs (e.g., "before thi...
Analyze contribution opportunities in the vllm-project/vllm repository for community developers. Given a module, feature, or model area, this skill collects information from open issues, recent PRs, GitHub discussions, code TODOs/FIXMEs, roadmap labels, and maintainer activity to generate a structured Markdown report of actionable tasks (feature development, model support, performance optimization, bug fixes, documentation, refactoring). Each task includes difficulty, prerequisite knowledge, ...
Design and implement vLLM features. Given user requirements (feature description, related PRs, reference materials), produces (1) core code implementation — NO test cases — and (2) a rich Markdown design document saved to the current project root. Use when the user asks to design a vLLM feature, implement a vLLM feature, architect a component for vLLM, generate a design doc for vLLM, or requests a feature design for ML inference systems. Triggered by phrases like "帮我设计vLLM的xxx功能", "design a v...
Generate comprehensive Chinese technical tutorial documents for vLLM features and modules. Produces deep-dive code walkthrough documents with Mermaid architecture/flow diagrams, comparison tables, code snippets, and performance analysis. Output is saved as Markdown to the skill's outputs/ directory. TRIGGER when: user asks to learn about a vLLM feature, module, or subsystem (e.g., "我想了解 vllm 中的 xxx", "帮我生成 vllm xxx 的教程", "generate a tutorial for vllm's xxx feature", "vllm xxx 特性分析"). DO NOT T...
Generate comprehensive Chinese technical tutorial documents for specific vLLM models (e.g., Qwen3-VL, DeepSeek-V3, Llama 4, InternVL3, etc.). Produces deep-dive model walkthrough documents with Mermaid architecture diagrams, comparison tables, input preprocessing flows, forward pass analysis, ViT computation (for VLMs), vLLM code implementation analysis, and technical principle explanations (MoE, MLA, Gated Attention, ViT, DiT, etc.). Output is saved as Markdown to the skill's outputs/ direct...
Fetch and organize multimodal-related open issues from vllm-project/vllm. Categorizes issues by problem type (Bug, Feature Request, Performance, CUDA Graph, EPD disaggregation, Prefix Caching, ViT/visual encoder, Video, Audio/Speech, specific VL models, etc.) and generates a structured Markdown report saved to the skill's ./outputs directory. Use when the user wants to collect, search, analyze, or summarize open issues related to multimodal features or models in the vllm repository. Triggered...
Generate a vLLM-style PR description (Purpose / Test Plan / Test Result) from a GitHub PR's code changes. Use when the user provides a vLLM PR link or number and asks to generate, write, or draft a PR description. Triggered by requests like "帮我生成PR描述", "generate PR description for vllm PR 12345", "帮我写vllm PR的描述", "draft a PR desc for https://github.com/vllm-project/vllm/pull/12345".
Fetch and analyze a Pull Request from the vllm-project/vllm GitHub repository, then generate a comprehensive Markdown report covering PR overview, code change analysis (with Mermaid architecture/flow diagrams), technical principles, discussion highlights, and risk assessment. Use when the user provides a vllm PR number and asks to summarize, analyze, review, or understand it. Triggered by requests like "帮我分析vllm的PR 12345", "总结一下vllm PR 10000", "vllm PR 9999 做了什么", "analyze vllm PR 12345".
Generate a vLLM-style RFC (Request for Comments) document based on user input. Use when the user wants to create an RFC for a major architectural change or design decision in vLLM. Triggered by requests like "帮我生成一个vLLM RFC", "create a vLLM RFC for ...", "写一个RFC关于...", "生成RFC文档".
Review AMD/ROCm-related pull requests from the vllm-project/vllm GitHub repository and produce a single combined Chinese report: a detailed PR summary (summary, background & motivation, code-change analysis with Mermaid diagrams, technical principles, discussion highlights, risk table) followed by severity-sorted, type-categorized ROCm review findings, a verdict, and copy-paste English review comments with file + diff line numbers ready for GitHub PR review. Use when the user provides a vllm ...
AI code review for aiter PRs. Catches perf regressions, silent correctness bugs, dispatch gate holes, and AI-generated code patterns. Invoke with a PR number; works through fetch → semantic understanding → rule checklist → verdict. Add new rules here as patterns emerge from real reviews.
AI code review for ATOM PRs. ATOM consumes aiter kernels and integrates with vLLM/SGLang plugins. Reviews check perf claims, aiter cross-repo deps, model coverage, dispatch correctness, and AI-generated code patterns. Invoke with a PR number.
Write or complete Chinese vLLM technical blog posts in the author's established Zhihu style. Use when the user provides a vLLM feature, model, architecture, optimization, or other topic and asks for a full blog post, or provides an existing Markdown outline/draft plus references and asks to research current vllm-project/vllm code and finish the missing sections. Produces concise technical diagrams and stores each new article with its images under this skill's outputs directory.
Generate test cases for the vllm-project/vllm repository (https://github.com/vllm-project/vllm). Use this skill when the user wants to write unit tests or integration/e2e tests for vllm code, functions, classes, or features. Triggered by requests like "帮我写XXX的测试用例", "生成XXX的单元测试", "为XXX功能写测试", "generate tests for XXX in vllm", "write a test for XXX vllm function". Do NOT use for vllm-ascend — use vllm-ascend-ut-generator instead.
Compare decode-phase kernel implementations of vLLM vs ATOM (ROCm/ATOM) running the same model with the same config, using torch profiler Chrome-trace JSONs from both engines. Samples one decode step, extracts one layer per layer type, produces per-layer-type comparison tables (which kernels each engine uses, fused or split, multi-stream, quantization differences, timing), derives a vLLM optimization TODO list, and writes a Markdown report (Chinese by default, English, or both) to the skill's...