Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Pytorch Pro

ASecurity

PyTorch guidance — tensors, autograd, training loops, DataLoader, distributed training, and production inference.

2 stars
0 votes
0 copies
0 views
Added 9/29/2026
ai-agentspythongoc++nodedebuggingapiperformance

Works with

cliapi

Security Analysis

A100/100

Scanned 9/29/2026

$npx -y skills add aicodedecode/awesome-muse-skills --skill pytorch-pro --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Pytorch Pro?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Pytorch Pro
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/aicodedecode-pytorch-pro/badge)](https://www.skillsdirectory.com/skills/aicodedecode-pytorch-pro)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: pytorch-pro
description: PyTorch guidance — tensors, autograd, training loops, DataLoader, distributed training, and production inference.
category: development
---

## Overview

PyTorch is the dominant framework for deep learning research and increasingly for production: dynamic computation graphs, Pythonic APIs, and an ecosystem (torchvision, torchaudio, Hugging Face, Lightning) that covers most use cases. Its imperative style makes debugging intuitive — you can print tensors mid-forward-pass like any Python code.

This skill covers the PyTorch fundamentals that transfer everywhere: tensors and autograd, correct training loops, efficient data loading, distributed training, and the path to production inference (TorchScript, ONNX, TensorRT, torch.compile).

## When to use

- Building or debugging PyTorch models.
- Writing correct training loops (train/eval modes, gradient handling).
- Optimizing data loading (DataLoader bottlenecks).
- Training on multiple GPUs (DDP).
- Exporting models for production inference.
- Choosing between PyTorch, TensorFlow, and JAX.

## Core concepts

- **Tensors.** The core data structure: `torch.tensor`, shapes/dtypes/devices. Think in shapes — most bugs are shape mismatches; print shapes liberally while developing. `device` discipline (`.to(device)`) avoids CPU/GPU mixing errors.
- **Autograd.** Automatic differentiation: operations on `requires_grad=True` tensors build a graph; `loss.backward()` computes gradients. `torch.no_grad()` for inference (saves memory, disables graph); `detach()` to cut graph edges.
- **nn.Module.** The building block: `__init__` defines layers, `forward` defines computation. `model.train()` vs `model.eval()` — dropout/BatchNorm behave differently; forgetting `eval()` at inference is a classic silent bug.
- **Training loop anatomy.** Zero grads → forward → loss → backward → optimizer step. `optimizer.zero_grad()` placement matters (gradients accumulate by default — a feature for accumulation, a bug when forgotten).
- **DataLoader.** `Dataset` + `DataLoader`: `num_workers` for parallel loading, `pin_memory` for GPU transfer speed, custom `collate_fn` for variable-length data. Data loading is the usual bottleneck — profile it (`nvidia-smi` showing idle GPU = starved pipeline).
- **Losses and optimizers.** CrossEntropyLoss (classification), MSE/BCE variants; Adam/AdamW defaults, SGD with momentum for some vision tasks; learning-rate schedulers (cosine, OneCycle, ReduceLROnPlateau). AdamW + cosine is the strong default.
- **Mixed precision.** `torch.cuda.amp` (autocast + GradScaler): ~2x speedup, half memory, with loss scaling for stability. Nearly free performance — use it unless you have a reason not to.
- **Gradient clipping/accumulation.** Clip (`clip_grad_norm_`) for RNNs/transformers stability; accumulation (`loss / k`, step every k batches) for effective large batches on limited VRAM.
- **Checkpointing.** Save `model.state_dict()` + `optimizer.state_dict()` + epoch + scheduler + RNG states — resumable training needs all of these, not just weights. Save best + latest.
- **Distributed (DDP).** `DistributedDataParallel` — one process per GPU, gradients averaged. `torchrun` launcher; `DistributedSampler` for the DataLoader; sync BatchNorm where needed. DDP (not DataParallel) is the correct multi-GPU approach.
- **torch.compile.** The 2.0+ compiler: often 30%+ speedup with one line (`torch.compile(model)`). Graph breaks from data-dependent control flow reduce gains — write compile-friendly code (avoid Python branching on tensor values in hot paths).
- **Inference optimization.** `model.eval()` + `torch.no_grad()`/`inference_mode()`; export paths: TorchScript (C++ runtime), ONNX (cross-framework), TensorRT (NVIDIA GPUs, max perf), quantization (INT8 for CPU/latency). Choose by deployment target.
- **Debugging.** NaN hunting (check loss per batch, `torch.autograd.set_detect_anomaly(True)` for the guilty op), shape printing, overfitting a single batch as a smoke test (if it can't overfit 10 samples, something's broken).
- **Ecosystem.** torchvision/torchaudio/huggingface for pretrained models; PyTorch Lightning for training-loop structure; `torchmetrics` for correct metric computation.

## Practical workflow

1. **Write the correct loop.** Train/eval modes, zero grads, no-grad inference — the template:
   ```python
   for epoch in range(epochs):
       model.train()
       for x, y in train_loader:
           x, y = x.to(device), y.to(device)
           optimizer.zero_grad()
           with torch.autocast(device_type="cuda", dtype=torch.float16):
               loss = criterion(model(x), y)
           scaler.scale(loss).backward()
           scaler.unscale_(optimizer)
           torch.nn.utils.clip_grad_norm_(model.parameters(), 1.0)
           scaler.step(optimizer); scaler.update()
       # validation
       model.eval()
       with torch.inference_mode():
           val_loss = sum(criterion(model(x.to(device)), y.to(device))
                          for x, y in val_loader) / len(val_loader)
   ```
2. **Smoke-test first.** Overfit a single batch — if loss doesn't → ~0, debug before full training (LR, data, labels, architecture).
3. **Fix data loading.** `num_workers=4+`, `pin_memory=True`, profile GPU utilization; move preprocessing (tokenization, augmentation) into the pipeline, not the training step.
4. **Scale with DDP.** `torchrun --nproc_per_node=4 train.py`; DDP wrap; DistributedSampler; scale LR with batch size (linear scaling rule as starting point).
5. **Checkpoint completely.** state_dicts for model+optimizer+scheduler, epoch, RNG states; save best-by-val and latest; test resume once.
6. **Compile and export.** `torch.compile` for training/inference speed; export to ONNX/TensorRT/quantized for deployment targets; benchmark latency/throughput per export.
7. **Monitor training.** Loss curves, LR schedule, gradient norms, GPU util/memory — W&B/MLflow from run one. Diverging loss? Check LR, data normalization, and label correctness first.
8. **Serve efficiently.** Batch inference requests, `inference_mode()`, right-sized instances; the deployment format (ONNX/TensorRT) chosen by latency requirements.

## Common pitfalls

- **Forgetting `model.eval()`** — dropout/BatchNorm active at inference; silent quality drop.
- **Forgetting `zero_grad()`** — gradients accumulating unintentionally; loss behaves strangely.
- **No `torch.no_grad()` at inference** — graph building wasting memory; `inference_mode()` is stricter.
- **DataLoader bottleneck** — GPU idle; `num_workers`, `pin_memory`, pipeline preprocessing.
- **Using DataParallel** — deprecated, slower; DDP with `torchrun`.
- **NaNs unchecked** — training diverging silently; anomaly detection + per-batch loss logging.
- **Wrong device mixing** — CPU tensor + CUDA tensor errors; consistent `.to(device)`.
- **Saving whole model instead of state_dict** — pickle fragility; `state_dict` + architecture code.
- **Incomplete checkpoints** — weights only, can't resume optimizer/scheduler; save everything.
- **LR not scaled with batch size** — DDP batch growth without LR adjustment; linear scaling starting point.
- **Compile-unfriendly code** — data-dependent Python branching killing `torch.compile` gains; restructure hot paths.
- **No smoke test** — full training on a broken setup; overfit-one-batch first.
- **Metric bugs** — wrong averaging across batches; `torchmetrics` or careful accumulation.

Attribution

aicodedecodeaicodedecode
View sourceSee grades on GitHubMore from aicodedecode →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Terse caveman voice: answer first, fluff gone, every technical fact kept. Use for /caveman, "caveman mode", "talk like caveman", "be brief", "less tokens". Stays on until "stop caveman" or "normal mode".

1100021 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

698621 votes

Writing Skills

Create and manage Claude Code skills in HASH repository following Anthropic best practices. Use when creating new skills, modifying skill-rules.json, understanding trigger patterns, working with hooks, debugging skill activation, or implementing progressive disclosure. Covers skill structure, YAML frontmatter, trigger types (keywords, intent patterns), UserPromptSubmit hook, and the 500-line rule. Includes validation and debugging with SKILL_DEBUG. Examples include rust-error-stack, cargo-dep...

3931 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3421 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Amp, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Grok Build, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

741 votes
View all in ai-agents →