Use when choosing or reviewing GPU targets for NPA workbench tools, training, rendering, inference, or workflow YAML resources.
Scanned 9/8/2026
Install to Claude Code
npx -y skills add nebius/nebius-physical-ai --skill gpu-selection --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Gpu Selection?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/nebius-gpu-selection)More formats (shields.io, HTML) on the badges page.
---
name: gpu-selection
description: Use when choosing or reviewing GPU targets for NPA workbench tools, training, rendering, inference, or workflow YAML resources.
---
# GPU Selection
## When To Use
Use this skill when a task asks which GPU family to use, changes workflow
resources, updates image routing, or reviews render/training placement.
## Procedure
1. Identify whether the workload needs RT cores, tensor throughput, multi-GPU
scaling, or only CPU resources.
2. Check the tool-specific skill for hard constraints.
3. Encode the choice in CLI flags, SDK config, or workflow YAML env/resources.
4. Keep image variants aligned with GPU selection.
5. For direct-Kubernetes Jobs, discover `nvidia.com/gpu.product` labels and
construct an ordered, compatible candidate list. Move to the next product
only for concrete scheduler evidence (`Unschedulable`, insufficient GPU
resource, or no matching product/affinity); runtime, pull, credential,
checkpoint, and application failures are not placement failures.
## Three-Tier Contract
- CLI: commands expose GPU choices through flags such as `--gpu-type`,
`--gpu-preset`, `--runtime`, or tool-specific image variant options.
- SDK: runtime config and request builders should carry GPU type/count rather
than deriving it from private environment names.
- YAML: workflow resources and env vars must express the GPU target explicitly
enough for reviewers to validate routing.
## Current Defaults
- H100: general training, CLIP embedding, detection, MJLab, Cosmos inference,
LeRobot training smoke, and non-render throughput.
- L40S: Isaac Lab and SONIC render validation on VM hosts.
- RTX PRO 6000 Blackwell: Isaac Lab and SONIC render validation on Kubernetes
with mounted NVIDIA GPU Operator drivers.
- B200 / B300: headless, state-based training and inference only.
- CPU: Retargeting and many dataset curation/import steps.
## Blackwell Is Two Different Targets
"Blackwell" spans two CUDA majors, and their binaries are mutually incompatible:
| GPU | Compute capability | SM | Nebius platform |
|---|---|---|---|
| RTX PRO 6000 Blackwell | 12.0 | `sm_120` | `gpu-rtx6000` |
| B200 | 10.0 | `sm_100` | `gpu-b200-sxm` (us-central1) |
| B300 (Blackwell Ultra) | 10.3 | `sm_103` | `gpu-b300-sxm` |
A green smoke on RTX PRO 6000 does not prove B200/B300. Within major 10,
forward compatibility holds, so `sm_100` SASS runs on `sm_103`: target B200
first, then confirm on B300. See
`docs/workbench/blackwell-datacenter-image-compatibility.md` and the per-image
verdicts in `npa/docker/workbench/blackwell-dc-images.json`.
## Gotchas
- H100, H200, and datacenter Blackwell (B200/B300) lack RT cores; do not route
Isaac Lab or SONIC render validation there. `npa.workbench.sonic.routing`
classifies these as `datacenter-headless` and rejects render workloads.
- L40S capacity can be constrained; if the task only needs non-render training,
H100 may be the pragmatic target.
- Preemptible GPU placement does not change any boot-disk allocation. Preserve
the identical `compute.disk.count` and `compute.disk.size.network-ssd` byte
requirements in quota plans.
- For repeated typed on-demand placement failures and consent-gated pool
switching, load `skills/atomic/gpu-allocation-fallback/SKILL.md`.
- B200/B300 enablement depends on upstream library support per tool. Treat it as
vendor-paced unless current tests prove the path. The 2026-08-03 final
Genesis/Sim2Real tags passed real kernel compilation and physics smokes on
both B200 and B300; the NVIDIA Isaac vendor stacks and the per-image Cosmos
blockers in `blackwell-dc-images.json` remain separate constraints.
- Sim2Real Isaac candidates are only L40S and RTX PRO 6000 label variants.
Never add H100/H200 as an Isaac capacity fallback. Non-Isaac Cosmos candidates
may use H100/H200 only when the selected image advertises a compatible SM and
the component's VRAM/model rules allow it.
- Record candidate order, skipped/attempted products and scheduler reasons,
selected product/node, allocated resource/count, Job name, and runtime image
digest in the component provenance. Exhaustion is a blocker, not permission
to change tier, backend, image semantics, or execution mode.
- Terraform's canonical compute outputs are `platform` and `preset`, with
`cpu_platform`/`cpu_preset` for CPU-only instances. Deprecated
`gpu_platform`/`gpu_preset` aliases are GPU-only and return null for CPU
instances; do not interpret a historical CPU value under those aliases as GPU
placement.
## Verify
```bash
npa/.venv/bin/python -m pytest npa/tests/guardrails/test_skills_index.py -q
```
The smoke test invokes help for GPU-sensitive training commands and parses the
workflow YAML resources referenced by the manifest.
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!