Practical guide for using the GWDG / NHR-Nord@Göttingen HPC systems — one unified cluster of Emmy (CPU) and Grete (GPU) islands shared by SCC, KISSKI, NHR, and REACT accounts. Use whenever the user mentions SCC (Scientific Compute Cluster), KISSKI, NHR, HLRN, Emmy, Grete, glogin/glogin-gpu, gwdg-storage, or "the GWDG cluster"; or wants to connect via SSH, write or debug a Slurm script (sbatch/srun/salloc), pick a partition (scc-cpu, scc-gpu, kisski, kisski-h100, grete, standard96s), request G...
Installs into .claude/skills of the current project.
Are you the author of Gwdg Hpc?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/lkaesberg-gwdg-hpc)
---
name: gwdg-hpc
description: >-
Practical guide for using the GWDG / NHR-Nord@Göttingen HPC systems — one
unified cluster of Emmy (CPU) and Grete (GPU) islands shared by SCC, KISSKI,
NHR, and REACT accounts. Use whenever the user mentions SCC (Scientific Compute
Cluster), KISSKI, NHR, HLRN, Emmy, Grete, glogin/glogin-gpu, gwdg-storage, or
"the GWDG cluster"; or wants to connect via SSH, write or debug a Slurm script
(sbatch/srun/salloc), pick a partition (scc-cpu, scc-gpu, kisski, kisski-h100,
grete, standard96s), request GPUs (A100/H100/V100 or MIG slices), manage storage
($HOME/$WORK/$PROJECT/workspaces/show-quota/ws_allocate), load modules (Lmod),
use conda/miniforge or Apptainer, run Hugging Face models, or move data to/from
these systems. Also for login nodes, account types, core-hour accounting, QoS,
or job troubleshooting — even when the cluster is not named but the context
(GWDG, Göttingen) is clear. Prefer this over generic Slurm knowledge; the
partitions, login nodes, and filesystems are site-specific.
---
# GWDG HPC: Emmy & Grete (SCC / KISSKI / NHR / REACT)
## The mental model (read this first)
There is **one unified HPC system** in Göttingen, operated by GWDG. It is not
three separate clusters. It is made of compute "islands":
- **Emmy** — the CPU islands (Emmy Phase 1/2/3), e.g. Intel Sapphire Rapids and Cascade Lake nodes.
- **Grete** — the GPU islands (Grete Phase 1/2/3), with Nvidia V100 / A100 / H100 GPUs.
- Plus the older **SCC Legacy** island and a few group-restricted islands (CIDBN, FG, SOE).
**SCC, KISSKI, NHR, and REACT are not different machines — they are different
*account types* on this shared system.** The account type is the single most
important fact, because it determines three things at once:
1. which **login nodes** you can reach,
2. which **filesystems** ($HOME, scratch) you land on, and
3. which **Slurm partitions** you are allowed to submit to.
So the first job is always to figure out *what kind of account the user has*,
and then read everything else off that. Do not assume; a user may have several
accounts (e.g. an SCC account and an NHR project).
The default software environment everywhere is the **`gwdg-lmod`** Lmod module stack.
---
## Step 1 — Establish the account type
If it isn't already clear from the conversation, determine it. Signals:
- **Username shape.** A Project-Portal username looks like **`u12345`** (a `u`
followed by five digits) and is *project-specific* — each project the user
belongs to is a separate account with its own storage and allocation. The
user's AcademicID (the name they use for other GWDG services) **cannot** be
used for SSH (the only exception is the storage transfer nodes).
- **Legacy usernames** (being phased out, must be migrated): a name-based login
like `jdoe`, `john.doe2`, `doe15` is a legacy **SCC** user; a login like
`nimjdoe` / `hbbmustr` (3 letters, 2 of them the German state code, then 5
chosen characters) is a legacy **NHR/HLRN** user.
- **Project Portal tree.** In <https://hpcproject.gwdg.de>, the project's path
reveals the type: `Projects / Scientific Compute Cluster (SCC)` → SCC;
`Projects / Extern / KISSKI` → KISSKI; `Projects / Extern / NHR-NORD@Göttingen` → NHR;
`Projects / Extern / EFRE-REACT …` → REACT (CIDAS and Research Units are treated like NHR).
- **On the cluster itself**, `show-quota` lists the data stores assigned to the
current user, and the login banner shows which partitions are available.
### The master table — everything keys off this
| Account | Login node(s) (SSH alias) | $HOME filesystem | scratch / $WORK | Partitions you can use |
| --- | --- | --- | --- | --- |
| **NHR** | `glogin` (Emmy), `glogin-gpu` (Grete) | vast-nhr | lustre-mdc, lustre-grete | `medium96s*`, `standard96(s)*`, `large96(s)*`, `huge96(s)*`, `grete*`, `jupyter*` |
| **SCC** | `glogin` / `glogin-p3` (Emmy), `glogin-gpu` (Grete), `login-mdc` (SCC Legacy) | vast-standard | scratch-scc | `scc-cpu`, `scc-gpu`, `medium`, `jupyter*` |
| **KISSKI** | `glogin-gpu` (Grete) | vast-kisski | vast-kisski | `kisski`, `kisski-h100`, `grete:interactive`, `jupyter*` |
| **REACT** | `glogin-gpu` (Grete) | vast-react | vast-react | `react`, `grete:interactive`, `jupyter*` |
Notes that trip people up:
- **SCC users doing CPU work should log in to `glogin-p3` (Emmy Phase 3)** — the `scc-cpu` partition lives on that island.
- **KISSKI is GPU-focused**: it logs in via `glogin-gpu` and has no general CPU partition of its own.
- **The `kisski` and `kisski-h100` partitions are effectively shared** — request only the GPUs/cores you need (there is no separate `:shared` variant). `kisski-h100` often has a shorter queue than the A100 `kisski`.
- **The `grete*` partitions are NHR-only.** KISSKI/REACT accounts are *denied* there (they get `kisski*` / `react` instead, plus `grete:interactive`). Don't hand a KISSKI user a `grete:shared` script.
- The `jupyter` partition (for JupyterHPC sessions) is available to everyone.
For full detail on partitions and hardware, read `references/partitions.md`.
---
## Step 2 — Connecting (SSH)
Pick the login node for the island you'll actually use (closest hardware match,
correct default software stack, right filesystems mounted). The recommended
`~/.ssh/config` aliases:
```sshconfig
# NHR + SCC (CPU islands)
Host Emmy
Hostname glogin.hpc.gwdg.de
User u12345
IdentityFile ~/.ssh/id_ed25519
Host Emmy-p3
Hostname glogin-p3.hpc.gwdg.de
User u12345
IdentityFile ~/.ssh/id_ed25519
# Grete (GPU island) — also the KISSKI / REACT login
Host Grete
Hostname glogin-gpu.hpc.gwdg.de
User u12345
IdentityFile ~/.ssh/id_ed25519
```
Then `ssh Emmy-p3` (SCC CPU), `ssh Grete` (any GPU work), etc. To connect
without a config entry: `ssh u12345@glogin-gpu.hpc.gwdg.de -i ~/.ssh/id_ed25519`.
- **GPU work → always use `glogin-gpu`.** Many GPU software modules are not even
visible on the other login nodes, and the CPU architecture matches the Grete nodes.
- **Legacy SCC login nodes (`login-mdc` / `gwdu101-102`) are not reachable from
outside GÖNET** without VPN. Use a general login node as a jumphost:
`ssh u12345@login-mdc.hpc.gwdg.de -J u12345@glogin.hpc.gwdg.de` (or a `ProxyJump`
entry in the config).
- On first connect SSH will ask you to verify the host key. The ed25519
fingerprint for all login nodes is
`SHA256:PPK0aO2QZ/k4duUx18Pp5AOKG/gFEBHgw/bl8vg9oJk`. If a key *changed*, clear
the old one with e.g. `ssh-keygen -R glogin.hpc.gwdg.de`.
Full config blocks (including legacy SCC, storage transfer nodes, and advanced
multi-account setups) and troubleshooting are in `references/connecting.md`.
---
## Step 3 — Software (the module system)
Software is provided through **Lmod**. The essentials:
```bash
module avail # what can I load right now
module avail STRING # filter by name
module spider STRING # search everything, incl. modules hidden behind dependencies
module load NAME/VERSION # make software available (omit /VERSION for the default)
module list # what's loaded
module purge # unload everything
```
The module tree is **hierarchical**: some modules only appear after a
prerequisite is loaded. The classic example is CUDA, which requires a compiler
first:
```bash
module load gcc cuda # gcc must come before cuda; default cuda is 12.6.2
# use `module spider cuda/<version>` to see which gcc a given cuda needs
```
Avoid putting `module load` lines in `~/.bashrc` — modules differ between
islands and this causes hard-to-debug failures. Load them in your shell or, for
jobs, inside the job script. More (software stacks `gwdg-lmod` / `nhr-lmod` /
`scc-lmod`, conda/miniforge, Apptainer, MPI) is in `references/software.md`.
---
## Step 4 — Storage (know where to put things)
Different filesystems for different jobs — this is exposed on purpose, for
performance. Run **`show-quota`** on a login node to see the data stores and
quotas for the current user. Key locations (paths are provided via environment
variables set at login):
| Variable / store | Typical limit | Use it for | Notes |
| --- | --- | --- | --- |
| `$HOME` | 60 GiB | configs, scripts, small installs | backed up; do **not** run jobs out of here |
| `$WORK` / workspace | 10–40 TiB | active job data, parallel I/O | fast (Lustre/BeeGFS/VAST); **no backups**, has expiry |
| `$TMPDIR` (job-local) | — | scratch *within* a single job | **deleted when the job ends** — copy results out |
| `$PROJECT` | 3 TiB | shared project data, larger installs/conda envs | project-wide |
| `$COLD` | 12 TiB | data not actively used by jobs | **NHR only** |
| Tape | 8 TiB | archive | NHR only; not a 10-year archive |
Rule of thumb: stage data into `$WORK`/`$TMPDIR` for the job, write results
there, then copy anything you want to keep to `$HOME` or `$PROJECT` before the
job (and thus `$TMPDIR`) disappears. For a series of jobs that share a fast
filesystem, allocate a **workspace** — and remember to name the data store
explicitly, since the listed default (`DONT_USE`) won't allocate:
```bash
ws_list -l # what stores are available on this node
ws_allocate -F lustre-rzg my_run 30 # 30-day workspace (lustre-rzg = fast on Grete/Emmy P3)
```
Details, the full `ws_*` lifecycle, and the per-account filesystem map are in
`references/storage.md`.
---
## Step 5 — Running jobs with Slurm
Work runs as **batch jobs** submitted with `sbatch`, or interactively with
`srun` / `salloc`. A job script is a shell script whose first lines are
`#SBATCH` directives. Those directives **must** sit at the very top — Slurm
stops reading them at the first real command.
### CPU batch job template
```bash
#!/bin/bash
#SBATCH --job-name=my-cpu-job
#SBATCH -p scc-cpu # partition — pick one your account can use (see master table)
#SBATCH -N 1 # nodes
#SBATCH -n 16 # tasks (use -c for threads-per-task instead, for OpenMP)
#SBATCH -t 12:00:00 # walltime hh:mm:ss — ALWAYS set this
#SBATCH -o slurm-%j.out # %j = job id
module purge
module load openmpi # whatever your code needs
srun ./my_program # srun for MPI; just call the program directly otherwise
```
### GPU batch job template (Grete / KISSKI)
```bash
#!/bin/bash
#SBATCH --job-name=train-gpu
#SBATCH -p grete:shared # shared GPU partition (or kisski, scc-gpu, ...)
#SBATCH -G A100:2 # request 2 GPUs; -G <type>:<count>
#SBATCH -t 05:00:00
#SBATCH --mail-type=all
module load miniforge3 gcc cuda
source activate my-env # your conda/miniforge environment
python -u train.py
```
GPU specifics worth getting right:
- **Request GPUs with `-G <type>:<count>`** (e.g. `-G A100:2`, `-G H100:1`).
- On a **shared** partition (`grete:shared`, `kisski`, `scc-gpu`, …) you choose
how many GPUs you want. On a **non-shared/exclusive** partition (`grete`,
`grete-h100`, …) you always get the whole node (4 GPUs) and are **billed for
all 4 regardless of how many you use** — so request multiples of 4 there, and
use a shared partition if you need fewer.
- Prefer **not** to hand-tune `--mem` / `-c`; Slurm assigns a fair share
proportional to the GPUs you requested. If you do set memory, remember usable
RAM is ~30 GiB less than the hardware figure (OS + ECC overhead).
- For quick debugging/interactive work, most accounts can use
**`grete:interactive`**, which serves **MIG GPU slices** requested as
`-G 1g.10gb` or `-G 2g.10gb`, etc.
### Submit and monitor
```bash
sbatch job.sh # submit; prints the job id
squeue --me # your queued/running jobs
squeue --start -j <jobid> # estimated start time
scontrol show job <jobid> # full detail (works after the job, too)
sinfo -p <partition> # node availability in a partition
scancel <jobid> # cancel a job
scancel -i --me # cancel all your jobs, asking about each
```
The complete parameter table, job arrays, interactive jobs, QoS / longer
walltimes, and enabling internet access inside a job are in `references/slurm.md`.
---
## Critical gotchas (the things that waste people's time)
- **Always set a walltime (`-t`).** Most partitions default to 12 h, and a job
that fails after running longer than 12 h is only refunded for 12 h. Request
close to the real runtime plus a buffer — shorter jobs also schedule sooner.
- **Exclusive GPU partitions bill for the whole node.** Use a `:shared`
partition when you need fewer than 4 GPUs.
- **Be a good neighbour on shared nodes.** Don't reserve all the RAM with a few
cores (or vice versa) — it strands the rest of the node for everyone else.
- **Compute nodes have no internet by default.** For jobs that must fetch
something, add `-C inet` / `--constraint=inet` (this routes http/https/ftp
through the GWDG proxy `http://www-cache.gwdg.de:3128`). Download data on a
login node beforehand when you can. For Hugging Face specifically, download on
a login node and run the job with `HF_HUB_OFFLINE=1` / `TRANSFORMERS_OFFLINE=1`
so a started job can't die on a network hiccup.
- **Keep Hugging Face caches off `$HOME`.** Point `HF_HOME` at `$PROJECT` or a
workspace — checkpoints blow past the 60 GiB / 7 M-file home quota fast. For
caching, offline use, fast downloads, and shared group weights, see
`references/huggingface_models.md`.
- **Short job? Use `--qos=2h`** (with `--time` ≤ 2 h) for a big priority boost —
but avoid flooding the scheduler with many sub-1h jobs (combine them or use a
job array).
- **`$TMPDIR` and workspaces are not permanent.** Copy results to `$HOME` /
`$PROJECT` before the job ends or the workspace expires. Workspaces have no backup.
- **Don't `module load` in `~/.bashrc`** — see Step 3.
---
## When details might be stale
The partition lists, hardware specs, quotas, and module versions in this skill
reflect the GWDG documentation as captured, but the cluster changes (new
hardware phases, retired partitions, adjusted quotas). When a *current* value
matters, prefer to confirm it live rather than asserting from memory:
- partitions / nodes → `sinfo`, `scontrol show partition <name>`
- a partition's default/max walltime → `scontrol show partition <name>` (look at `DefaultTime` / `MaxTime`)
- the user's storage and quotas → `show-quota`
- remaining compute-time budget → `sbalance`
- which workspace data stores exist on a node → `ws_list -l`
- available software → `module avail` / `module spider <name>`
The authoritative source is **<https://docs.hpc.gwdg.de/>**. If you can't verify
and aren't sure, say so rather than guessing a partition name or flag.
---
## Getting help
Point users at the right support address for their account type:
- **SCC / general HPC:** hpc-support@gwdg.de
- **NHR:** nhr-support@gwdg.de
- **KISSKI / other:** see the support page — <https://docs.hpc.gwdg.de/support/> —
for the address matching the account type.
---
## Reference map (read the file that matches the task)
- **`references/connecting.md`** — SSH config (simple + advanced, multi-account),
login-node DNS names and aliases, host-key fingerprints, legacy-SCC jumphost,
storage transfer nodes, VS Code / rsync over SSH, SSH troubleshooting.
- **`references/slurm.md`** — full `#SBATCH` parameter reference, all job-control
commands, walltime & QoS, shared-node etiquette, job arrays, interactive jobs,
internet-in-jobs, multiple-programs-per-node.
- **`references/partitions.md`** — complete CPU and GPU partition tables (by
account type *and* by hardware), CPU/GPU `--constraint` flags, MIG slices,
hardware totals, and how to choose a partition.
- **`references/storage.md`** — every data store, the environment variables,
`show-quota`, workspaces, job-local storage, per-account filesystems, and
moving data to/from non-HPC GWDG storage.
- **`references/software.md`** — Lmod in depth, the `gwdg-lmod` / `nhr-lmod` /
`scc-lmod` stacks, CUDA & the hierarchical module system, conda/miniforge,
Apptainer containers, MPI, and writing your own module files.
- **`references/huggingface_models.md`** — running Hugging Face on the cluster:
cache placement (`HF_HOME` off `$HOME`), offline use on compute nodes, fast
downloads (`hf_transfer`), and the **AG GIPP shared model weights** directory
on the SCC — conventions plus how to wire paths so models resolve automatically.