Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Implement And Scale Bioinformatics Pipeline

ASecurity

Implement a designed genomics workflow in the chosen engine, containerize and pin every tool for reproducibility, scale it with scatter/gather on HPC Slurm or cloud Batch (spot on the fault-tolerant steps), and validate it against a GIAB/GA4GH hap.py truth set — then produce a pipeline-validation report. Reach for this when the user asks "build this pipeline in Nextflow/Snakemake/WDL", "containerize and pin it for reproducibility", "scale/cost-optimize this on Slurm or cloud", or "benchmark o...

7 stars
0 votes
0 copies
0 views
Added 9/23/2026
ai-agentsrustgonoderailsdockeraws

Security Analysis

A100/100

Scanned 9/23/2026

$npx -y skills add mcorbett51090/RavenClaude --skill implement-and-scale-bioinformatics-pipeline --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Implement And Scale Bioinformatics Pipeline?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Implement And Scale Bioinformatics Pipeline
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/mcorbett51090-implement-and-scale-bioinformatics-pipeline/badge)](https://www.skillsdirectory.com/skills/mcorbett51090-implement-and-scale-bioinformatics-pipeline)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: implement-and-scale-bioinformatics-pipeline
description: Implement a designed genomics workflow in the chosen engine, containerize and pin every tool for reproducibility, scale it with scatter/gather on HPC Slurm or cloud Batch (spot on the fault-tolerant steps), and validate it against a GIAB/GA4GH hap.py truth set — then produce a pipeline-validation report. Reach for this when the user asks "build this pipeline in Nextflow/Snakemake/WDL", "containerize and pin it for reproducibility", "scale/cost-optimize this on Slurm or cloud", or "benchmark our variant calls against GIAB". Used by `genomics-pipeline-engineer` (primary).
---

# Skill: implement-and-scale-bioinformatics-pipeline

> **Invoked by:** `genomics-pipeline-engineer` (primary). Also consulted by `bioinformatics-workflow-architect` to confirm the chosen engine/compute can actually run the workflow before finalizing the architecture.
>
> **When to invoke:** "Build the pipeline in <engine>"; "containerize + pin it"; "scale / cost-optimize on Slurm or cloud"; "benchmark against GIAB/hap.py"; any move from a designed workflow to a running, reproducible, validated pipeline.
>
> **Output:** the implemented workflow (per-step processes, containers, resource labels, params/config), the reproducibility manifest (pinned versions + lockfile + provenance), the scaling/cost plan, and a truth-set concordance report captured in the validation-report template.

## Procedure

1. **Start from the analysis plan.** Take the step graph, reference build, accessory files, and truth set from [`design-genomics-analysis-workflow`](../design-genomics-analysis-workflow/SKILL.md) / [`../../templates/analysis-plan-spec.md`](../../templates/analysis-plan-spec.md). Don't implement against a vague ask — an unspecified step graph produces an unmaintainable pipeline.
2. **Implement one process per step.** Each step (QC, trim, align, dedup, BQSR, call, joint-genotype / quantify) is a discrete process/rule with its own **pinned container**, a **resource label** (CPU/RAM/time), and typed inputs/outputs. Never a mega-script — the process graph is the unit of resume and reuse. Add a **MultiQC** roll-up at the end.
3. **Containerize + pin for reproducibility.** Per-process containers (Docker in dev → **Apptainer/Singularity** on HPC) or locked Conda/Bioconda envs; every tool at an **exact version**; commit the **lockfile** and a params file. Record the reference build + accessory-file **checksums**. Capture provenance (engine `-with-trace/-with-report` / Snakemake report, seeds).
4. **Scale with scatter/gather + right-sized resources.** Scatter per-sample and per genomic interval for alignment/calling; gather at joint-genotyping. Right-size resources from **profiling a representative sample**, not guesses. Use the engine's **resume** so a failure restarts from the last good step.
5. **Cost-optimize the compute placement.** On cloud, put the **short, fault-tolerant, checkpointed** steps on **spot/preemptible** with retries; keep **long single-shot** steps (a big joint-genotype gather) on-demand. Prefer CRAM over BAM, delete intermediates, keep data close to compute (watch egress). Record a **before/after runtime + cost**.
6. **Validate against the truth set.** For variant calling, run **GA4GH hap.py** against the **GIAB** truth VCF + confident-region BED and report **precision / recall / F1 by variant type (SNV/indel)** *in the confident regions*. For RNA/single-cell, check the expected QC signatures + spike-ins/markers. Regression-gate: a tool-version bump that degrades concordance must fail here, not in production.
7. **Capture it** in [`../../templates/pipeline-validation-report.md`](../../templates/pipeline-validation-report.md) — the engine/version, the pinned environment, the scaling/cost result, and the concordance numbers in one auditable page.

## Worked example

> User: "Build the WGS germline pipeline we designed in Nextflow, make it reproducible, run it on AWS Batch cheaply, and prove it works."

- **Implement:** DSL2 processes — `fastqc` → `multiqc`, `fastp`, `bwamem2_align`, `markduplicates`, `bqsr`, `haplotypecaller` (scattered over interval BED, gVCF), `genomicsdb_import` + `genotypegvcfs` (gather), `vqsr`, `vep`. Each with a pinned container and a `resourceLabel`.
- **Reproducibility:** each process pins its Biocontainer to an exact tag; a `nextflow.config` with a params file; `-with-report -with-trace -with-timeline` for provenance; reference (GRCh38 analysis set) + dbSNP checksums recorded; the container digests committed.
- **Scale on AWS Batch:** the AWS Batch executor; **scatter** HaplotypeCaller over ~can-be-many interval shards; the alignment shards + per-interval calling run on **spot** with `errorStrategy 'retry'` (reclaim = a cheap retry); the single-shot joint-genotype gather runs **on-demand**. CRAM output, intermediates cleaned. Before/after: record the spot-vs-on-demand cost delta.
- **Validate:** run the pipeline on **GIAB HG002**, then `hap.py calls.vcf.gz` against the HG002 truth VCF + confident BED → report SNV and **indel** precision/recall/F1 in the confident regions. Gate the pipeline on meeting the target before it's trusted for real samples.
- **Report:** all of it in the pipeline-validation report — engine version, pinned digests, scatter config, spot/cost result, and the hap.py concordance table.

## Guardrails

- Implement against a **captured analysis plan**, never a vague ask.
- **One process per step, one pinned container per process** — mega-scripts kill resume, reuse, and reproducibility.
- **Pin every version + commit a lockfile + record reference checksums** — a floating solve or a `latest` tag is a non-reproducible result.
- Scatter/gather + **right-size from profiling** is the primary scaling lever, not bigger nodes.
- **Spot only on fault-tolerant, retryable steps**; keep long single-shot steps on-demand.
- **Validate against GIAB/hap.py before claiming accuracy** — report precision/recall/F1 by variant type in the confident regions, never "looks fine."
- Never mix reference builds/coordinates — the build + its matching accessory files are one set.
- Volatile facts (tool versions, GIAB truth-set versions, cloud pricing) carry a **retrieval date**; re-verify before shipping. See [`../../knowledge/genomics-workflow-patterns-2026.md`](../../knowledge/genomics-workflow-patterns-2026.md).

Attribution

mcorbett51090mcorbett51090
View sourceSee grades on GitHubMore from mcorbett51090 →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Terse caveman voice: answer first, fluff gone, every technical fact kept. Use for /caveman, "caveman mode", "talk like caveman", "be brief", "less tokens". Stays on until "stop caveman" or "normal mode".

1100021 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

698461 votes

Writing Skills

Create and manage Claude Code skills in HASH repository following Anthropic best practices. Use when creating new skills, modifying skill-rules.json, understanding trigger patterns, working with hooks, debugging skill activation, or implementing progressive disclosure. Covers skill structure, YAML frontmatter, trigger types (keywords, intent patterns), UserPromptSubmit hook, and the 500-line rule. Includes validation and debugging with SKILL_DEBUG. Examples include rust-error-stack, cargo-dep...

3931 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3421 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Amp, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Grok Build, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

741 votes
View all in ai-agents →