Skip to content
Back to skills

Genomics Pro

ASecurity

Genome-scale analysis — variant calling, annotation, GWAS, and structural variation for research and clinical genomics.

  • 2 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added September 29, 2026
ai-agentsgobashtestingdatabase

Works with

  • cli

Security analysis

A100/100

Scanned September 29, 2026

npx -y skills add aicodedecode/awesome-muse-skills --skill genomics-pro --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Genomics Pro?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Genomics Pro
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/aicodedecode-genomics-pro/badge)](https://www.skillsdirectory.com/skills/aicodedecode-genomics-pro)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: genomics-pro
description: Genome-scale analysis — variant calling, annotation, GWAS, and structural variation for research and clinical genomics.
category: scientific
---

## Overview

genomics-pro covers analysis at genome scale: whole-genome and whole-exome sequencing, variant
discovery and interpretation, population-scale association studies, and structural variation. Where
bioinformatics-pro handles the general NGS mechanics, this skill goes deep on genomic questions —
what variants exist, what they do, and how they associate with traits or disease.

It spans germline and somatic analysis, single-sample and cohort workflows, and the annotation
resources (ClinVar, gnomAD, COSMIC, dbSNP) that turn a VCF of millions of variants into a shortlist
worth investigating.

## When to use

- Designing a WGS/WES study: coverage targets (30x germline WGS, 100x+ tumor), capture kits,
  trio vs singleton design.
- Germline variant calling: GATK best practices, joint genotyping across cohorts.
- Somatic analysis: tumor-normal pairs, mutational signatures, copy-number and LOH.
- Variant annotation and filtering: population frequency, predicted consequence, clinical databases.
- GWAS: QC, association testing, fine-mapping, polygenic scores.
- Structural variants and CNVs: detection from short reads, long-read validation.
- Interpreting a clinical-grade report: ACMG classification, incidental findings policy.

## Core concepts

- **Germline vs somatic.** Germline variants are inherited, ~diploid, and called per-sample then
  joint-genotyped; somatic variants are acquired, often subclonal, and need tumor-normal comparison
  with allele-frequency-aware callers (Mutect2, Strelka2). Mixing the two workflows corrupts both.
- **The VCF contract.** CHROM/POS/REF/ALT plus genotype fields (GT, AD, DP, GQ, PL). Understand
  normalization (left-alignment, parsimony) — the same variant can be written multiple ways, which
  breaks naive comparisons; always normalize before intersecting callsets.
- **Filtering strategy.** Common variants are rarely causal for rare disease: filter by population
  allele frequency (gnomAD AF, typically <1% or <0.1% for rare disease), consequence (VEP: stop-gain,
  frameshift, missense with CADD/REVEL), inheritance model, and phenotype match (HPO terms).
- **Cohort QC for GWAS.** Sample call rate, heterozygosity outliers, sex checks, relatedness
  (IBD/kinship), population stratification (PCA against 1000 Genomes, include PCs as covariates).
  Genomic inflation factor λ should be near 1; λ >> 1 means uncontrolled structure.
- **Association testing.** Linear/logistic regression per variant with covariates, or mixed models
  (BOLT-LMM, SAIGE) for relatedness and case-control imbalance. Genome-wide significance: p < 5e-8.
- **Structural variation.** Deletions, duplications, inversions, translocations — detected via
  discordant pairs, split reads, and read depth (Manta, Delly, LUMPY). Short reads miss many SVs;
  long reads (PacBio/ONT) are the gold standard for complex regions.
- **ACMG variant classification.** Pathogenic / likely pathogenic / VUS / likely benign / benign —
  based on population, computational, functional, segregation, and de novo evidence. A VUS is not a
  diagnosis; report it as uncertainty, not a finding.
- **Reference matters more here.** GRCh38 vs T2T-CHM13 changes variant coordinates and resolves
  previously "dark" regions. Liftover between builds is lossy — realign when it matters.

## Practical workflow

1. **Design.** Trio WES for rare disease (de novo detection), 30x WGS for comprehensive SVs,
   tumor-normal 100x/30x for somatic. Record capture kit and reference build.
2. **Process.** Align → MarkDuplicates → BQSR → HaplotypeCaller (germline) or Mutect2 (somatic).
   For cohorts: joint genotyping with GenomicsDBImport + GenotypeGVCFs.
3. **QC.** Ti/Tv ratio (~2.0-2.1 WGS, ~3.0 WES), het/hom ratio, call rate, contamination
   (VerifyBamID), sex concordance. Outliers get investigated, not deleted silently.
4. **Annotate.** VEP with gnomAD frequencies, ClinVar, CADD/REVEL, splice predictors (SpliceAI).
   Keep the full annotated VCF; filtering is a view, not a deletion.
5. **Filter/interpret.** Frequency → consequence → inheritance → phenotype → literature. For GWAS:
   association → clumping/fine-mapping → colocalization with eQTLs → polygenic scoring.
6. **Validate.** Sanger or long-read confirmation for clinical-grade calls; replication cohort for
   GWAS hits. Report with ACMG terms and explicit limitations.

Example command sketch:
```bash
gatk HaplotypeCaller -R ref.fa -I sample.bam -O sample.g.vcf.gz -ERC GVCF
gatk GenomicsDBImport --genomicsdb-workspace-path db -V s1.g.vcf.gz -V s2.g.vcf.gz -L chr1
vep -i variants.vcf -o annotated.vcf --cache --af_gnomad --plugin CADD
plink2 --bfile cohort --glm --covar pcs.txt --out assoc
```

## Common pitfalls

- Calling somatic variants without a matched normal — germline contamination of the "somatic" list.
- Comparing VCFs without normalization; same variant, different representation.
- Population stratification in GWAS mistaken for signal (always plot the QQ plot).
- Over-interpreting VUS or in-silico predictors as clinical evidence.
- Ignoring sex chromosomes and mitochondrial DNA in "genome-wide" analyses.
- Using hg19 coordinates with GRCh38 annotations (or vice versa).
- Treating a polygenic risk score as diagnostic — it is probabilistic and ancestry-sensitive.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…