Skip to content
Back to skills

Genome Assembly

ASecurity

Genome assembly fundamentals — long-read assembly, polishing, QC metrics, and annotation.

  • 2 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 29, 2026
ai-agentsgobash

Security analysis

A100/100

Scanned September 29, 2026

npx -y skills add aicodedecode/awesome-muse-skills --skill genome-assembly --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Genome Assembly?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Genome Assembly
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/aicodedecode-genome-assembly/badge)](https://www.skillsdirectory.com/skills/aicodedecode-genome-assembly)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: genome-assembly
description: Genome assembly fundamentals — long-read assembly, polishing, QC metrics, and annotation.
category: scientific
---

## Overview

genome-assembly covers de novo genome assembly: data requirements, long-read assemblers
(HiFi and ONT), assembly QC (N50, BUSCO, k-mer completeness), scaffolding (Hi-C), polishing,
and annotation. Assembly quality determines every downstream analysis — a fragmented,
contaminated assembly poisons variant calling, annotation, and comparative genomics.

## When to use

- Planning sequencing: coverage, read type (HiFi vs ONT vs short-read), Hi-C needs.
- Assembling: hifiasm, Flye, Verkko — choosing by data type.
- QC: N50/L50, BUSCO, Merqury k-mer spectra, contamination screening.
- Scaffolding: Hi-C (YaHS, SALSA), optical maps.
- Polishing: when needed, with what reads.
- Haplotype-resolved assembly: trio, Hi-C phasing.
- Annotation: BRAKER, liftoff, functional annotation basics.

## Core concepts

- **Read type determines outcome.** PacBio HiFi (99.9% accurate, 15-20 kb): the current
  gold standard — accurate enough to skip polishing. ONT (100+ kb reads, ~99% accurate):
  best for the hardest repeats, needs polishing. Short reads alone: fragmented assemblies
  of complex genomes — acceptable for bacteria, inadequate for most eukaryotes.
- **Coverage.** 30x HiFi for a good human-scale assembly; 60x+ for the hardest regions;
  ONT 30-50x. More coverage helps to a point — beyond ~60x HiFi, returns diminish and
  compute costs balloon.
- **Assemblers.** hifiasm (HiFi — best-in-class for phased assembly), Flye (ONT and
  metagenomic), Verkko (HiFi + ONT ultra-long integration, telomere-to-telomere attempts).
  Match assembler to data; don't run a short-read assembler on long reads.
- **Haplotype resolution.** Diploid genomes: collapsed (mosaic), primary/alternate, or
  fully phased (trio data or Hi-C). Collapsed assemblies create false duplications and
  break variant calling — phase when the biology needs it (heterozygous organisms always
  benefit).
- **QC metrics.** N50/L50 (contiguity — N50 of 50 Mb vs 50 kb is a different universe);
  BUSCO (gene completeness vs lineage expectations — >95% for good assemblies);
  Merqury (k-mer completeness and consensus quality QV — reference-free truth);
  assembly size vs expected genome size (too big = uncollapsed haplotypes/contamination,
  too small = missing sequence). Check all four, not just N50.
- **Contamination screening.** BlobTools (coverage vs GC vs taxonomy): symbionts, food,
  and lab contaminants assemble alongside your organism. Screen before publishing —
  contaminated "genome papers" are embarrassing and common.
- **Scaffolding.** Hi-C contact data orders contigs into chromosomes (YaHS); validate with
  contact maps (suspicious joins show as off-diagonal breaks). Scaffolding doesn't fix a bad
  contig assembly — it organizes a good one.
- **Polishing.** Needed for ONT (not HiFi): Racon/Medaka with long reads, then short-read
  polishing (Pilon/NextPolish) for residual errors. Over-polishing a good assembly can
  introduce errors — verify with Merqury QV before/after.
- **Annotation.** BRAKER (RNA-seq + protein evidence → gene models); liftoff (transfer
  annotation from a related genome — fast, reference-biased); functional annotation
  (InterProScan, eggNOG). Annotation quality limits every downstream analysis — budget
  real effort here.

## Practical workflow

1. **Plan.** Genome size estimate (flow cytometry/k-mers), heterozygosity, repeat content →
   read type + coverage + Hi-C decision.
2. **QC reads.** Read length/N50 distributions, quality; screen for contamination early.
3. **Assemble.** hifiasm (HiFi) / Flye (ONT) with appropriate parameters; try defaults
   first — they're good now.
4. **QC assembly.** N50, BUSCO, Merqury QV/completeness, size sanity, BlobTools
   decontamination.
5. **Scaffold (if Hi-C).** YaHS; inspect contact maps; break misjoins.
6. **Polish (if ONT).** Long-read then short-read; verify QV improvement.
7. **Annotate.** BRAKER with RNA-seq evidence; functional annotation; validate gene count
   vs relatives.
8. **Deposit.** INSDC submission with metadata; assembly QC report alongside.

Example command sketch:
```bash
hifiasm -o asm -t 32 --h1 r1.fq --h2 r2.fq hifi.fq   # trio-phased
busco -i asm.bp.p_ctg.fa -l mammalia_odb10 -m genome
merqury.sh reads.meryl asm.bp.p_ctg.fa merqury_out/
```

## Common pitfalls

- Short-read-only assembly of a complex eukaryote (hopeless fragmentation).
- N50 worship without BUSCO/Merqury/contamination checks.
- Uncollapsed haplotypes doubling the assembly size.
- Contamination published as novel sequence (no BlobTools screening).
- Scaffolding a bad contig assembly (garbage organized into chromosomes).
- Over-polishing degrading a good HiFi assembly.
- Annotation as an afterthought (bad gene models poisoning downstream work).

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…