Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Alterlab Gtars

ASecurity

Runs high-performance genomic interval analysis with gtars (databio), a Rust toolkit with Python bindings — the performance-critical backend for the geniml ML library. Use when computing overlaps/jaccard/coverage between BED region sets, indexing intervals with IGD, generating uniwig accumulation/coverage tracks, tokenizing genomic regions for ML, splitting single-cell fragments into pseudobulks, or computing GA4GH refget sequence digests. NOT for training region embeddings (use alterlab-geni...

68 stars
0 votes
0 copies
0 views
Added 5/28/2026
ai-agentspythonrustgoshellbashapidatabasebackendperformancedocumentation

Works with

cliapi

Security Analysis

A96/100
mediumInstalls packages at runtime which could introduce malicious dependencies

Pro scans all 7 files and shows the line behind each finding

Scanned 9/23/2026

$npx -y skills add AlterLab-IEU/AlterLab-Academic-Skills --skill alterlab-gtars --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Alterlab Gtars?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Alterlab Gtars
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/alterlab-ieu-alterlab-gtars/badge)](https://www.skillsdirectory.com/skills/alterlab-ieu-alterlab-gtars)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: alterlab-gtars
description: Runs high-performance genomic interval analysis with gtars (databio), a Rust toolkit with Python bindings — the performance-critical backend for the geniml ML library. Use when computing overlaps/jaccard/coverage between BED region sets, indexing intervals with IGD, generating uniwig accumulation/coverage tracks, tokenizing genomic regions for ML, splitting single-cell fragments into pseudobulks, or computing GA4GH refget sequence digests. NOT for training region embeddings (use alterlab-geniml) or non-genomic spatial joins (use alterlab-geopandas). Part of the AlterLab Academic Skills suite.
license: MIT
allowed-tools: Read Write Edit Bash(uv:*) Bash(python:*) Bash(gtars:*) Bash(cargo:*)
compatibility: No API key required. Python API runs locally via `uv run python` with the `gtars` package (PyPI, verified 0.10.0). The CLI is a separate Rust binary (`gtars-cli` 0.10, install via cargo).
metadata:
    skill-author: AlterLab
    version: "1.1.1"
    last_updated: "2026-09-23"
---

# Gtars: Genomic Tools and Algorithms in Rust

## Overview

Gtars (from databio, the lab behind `geniml`) is a high-performance Rust toolkit for manipulating, analyzing, and processing genomic interval data. Its primary purpose is to be the performance-critical backend for `geniml`, a Python library for machine learning on genomic intervals. It provides overlap/set operations, IGD overlap indexing, coverage (uniwig) tracks, region tokenization for ML, single-cell fragment pseudobulking, and GA4GH refget sequence-collection management.

## When to Use This Skill

Use this skill when working with:
- Genomic interval files (BED) — overlaps, jaccard, set ops, coverage
- IGD indexing for fast overlap queries over large interval databases
- Coverage / accumulation tracks via uniwig
- Genomic ML preprocessing and region tokenization
- Single-cell fragment files (split into pseudobulks by cluster)
- Reference sequence digests and retrieval (refget)

### Does NOT Trigger

| Scenario | Use Instead |
|----------|-------------|
| Training region / single-cell embeddings (Region2Vec, scEmbed, BEDspace) | `alterlab-geniml` |
| Per-read BAM/CRAM/VCF access (CIGAR, MAPQ, pileups) | `alterlab-pysam` |
| Normalized bigWig coverage from BAM, TSS heatmaps/profiles | `alterlab-deeptools` |
| Geographic (non-genomic) spatial joins and overlaps | `alterlab-geopandas` |

> Version note: examples are verified against the **`gtars` Python package 0.10.0** (PyPI, 2026-09; first written for 0.8 — the calls below are unchanged). The Python API is exposed through submodules — `gtars.models`, `gtars.tokenizers`, `gtars.refget`, `gtars.utils`, plus the newer `gtars.genomic_distributions`, `gtars.lola`, `gtars.vrs` — NOT as flat top-level functions. There is no `gtars.igd` or `gtars.uniwig` Python submodule; IGD building and uniwig track generation are CLI-only.

## Installation

### Python package

```bash
uv pip install gtars   # or: uv add gtars
```

Import surface (verified, 0.10.0):

```python
from gtars.models import RegionSet, Region, RegionSetList
from gtars.tokenizers import Tokenizer, tokenize_fragment_file
from gtars import refget          # RefgetStore, digest_fasta, sha512t24u_digest, ...
from gtars import utils           # read/write .gtok token files
```

### CLI (separate Rust binary)

The CLI ships as the `gtars-cli` crate (binary name `gtars`) and is installed with Cargo. Most subcommands are behind feature flags:

```bash
# All commonly used commands
cargo install gtars-cli --features "uniwig overlaprs igd bbcache scoring fragsplit genomicdist"

# Or a subset
cargo install gtars-cli --features "uniwig igd"
```

Available CLI subcommands: `igd`, `overlaprs`, `uniwig`, `bbcache`, `pb` (fragment pseudobulking), `scoring`, `genomicdist`, `ranges`, `consensus`, `prep`. Flag sets differ per subcommand and evolve across versions — always confirm with `gtars <command> --help`.

## Core Capabilities

Gtars is organized into specialized modules, each focused on specific genomic analysis tasks:

### 1. Overlap Detection and Set Operations

Detect overlaps and compute set operations / similarity between region sets with `RegionSet` (Python), or index a large interval database with IGD (CLI).

**When to use:**
- Finding overlapping regulatory elements, comparing ChIP-seq peaks
- Variant annotation; identifying shared genomic features
- Jaccard / overlap-coefficient similarity between BED files

**Quick example (Python):**
```python
from gtars.models import RegionSet

peaks = RegionSet("chip_peaks.bed")
promoters = RegionSet("promoters.bed")

# Regions in peaks that overlap a promoter (the real method is subset_by_overlaps)
in_promoters = peaks.subset_by_overlaps(promoters)
in_promoters.to_bed("peaks_in_promoters.bed")

print(peaks.count_overlaps(promoters))  # per-region overlap counts
print(peaks.jaccard(promoters))         # similarity score
```

For querying a large reference database many times, build an IGD index once via the CLI (`gtars igd create ...`) and search it (`gtars igd search ...`). See `references/overlap.md`.

### 2. Coverage / Accumulation Tracks (uniwig, CLI)

Generate coverage / accumulation tracks from a BED or BAM file with the uniwig CLI subcommand.

**When to use:**
- ATAC-seq accessibility profiles, ChIP-seq coverage, RNA-seq read coverage

**Quick example (CLI):**
```bash
# uniwig reads a sorted BED/BAM and writes accumulation tracks.
# Flags differ by version; confirm with `gtars uniwig --help`.
gtars uniwig --file fragments.bed --filetype bed \
             --fileheader coverage --outputtype bw
```

See `references/coverage.md` for verified flags and `RegionSet.coverage()` for an in-memory alternative.

### 3. Genomic Tokenization

Convert genomic regions into discrete tokens for ML (the preprocessing layer `geniml` builds on).

**When to use:**
- Preprocessing peaks/regions into a fixed vocabulary for genomic ML models
- Feeding token IDs to geniml or custom transformer models

**Quick example (Python):**
```python
from gtars.tokenizers import Tokenizer
from gtars.models import Region

tokenizer = Tokenizer.from_bed("universe.bed")   # vocab = the universe BED
tokens = tokenizer.tokenize([Region("chr1", 1000, 2000, None)])  # -> ['chr1:1000-2000']
ids = tokenizer.convert_tokens_to_ids(tokens)                     # -> [<int>]
```

See `references/tokenizers.md`. Note: the class is `Tokenizer` (there is no `TreeTokenizer`).

### 4. Reference Sequence Management (refget)

Compute GA4GH refget digests and manage/retrieve reference sequences.

**When to use:**
- Validating reference genome integrity via sequence digests
- Building a local sequence-collection store and extracting subsequences

**Quick example (Python):**
```python
from gtars import refget

# Digest a FASTA into a GA4GH SequenceCollection (no sequence data loaded)
collection = refget.digest_fasta("hg38.fa")

# Or a one-off sequence digest
d = refget.sha512t24u_digest("ACGTACGT")   # 32-char base64url GA4GH sha512t24u digest
```

See `references/refget.md` for `RefgetStore` (load, store, and `get_substring`).

### 5. Fragment Pseudobulking (pb, CLI)

Split a single-cell fragment file into pseudobulks based on a cluster/cell-group mapping.

**When to use:**
- Processing single-cell ATAC-seq; cluster-based fragment aggregation

**Quick example (CLI):**
```bash
# The fragsplit feature exposes the `pb` (pseudobulk) subcommand.
gtars pb --fragments fragments.bed.gz --mapping cluster_mapping.tsv
```

The Python side also offers `gtars.tokenizers.tokenize_fragment_file(...)` for tokenizing fragments directly. See `references/cli.md`.

### 6. Fragment / Region Scoring

Score region/fragment files against reference datasets with the `scoring` CLI subcommand.

**When to use:**
- Evaluating enrichment of regions against a reference universe
- Batch quality-metric computation across samples

```bash
gtars scoring --help   # confirm subcommands and flags for your installed version
```

## Common Workflows

### Workflow 1: Peak Overlap Analysis

Identify peaks overlapping promoters (Python):

```python
from gtars.models import RegionSet

peaks = RegionSet("chip_peaks.bed")
promoters = RegionSet("promoters.bed")

overlapping_peaks = peaks.subset_by_overlaps(promoters)
overlapping_peaks.to_bed("peaks_in_promoters.bed")

# Iterate results (regions expose .chr/.start/.end)
for r in overlapping_peaks:
    print(r.chr, r.start, r.end)
```

### Workflow 2: IGD index + repeated overlap queries (CLI)

Index a large reference database once, then search it many times:

```bash
# Build the IGD database from a directory or list of BED files
gtars igd create --help     # confirm the exact input/output flags for your version

# Search the database with query regions
gtars igd search --help
```

### Workflow 3: ML Preprocessing (tokenization)

Prepare genomic regions for an ML model:

```python
from gtars.tokenizers import Tokenizer
from gtars.models import RegionSet

# Step 1: Build a tokenizer from the universe BED (defines the vocabulary)
tokenizer = Tokenizer.from_bed("universe.bed")

# Step 2: Tokenize a region set (tokenize accepts a RegionSet or a list of Region)
regions = RegionSet("training_peaks.bed")
tokens = tokenizer.tokenize(regions)                      # list of 'chr:start-end' strings
ids = tokenizer.convert_tokens_to_ids(tokens)             # integer IDs for the model

# Step 3: feed `ids` to geniml or a custom model (see alterlab-geniml for training)
```

## Python vs CLI Usage

**Use Python API when:**
- Integrating with analysis pipelines
- Need programmatic control
- Working with NumPy/Pandas
- Building custom workflows

**Use CLI when:**
- Quick one-off analyses
- Shell scripting
- Batch processing files
- Prototyping workflows

## Reference Documentation

- **`references/python-api.md`** — `RegionSet` / `Region` operations, set ops, overlaps, export
- **`references/overlap.md`** — overlap detection and IGD indexing
- **`references/coverage.md`** — uniwig coverage tracks
- **`references/tokenizers.md`** — region tokenization for ML
- **`references/refget.md`** — refget digests and `RefgetStore`
- **`references/cli.md`** — CLI subcommand overview

## Relationship to geniml

gtars is the Rust performance backend for `geniml` (databio's ML-on-genomic-intervals library). Use gtars for the heavy interval ops and tokenization; use `geniml` (skill `alterlab-geniml`) for embedding/model training built on top of those tokens.

## Data Formats

- **BED**: genomic intervals (3-column or extended) — the core input
- **BigWig / WIG**: coverage / accumulation tracks (uniwig output)
- **FASTA**: reference sequences (refget)
- **Fragment files**: single-cell fragments (often `.bed.gz`), with a separate cluster-mapping file for `pb`

## Verifying the API

This package's surface differs between releases and the Python and CLI APIs are NOT mirror images. Before relying on an unfamiliar method or flag, confirm against the installed version:

```bash
uv run python -c "from gtars import models, tokenizers, refget; print(dir(models), dir(tokenizers), dir(refget))"
gtars <command> --help
```

Part of the AlterLab Academic Skills suite.

Attribution

AlterLab-IEUAlterLab-IEU
View sourceSee grades on GitHubMore from AlterLab-IEU →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Terse caveman voice: answer first, fluff gone, every technical fact kept. Use for /caveman, "caveman mode", "talk like caveman", "be brief", "less tokens". Stays on until "stop caveman" or "normal mode".

1100021 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

698621 votes

Writing Skills

Create and manage Claude Code skills in HASH repository following Anthropic best practices. Use when creating new skills, modifying skill-rules.json, understanding trigger patterns, working with hooks, debugging skill activation, or implementing progressive disclosure. Covers skill structure, YAML frontmatter, trigger types (keywords, intent patterns), UserPromptSubmit hook, and the 500-line rule. Includes validation and debugging with SKILL_DEBUG. Examples include rust-error-stack, cargo-dep...

3931 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3421 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Amp, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Grok Build, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

741 votes
View all in ai-agents →