
Claude Skills by HolobiomicsLab
github.com/HolobiomicsLabUse when after executing the Juicer pipeline on raw Hi-C FASTQ files, to confirm that the pipeline has generated the expected .hic output artifact and that the contact matrix construction and normalization steps completed without error.
Use when before running HiC-Pro's normalization stage on aligned Hi-C BAM files. Specifically, when you have SAM/BAM-formatted aligned Hi-C reads that need bias correction and matrix balancing to produce normalized contact maps suitable for downstream chromatin structure analysis.
Use when after generating raw Hi-C contact matrices from aligned reads (post-merge, pre-analysis).
Use when you have completed Hi-C map generation (producing .hic files from aligned reads) and need to detect and annotate topological features such as chromatin loops, topologically associating domains (TADs), or interaction peaks.
Use when after running the ENCODE Hi-C uniform processing pipeline or Juicer on FASTQ input data and generating a .hic output file.
Use when you have a methylBase object containing aligned methylation calls across multiple samples and need to verify whether samples cluster by expected experimental condition (e.g., test vs. control) or identify unexpected sample relationships.
Use when when installing HiC-Pro on a shared HPC cluster or multi-node computing environment where job submission must be routed through a scheduler rather than running locally.
Use when you have raw .idat files or a beta-valued matrix from an Illumina HumanMethylation450 or EPIC array experiment and need to import the full probe set into R for downstream quality control, normalization, and differential methylation analysis.
Use when you have raw Illumina EPIC or 450k methylation array data (.idat files or beta-valued matrices) and need to perform comprehensive quality assessment, probe correction, batch effect adjustment, and identification of differentially methylated regions or blocks across sample groups.
Use when your input is raw .idat files or a beta-valued matrix from Illumina HumanMethylation450 or EPIC arrays, and you need to remove unreliable probes (those with detection p-value > 0.01 or insufficient bead counts) before performing differential methylation or other downstream analyses.
Use when after running ChAMP detection functions (champ.
Use when when you have aligned paired scATAC-seq and scRNA-seq data from the same cells (multiome data) and need to create a single reduced-dimension coordinate space that integrates both chromatin accessibility and gene expression signals for joint clustering, trajectory analysis, or visualization.
Use when you have a pre-generated .hic contact map file (from Juicer pipeline or external source) and need to systematically call chromatin loops, detect topologically associating domains, or annotate other structural features without re-running the full alignment and contact matrix construction.
Use when you have raw Hi-C FASTQ files from a high-throughput chromatin conformation capture experiment and need to generate a normalized contact matrix (.hic file) for downstream genomic analysis.
Use when you have paired-end Hi-C FASTQ files from a public repository (NCBI SRA, GEO, or ENCODE-deposited) and need to produce standardized .hic binary contact maps that conform to ENCODE reference formats and integrity standards for downstream 3D genome analysis.
Use when you have raw Hi-C FASTQ data and need to generate contact maps at kilobase resolution, or you have pre-generated .hic files and need to annotate structural features (loops, domains) for downstream 3D genome analysis.
Use when you have filtered peak counts from ATAC or DNase-seq data (with GC bias correction and sample/peak filtering applied) and want to annotate peaks by k-mer content rather than known transcription factor motifs—particularly when comparing how k-mer size affects the magnitude of chromatin.
Use when after computing deviations for both motif and kmer annotations on the same chromVAR dataset, when you need to determine whether kmers and motifs are redundant predictors of chromatin accessibility variability or provide complementary information for downstream clustering, annotation, or.
Use when you have a single-cell count matrix with 10 million or more cells that must be processed through dimension reduction, clustering, or integration pipelines. Use it specifically before executing matrix-free spectral embedding (tl.
Use when you have performed spectral dimension reduction on single-cell omics count matrices and wish to partition cells into discrete populations.
Use when you are building or extending a multi-module Python library for scientific computation (e.
Use when you have paired ChIP and control BED/BEDPE files and need to account for local sequencing bias before peak calling. Use it specifically when control signal varies across genomic regions at multiple spatial scales (e.
Use when when performing ChIP-Seq peak calling with MACS3, after duplicate filtering and fragment length prediction (d), to construct the background model that will be compared against ChIP signal.
Use when when you have aligned ChIP-Seq reads (BED or BEDPE format) and a corresponding control sample, and you need explicit control over peak-calling parameters—including fragment-length prediction, local bias windows (d, slocal=1kb, llocal=10kb), background scaling, and score-cutoff.
Use when when deploying a complex bioinformatics pipeline (e.g., HiC-Pro) that depends on multiple external tools with version constraints (samtools ≥1.9, bowtie2, R packages, Python libraries) and you need to verify their availability and configure their paths before running the analysis.
Use when when benchmarking or validating the scalability of single-cell algorithms that claim linear or sublinear space complexity, particularly when processing datasets with ≥10 million cells.
Use when you have extracted clustering or classification accuracy metrics (NMI, ARI, purity scores) for two or more competing methods evaluated on multiple datasets, and need to determine which method performs overall rather than on individual datasets alone.
Use when you have loaded a normalized beta-valued methylation matrix (e.
Use when after merging methylation call files across all samples into a unified methylBase object (via unite()), when you need to assess whether biological replicates cluster together, identify unexpected sample groupings, or visualize global methylation similarity relationships before proceeding.
Use when after calculating differential methylation across samples using calculateDiffMeth(), when you need to separately enumerate and extract hyper-methylated (increased methylation) versus hypo-methylated (decreased methylation) bases that meet both statistical significance (q-value < 0.
Use when you have loaded individual methylation call files as methylRawList objects from bisulfite sequencing experiments (via methRead()) and need to perform base-level comparative analysis across two or more samples.
Use when immediately after loading raw methylation array data using champ.load() or champ.import() to verify data integrity.
Use when after identifying differentially methylated bases or regions (via calculateDiffMeth() and getMethylDiff()), when you need to characterize WHERE these methylation changes occur relative to gene structure and CpG density landscapes.
Use when you have loaded normalized methylation beta-value matrices from Illumina EPIC or 450k arrays and need to move beyond single-CpG differential methylation testing to identify multi-CpG regions with coordinated differential methylation signals.
Use when after reading in per-sample methylation call files with methRead() and obtaining methylRawList objects, but before calculating differential methylation or performing annotation.
Use when when analyzing DNA methylation data from bisulfite sequencing (RRBS, target-capture, or whole-genome) and the dataset is too large to fit comfortably in memory, or when you need to process multiple large samples sequentially without reloading data.
Use when before invoking any Python module in a multi-step Hi-C processing pipeline, or when a dependency has been freshly installed or reinstalled.
Use when you have an ArchR project object with dimensionality reduction results (LSI or combined dimensions from scATAC-seq ± scRNA-seq) and want to infer pseudotime trajectories and cell-state transitions.
Use when you have a chromVARDeviations object with multiple annotation sets (such as JASPAR motifs and kmers) and need to determine which annotation pairs are redundant (high correlation) versus synergistic (high synergy z-scores).
Use when you have a set of differentially accessible peaks (output from differential accessibility testing, e.g., tl.
Use when after identifying a set of differentially accessible peaks (via tl.diff_test or equivalent), when you need to infer which transcription factors may regulate the observed chromatin state changes.
Use when you have a filtered set of non-overlapping peaks from ATAC-seq data and a collection of motifs (typically from JASPAR or similar databases), and you need to identify which peaks contain matches to which motifs as a prerequisite for computing motif-based deviation scores across samples.
Use when you have bias-corrected ATAC-seq footprint signals (BigWig files) from two or more distinct conditions (e.
Use when you have independently generated or received both scATAC-seq peak count matrices and scRNA-seq gene expression matrices from the same set of cells (multiome experiment), and you need to perform joint analysis such as co-clustering, trajectory inference, or regulatory inference that.
Use when after generating a q-value bedgraph track from ChIP-Seq pileup versus local lambda comparison, and you need to identify statistically significant narrow peaks with defined boundaries.
Use when after running macs3 callpeak with the -f BEDPE flag on paired-end ChIP-Seq data (e.g., CTCF_PE_ChIP_chr22_50k.bedpe.
Use when you have aligned ATAC-seq BAM files and want to discriminate between transcription factor binding sites that are actually occupied by protein versus sites with matching sequence motifs that are unbound.
Use when after applying a quantitative analysis function (e.g., cooltools.insulation, contact frequency calculations) to Hi-C cooler files or other genomic datasets, validate that the output numeric columns contain values within plausible ranges (e.
Use when analyzing differential methylation from bisulfite sequencing data where you suspect overdispersion (variance exceeds binomial expectations), or when comparing uncorrected and corrected statistical tests to determine whether more stringent thresholds are justified by the data.
Use when you have loaded fragment data from single-cell ATAC-seq experiments into a backed AnnData object (with fragments stored in .obsm['fragment_paired'] or .