All authors
lilinji avatar

Claude Skills by lilinji

github.com/lilinji
193 skillsA× 1930 installs367 views
Bio Machine Learning Survival AnalysisA

Builds and validates predictive time-to-event models on clinical and omics data with penalized Cox, random survival forests, gradient-boosted and deep survival models, and prediction-grade evaluation (Uno's C, time-dependent AUC, integrated Brier, calibration, competing risks). Use when building an individualized risk predictor or prognostic omics signature, choosing a survival model, or evaluating one beyond the C-index. For Kaplan-Meier, log-rank, and classical Cox hazard-ratio inference in...

datapythongo
0
8
Bio Metabolomics Statistical AnalysisA

Decision-grade statistical analysis for metabolomics intensity tables. Covers transformation and scaling (Pareto vs unit-variance as a hidden hypothesis), unsupervised structure (PCA/HCA for QC), permutation-validated PLS-DA/OPLS-DA (R2 vs Q2, double CV, VIP as heuristic), univariate testing (Welch/Mann-Whitney/ANOVA/LMM with covariate adjustment), and dependence-aware multiple testing. Use when testing which metabolites differ, building or validating a discriminant model, choosing a scaling,...

datapythonrust
0
8
Bio Methylation Array Qc FilteringA

Performs probe filtering and sample-level QC on Illumina Infinium methylation arrays (450K / EPIC / EPICv2) to decide which probes and samples to trust. Drops detection-p-failed and low-bead-count probes, removes cross-reactive/non-specific probes (Chen 2013 / Pidsley 2016 lists via maxprobes), excludes SNP-overlapping probes with dropLociWithSnps, and handles sex-chromosome probes. Collapses EPICv2 replicate probes with betasCollapseToPfx and harmonizes across array versions (EPICv2 hg38 vs ...

researchrustgo
0
8
Bio Methylation Epigenetic ClocksA

Computes DNA methylation age (DNAm age) and pace of aging by applying frozen elastic-net epigenetic clocks to a clean beta matrix with methylclock, dnaMethyAge, or methylCIPHER. Covers the clock menu by question (chronological Horvath/Hannum/skin&blood; health-mortality PhenoAge/GrimAge; DunedinPACE pace; pediatric/gestational; mitotic epiTOC), age acceleration (EAA/IEAA/EEAA) as the real endpoint, the principal-component (PC) clock fix for the per-CpG reliability crisis, and EPICv2 clock-CpG...

testinggotesting
0
8
Bio Methylation MethylkitA

Imports Bismark coverage or cytosine-report files into the methylKit object model, then runs the import-to-results spine - filterByCoverage, normalizeCoverage, unite/destrand, calculateDiffMeth, getMethylDiff - for both per-CpG (DMC) and fixed-tile (DMR) differential methylation, plus tileMethylCounts, PCA/correlation/clustering QC, and assocComp/removeComp batch handling. Covers the silent default traps that shape the false-positive rate: overdispersion='none' does no correction while 'MN' f...

datarustgo
0
8
Bio Ml Docking RescoringA

Performs ML-based protein-ligand pose prediction and scoring using DiffDock-L (diffusion-based), Boltz-1 / Boltz-2 (foundation model with affinity), Chai-1, AlphaFold3 ligand, EquiBind, TANKBind, NeuralPLexer, and hybrid workflows (DiffDock pose + GNINA rescore + PoseBusters QC). Explicit handling of when ML beats classical docking, when classical beats ML, the PB-invalid pose problem, and rescoring as the standard production hybrid. Use when modern docking is needed: foundation-model ligand-...

datapythongo
0
8
Bio Population Genetics Association TestingA

Single-variant common-variant GWAS with plink2 --glm (linear/logistic, Firth) and the linear mixed models GEMMA, BOLT-LMM, SAIGE, regenie (SPA). A GWAS statistic is valid only when genotype is independent of unmodeled phenotype drivers after the chosen covariates and random effects, so the engine follows sample structure and case:control imbalance, not taste: PC covariates absorb continuous ancestry but cannot remove relatedness (a covariance structure needing an LMM), genomic inflation above...

datapythongo
0
8
Bio Population Genetics Plink BasicsA

Manages PLINK genotype filesets - format conversion (VCF, BED/BIM/FAM, PED/MAP, pgen/pvar/psam) and sample/variant QC (missingness, MAF, HWE, sex check, heterozygosity, KING relatedness) with PLINK 1.9 and 2.0. PLINK rewrites allele bookkeeping: PLINK 1.x A1 defaults to the minor allele and is recomputed every load, silently flipping effect-allele meaning unless --keep-allele-order, while PLINK 2.0 tracks explicit REF/ALT. QC order matters (variant before sample missingness), HWE is controls-...

datapythonrust
0
8
Bio Population Genetics Population StructureA

Infers and describes population structure with PCA (plink2 --pca, smartpca/EIGENSOFT, FlashPCA2), model-based clustering (ADMIXTURE, fastSTRUCTURE), FST estimators (Weir-Cockerham vs Hudson), and f-statistics (f3/f4/D via AdmixTools/admixr), plus Python plotting of PCs and Q barplots. Every output is a model-conditioned description of variance, not truth: PCs conflate ancestry with LD/inversions/relatedness/batch, ADMIXTURE Q-values are panel- and K-dependent artifacts, and CV-minimum K is a ...

datapythongo
0
8
Bio Population Genetics Scikit Allel AnalysisA

In-memory Python population genetics with scikit-allel - GenotypeArray/HaplotypeArray/AlleleCountsArray, diversity (pi, theta, Tajima's D), SFS, FST (Weir-Cockerham, Hudson, Patterson), f3/D admixture stats, LD pruning, PCA, and selection scans (iHS, XP-EHH, nSL, Garud H). Nearly every statistic is a ratio or density with one silent denominator bug in two faces: omit is_accessible= and per-base pi/theta divide by total span not accessible bp (deflated 2-5x); average per-SNP FST instead of sum...

datapythongo
0
8
Bio Single Cell Differential AbundanceA

Test whether cell-type proportions or composition changed between conditions in single-cell data using Milo (miloR), scCODA, sccomp, and propeller. Use when comparing cell-type proportions / composition between conditions, asking which populations expanded or contracted with treatment or disease, running neighborhood-level (cluster-free) abundance testing, or guarding against compositional shifts that masquerade as differential expression.

datapythongo
0
8
Bio Single Cell Doublet DetectionA

Detect and remove doublets (two or more cells in one droplet) from single-cell RNA-seq using scDblFinder (R), Scrublet (Python), and DoubletFinder (R). Use when flagging artificial intermediate populations before clustering, setting the expected doublet rate from recovered-cell counts, running detection per sample before integration, choosing between simulate-and-score methods, or interpreting a non-bimodal score histogram.

datapythongo
0
8
Bio Tcr Bcr Analysis Mixcr AnalysisA

Align V(D)J reads and assemble TCR/BCR clonotypes with MiXCR, driven by a chemistry-matched preset. Use when choosing/auditing the preset for a library (5'RACE/template-switch vs multiplex-primer amplicon -> rigid vs floating boundaries; RNA vs gDNA -> --rna/--dna; bulk vs 10x single-cell; UMI vs no-UMI -> tag pattern and barcode collapse; kit presets Takara/NEBNext/QIAseq/BD/MiLaboratory); assembling clonotypes by CDR3 vs VDJRegion; setting the reads-vs-UMI-vs-cell quantitation denominator; ...

datarustgo
0
8
Bio Temporal Genomics Trajectory ModelingA

Models continuous temporal trajectories from BULK or time-resolved omics where the x-axis is measured experimental time: penalized GAMs (mgcv) for smooth trends and changepoint detection (segmented, ruptures) for abrupt regime shifts. Use when deciding between a smooth GAM and a changepoint model; choosing the GAM distribution (nb() plus a library-size offset for raw counts vs Gaussian on vst/log-CPM); setting the basis-dimension ceiling k below the number of timepoints and letting REML pick ...

developmentpythonrust
0
8
Bio Tumor Fraction EstimationA

Estimates tumor fraction (the genome-wide proportion of cfDNA molecules that are tumor-derived, the cfDNA analogue of bulk-tumor purity) from shallow whole-genome sequencing with ichorCNA, an HMM over 1 Mb bins that jointly EM-estimates tumor fraction, ploidy, and subclonal prevalence over a normal/ploidy grid. Encodes the load-bearing reframes: tumor fraction is the quantity that travels across assays and is NOT mutation VAF (clonal-het VAF approximately TF/2), CNA-based estimation has a har...

testingpythonrust
0
8
Bio Workflows Cnv PipelineA

Orchestrates the copy-number pipeline from BAM to segmented, integer-called, annotated CNVs, forking on germline-vs-somatic - CNVkit (somatic exome/panel: coverage -> assay-matched reference/PoN -> fix -> segment -> purity/ploidy-aware call), GATK gCNV (germline rare-CNV cohort), and allele-specific callers (ASCAT/FACETS/PURPLE) for purity/ploidy. Use when committing the build + target/access BED + PoN once (assay-matched), building the reference from normals BEFORE segmenting, fitting purity...

datagobash
0
8
Bio Workflows Outbreak PipelineA

Orchestrates genomic-epidemiology outbreak investigation from pathogen isolates to transmission networks, forking bacterial (snippy -> Gubbins recombination-masking -> IQ-TREE -> TreeTime -> TransPhylo) vs viral (Nextstrain/augur), with parallel MLST typing (cgMLST delegated to epidemiological-genomics/pathogen-typing) and AMR surveillance. Use when committing ONE reference genome for SNP calling (every isolate and distance inherits its coordinates), applying MANDATORY Gubbins recombination-m...

datapythonrust
0
8
Bio Workflows Rnaseq To DeA

Orchestrates the end-to-end bulk RNA-seq differential-expression pipeline from FASTQ to an annotated DE gene table, chaining fastp QC/trim, Salmon (decoy-aware) or STAR+featureCounts quantification, tximport gene-level collapse, DESeq2/edgeR/limma-voom testing, apeglm shrinkage, and VST-based visualization. Use when committing the reference release and gene-ID namespace once for the whole run, sequencing steps in the defensible order (tximport before DE, raw counts into the model, VST only fo...

datagobash
0
8
Biorxiv SearchA

Search bioRxiv biology preprints with natural language queries. Semantic search powered by Valyu.

researchtypescriptpython
0
8
Bulkrna Batch CorrectionA

Load when removing batch effects from a multi-cohort bulk RNA-seq dataset using ComBat (R

datapythonrust
0
8
Bulkrna CoexpressionA

Load when discovering gene co-expression modules and hub genes in a bulk RNA-seq cohort via

datapythongo
0
8
Bulkrna DeA

Load when comparing gene expression between two conditions in bulk RNA-seq count data. Skip

datapythongo
0
8
Bulkrna DeconvolutionA

Load when estimating cell-type proportions in bulk RNA-seq samples from a single-cell or

datapythongo
0
8
Bulkrna EnrichmentA

Load when running pathway / GO term enrichment on a bulk RNA-seq DE result list. Skip when

datapythongo
0
8
Bulkrna Geneid MappingA

Load when converting gene identifiers between Ensembl, Entrez, and HGNC symbol in a bulk

datapythongo
0
8
Bulkrna Ppi NetworkA

Load when querying STRING for the protein-protein interaction subgraph induced by a bulk

datapythongo
0
8
Bulkrna QcA

Load when checking a bulk RNA-seq count matrix for library-size outliers, gene detection

datapythonrust
0
8
Bulkrna Read AlignmentA

Load when summarising STAR / HISAT2 / Salmon alignment-rate logs in bulk RNA-seq. Skip when

datapythonrust
0
8
Bulkrna Read QcA

Load when checking raw FASTQ quality (Phred / GC / adapter / Q20-Q30) before alignment in

datapythonrust
0
8
Bulkrna SplicingA

Load when summarising rMATS / SUPPA2 alternative-splicing output and identifying significant

datapythongo
0
8
Bulkrna SurvivalA

Load when stratifying patients by gene expression and testing for survival differences (Kaplan-Meier

datapythongo
0
8
Bulkrna TrajblendA

Load when placing bulk RNA-seq samples on a single-cell reference's pseudotime axis (NNLS

datapythongo
0
8
Consensus DomainsA

Load when you want a verified multi-method consensus over spatial tissue domains on a preprocessed

datapythonrust
0
8
Consensus InterpretA

Load when biologically interpreting a finished verified consensus run (consensus-domains

datapythongo
0
8
Genomics AlignmentA

Load when computing alignment QC metrics (mapping rate, MAPQ distribution, insert size, duplicate

datapythongo
0
8
Genomics AssemblyA

Load when computing genome-assembly QC metrics — N50/N90, L50/L90, total length, contig count,

datapythongo
0
8
Genomics Cnv CallingA

Load when calling CNV segments via CBS-style segmentation on a bin-level log2-ratio CSV from

datapythongo
0
8
Genomics EpigenomicsA

Load when summarising a peak file (BED / narrowPeak) from ATAC-seq / ChIP-seq / CUT&Tag —

datapythongo
0
8
Genomics PhasingA

Load when summarising a phased VCF (output of WhatsHap / SHAPEIT5 / Eagle2) — phased fraction

datapythongo
0
8
Genomics QcA

Load when running pre-alignment FASTQ quality control — Phred quality scores, Q20/Q30 rates,

datapythongo
0
8
Genomics Sv DetectionA

Load when summarising structural variants from an SV VCF (DEL / DUP / INV / TRA) — BND-notation

datapythongo
0
8
Genomics Variant AnnotationA

Load when summarising functional impact of an annotated variant CSV — per-IMPACT counts (HIGH

datapythongo
0
8
Genomics Variant CallingA

Load when summarising small variants (SNVs / indels) from a VCF or computing demo-pattern

datapythongo
0
8
Genomics Vcf OperationsA

Load when summarising / filtering a VCF — variant classification (SNP / MNP / INS / DEL /

datapythongo
0
8
Hermes Tweet XquikA

Use Hermes Tweet and Xquik for X/Twitter agent workflows. Plan social listening, account and follower analysis, post research, monitors, webhook alerts, REST API calls, MCP client setup, and safe tweet actions.

toolsbashgit
0
8
LiteratureA

Load when extracting GEO accessions, dataset metadata, and downloadable references from a

datapythongo
0
8
Medical Entity ExtractorA

Extract medical entities (symptoms, medications, lab values, diagnoses) from patient messages.

ai-agentstypescriptgo
0
8
Medical Research ToolkitA

Query 14+ biomedical databases for drug repurposing, target discovery, clinical trials, and literature research. Access ChEMBL, PubMed, ClinicalTrials.gov, OpenTargets, OpenFDA, OMIM, Reactome, KEGG, UniProt, and more through a unified MCP endpoint. Use when researching disease targets, finding approved/investigational drugs, searching clinical evidence, discovering genetic associations, or analyzing compound bioactivity data.

researchgobash
0
8
Memory ManagementA

Ensure deterministic and safe memory usage: static allocation, pools, stack sizing, MPU usage, and corruption detection for medical device firmware.

datac++security
0
8
Metabolomics AnnotationA

Load when annotating LC-MS features against a built-in 15-metabolite HMDB demo dictionary

datapythongo
0
8