
Claude Skills by HolobiomicsLab
github.com/HolobiomicsLabUse when you have raw MS/MS spectra (in formats like mzML, json, mgf, msp, mzxml) that contain background noise or numerous low-intensity peaks before running MS2Query library matching.
Use when after executing feature detection and quantification on raw LC-MS data (mzML or NetCDF format) using an automated pipeline such as MetaboAnalystR 4.0, and before proceeding to downstream normalization, scaling, or functional analysis.
Use when after peak filtering (by m/z, isotopic presence, formula assignment error, and sample prevalence) and before multivariate analysis (PCA, NMDS, PERMANOVA) when comparing peak abundance patterns across samples with potential differences in ionization efficiency, ion suppression, or total ion.
Use when when training Word2Vec embeddings on mass spectra represented as peak-word documents, and you need to preserve the quantitative intensity relationships between fragments without allowing a single dominant peak to overwhelm the learned word associations.
Use when you have loaded raw mass spectrometry spectral data (in MGF, MSP, mzML, or mzXML format) and need to decide which intensity threshold(s) to use for filtering out noise and low-abundance peaks.
Use when when you have loaded mass spectrometry data (from mzML or Bruker .
Use when when you have raw mzML files and a corresponding feature table (CSV format, e.g., from mzmine) and need to generate peak matrices with fixed dimensions (e.g., 2 × 120) that encode margin vs. peak signal regions for training a neural network classifier to filter false positive LCMS peaks.
Use when after molecular formula assignment and peak filtering are complete, when you have a filtered peak list (m/z values and molecular formulas) and want to discover biochemical transformations without prior knowledge of reaction networks.
Use when after elution peaks have been detected on composite mass tracks using local maxima and prominence thresholds, and before mapping detected features back to individual samples or performing pre-annotation.
Use when when you have manually labeled LC-MS peaks as 'High quality' or 'Low quality' using NeatMS's annotation tool and need to create training/validation/test batches.
Use when after composite-map peak detection (scipy.signal.find_peaks) has identified candidate peaks on aligned mass tracks, but before compiling the final feature table.
Use when when validating a metabolomics pathway analysis method (particularly decomposition-based approaches like PLAGE) against data quality degradation, or when comparing robustness across methods (PLAGE vs. ORA vs. GSEA).
Use when you have a table of detected chromatographic peaks (e.g., from CentWave peak detection in xcms) and need to isolate a single target m/z (e.g., m/z 304.1131 for a pesticide) or a narrow m/z range, or when you must restrict analysis to a known retention time window (e.
Use when when using mpactr filter functions (e.g., filter_mispicked_ions, filter_group, filter_cv) with R6 reference semantics and uncertain whether the copy_object parameter controls deep copying or in-place modification.
Use when after converting peak-picker output (from MZmine, XCMS, MS-DIAL, or Compound Discoverer) into LipidMatch-compatible format. Use this skill when you need to verify that the converted file will be successfully read by LipidMatch before proceeding to lipid identification;
Use when when you have mass spectrometry data organized in a Pandas DataFrame with m/z values, retention time (RT), and intensity measurements, and you want to visualize the joint distribution and correlation of these three dimensions to identify peaks, assess separation, and detect patterns across.
Use when when you have a tandem mass spectrum (MSMS) with known peptide sequence and wish to assess whether enabling neutral loss annotation (e.g., NH3: −17.026549, H2O: −18.010565) increases the proportion of observed m/z peaks that can be matched to predicted fragment ions.
Use when when you have a tandem mass spectrometry spectrum with a known or inferred peptide sequence that may contain post-translational modifications (phosphorylation, glycosylation, cross-links), and you need to annotate which observed m/z peaks correspond to specific fragment ion types (b, y, a.
Use when you have raw MS2 spectra from a sample and need to collapse them into a single sample-level representation for comparison across multiple samples, particularly when samples have poor feature overlap, strong retention time shifts between LC methods, or were acquired on different mass.
Use when you have paired genomic-metabolomic link scores (e.
Use when you have a pretrained model with documented performance on a bounded input domain (e.g., molecules ≤19 heavy atoms, sequences <1000 bp) and you need to establish whether and how much accuracy drops on held-out test cases outside that domain boundary.
Use when when you have access to a set of gallery or benchmark scripts executed across multiple plotting backends and need to quantify which backend delivers the fastest median execution time for specific mass spectrometry plot types (chromatogram, mobilogram, peakmap, peakmap-marginals, spectrum.
Use when you have normalized peak intensities or abundance matrices from mass spectrometry (e.
Use when after computing Multi-Block Variable Importance in Projection (MB-VIP) scores on a fitted MB-PLS discriminant model, when you need to distinguish signal features from noise by establishing empirical significance thresholds rather than relying on parametric assumptions.
Use when you have centroided MS2 spectra (ddMS2 data in mzML format) from HRMS analysis and need to identify potential PFAS compounds among thousands of features.
Use when you have detected features in LC- or GC-HRMS data (via pyOpenMS or custom feature tables) and need to systematically rank them for likelihood of being PFAS compounds.
Use when you have an m/z-resolved feature list from LC- or GC-HRMS analysis (either detected by pyOpenMS or provided as a custom Excel table) and need to prioritize potential PFAS compounds by identifying clusters of homologous structures.
Use when after generating a Chemical Feature Tree from q2-qemistree (or any tree artifact) and before using it for alpha-diversity or beta-diversity phylogenetic analyses.
Use when you have a published computational pipeline with deposited code and validation data, and you need to verify that the pipeline can be executed end-to-end to reproduce reported validation metrics (annotation accuracy, coverage, or equivalent performance benchmarks).
Use when after community-dependent gap-filling has proposed reactions to fill metabolic gaps in individual consensus reconstructions.
Use when a Shiny application or R package currently runs only on Windows and you need to enable deployment on Linux or macOS.
Use when you have mass spectrometry data (m/z, retention time, intensity) loaded into a Pandas DataFrame and need to explore the full 3D structure of a peak map interactively, particularly when static 2D heatmaps obscure important intensity relationships or when stakeholders require browser-based.
Use when augmenting mass spectrometry ion images for contrastive learning, particularly when the model must generalize across different detector conditions or signal-to-noise ratios.
Use when when performing targeted peak detection on LC-MS data where compounds have been assigned expected ionization polarities (positive or negative mode) in the target list, and you want to prevent false peak assignments from the opposite polarity and avoid manual pre-filtering of raw data by.
Use when you have a comprehensive target list (containing compounds from both positive and negative ionization modes) but need to screen or detect peaks in a single LC-MS run acquired in a specific polarity mode.
Use when when you have multiple file format variants (compressed indexed gzip, standard gzip, SQLite database, uncompressed mzML) that all need to be read via a unified interface, and you want to avoid a long chain of conditional logic in client code.
Use when after fitting a polynomial calibration model to tunemix reference data in DEIMoS, assess whether the model explains sufficient variance in the m/z–drift-time–CCS relationship.
Use when when you have validated link annotations from multiple independent datasets (≥2), individual scoring functions with per-dataset enrichment p-values, and you want to test whether a combined scoring strategy (e.
Use when preparing mass spectrum input tensors for transformer encoder layers in IDSL_MINT.
Use when after training a GNN model on molecular structures with continuous targets (e.g., CCS values), when you need to understand which node-level (atom) or edge-level (bond) features contribute most to individual or aggregate predictions.
Use when when you have tandem mass spectrometry data (LC-MS/MS in MGF, mzML, mzXML, or mzData format) paired with either genome sequences or precursor peptide predictions, and you want to confirm the presence and identity of modified ribosomally synthesized and post-translationally modified.
Use when you have centroided LC-MS/MS spectra (in MGF, mzXML, mzML, or mzData format) and genomically-predicted precursor peptide sequences, and you need to identify which predicted RiPPs are actually expressed and modified in the sample.
Use when after molecular formula assignment has been performed on FT-ICR MS peaks and you need to remove assignments with unacceptable mass error before proceeding to chemodiversity analysis, transformation network generation, or multivariate statistics.
Use when a mass spectrum calibration procedure initialized with a narrow ppm window (e.g., ±1.0 or ±5.0 ppm) finds fewer than 5 reference m/z matches.
Use when when you have computed similarity scores (e.g., MS2DeepScore, Spec2Vec, modified Cosine) between pairs of spectra or compounds and want to compare their ability to retrieve chemically related pairs. Apply this skill if you have ground-truth structural similarity labels (e.
Use when you have extracted fragmentation patterns from a collection of MS/MS spectra (using mineMS2) and have partitioned spectra into components via GNPS molecular networking (e.g., connected components, cliques, or high-similarity pairs with cosine > threshold).
Use when after retrieving top-scoring library candidates from a full MS2Deepscore comparison, but before or during final re-ranking.
Use when after loading an MsmsSpectrum object but before intensity filtering or spectral annotation.
Use when you have assembled genome FASTA sequences (from SPAdes, metaSPAdes, or antiSMASH output) and need to systematically identify precursor peptides corresponding to a target RiPP class before constructing the structure database for dereplication.
Use when you have centroided MS2 spectra from data-dependent acquisition (ddMS2) in mzML format and seek to prioritize potential PFAS features by detecting diagnostic fragment masses.