
Claude Skills by HolobiomicsLab
github.com/HolobiomicsLabUse when when setting up a ViMMS chemical sampling environment and you need to restrict the chemical search space to a specific m/z range (e.g., 100–1000) and MS level (e.g., MS1 only) before generating virtual LC-MS/MS data.
Use when you have a collection of chemical structures in SMILES format and need to create paired structure–spectrum training data for a generative model, but do not have experimental MS/MS spectra available.
Use when you have downloaded and extracted a GNPS archive (from METABOLOMICS-SNETS, METABOLOMICS-SNETS-V2, FEATURE-BASED-MOLECULAR-NETWORKING for GNPS1, or classical_networking_workflow/feature_based_molecular_networking_workflow for GNPS2) and need to construct a queryable molecular family graph.
Use when you have untargeted metabolomics peak intensity data and spectral groupings (Molecular Families or Mass2Motifs) but lack confident chemical annotations or want to avoid pathway database dependency.
Use when you have annotated metabolite structures (with SMILES strings) from a reference library (e.
Use when when you have paired tandem MS spectra and corresponding molecular structures (as SMILES strings or InChI keys) and need to train a model that jointly embeds spectra and structures for structure annotation by database lookup.
Use when when you have annotated chemical structures (SMILES or InChI strings) from a curated MS/MS dataset and need to compute pairwise structural similarity scores (Tanimoto or other metrics) as training labels, or when preparing molecular representations for comparison against mass spectral data.
Use when you have received a JSON response from the CSI:FingerID web service endpoint after submitting a fragmentation tree or tandem mass spectrum query, and you need to extract the predicted molecular fingerprint representation and associated scoring metrics for compound identification or CANOPUS.
Use when you have tandem MS/MS spectra paired with known molecular structures (for training) or unknown spectra requiring structure identification, and you want to predict dense molecular fingerprint vectors that encode structural similarity.
Use when you have acquired MS/MS spectra (in MGF format with required fields: TITLE, PRECURSOR_MZ, PRECURSOR_TYPE, COLLISION_ENERGY) from known or unknown compounds and need to predict their molecular formulas with ranked candidates and confidence scores.
Use when you have user-specified lipid class constraints (e.
Use when processing tandem MS/MS libraries in mgf format (such as GNPS) that lack a Molecular Formula (MF) field but contain valid SMILES strings. The computed formulas are required before combining libraries or writing them to msp format for MS-DIAL compatibility.
Use when when you need to constrain a large metabolite database to a specific instrumental range (e.g., m/z 100–1000) before generating virtual chemical mixtures for LC-MS/MS simulation, or when you need to verify that a reported filtered database count can be reproduced from raw database files.
Use when when you have grouped features consolidated into empirical compounds (EmpCpds) with inferred adduct assignments from khipu, and need to perform MS1-level annotation by matching neutral formulas against JMS-compliant reference libraries (HMDB, LMSD).
Use when you have detected peaks from untargeted LC/HRMS analysis (via IDSL.IPA or equivalent peak picker) with m/z and retention time values, and you need to assign molecular formulas to those peaks.
Use when you have a query MS/MS spectrum with a SMILES string and adduct type (e.g., '[M+H]+', '[M+Na]+'), and you need to validate fragment ions against chemically plausible losses from the parent compound.
Use when you have MS/MS fragmentation spectra (from Orbitrap or Q-TOF instruments) in MGF format with known precursor m/z, adduct type, and collision energy, and you need to generate ranked molecular formula candidates.
Use when immediately after formula assignment from raw FT-ICR MS peak detection, when you have a peak intensity matrix with assigned molecular formulas and need to remove spurious or low-confidence assignments before calculating thermodynamic indices, determining compound classes, or performing.
Use when when you have a feature list from HRMS data with molecular formula annotations (inferred or assigned) and need to compute per-carbon mass defect ratios (MD/C, m/C) as part of PFAS candidate prioritization.
Use when when you have a feature list from HRMS with tentatively assigned molecular formulas (from in silico tools or databases) and need to assess formula plausibility before applying downstream PFAS-specific filters.
Use when when you have an experimental tandem mass spectrum (m/z peaks and intensities) and a chemical formula, and need to identify the true molecular structure from a candidate library (e.g., PubChem).
Use when when you have a molecular structure in XYZ or similar coordinate format and need to initialize QCxMS2 or related workflows for EI mass spectrum calculation.
Use when you have an in-house collection of liquid chromatography spectra and retention time measurements for small molecules, and you want to improve structural identification accuracy by predicting retention times.
Use when when you have a collection of molecular structures (as InChI strings, SMILES, or RDKit Mol objects) and need to feed them into a pretrained or transfer-learning neural network that expects both molecular graph topology and structural fingerprints as inputs.
Use when when you have molecular structures (SMILES or chemical graphs) from a database like PubChem and need to predict molecular properties (e.
Use when you have a GNPS mass spectral molecular network (classical or feature-based) and MS2LDA-derived Mass2Motif data, and you need to annotate network nodes with both chemical class information from GNPS library matches and substructural motifs from MS2LDA to enable joint interpretation of.
Use when you have a GNPS mass spectral molecular network (in .graphml or Cytoscape format) and wish to annotate its nodes and edges with chemical class assignments from the GNPS library and/or MS2LDA motif probabilities from an independent LDA experiment.
Use when after generating candidate transformed structures from biotransformation rules and when you have MS/MS spectral feature data that you wish to organize into putative molecular families.
Use when you have untargeted metabolomics data (e.g., LC-MS/MS spectra) and need to organize compounds by structural relatedness to enable structure discovery for unknown metabolites.
Use when when you have both (1) a molecular network graph from GNPS with MS/MS feature nodes and edges, and (2) a quantitative bioassay matrix (fractions × bioactivity measurements) from parallel LC-MS/MS fractionation of the same sample extract.
Use when after completing dereplication and cosine similarity clustering in the MolNotator pipeline, when you have merged deduplicated molecular predictions and ion annotations and need to construct the final network representation for visualization and compound identification.
Use when you have a GNPS-generated classical or feature-based molecular network (in graphml or JSON format) and corresponding MS2LDA or chemical class assignment data, and you need to embed substructural motif identifiers, confidence scores, or chemical class labels as node/edge attributes for.
Use when when you have a GNPS molecular networking task ID (from GNPS1 or GNPS2 workflows: METABOLOMICS-SNETS, METABOLOMICS-SNETS-V2, FEATURE-BASED-MOLECULAR-NETWORKING, classical_networking_workflow, or feature_based_molecular_networking_workflow) and need to prepare the job archive for NPLinker.
Use when you have a GNPS-generated molecular network (classical or feature-based workflow) and corresponding MS2LDA LDA experiment results (Mass2Motif assignments and/or chemical class predictions), and you need to propagate those annotations to individual network nodes to support visual and.
Use when you have LC-MS/MS DDA data from one or more samples and need to organize fragmentation spectra by similarity relationships to support compound annotation, enable cross-sample comparisons, and identify known and unknown metabolites sharing structural features.
Use when after preparing a feature table and MS/MS spectral data (mzML or MGF format with precursor m/z, retention time, and MS/MS spectra) and before submitting to GNPS for molecular network generation.
Use when you have a CSV file with rows of molecule definitions (chemical formulas, retention times, intensities, or other peak properties) and need to feed them into SMITER's simulation workflow.
Use when when you have a molecular target compound defined by SMILES, InChI, or chemical formula and need to feed it into a pretrained spectrum prediction model (ICEBERG or SCARF) to generate tandem mass spectra or conduct structural elucidation.
Use when when you have a set of chemical structures (SMILES strings or SDF files) that need to be processed for training a graph neural network model on molecular property prediction tasks, specifically when the target property (e.
Use when when you have a molecular structure in any representation (drawn structure, PDB file, common name) and need to input it into mass spectrum prediction tools like ICEBERG or SCARF, or when screening candidates from chemical databases like PubChem.
Use when you have a pretrained encoder that produces fixed-size embeddings from MS/MS spectra (or other molecular data modalities) and you need to recover the corresponding molecular structure as a SMILES string.
Use when after RAMClustR clustering of XCMS-detected features and prior to final compound annotation, when you need to verify the robustness of molecular weight inference or when findMain and RAMClustR predictions are available for the same compound clusters and you want to assess concordance or.
Use when when you have loaded a MoNA mass spectral library (GC-MS or LC-MS/MS) in MSP format and observe that SMILES strings are present in the Comment field rather than in a dedicated SMILES metadata field.
Use when when a trained Siamese neural network model makes predictions on new spectrum pairs and you need to identify and exclude high-uncertainty predictions to improve RMSE.
Use when after LDA has converged and inferred Mass2Motifs from preprocessed mass spectrometry spectral data, when the raw motif-fragment distributions contain noise or low-confidence associations that obscure the dominant fragmentation patterns.
Use when after executing MassQL queries against a MotifDB reference database and retrieving ranked motif matches, when you need to determine which database entries represent true structural correspondence versus spurious matches, and to decide whether a motif's top-ranking hit is sufficiently.
Use when you have a GNPS-generated classical or feature-based mass spectral molecular network (graphml or JSON format) and a corresponding MS2LDA experiment with Mass2Motif assignments on the same spectra, and you want to visualize and quantify which structural motifs are shared within and across.
Use when after Mass2Motifs have been inferred from tandem MS/MS spectra via LDA topic modeling and you need to assign putative substructure annotations to those motifs.
Use when when chaining multiple mpactr filters on a peak table and you need to decide whether to preserve intermediate filtered objects or accept in-place mutation for memory efficiency.
Use when you have a preprocessed peak table from tandem MS/MS data (e.