
Claude Skills by HolobiomicsLab
github.com/HolobiomicsLabUse when after computing a scoring function over all possible genomic-metabolomic candidate pairs (e.g., all 2966 MIBiG-GNPS BGC-spectrum pairs), when you have a subset of known validated links and need to assess whether the scoring function ranks them significantly higher than expected by chance.
Use when after computing link scores (e.g., strain correlation, IOKR, or combined scores) across GCF-MF pairs, use this skill to assess whether scores achieve sufficient separation between validated links and the background population.
Use when you have two or more complementary scoring functions (e.g., strain correlation and IOKR scores) that you wish to combine, and you need to determine which combination strategy and parameters maximize enrichment of known true links in a validation set.
Use when after implementing or modifying the scoring module that computes average InChIKey scores and neighbourhood scores for candidate matches, or when integrating new scoring logic into an existing MS2Query pipeline.
Use when you have a set of gallery or example scripts that must run consistently across multiple backend implementations (e.g., matplotlib, Bokeh, Plotly), and you need to verify that reported execution times are accurate or detect performance changes.
Use when when you have a calibrated FT-ICR mass spectrum (e.g., ESI-NEG mode) and need to decide between rapid single-assignment (first_hit=True) and exhaustive multi-assignment (first_hit=False) modes.
Use when you have executed batch searches across two or more domain-specific MASST tools and obtained separate output files (_microbe.json, _plant.json, _tissue.
Use when when you need to confirm that a generated or retrieved release artifact from a version control system (e.g., git tag v1.0.0) produces byte-for-byte or functionally equivalent outputs to the official release published on a platform (e.g., GitHub Releases) on a specific date.
Use when you have a trained neural network model and a labelled validation dataset (with high-quality and low-quality peak annotations), and you need to determine the optimal probability threshold that maximizes the difference between true positive rate and false positive rate for classifying MS1.
Use when when evaluating whether an MS data processing platform (such as mzmine) supports the full range of separation/ionization techniques your laboratory uses, or when assessing whether gaps exist in the software architecture that would require external pre- or post-processing for specific.
Use when when you have raw LC-HRMS metabolomics data in .mzML or .
Use when when deploying a multi-service microarchitecture (such as MAGMa''s four distinct subprojects: magmaweb, joblauncher, job, and pubchem) via Docker Compose and you need to ensure that service startup order respects true readiness rather than container existence—particularly when services have.
Use when after you have processed raw LC-MS/MS spectral data through the specXplore importing pipeline in a Jupyter notebook and produced an in-memory specXplore session data object containing t-SNE embeddings (based on ms2deepscore similarity scores) and associated spectral metadata.
Use when when you need to verify that a GitHub Actions workflow (e.g., main.yml) executes successfully on your local machine, reproduce a reported passing or failing CI build status, or debug why a workflow badge reports success/failure.
Use when when a Shiny application is documented or observed to run only on Windows, blocking deployment to Linux or macOS users. Typical triggers include hardcoded Windows path separators, unavailable packages on non-Windows systems, or system calls specific to the Windows API.
Use when after a TCN-based formula prediction model has generated initial formula candidates from MS/MS spectra, apply this skill to rescore and refine those candidates when you need to improve ranking accuracy.
Use when you have preprocessed MS/MS spectra binned into 10,000 equally-sized m/z bins (10–1000 m/z range) with square-root-transformed intensities, and you need to generate 200-dimensional spectral embeddings for structural similarity prediction, visualization via dimensionality reduction (e.
Use when you have a collection of preprocessed tandem mass spectra (binned into 10,000 equally-sized m/z bins, intensities square-root transformed, top 1,000 peaks retained), a trained MS2DeepScore Siamese model, and you need to predict structural similarity scores (Tanimoto or Dice) for all or a.
Use when you have QCpool (pooled quality control) samples measured at regular intervals across one or more LC-MS/MS sequences and need to detect whether instrument performance degrades, drifts, or destabilizes during the analytical run.
Use when when you have a pre-computed hierarchical dendrogram from structural clustering (e.g., of LC-MS features based on m/z and retention time) and want to compare or validate the cluster assignments produced by a fixed constant-threshold method.
Use when after computing pairwise similarity scores across a collection of preprocessed mass spectra using matchms similarity measures (Cosine-related, molecular fingerprint-based, or metadata-related assessments), use this skill to persist the resulting similarity matrix to a named output file.
Use when when you have computed Spec2Vec similarity scores (typically cosine similarity in [0, 1] range) between discovered Mass2Motifs and a spectral library, and need to decide which matches are sufficiently confident to include in per-motif annotation output.
Use when you have a set of metabolites or chemical formulas to analyze and want to evaluate how different MS/MS fragmentation strategies (e.g., TopN, exclusion lists, dynamic window selection) would perform without access to real instrument time.
Use when after executing a reproducible simulation pipeline (particularly for Over-representation Analysis in metabolomics), compare the newly generated outputs against reference results to confirm that the simulation was correctly implemented and that the computational environment did not.
Use when when you have a computational simulation framework (e.
Use when when you have a log2-normalized, zero-mean, unit-variance intensity matrix (rows=metabolites, columns=samples) and a curated metabolite set database (e.
Use when your LC-HRMS metabolomics analysis must run on a high-performance computing cluster (e.g., HiPerGator, SLURM-managed systems) that lacks Docker support or prefers Singularity for security and portability. You have .mzML or .
Use when building a transformer-based neural network for chemical formula ranking or classification from mass spectrometry spectra, and you need to encode categorical chemical formulas (e.
Use when when you have processed LC-MS/MS data with precursor m/z, ionization mode, collision energy (if available), and fragment peak lists (m/z and intensity pairs), and need to query CSI:FingerID for molecular fingerprint predictions as part of an automated metabolite identification workflow.
Use when when preprocessing raw SMILES strings from external chemistry databases (e.g., CCSBase, METLIN-CCS, or custom compound libraries) that may contain multiple valid but non-canonical notations for the same molecular structure.
Use when you have implemented a new RDKit-based ComputeConverter for SMILES↔InChI conversions and need to verify that the conversion methods preserve molecular structure integrity across round-trip transformations (SMILES → InChI → SMILES or vice versa) before registering it in the MSMetaEnhancer.
Use when when processing downloaded mass spectral libraries (particularly MoNA EI or MS2 libraries) where SMILES information exists but is embedded in unstructured Comment fields rather than a dedicated SMILES field, or when assigning SMILES from external structure databases (SDF files) to library.
Use when you have SMILES strings for candidate novel psychoactive substance structures and need to convert them into a machine-readable molecular representation before computing descriptors, generating mass spectra, or calculating chemical fingerprints.
Use when when you have a dataset of molecular structures encoded as SMILES strings that will be processed downstream (e.
Use when you have raw molecular structures in SMILES or SDF format that will feed into BitterPredict.m or other structure-based classifiers.
Use when you have a set of chemical structures (as SMILES strings or convertible to SMILES) and need to submit them programmatically to the NP Classifier /classify endpoint for batch or automated classification.
Use when you have raw SMILES strings from multiple external database sources (e.g., PubChem, ChEMBL, vendor databases) that need to be integrated into a unified structure registry. Indicators include: (1) raw SMILES table exists at a known interim input path (e.
Use when you have a mass spectral library in MSP format (e.g., from NIST, SWGDRUG, or other sources) exported alongside a folder of MOL files, and you need to populate the SMILES field in each library record to enable structure-based filtering, annotation, or downstream MS-DIAL analysis.
Use when after composite map peak detection has generated a full unfiltered peak list with SNR values computed for each candidate peak.
Use when analyzing MALDI-mass spectrometry imaging data in which sodium or other alkali metal contamination is suspected, or when peak lists show unexplained mass differences in the range of ~20–25 Da (characteristic of Na adducts).
Use when you need to verify the scope and completeness of a software platform's analytical capabilities—particularly when the project claims to support multiple input modalities (e.
Use when when you need to understand the modular structure of a multi-component research software project—particularly when integrating, documenting, or extending a system whose architecture is not immediately obvious from high-level descriptions.
Use when you need to assess whether a newly developed FT-ICR MS pipeline (e.
Use when when a tool claims to be 'scalable' or 'performance-conscious' but lacks published performance benchmarks, or when you need to confirm that runtime and memory scale linearly (or predictably) with sample count before deploying the tool on large LC-MS datasets (e.g., >100 samples).
Use when when you have a scientific software tool (e.g., Met-ID) that is architected to support plugins or configuration-driven modules, and you need to register and apply a novel reagent, derivatizing matrix, or analytical method (e.
Use when you need to determine the full scope of hardware and methodological compatibility for a bioinformatics tool before designing an analytical workflow.
Use when after making code modifications (bug fixes, new features, or refactoring) to the MS2Query codebase, or when contributing changes via pull request. The skill is essential before pushing feature branches to the repository or merging changes into master.
Use when when you have a GitHub-hosted Python project (or other supported language) with an existing test suite and want to gate code contributions on multiple quality dimensions beyond unit tests—specifically when you need automated reporting of code coverage, technical debt, security issues, and.
Use when when you need to reverse-engineer or formally document the computational steps within a closed or under-documented scientific software module—particularly when the software performs in silico generation, enumeration, or filtering of candidate molecular structures and the published paper or.
Use when when you need to verify that a specific data transformation (e.g., precursor m/z zeroing, feature scaling, or field masking) is applied consistently across multiple execution workflows (training, evaluation, inference) in a codebase.