Category

Data & Analytics

Data analysis, BI, visualization, datasets, statistics, and ML workflows

13,285
skills in category
554
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse data & analytics skills

Showing 6,721–6,744 of 13,285 skills

Target PredictionA

--> --- name: bio-small-rna-seq-target-prediction description: Predict miRNA target genes using sequence-based algorithms and database lookups. Use when identifying potential mRNA targets of differentially expressed or functionally important miRNAs. tool_type: mixed primary_tool: miRanda measurable_outcome: Execute skill workflow successfully with valid output within 15 minutes. allowed-tools: - read_file - run_shell_command ---

datapythongo
0
6
Differential MirnaA

--> --- name: bio-small-rna-seq-differential-mirna description: Perform differential expression analysis of miRNAs between conditions using DESeq2 or edgeR with small RNA-specific considerations. Use when identifying miRNAs that change between treatment groups, disease states, or developmental stages. tool_type: r primary_tool: DESeq2 measurable_outcome: Execute skill workflow successfully with valid output within 15 minutes. allowed-tools: - read_file - run_shell_command ---

datashellexpress
0
6
Tximport WorkflowA

--> --- name: bio-rna-quantification-tximport-workflow description: Import transcript-level quantifications from Salmon/kallisto into R for gene-level analysis with DESeq2/edgeR using tximport or tximeta. Use when importing transcript counts into R for DESeq2/edgeR. tool_type: r primary_tool: tximport measurable_outcome: Execute skill workflow successfully with valid output within 15 minutes. allowed-tools: - read_file - run_shell_command --- Import transcript-level estimates from Salmon, kal...

datashellexpress
0
6
Count Matrix QcA

--> --- name: bio-rna-quantification-count-matrix-qc description: Quality control and exploration of RNA-seq count matrices before differential expression. Check for outliers, batch effects, and sample relationships. Use when assessing count matrix quality before DE analysis. tool_type: mixed primary_tool: DESeq2 measurable_outcome: Execute skill workflow successfully with valid output within 15 minutes. allowed-tools: - read_file - run_shell_command --- Quality control and exploratory analys...

datapythonshell
0
6
Orf DetectionA

--> --- name: bio-ribo-seq-orf-detection description: Detect and quantify translated ORFs from Ribo-seq data including uORFs and novel ORFs using RiboCode and ORFquant. Use when identifying translated regions beyond annotated coding sequences or quantifying ORF-level translation. tool_type: mixed primary_tool: RiboCode measurable_outcome: Execute skill workflow successfully with valid output within 15 minutes. allowed-tools: - read_file - run_shell_command ---

datapythongo
0
6
Metadata JoinsA

--> --- name: bio-expression-matrix-metadata-joins description: Merge sample metadata with count matrices and add gene annotations. Use when preparing data for differential expression analysis or visualization. tool_type: mixed primary_tool: pandas measurable_outcome: Execute skill workflow successfully with valid output within 15 minutes. allowed-tools: - read_file - run_shell_command ---

datapythonshell
0
6
Timeseries DeA

--> --- name: bio-differential-expression-timeseries-de description: Analyze time-series RNA-seq data using limma voom with splines, maSigPro, and ImpulseDE2. Identify genes with dynamic expression patterns. Use when analyzing time-series or longitudinal expression data. tool_type: r primary_tool: limma measurable_outcome: Execute skill workflow successfully with valid output within 15 minutes. allowed-tools: - read_file - run_shell_command --- Identify genes with significant temporal express...

datagoshell
0
6
Edger BasicsA

--> --- name: bio-de-edger-basics description: Perform differential expression analysis using edgeR in R/Bioconductor. Use for analyzing RNA-seq count data with the quasi-likelihood F-test framework, creating DGEList objects, normalization, dispersion estimation, and statistical testing. Use when performing DE analysis with edgeR. tool_type: r primary_tool: edgeR measurable_outcome: Execute skill workflow successfully with valid output within 15 minutes. allowed-tools: - read_file - run_shell...

datagoshell
0
6
Deseq2 BasicsA

--> --- name: bio-de-deseq2-basics description: Perform differential expression analysis using DESeq2 in R/Bioconductor. Use for analyzing RNA-seq count data, creating DESeqDataSet objects, running the DESeq workflow, and extracting results with log fold change shrinkage. Use when performing DE analysis with DESeq2. tool_type: r primary_tool: DESeq2 measurable_outcome: Execute skill workflow successfully with valid output within 15 minutes. allowed-tools: - read_file - run_shell_command --- D...

datashellexpress
0
6
De ResultsA

--> --- name: bio-de-results description: Extract, filter, annotate, and export differential expression results from DESeq2 or edgeR. Use for identifying significant genes, applying multiple testing corrections, adding gene annotations, and preparing results for downstream analysis. Use when filtering and exporting DE analysis results. tool_type: r primary_tool: DESeq2 measurable_outcome: Execute skill workflow successfully with valid output within 15 minutes. allowed-tools: - read_file - run...

datagoshell
0
6
Batch CorrectionA

--> --- name: bio-differential-expression-batch-correction description: Remove batch effects from RNA-seq data using ComBat, ComBat-Seq, limma removeBatchEffect, and SVA for unknown batch variables. Use when correcting batch effects in expression data. tool_type: r primary_tool: sva measurable_outcome: Execute skill workflow successfully with valid output within 15 minutes. allowed-tools: - read_file - run_shell_command ---

datashellexpress
0
6
Single Cell SplicingA

--> --- name: bio-single-cell-splicing description: Analyzes alternative splicing at single-cell resolution using BRIE2 for probabilistic PSI estimation or leafcutter2 for cluster-based analysis with NMD detection. Identifies cell-type-specific splicing patterns. Use when analyzing isoform usage in scRNA-seq or finding splicing differences between cell populations. tool_type: python primary_tool: BRIE2 measurable_outcome: Execute skill workflow successfully with valid output within 15 minutes...

datapythonshell
0
6
Differential SplicingA

--> --- name: bio-differential-splicing description: Detects differential alternative splicing between conditions using rMATS-turbo (BAM-based) or SUPPA2 diffSplice (TPM-based). Reports events with FDR-corrected significance and delta PSI effect sizes. Use when comparing splicing patterns between treatment groups, tissues, or disease states. tool_type: mixed primary_tool: rMATS-turbo measurable_outcome: Execute skill workflow successfully with valid output within 15 minutes. allowed-tools: - ...

datapythonshell
0
6
Context Specific ModelsA

--> --- name: bio-systems-biology-context-specific-models description: Build tissue and condition-specific metabolic models using GIMME, iMAT, and INIT algorithms with expression data constraints. Create models that reflect cell-type specific metabolism. Use when building tissue-specific metabolic models or integrating transcriptomics with FBA. tool_type: python primary_tool: cobrapy measurable_outcome: Execute skill workflow successfully with valid output within 15 minutes. allowed-tools: - ...

datapythongo
0
6
Python Pandas Best PracticesA

--> --- name: 'pandas-best-practices' description: 'Standards for efficient, readable, and performant data manipulation using Python''s Pandas library.' measurable_outcome: Execute skill workflow successfully with valid output within 15 minutes. allowed-tools: - read_file - run_shell_command - write_file --- This skill provides guidelines for working with tabular data in Python. It focuses on vectorization, memory management, and method chaining to write "Modern Pandas" code.

datapythongo
0
6
Rmarkdown ReportsA

--> --- name: bio-reporting-rmarkdown-reports description: Create reproducible bioinformatics analysis reports with R Markdown including code, results, and visualizations in HTML, PDF, or Word format. Use when generating analysis reports with RMarkdown. tool_type: r primary_tool: rmarkdown measurable_outcome: Execute skill workflow successfully with valid output within 15 minutes. allowed-tools: - read_file - run_shell_command ---

datashellexpress
0
6
Quarto ReportsA

--> --- name: bio-reporting-quarto-reports description: Build reproducible scientific documents, presentations, and websites with Quarto supporting R, Python, Julia, and Observable JS. Use when creating reproducible reports with Quarto. tool_type: mixed primary_tool: Quarto measurable_outcome: Execute skill workflow successfully with valid output within 15 minutes. allowed-tools: - read_file - run_shell_command ---

datapythonshell
0
6
Jupyter ReportsA

--> --- name: bio-reporting-jupyter-reports description: Creates reproducible Jupyter notebooks for bioinformatics analysis with parameterization using papermill. Use when generating automated analysis reports, running notebook-based pipelines, or creating shareable computational notebooks. tool_type: python primary_tool: papermill measurable_outcome: Execute skill workflow successfully with valid output within 15 minutes. allowed-tools: - read_file - run_shell_command ---

datapythonshell
0
6
GseaA

--> --- name: bio-pathway-gsea description: Gene Set Enrichment Analysis using clusterProfiler gseGO and gseKEGG. Use when analyzing ranked gene lists to find coordinated expression changes in gene sets without arbitrary significance cutoffs. Detects subtle but coordinated expression changes. tool_type: r primary_tool: clusterProfiler measurable_outcome: Execute skill workflow successfully with valid output within 15 minutes. allowed-tools: - read_file - run_shell_command ---

datagoshell
0
6
Go EnrichmentA

--> --- name: bio-pathway-go-enrichment description: Gene Ontology over-representation analysis using clusterProfiler enrichGO. Use when identifying biological functions enriched in a gene list from differential expression or other analyses. Supports all three ontologies (BP, MF, CC), multiple ID types, and customizable statistical thresholds. tool_type: r primary_tool: clusterProfiler measurable_outcome: Execute skill workflow successfully with valid output within 15 minutes. allowed-tools: ...

datagoshell
0
6
PaperBananaA

--> --- name: paper-banana description: Agentic framework for automating the generation of publication-ready academic illustrations and statistical plots. license: CC-BY-SA-4.0 metadata: author: Peking University & Google Cloud AI Research version: "1.0.0" compatibility: - system: Python 3.9+ allowed-tools: - run_shell_command - read_file measurable_outcome: Execute skill workflow successfully with valid output within 15 minutes. --- PaperBanana is an advanced agentic framework designed to au...

datapythongo
0
6
Data AnalysisA

--> --- name: biomedical-data-analysis description: Omics data forge keywords: - pandas - R-tidyverse - SQL - visualization - reproducible measurable_outcome: Deliver a cleaned dataset + statistical summary + at least one visualization or dashboard spec for each request within 1 working session (≤30 minutes). license: MIT metadata: author: BioSkills Team version: "1.0.0" compatibility: - system: Python 3.9+ / R 4.0+ allowed-tools: - run_shell_command - read_file - python_repl --- Run the cros...

datapythonshell
0
6
Spectral LibrariesA

--> --- name: bio-proteomics-spectral-libraries description: Build, manage, and search spectral libraries for proteomics. Use when creating or working with spectral libraries for DIA analysis. Covers DDA-based library generation, predicted libraries (Prosit, DeepLC), and library formats. tool_type: mixed primary_tool: encyclopedia measurable_outcome: Execute skill workflow successfully with valid output within 15 minutes. allowed-tools: - read_file - run_shell_command ---

datapythonshell
0
6
QuantificationA

--> --- name: bio-proteomics-quantification description: Protein quantification from mass spectrometry data including label-free (LFQ, intensity-based), isobaric labeling (TMT, iTRAQ), and metabolic labeling (SILAC) approaches. Use when extracting protein abundances from MS data for differential analysis. tool_type: mixed primary_tool: MSstats measurable_outcome: Execute skill workflow successfully with valid output within 15 minutes. allowed-tools: - read_file - run_shell_command ---

datapythongo
0
6