Category

Data & Analytics

Data analysis, BI, visualization, datasets, statistics, and ML workflows

13,068
skills in category
545
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse data & analytics skills

Showing 2,2812,304 of 13,068 skills

Fairness Label Noise EvalA

This evaluation probes the ability of label noise correction methods to mitigate group-dependent label noise while preserving predictive performance and improving algorithmic fairness. It measures how well pre-processing techniques remove underlying discrimination from training data before classifier training. Use when the user wants to benchmark on OpenML (9 datasets), or asks about evaluating this task. Reports AUC.

datapythongo
0
3
Evasive Acceleration EvalA

Evaluates whether a two-dimensional risk metric (Evasive Acceleration) can statistically distinguish crash precursors from routine non-crash traffic conflicts at varying lead times before impact. It tests early-warning timeliness and discrimination capability under realistic false-alarm constraints. Use when the user has predictions and gold and needs to compute AUPRC.

datapythongo
0
3
Eth Phishing Detection EvalA

Evaluates a model's ability to detect phishing addresses on the Ethereum blockchain by analyzing temporal transaction dynamics and graph topology. It probes whether the model can effectively fuse edge-level temporal patterns with node-level structural and statistical features to distinguish malicious accounts from legitimate ones. Use when the user wants to benchmark on $D_1$, $D_2$, $D_3$, or asks about evaluating this task. Reports AUC.

datapythonnode
0
3
Ens 10 EvalA

Probes the ability of deep learning and statistical models to correct biases in long-term ensemble weather forecasts. It evaluates how well models can post-process raw ensemble members to produce calibrated predictive distributions for surface and atmospheric variables. Use when the user wants to benchmark on ENS-10, or asks about evaluating this task. Reports CRPS.

datapythonperformance
0
3
Ember EvalA

Evaluates machine learning models on static malware classification of Windows PE binaries. It probes the effectiveness of engineered static features versus raw binary inputs for distinguishing malicious from benign software. Use when the user wants to benchmark on EMBER, or asks about evaluating this task. Reports ROC AUC.

datapythongit
0
3
Ema Auditing EvalA

This evaluation probes a trained model's data memorization by auditing whether specific query images were included in its training set. It measures the ability to correctly distinguish memorized data from non-memorized data under varying calibration set qualities and query dataset constraints. Use when the user has predictions and gold and needs to compute auditing score (ρ_EMA / ρ_KS).

datapythongo
0
3
Elliptic Fraud EvalA

Evaluates the utility, robustness, and interpretability of graph-derived signals for tabular machine learning on a binary node classification task. It compares graph-augmented models against tabular baselines using statistical hypothesis testing and graph perturbation analysis to ensure reproducibility. Use when the user wants to benchmark on Elliptic, or asks about evaluating this task. Reports F1-score.

datapythonnode
0
3
Ecg Ssl EvalA

Evaluates the transferability of self-supervised representations learned from single-lead ECG signals to downstream clinical and activity recognition tasks. It probes the model's ability to extract robust cardiac and physiological features by training linear probes on frozen encoder outputs across classification and regression benchmarks. Use when the user wants to benchmark on PhysioNet 2017, PTB-XL, Human Activity Recognition (HAR), or asks about evaluating this task. Reports Accuracy.

datapythonperformance
0
3
Ecg Classification EvalA

Evaluates the ability of deep learning architectures to accurately classify electrocardiogram (ECG) recordings into predefined physiological or pathological categories. The benchmark probes joint time-frequency feature extraction capabilities by comparing models that embed Fourier analysis directly into convolutional layers against traditional signal processing and baseline CNN approaches. Use when the user wants to benchmark on MIT-BIH, ECG-ID, Apnea-ECG, or asks about evaluating this task. ...

datapythongo
0
3
Dwrf Weather Forecast EvalA

Evaluates the accuracy, uncertainty quantification, and physical consistency of high-resolution ensemble weather forecasts for renewable energy applications. It probes a model's ability to downscale coarse atmospheric data to 1 km resolution while preserving multi-scale turbulence, thermodynamic constraints, and extreme event probabilities. Use when the user wants to benchmark on Northwestern Gobi Desert Wind Farm & ERA5 Reanalysis, or asks about evaluating this task. Reports RMSE, CRPS.

datapythongo
0
3
Dvcs Cff Extraction EvalA

This benchmark evaluates machine learning models' ability to extract Compton Form Factors (CFFs) from deeply virtual Compton scattering (DVCS) cross-section data. It specifically probes how well models adhere to quantum chromodynamics (QCD) constraints, generalize across kinematic regions, and accurately quantify both aleatoric and epistemic uncertainties during the extraction process. Use when the user wants to benchmark on DVCS unpolarized proton target data, or asks about evaluating this t...

datapythongo
0
3
Dutch Medical Dialogue EvalA

Evaluates the quality of synthetically generated Dutch medical dialogues across structural, lexical, and qualitative dimensions to assess conversational naturalness and domain-specific accuracy. Use when the user wants to benchmark on Synthetic Dutch Medical Dialogues, or asks about evaluating this task. Reports MSTTR.

datapythongo
0
3
Dutch Llm Bench EvalA

Evaluates Dutch LLMs on reasoning, sentiment analysis, linguistic acceptability, world knowledge, and word sense disambiguation using zero-shot multiple-choice and binary classification tasks. Use when the user wants to benchmark on ARC (Dutch), DBRD, Dutch CoLA, Global MMLU (Dutch), XLWIC-NL, or asks about evaluating this task. Reports accuracy.

datapythongo
0
3
Duq Weather Forecasting EvalA

Evaluates a deep learning model's capability to perform spatio-temporal weather forecasting and quantify predictive uncertainty. It tests the model's ability to fuse historical observations with numerical weather prediction (NWP) data to generate accurate point forecasts and reliable 90% prediction intervals over a 37-hour horizon. Use when the user wants to benchmark on Beijing weather dataset, or asks about evaluating this task. Reports SS_avg.

datapython
0
3
Dsp Toxicity Prediction EvalA

Evaluates machine learning models' ability to predict diarrhetic shellfish poisoning (DSP) toxicity events in mussels using long-term environmental and phytoplankton monitoring data. It probes the model's capacity to integrate biological indicators (toxic species abundance) with abiotic drivers (salinity, river flow, temperature) for binary hazard forecasting. Use when the user wants to benchmark on Gulf of Trieste HAB monitoring dataset, or asks about evaluating this task. Reports F1 score.

datapythongo
0
3
Droughtset EvalA

Evaluates spatiotemporal forecasting models on predicting three drought indices (soil moisture, evaporative stress index, and solar-induced chlorophyll fluorescence) across the U.S. CONUS using weekly climate and vegetation data. It also assesses the models' ability to classify drought events based on soil moisture percentiles. Use when the user wants to benchmark on DroughtSet, or asks about evaluating this task. Reports MAE.

datapythontesting
0
3
Discovery SensitivityA

Evaluates the ability of machine learning classifiers to distinguish rare signal events from dominant Standard Model backgrounds in simulated high-energy physics collisions, and their capacity to accurately estimate signal fractions and discovery sensitivity via unbinned template fits. Use when the user has predictions and gold and needs to compute discovery-sensitivity.

datapythongo
0
3
Direct Evidence ScoreA

Assesses machine translation quality by measuring the proportion of source words that have strong lexical co-occurrence evidence in the training corpus. It evaluates whether data-driven lexical transfer fidelity correlates with standard translation quality metrics like BLEU. Use when the user has predictions and gold and needs to compute DE Score.

datapythongo
0
3
Diffusion Rep EvalA

Evaluates whether conditional diffusion models learn semantically meaningful and factorized representations by measuring generation accuracy against ground truth latent coordinates and the predictive power of internal model embeddings over those coordinates. Use when the user wants to benchmark on Synthetic 2D Gaussian Bump Dataset, or asks about evaluating this task. Reports predicted label accuracy.

datapythongo
0
3
Dejavu Forecasting EvalA

Evaluates time series forecasting accuracy and prediction interval calibration using a data-centric cross-similarity approach. It probes the model's ability to aggregate future paths from similar historical reference series to generate point forecasts and uncertainty bounds across different frequencies and historical sample lengths. Use when the user wants to benchmark on M1 and M3 forecasting competitions, or asks about evaluating this task. Reports MASE.

datapython
0
3
Dcg Bench EvalA

Evaluates multimodal large language models' ability to generate executable HTML/JavaScript code for dynamic chart animations from text or video prompts. It probes instruction following, code executability, and fine-grained semantic alignment between generated visualizations and input specifications. Use when the user wants to benchmark on DCG-8K, or asks about evaluating this task. Reports Execution Pass Rate.

datajavascriptpython
0
3
Data Poisoning EvalA

This evaluation probes the robustness of machine learning models against data poisoning attacks by measuring classification accuracy degradation and recovery under label flipping and image replacement attacks. It assesses how well statistical anomaly detection, adversarial training, and ensemble learning defenses mitigate performance drops and false prediction rates. Use when the user wants to benchmark on CIFAR-10, Insurance Claims, or asks about evaluating this task. Reports classification ...

datapythongo
0
3
Dark Machines Anomaly Score EvalA

Evaluates the ability of unsupervised machine learning models to detect deviations from Standard Model physics in high-energy collider data without assuming specific new physics signatures. It probes model-agnostic anomaly detection by measuring how well density estimation and reconstruction-based methods separate background events from potential signal events. Use when the user wants to benchmark on Dark Machines Anomaly Score Challenge Dataset, or asks about evaluating this task. Reports re...

datapythongo
0
3
Critical Icu Prediction EvalA

Evaluates traditional machine learning and deep learning models on a large-scale, multi-institutional OMOP CDM dataset for ICU clinical prediction. It probes the ability of models to forecast patient outcomes (mortality, length of stay, readmission, sepsis) using early admission temporal features. Use when the user wants to benchmark on CRITICAL, or asks about evaluating this task. Reports AUROC.

datapythontesting
0
3