Data & Analytics
Data analysis, BI, visualization, datasets, statistics, and ML workflows
Browse data & analytics skills
Showing 2,233–2,256 of 13,064 skills
Evaluates end-to-end performance of HPC systems for scientific machine learning, focusing on data staging, I/O efficiency, and model convergence under massive dataset constraints. It measures how well systems handle large-scale volumetric and high-resolution image workloads while meeting strict accuracy targets. Use when the user wants to benchmark on CosmoFlow, DeepCAM, or asks about evaluating this task. Reports MAE, IOU.
Evaluates machine learning surrogates for 2D airfoil aerodynamics on prediction accuracy, computational speed-up, physical consistency, and out-of-distribution generalization. The benchmark compares learned models against a standard CFD solver (OpenFOAM) across in-distribution and novel geometric configurations. Use when the user wants to benchmark on ML4CFD Competition Dataset, or asks about evaluating this task. Reports Global Score.
This benchmark evaluates AI agents' ability to execute end-to-end machine learning development workflows. It probes capabilities across six categories: dataset handling, model training, debugging, model implementation, API integration, and performance improvement. Success is measured by whether agents can produce fully functional, error-free code and configurations that satisfy the task specifications. Use when the user wants to benchmark on ML-Dev-Bench, or asks about evaluating this task. R...
Evaluates large language models' ability to reason over real-world spreadsheet data, including complex headers, multi-sheet files, and cross-file contexts. It probes capabilities across six meta operations: lookup, edit, calculate, compare, visualize, and reasoning. Use when the user wants to benchmark on MiMoTable, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates deep learning and traditional machine learning models on critical care prediction tasks using the MIMIC-III dataset. It probes a model's ability to predict patient mortality across multiple time horizons, classify ICD-9 diagnosis groups, and forecast hospital length of stay from raw clinical time series and tabular data. Use when the user wants to benchmark on MIMIC-III, or asks about evaluating this task. Reports binary classification.
Evaluates the generalization capability and scaling efficiency of self-supervised vision foundation models on a diverse suite of 12 biomedical image classification tasks. It probes how model capacity, data diversity, and pretraining objectives affect downstream diagnostic performance when using a frozen feature extractor. Use when the user wants to benchmark on MedMNIST (12 benchmarks), or asks about evaluating this task. Reports MCC.
Evaluates the short-term traffic flow forecasting capability of various machine learning and deep learning models on real-world urban road networks. It specifically probes how well models generalize across different traffic profiles and prediction horizons while maintaining computational efficiency. Use when the user wants to benchmark on Madrid Traffic Dataset, or asks about evaluating this task. Reports R^2 (Coefficient of determination).
Evaluates a model's ability to detect out-of-distribution (OOD) malware variants and classify known malware families without using OOD samples during training. It probes both classification accuracy on in-distribution data and the statistical separation capability between known and novel threats using cluster-driven decision boundaries. Use when the user wants to benchmark on Unspecified malware dataset (25 families), or asks about evaluating this task. Reports AUROC.
Evaluates AI-driven log analytics systems on three core tasks: parsing unstructured log messages into event templates, compressing log data efficiently, and detecting system anomalies using supervised or unsupervised models. It probes how well algorithms generalize across diverse, real-world system logs ranging from distributed systems to mobile apps. Use when the user wants to benchmark on Loghub, or asks about evaluating this task. Reports Parsing Accuracy (PA).
Evaluates LLMs across 44 existing tasks to uncover latent cognitive skills using psychometric factor analysis, rather than relying on aggregated benchmark scores. It probes whether models possess coherent, interpretable skill profiles across diverse domains like reading comprehension, mathematical reasoning, and ethical judgment. Use when the user wants to benchmark on SQuAD, GSM8K, GPQA, TriviaQA, XSum, MNLI (textual entailment), Ethical/Social Judgment datasets, or asks about evaluating thi...
Evaluates Large Language Models' proficiency in Knowledge Graph Engineering tasks, specifically focusing on RDF syntax repair, SPARQL query generation and semantics, and data serialization format handling across multiple graph structures. Use when the user wants to benchmark on LLM-KG-Bench 3.0, or asks about evaluating this task. Reports capability compass.
Evaluates machine learning models' ability to classify lightning-ignited wildfire occurrences versus non-occurrences using meteorological, vegetation, and spatio-temporal features. It probes the models' generalization across different feature configurations and geographic regions, highlighting the necessity of separate models for lightning versus anthropogenic fires. Use when the user wants to benchmark on Global Lightning-Ignited Wildfire Dataset, or asks about evaluating this task. Reports ...
Evaluates text classification performance and environmental impact (energy, cost, emissions) of NLP models on legal domain datasets. It compares traditional machine learning approaches against transformer-based models across multiple legal benchmarks. Use when the user wants to benchmark on LexGLUE, or asks about evaluating this task. Reports mF1.
Evaluates a multimodal machine learning workflow's ability to predict 3D subsurface geological, hydrogeological, and geophysical features from sparse, heterogeneous field data. It tests the model's generalization capability using transductive learning and mutual information maximization across five cross-validation splits. Use when the user wants to benchmark on Lana'i 3D Subsurface Grid, or asks about evaluating this task. Reports R-squared.
Evaluates tabular machine learning models in data lake environments by leveraging auxiliary tables to improve prediction on a target table. It probes two integration paradigms: table unionability (vertical concatenation to increase training samples) and table joinability (horizontal enrichment to add features). Use when the user wants to benchmark on LakeMLB, or asks about evaluating this task. Reports predictive performance.
Evaluates an automated video classification pipeline for multi-species wildlife behavior monitoring by comparing machine learning predictions against expert manual annotations and traditional ground-based sampling methods. Use when the user wants to benchmark on KABR Drone Behavioral Dataset (custom), or asks about evaluating this task. Reports accuracy.
This evaluation probes the robustness of supervised machine learning models for IoT intrusion detection when their training data is corrupted by adversarial poisoning attacks. It measures how different model architectures degrade in detection capability under label manipulation, outlier injection, and feature impersonation. Use when the user wants to benchmark on CICIoT2023, Edge-IIoTset, N-BaIoT, or asks about evaluating this task. Reports Accuracy.
This benchmark probes instruction-following capabilities by testing models on 20 carefully designed prompts that enforce format compliance, content constraints, logical sequencing, and multi-step execution. It measures whether models can adhere to verifiable, unambiguous constraints rather than relying on superficial pattern matching or memorized benchmark performance. Use when the user wants to benchmark on Instruction Adherence Diagnostic Prompts, or asks about evaluating this task. Reports...
Evaluates the ability of hybrid time-series models to forecast urban air quality indices and specific pollutant concentrations (PM2.5, O3, CO, NOx) using historical environmental data and temporal features. Use when the user wants to benchmark on CPCB Indian Air Quality Dataset, or asks about evaluating this task. Reports RMSE.
Evaluates image classification accuracy on ultra-constrained microcontrollers (MCUs) with strict SRAM and Flash limits, measuring the trade-off between model quantization, memory footprint, and performance. Use when the user wants to benchmark on ImageNet, or asks about evaluating this task. Reports Top-1 accuracy.
Evaluates the ability of individual machine learning classifiers and ensemble strategies to detect network intrusions and classify traffic types. It probes model robustness, precision-recall trade-offs, and computational efficiency across diverse real-world network traffic datasets with varying attack profiles. Use when the user wants to benchmark on RoEduNet-SIMARGL2021, CICIDS-2017, or asks about evaluating this task. Reports F1 Score.
Evaluates machine learning models' ability to detect and classify cyberattacks in Industrial Control Systems (ICS) network traffic. It probes the capability to distinguish between normal operations and specific attack types like DDoS, IP-Scan, MitM, Port-Scan, and Replay using flow-level features. Use when the user wants to benchmark on ICS-Flow, or asks about evaluating this task. Reports F1-score.
Probes the causal impact of LLM-assisted peer reviews on paper scoring and acceptance outcomes at a major machine learning conference. It measures whether AI-assisted reviews systematically inflate scores and increase acceptance probabilities, particularly for borderline submissions. Use when the user wants to benchmark on ICLR Conference Reviews (2018-2024), or asks about evaluating this task. Reports acceptance_rate_difference.
Evaluates machine learning models' ability to classify electron antineutrino (IBD) events from background accidents in a liquid scintillator detector. It measures how well the models preserve signal efficiency while controlling background contamination compared to traditional cut-based selection. Use when the user wants to benchmark on JUNO IBD/Accident Dataset, or asks about evaluating this task. Reports efficiency.