Data & Analytics
Data analysis, BI, visualization, datasets, statistics, and ML workflows
Browse data & analytics skills
Showing 2,185–2,208 of 13,064 skills
This evaluation probes the trade-off between total system throughput and user fairness in UAV-enabled wireless networks. It measures how effectively a resource allocation and trajectory design scheme balances maximizing aggregate data rates against ensuring equitable service across users with varying channel conditions. Use when the user has predictions and gold and needs to compute system throughput.
Evaluates the ability of machine learning and dynamical models to forecast subseasonal temperature and precipitation over the contiguous United States at 3–4 and 5–6 week lead times. It benchmarks predictive accuracy against operational baselines and tests robustness to high noise and spatial heterogeneity in climate data. Use when the user wants to benchmark on SubseasonalClimateUSA, or asks about evaluating this task. Reports % IMPROVEMENT OVER MEAN DEB. CFSV2 RMSE.
Evaluates the precision of statistical estimators (Empirical Bayes, Synthetic Regression, Direct Training) for estimating model performance on data subgroups with limited observations. It probes how well these methods reduce mean squared error and produce reliable confidence intervals when benchmarking LLMs, vision models, and tabular classifiers on niche tasks. Use when the user wants to benchmark on LLM MC QA tasks, Computer Vision tasks (LAION CLIP benchmark), COCO Captions, Tabular Fairne...
Evaluates a model's ability to extrapolate daily soil moisture dynamics across three depth layers using in-situ ground measurements and meteorological forcing. It probes temporal fidelity and absolute accuracy against independent station data. Use when the user wants to benchmark on ISMN & CEMADEN in-situ soil moisture measurements, or asks about evaluating this task. Reports NRMSE.
Evaluates the accuracy and spatial-temporal fidelity of a machine learning-derived daily soil moisture product for Europe. It probes the model's ability to generalize across diverse climates, capture drought dynamics, and outperform existing reanalysis and satellite-based soil moisture datasets. Use when the user wants to benchmark on SoMo.ml-EU, or asks about evaluating this task. Reports uRMSD.
This benchmark evaluates the capability of machine learning models to detect offensive language in Sinhala text. It probes binary text classification performance on a highly imbalanced dataset of Sinhala tweets, measuring how well models distinguish between offensive and non-offensive content. Use when the user wants to benchmark on SOLD, or asks about evaluating this task. Reports macro-averaged F1-score.
Evaluates the predictive accuracy of ensemble machine learning models for forecasting solar power generation using meteorological parameters. It probes regression performance under varying feature sets and ensemble aggregation strategies. Use when the user wants to benchmark on SRRA dataset, or asks about evaluating this task. Reports RMSE.
Evaluates tree-based machine learning models for day-ahead solar power generation forecasting at hourly resolution. It probes the models' ability to capture spatial heterogeneity in meteorological features and their robustness under varying weather conditions. Use when the user wants to benchmark on Belgian Solar Power Generation Dataset, or asks about evaluating this task. Reports RMSE.
Evaluates the ability of multilayer perceptrons to predict vegetation phenology parameters (start of season, peak of season, peak NDVI value) from soil temperature and meteorological variables in subarctic grasslands. It probes how well non-linear models capture complex, non-linear interactions between climate drivers and vegetation dynamics compared to simple linear baselines. Use when the user wants to benchmark on Subarctic grassland phenology dataset (Iceland, 2014-2019), or asks about ev...
This metric quantifies the degree of overlap between different LLM benchmarks by comparing their token-level perplexity signatures derived from in-the-wild pretraining corpora, revealing whether performance correlations stem from shared latent capacity familiarity or benchmark-orthogonal factors like question format. Use when the user has predictions and gold and needs to compute signature_overlap.
Evaluates a model's ability to classify text sentiment into binary or fine-grained polarity categories. Specifically probes how well the model handles negation scope and polarity disentanglement through multi-task learning. Use when the user wants to benchmark on SST-binary, SST-fine, SemEval-binary, SemEval-fine, or asks about evaluating this task. Reports accuracy.
Evaluates unsupervised lossy compression of irregular sensor data on radiation-hardened edge ASICs. It probes the ability to reconstruct high-granularity calorimeter images under extreme bandwidth and latency constraints. Use when the user wants to benchmark on CMS HGCal Trigger Data, or asks about evaluating this task. Reports Energy Mover's Distance (EMD).
Evaluates machine learning models for seismic wavefield forecasting, reconstruction, and generalization under realistic constraints like noise and limited data. It probes model robustness and dynamic learning by comparing performance across multiple tasks against naive baselines. Use when the user wants to benchmark on global wavefields, DAS, synthetic 3D crustal wavefields, or asks about evaluating this task. Reports multi-metric scoring.
Evaluates a seismology foundation model's ability to classify seismic event types, localize epicenters and depths, and determine focal mechanisms using multi-modal seismic data. It probes cross-dataset generalization and compares fine-tuned, frozen, and scratch-trained variants against spectrum-based baselines. Use when the user wants to benchmark on PNW dataset, SCSN dataset, or asks about evaluating this task. Reports AUC.
Evaluates an autonomous agent's capability to generate executable scientific and data-science code by measuring execution reliability and task-goal satisfaction. It probes the model's ability to navigate constrained search budgets while producing outputs that match predefined success criteria across diverse disciplinary benchmarks. Use when the user wants to benchmark on ScienceAgentBench, DA-Code, or asks about evaluating this task. Reports Success Rate (SR).
Evaluates the ease of implementation and qualitative complexity of five big-data systems when executing real-world scientific image analytics workloads in astronomy and neuroscience. It probes how well each system handles multidimensional array operations, Python integration, and native support for scientific image formats. Use when the user wants to benchmark on Neuroscience Use Case, Astronomy Use Case, or asks about evaluating this task. Reports Lines of Code (LoC).
Evaluates a model's ability to segment specific objects in remote sensing images guided by natural language expressions, with a focus on accurately localizing small, scattered targets that are characteristic of aerial and satellite imagery. Use when the user wants to benchmark on RefSegRS, or asks about evaluating this task. Reports mIoU.
Evaluates scientific machine learning models on sim-to-real transfer for complex physical systems. It probes a model's ability to predict spatiotemporal dynamics from real-world measurements, leveraging simulated pretraining, and assesses long-term prediction stability under autoregressive rollout. Use when the user wants to benchmark on Cylinder, ControlledCylinder, FSI, Foil, Combustion, or asks about evaluating this task. Reports RMSE.
Evaluates multi-modal large language models on weather radar forecast quality analysis, specifically testing their ability to perform quantitative rating of radar frames/sequences and generate qualitative assessment reports. It probes domain-specific meteorological understanding, temporal pattern evolution tracking, and alignment with expert judgment. Use when the user wants to benchmark on RQA-70K, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates hybrid quantum-classical neural networks (QMLP and QCNN) on malware classification tasks. It probes the ability of NISQ-era quantum circuits to encode high-dimensional static feature vectors and learn discriminative patterns for binary and multiclass malicious software detection. Use when the user wants to benchmark on API-Graph, EMBER-Domain, AZ-Domain, EMBER-Class, AZ-Class, or asks about evaluating this task. Reports Accuracy.
Evaluates machine learning models for predicting photovoltaic (PV) output power in smart buildings across multiple time frames (30 min, 1 hour, 4 hours) using location-specific meteorological and temporal features. Use when the user wants to benchmark on Tartu, Estonia PV dataset, or asks about evaluating this task. Reports MAPE.
This benchmark evaluates multimodal large language models' ability to interpret synthetic data visualizations. It probes their capacity to extract quantitative features, identify statistical properties, and answer specific questions about plots like histograms, time series, boxplots, and violin plots without relying on real-world data contamination. Use when the user wants to benchmark on PUB Synthetic Plot Dataset, or asks about evaluating this task. Reports overall score.
This benchmark evaluates the discriminative performance of machine learning models on tabular classification tasks under data-scarce conditions (sample sizes ≤500). It specifically probes whether complex AutoML and deep learning approaches can consistently outperform simple baselines like logistic regression when training data is limited. Use when the user wants to benchmark on PMLBmini, or asks about evaluating this task. Reports AUC.
Evaluates the sharpness and calibration of probabilistic net-load forecasts. It measures how closely predicted quantiles align with actual observations and how narrow the prediction intervals are while maintaining statistical reliability. Use when the user has predictions and gold and needs to compute Pinball Score.