Data & Analytics
Data analysis, BI, visualization, datasets, statistics, and ML workflows
Browse data & analytics skills
Showing 2,161–2,184 of 13,064 skills
Evaluates the ability of a Gaussian linear state-space model to accurately forecast short-term wind speeds in the North-East Atlantic using historical observations. It also assesses the model's capacity to reproduce realistic spatiotemporal wind statistics and compares parameter estimation methods (GMM vs ML). Use when the user wants to benchmark on ERA Interim reanalysis data (North-East Atlantic), or asks about evaluating this task. Reports MSPE.
Evaluates machine learning and deep learning models' ability to predict wildfire occurrences in Morocco using integrated spatio-temporal environmental and meteorological features. It specifically tests temporal generalization by training on pre-2022 data and validating on post-2022 data. Use when the user wants to benchmark on Morocco Wildfire Dataset, or asks about evaluating this task. Reports accuracy.
Evaluates social bias in Spoken Language Models (SLMs) by isolating content-induced bias and acoustic bias (gender/accent) using a synthesized speech version of the BBQ dataset. It measures how architectural differences in speech encoders affect bias propagation and whether acoustic cues override textual context. Use when the user wants to benchmark on VoiceBBQ, or asks about evaluating this task. Reports bias_score.
Evaluates the ability of coding agents to generate executable visualization code across multiple programming languages and chart families, including iterative self-debugging capabilities. It probes both initial code generation fidelity and the model's capacity to recover from execution errors using feedback logs. Use when the user wants to benchmark on VisPlotBench, or asks about evaluating this task. Reports Execution Pass Rate.
This evaluation probes the semantic, epistemic, and verifiability characteristics of real-world user-submitted fact-checking requests. It measures how public demand for verification aligns with or diverges from synthetic benchmark distributions, highlighting gaps in current misinformation evaluation corpora. Use when the user wants to benchmark on User Fact-Checking Claims Dataset, or asks about evaluating this task. Reports Veracity Score.
Evaluates the classification accuracy and computational efficiency of machine learning models for network intrusion detection on a realistic dataset of contemporary traffic and synthetic attacks. It also assesses the privacy preservation and data utility of a Pearson Correlation Coefficient (PCC) feature selection and Least Squares Method (LSM) data distortion pipeline. Use when the user wants to benchmark on UNSW-NB15, or asks about evaluating this task. Reports Accuracy.
Evaluates the ability of generative models to synthesize realistic 10-meter hourly wind speed maps over the UK, focusing on statistical fidelity, spatial structure, and extreme event intensity distribution. Use when the user wants to benchmark on ERA5, or asks about evaluating this task. Reports FID.
Evaluates the predictive accuracy and computational efficiency of a user-behavior-based network traffic forecasting method against statistical and neural network baselines on real-world SMS traffic data. Use when the user wants to benchmark on Guangzhou and Milan SMS datasets, or asks about evaluating this task. Reports R2.
Evaluates LLMs' ability to perform multi-step, constraint-aware reasoning and inference on real-world time series data. It probes compositional reasoning, numerical precision, and the capacity to assemble complex analytical or forecasting workflows via executable code generation. Use when the user wants to benchmark on TSAIA, or asks about evaluating this task. Reports Success Rate.
Evaluates machine learning models for detecting fraudulent bank transfers by optimizing instance-dependent cost-sensitive objectives. It probes the model's ability to minimize financial losses and maximize expected savings under highly imbalanced transaction data. Use when the user wants to benchmark on Credit Card Transaction Data, Bank data set, or asks about evaluating this task. Reports Expected Savings.
Evaluates the scalability and efficiency of distributed machine learning systems by measuring how many training tokens each system can process per second across different hardware topologies and model parallelism strategies. Use when the user has predictions and gold and needs to compute training throughput (tokens/second).
Evaluates machine learning models for detecting tornadoes using full-resolution polarimetric weather radar imagery. It probes the ability of classifiers to distinguish tornadic signatures from non-tornadic weather patterns across varying difficulty levels and threshold settings. Use when the user wants to benchmark on TorNet, or asks about evaluating this task. Reports AUC (ROC).
Evaluates machine learning classifiers' ability to predict tornado occurrences up to five days in advance using historical meteorological grid data. It measures detection capability and false alarm rates under a strict temporal train-test split simulating real-world forecasting. Use when the user wants to benchmark on Custom Tornado Forecasting Dataset, or asks about evaluating this task. Reports POD.
This benchmark evaluates vision-language models' ability to reason over time series data across medical, financial, and meteorological domains. It probes pattern recognition, anomaly detection, and causal reasoning using synthetically generated multiple-choice questions derived from real-world datasets. Use when the user wants to benchmark on PTB-XL, MIT-BIH, MIMIC-IV Waveform, Yahoo Finance, WeatherBench 2, or asks about evaluating this task. Reports accuracy.
Evaluates the effectiveness of transfer learning on time series data by comparing pre-trained models against models trained from scratch across intra-domain and cross-domain settings. It probes whether shared temporal structures enable knowledge transfer and how dataset size and domain similarity affect predictive performance and training convergence. Use when the user wants to benchmark on LEN-DB, SPEECH, EMG, S&P 500, LOMAX, STEAD, or asks about evaluating this task. Reports MAE, weighted F...
Evaluates the relative contribution of network traffic features to intrusion detection models by measuring performance degradation when features are shuffled. Probes the model's reliance on specific packet-level and flow-level statistics for distinguishing benign from malicious traffic and classifying attack types. Use when the user wants to benchmark on TII-SSRC-23, or asks about evaluating this task. Reports Permutation Feature Importance (PFI).
Evaluates deep learning surrogate models on autoregressive time-series prediction across 16 diverse physics simulations. It probes the model's ability to forecast future spatiotemporal states from short historical snapshots and maintain stability over longer rollout horizons. Use when the user wants to benchmark on The Well, or asks about evaluating this task. Reports VRMSE.
Evaluates large language models on table question answering with advanced data analysis tasks, including forecasting and chart generation, as well as unclear queries that lack explicit parameters. Probes the model's ability to perform semantic parsing, infer missing parameters, generate executable analysis code, and reason about data visualization. Use when the user wants to benchmark on Text2Analysis, or asks about evaluating this task. Reports ECR, pass@1.
Evaluates time-series foundation models, statistical methods, and machine learning algorithms on univariate forecasting tasks. It probes their ability to handle diverse statistical properties like stationarity, seasonality, sparsity, and noise across real-world and synthetic datasets. Use when the user wants to benchmark on TempusBench Univariate Benchmark, or asks about evaluating this task. Reports MASE.
Predicts whether a patient will give positive feedback (thumbs-up) for a doctor's response in a Romanian telemedicine platform. It probes the model's ability to leverage clinical communication features, patient/doctor history, and metadata to forecast user satisfaction. Use when the user wants to benchmark on Romanian Telemedicine Platform Dataset, or asks about evaluating this task. Reports ROC-AUC.
This benchmark evaluates the ability of machine learning models to predict click-through rates (CTR) for advertisements on a large-scale e-commerce platform. It specifically probes how well models capture static user-ad interactions versus dynamic, temporal user behavior sequences to forecast future clicks. Use when the user wants to benchmark on Alibaba's Taobao Advertising Dataset, or asks about evaluating this task. Reports AUC.
Evaluates the fidelity, utility, and privacy of synthetic tabular data generated by language models compared to real data and other generative baselines. It measures how well the synthetic distribution matches the original across statistical, downstream utility, and privacy dimensions. Use when the user wants to benchmark on Adult, Default, Shoppers, Magic, Beijing, or asks about evaluating this task. Reports C2ST.
Evaluates the performance of traditional machine learning, deep learning, attention-based, and contrastive learning methods on tabular classification tasks. It probes how data characteristics (dimensionality, difficulty) influence the optimal learning strategy and compares different masking/filling strategies used in contrastive learning. Use when the user wants to benchmark on OpenML Tabular Benchmark, or asks about evaluating this task. Reports F1 score.
Evaluates the robustness of tabular machine learning models when the set of available features dynamically changes in open environments. It measures performance degradation across classification and regression tasks under varying degrees of feature removal (20% to 100%). Use when the user wants to benchmark on TabFSBench datasets, or asks about evaluating this task. Reports performance gap (Δ).