Data & Analytics
Data analysis, BI, visualization, datasets, statistics, and ML workflows
Browse data & analytics skills
Showing 2,257–2,280 of 13,068 skills
Evaluates the ability of individual machine learning classifiers and ensemble strategies to detect network intrusions and classify traffic types. It probes model robustness, precision-recall trade-offs, and computational efficiency across diverse real-world network traffic datasets with varying attack profiles. Use when the user wants to benchmark on RoEduNet-SIMARGL2021, CICIDS-2017, or asks about evaluating this task. Reports F1 Score.
Evaluates machine learning models' ability to detect and classify cyberattacks in Industrial Control Systems (ICS) network traffic. It probes the capability to distinguish between normal operations and specific attack types like DDoS, IP-Scan, MitM, Port-Scan, and Replay using flow-level features. Use when the user wants to benchmark on ICS-Flow, or asks about evaluating this task. Reports F1-score.
Probes the causal impact of LLM-assisted peer reviews on paper scoring and acceptance outcomes at a major machine learning conference. It measures whether AI-assisted reviews systematically inflate scores and increase acceptance probabilities, particularly for borderline submissions. Use when the user wants to benchmark on ICLR Conference Reviews (2018-2024), or asks about evaluating this task. Reports acceptance_rate_difference.
Evaluates machine learning models' ability to classify electron antineutrino (IBD) events from background accidents in a liquid scintillator detector. It measures how well the models preserve signal efficiency while controlling background contamination compared to traditional cut-based selection. Use when the user wants to benchmark on JUNO IBD/Accident Dataset, or asks about evaluating this task. Reports efficiency.
Evaluates the quality of the HPLT v2 multilingual corpus by training downstream models (masked language models, generative LMs, and MT systems) and measuring their performance on standard linguistic, natural language understanding, and machine translation benchmarks. Use when the user wants to benchmark on Universal Dependencies (UD) treebanks, WikiAnn, FLORES-200, or asks about evaluating this task. Reports BLEU.
Evaluates machine learning and deep learning models on predicting clinical outcomes from high-resolution ICU time-series data. It probes capabilities in handling class imbalance, long temporal dependencies, and varying data resolutions across stay-level and online monitoring tasks. Use when the user wants to benchmark on HiRID, or asks about evaluating this task. Reports AUPRC.
This benchmark evaluates a model's ability to perform Named Entity Recognition (NER) on Hindi text. It probes the model's capacity to identify and classify entity spans (e.g., Person, Location, Organization, and others) in a language characterized by free word order, lack of capitalization, and spelling variations. Use when the user wants to benchmark on HiNER, or asks about evaluating this task. Reports F1-Score.
Evaluates the ability of spatiotemporal graph neural networks to perform multistep-ahead forecasting on correlated time series while simultaneously learning hierarchical cluster structures end-to-end. It probes the model's capacity to leverage relational inductive biases and self-supervised aggregation for improved prediction accuracy. Use when the user wants to benchmark on METR-LA, PEMS-BAY, AQI, CER-E, or asks about evaluating this task. Reports MAE.
This evaluation framework quantifies two distinct types of hallucinations in large language models: factuality (truthfulness of generated information) and faithfulness (consistency with input context or instructions). It probes model performance across 15 diverse knowledge-intensive tasks, including closed-book QA, summarization, reading comprehension, and fact-checking, using zero-shot and few-shot in-context prompts. Use when the user wants to benchmark on FEVER, FaithDial, NQ-open, TriviaQ...
Evaluates a Convolutional Autoencoder's ability to compress and reconstruct geophysical turbulence fields while preserving high-order statistical moments. It specifically probes the model's capacity to capture non-Gaussian, intermittent structures like extreme vertical drafts without degrading point-wise accuracy. Use when the user wants to benchmark on Stratified turbulence simulation data, or asks about evaluating this task. Reports MAPE on kurtosis ($K_w$).
Evaluates the ability of a generative model to synthesize realistic 3-component broadband ground motion acceleration time histories conditioned on seismic parameters. It probes the model's capacity to match empirical spectral intensities (PSA, FAS, EAS) and capture aleatory variability across different frequency bands and tectonic settings. Use when the user wants to benchmark on BBP dataset, Kik-net dataset, or asks about evaluating this task. Reports Normalized model residual (epsilon).
Evaluates multi-document summarization performance and model explainability by comparing sentence vs. paragraph inputs and analyzing how attention weights correlate with reference summary similarity to reveal positional bias. Use when the user wants to benchmark on MultiNews, WikiSum, or asks about evaluating this task. Reports ROUGE-F (1/2/L).
Evaluates machine learning models on glycan analysis tasks, including taxonomic classification, immunogenicity prediction, glycosylation type prediction, and protein-glycan binding affinity estimation. It probes the ability of sequence-based and graph-based encoders to capture multi-relational glycan structures and benefit from multi-task learning. Use when the user wants to benchmark on GlycanML, or asks about evaluating this task. Reports Macro-F1.
Evaluates machine learning models' ability to predict particle-level dynamic propensity and dynamic heterogeneity from static amorphous structural configurations in glass-forming liquids. Use when the user wants to benchmark on GlassBench, or asks about evaluating this task. Reports Pearson correlation coefficient ($\rho_P$).
Evaluates a model's ability to generate images that accurately reflect complex, multidisciplinary textual prompts. It probes semantic correctness, visual plausibility (spelling, logical consistency, readability), and the integration of domain knowledge with reasoning during image generation. Use when the user wants to benchmark on GenExam, or asks about evaluating this task. Reports strict score.
Evaluates the accuracy and temporal consistency of 1-hourly global weather forecasts up to 90 hours lead time. It probes the model's ability to capture both large-scale atmospheric patterns and fine-scale variability across meteorological, energy, aviation, and marine variables. Use when the user wants to benchmark on ERA5 (2018 testing data), or asks about evaluating this task. Reports RMSE.
Evaluates the ability of a machine learning closure model to stabilize reduced-order models for turbulent geophysical fluid dynamics. Specifically, it probes whether an extreme learning machine can predict mode-dependent eddy viscosities to maintain long-time integration accuracy and statistical steady-state behavior in coarse-grained ocean circulation simulations. Use when the user wants to benchmark on Four-gyre barotropic circulation problem, or asks about evaluating this task. Reports L2-...
Evaluates the quality of generated Semantic Identifiers (SIDs) for generative retrieval in industrial recommendation and search systems. It measures how well SIDs capture item relationships and distribute usage fairly, and assesses their impact on downstream retrieval hitrate and online transaction metrics. Use when the user wants to benchmark on FORGE, or asks about evaluating this task. Reports HR@K.
This benchmark probes a model's ability to classify financial news sentences or phrases into positive, neutral, or negative semantic orientations. It specifically tests domain-specific sentiment analysis by evaluating how well models capture contextual cues, economic concepts, and directional event expectations in financial texts. Use when the user wants to benchmark on Financial PhraseBank, or asks about evaluating this task. Reports accuracy.
Evaluates the effectiveness, efficiency, and interpretability of explicit feature interaction models for click-through rate (CTR) prediction on large-scale, highly sparse advertising and recommendation datasets. Use when the user wants to benchmark on Avazu, Criteo, ML-1M, KDD12, iPinYou, KKBox, or asks about evaluating this task. Reports AUC.
Assesses whether multilingual LLM responses contain misinformation or fake content when queried on specific topics. The evaluation relies on an automated GPT-based judge to determine the presence of fake information in model generations. Use when the user has predictions and gold and needs to compute fake news occurrence.
This evaluation benchmarks pre-processing group fairness techniques on tabular datasets by measuring their impact on fairness metrics and downstream model performance. It probes whether bias mitigation methods improve group-level parity without significantly degrading predictive utility across varying decision thresholds. Use when the user wants to benchmark on Adult, COMPAS, German, or asks about evaluating this task. Reports disparate impact.
Evaluates a model's ability to remove the causal influence of protected attributes from predictions while maintaining predictive accuracy. It probes counterfactual fairness and causal effect removal on both synthetically generated causal graphs and real-world tabular datasets. Use when the user wants to benchmark on Synthetic Causal Case Studies, Law School Admissions, Adult Census Income, or asks about evaluating this task. Reports ATE.
Evaluates the ability of an AutoML-based fairness repair framework to mitigate bias in machine learning models while preserving predictive accuracy. It measures the trade-off between accuracy retention and bias reduction across multiple binary classification datasets and model architectures. Use when the user wants to benchmark on Adult Census (race), Bank Marketing (age), German Credit (sex), Titanic (sex), or asks about evaluating this task. Reports Accuracy difference.