Data & Analytics
Data analysis, BI, visualization, datasets, statistics, and ML workflows
Browse data & analytics skills
Showing 2,305–2,328 of 13,068 skills
Evaluates change point detection algorithms for identifying the onset of creative fatigue in digital advertising campaigns. It probes early warning capabilities, precision-recall trade-offs, and detection latency against ground-truth performance degradation events. Use when the user wants to benchmark on Synthetic Gradual Decline, Synthetic Sharp Decline, or asks about evaluating this task. Reports Delay (days).
Evaluates LLMs on Chinese public security domain tasks including text classification, information extraction, question answering, and text generation. Probes domain-specific accuracy, reliability, and contextual understanding in high-stakes law enforcement scenarios. Use when the user wants to benchmark on Weibo Sentiment Analysis, Rumor Detection, Telecommunication Fraud Detection, Drug-Related Case Reports, Public Security Case Reading Comprehension, Public Security Case Summary, or asks ab...
This metric quantifies the promotional or visionary framing of a machine learning paper's abstract, isolating rhetorical style from the underlying technical content. It uses counterfactual abstracts generated by diverse LLM personas and pairwise comparisons to produce a calibrated continuous score of rhetorical strength. Use when the user has predictions and gold and needs to compute Rhetorical Score.
Probes a model's ability to perform compositional visual reasoning and ground referring expressions in the presence of semantically similar distractors. It evaluates whether models can parse complex logical structures and distinguish fine-grained visual differences rather than relying on statistical biases. Use when the user wants to benchmark on Cops-Ref, or asks about evaluating this task. Reports accuracy.
Evaluates clinical prediction capabilities on unstructured notes and structured EHR data. It benchmarks zero-shot LLMs, finetuned BERTs, and conventional ML/DL models on mortality, readmission, and length-of-stay prediction tasks. The setup tests out-of-the-box prompting versus task-specific finetuning across diverse model families. Use when the user wants to benchmark on MIMIC-IV, TJH, or asks about evaluating this task. Reports AUROC.
Evaluates the out-of-distribution robustness of machine learning climate emulation models under time-domain shifts (training on historical data, testing on recent years) and source-domain shifts (training on one SSP scenario, testing on others). This protocol assesses how well models generalize to changing climate dynamics and divergent emission pathways. It specifically measures performance degradation when distribution shifts occur in temporal or scenario domains. Use when the user wants to...
Evaluates the computational efficiency and 3D reconstruction accuracy of a GPU-accelerated binary feature descriptor (CLATCH) compared to traditional and deep learning-based descriptors within a Structure-from-Motion pipeline. Use when the user wants to benchmark on Photogrammetry Image Sets (8 scenes), or asks about evaluating this task. Reports SfM Scene RMSE (pixels).
Evaluates the capability of machine learning models (traditional classifiers and CNNs on barcode-encoded features) to classify malware samples into benign or specific malware families. It probes how well structural patterns in 2D barcodes (QR and Aztec codes) capture executable features for downstream classification tasks. Use when the user wants to benchmark on CIC-MalMem-2022, or asks about evaluating this task. Reports accuracy.
Evaluates the ability of time series forecasting models to predict trajectories of low-dimensional chaotic dynamical systems. It probes how well models capture underlying deterministic chaos, smoothness, and multi-scale temporal dependencies without explicit trend or seasonality signals. Use when the user wants to benchmark on Chaotic Dynamical Systems Benchmark, or asks about evaluating this task. Reports sMAPE.
Evaluates multimodal large language models' ability to interpret financial charts, tables, and diagrams in Chinese, and answer domain-specific questions. It probes visual reasoning, statistical and structural analysis, and financial concept comprehension under zero-shot conditions. Use when the user wants to benchmark on CFBenchmark-MM, or asks about evaluating this task. Reports accuracy.
Evaluates Vision-Language Models' ability to analyze multi-scale candlestick charts (daily and weekly) and predict 30-day forward stock returns. It probes their capacity for visual technical analysis, trend recognition, and regression-based financial forecasting without relying on textual market data. Use when the user wants to benchmark on Multi-Scale Candlestick Stock Return Benchmark, or asks about evaluating this task. Reports IC (Information Coefficient).
Evaluates large language models' ability to perform structured reasoning over temporally segmented in-vehicle CAN traffic logs. It probes capabilities in temporal analysis, multi-condition inference, and behavioral interpretation for automotive cybersecurity forensics. Use when the user wants to benchmark on CAN-QA, or asks about evaluating this task. Reports accuracy.
Evaluates a multimodal language model's ability to process long-duration ECG signals for clinical forecasting, diagnostic classification, report generation, and statistical grounding. It probes temporal reasoning, multi-lead interpretation, and instruction-following in medical AI settings. Use when the user wants to benchmark on Icentia11k, PTB-XL, CSN, CODE-15%, CPSC-2018, HEEDB, Penn, MIMIC-IV-ECG, ECGBench / ECG-QA, ECG Grounding Benchmark, or asks about evaluating this task. Reports F1 sc...
This benchmark evaluates machine learning models on two distinct scientific tasks using the BubbleML dataset: predicting optical flow for bubble dynamics and solving multiphysics PDEs for temperature and velocity field propagation. It probes a model's ability to capture non-rigid object motion, sharp physical interfaces, and long-horizon temporal dynamics in phase-change simulations. Use when the user wants to benchmark on BubbleML, or asks about evaluating this task. Reports end-point error ...
Predicts county-level breast cancer screening rates using census-tract-level socioeconomic, demographic, and geospatial features. Evaluates the regression performance of Random Forest, Linear Regression, and Support Vector Machine models to identify which algorithm best captures underlying patterns in screening access disparities. Use when the user wants to benchmark on US Census Tracts Mammography Screening Dataset, or asks about evaluating this task. Reports R^2.
This evaluation probes the classification accuracy and prototype-based interpretability of deep learning models on mammography datasets. It measures how well models predict malignancy while ensuring that their learned prototypes align with domain-specific radiological features (e.g., mass/calcification types and BIRADS descriptors). Use when the user wants to benchmark on CBIS-DDSM, CMMD, VinDr-Mammo, or asks about evaluating this task. Reports F1.
Evaluates 3D medical image segmentation models on intracranial meningioma MRI scans. It probes volumetric accuracy and boundary sharpness across three tumor subregions (enhancing tumor, tumor core, whole tumor) under varying contrast and lesion size conditions. Use when the user wants to benchmark on BraTS 2023 Intracranial Meningioma Challenge, or asks about evaluating this task. Reports DSC.
Evaluates the ability of machine learning models to classify brain MRI images into four categories: glioma, meningioma, pituitary tumor, and no tumor. It tests feature extraction and decision fusion capabilities using both deep learning and traditional ML classifiers. Use when the user wants to benchmark on Kaggle Brain Tumor MRI dataset, Figshare Brain Tumor dataset, or asks about evaluating this task. Reports Accuracy.
Evaluates the out-of-distribution (OOD) generalization of machine learning models for molecular property prediction. It probes whether models trained on in-distribution (ID) molecules can accurately extrapolate to novel chemical spaces, highlighting the disconnect between ID accuracy and OOD robustness. Use when the user wants to benchmark on QM9, or asks about evaluating this task. Reports RMSE.
This benchmark evaluates a model's ability to answer naturally occurring yes/no questions based on a provided passage. It probes complex inferential reasoning and non-factoid inference, requiring the model to go beyond simple keyword matching or shallow statistical features to determine entailment or contradiction between the question and the passage. Use when the user wants to benchmark on BoolQ, or asks about evaluating this task. Reports accuracy.
Evaluates vision-language models on abductive and defeasible reasoning in videos depicting unpredictable events. It probes the ability to infer hidden causes from limited visual cues and revise hypotheses when new evidence emerges, testing reasoning beyond simple statistical recall. Use when the user wants to benchmark on BlackSwanSuite, or asks about evaluating this task. Reports accuracy.
Evaluates data-driven machine learning models for medium-range weather forecasting over India. It probes the ability of models to capture spatial and temporal atmospheric dynamics across diverse Indian microclimates for variables like geopotential height, temperature, and precipitation. Use when the user wants to benchmark on BharatBench, or asks about evaluating this task. Reports RMSE.
Evaluates the inherent trade-off between diversity (agreement of model rankings across tasks) and stability (sensitivity of final rankings to label noise) in multi-task machine learning benchmarks. It quantifies how much a benchmark's leaderboard ranking changes when trivial label noise is injected, and how diverse the rankings are across its constituent tasks. Use when the user wants to benchmark on GLUE, SuperGLUE, MTEB, BigBenchHard, MMLU, OpenLLM, VTAB, ImageNet, or asks about evaluating ...
This benchmark evaluates machine learning models on bioacoustic animal sound recognition across 12 diverse datasets spanning birds, mammals, anurans, and insects. It probes two core capabilities: multi-label species classification and temporal sound detection, testing models' ability to generalize across species and handle varying recording conditions and class imbalances. Use when the user wants to benchmark on wtkn, bat, cbi, hbdb, dogs, dcase, enabirds, hiceas, rfcx, hainan-gibbons, esc, s...