Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 8,113–8,136 of 20,861 skills
This benchmark evaluates a model's ability to perform multitask learning across ten diverse natural language processing tasks by framing them as a unified question-answering problem. It probes zero-shot generalization, domain adaptation, and the effectiveness of anti-curriculum training strategies without relying on task-specific modules. Use when the user wants to benchmark on decaNLP, or asks about evaluating this task. Reports decaScore.
Evaluates transformer-based models on word-level extractive summarization for policy debate evidence. It measures how well models can identify and extract relevant tokens to form summaries of debate arguments. Use when the user wants to benchmark on DebateSum, or asks about evaluating this task. Reports ROUGE F1.
Evaluates object detection and pose classification capabilities on historical European paintings. Probes a model's ability to recognize culturally heritage-specific entities and human-like poses in artistic contexts rather than natural photographs. Use when the user wants to benchmark on DEArt, or asks about evaluating this task. Reports mAP@0.5.
Evaluates the capability of neural architectures to perform binary emotion recognition (valence, arousal, dominance) directly from raw, multi-channel EEG time-series data without hand-crafted features. It measures how well a model generalizes across subjects using a standard cross-validation protocol. Use when the user wants to benchmark on DEAP, or asks about evaluating this task. Reports Accuracy.
This benchmark probes the acoustic faithfulness of Audio Multimodal Large Language Models (Audio MLLMs) by measuring how reliably they attend to acoustic cues (emotional prosody, background sounds, speaker identity) when faced with conflicting textual semantics or misleading prompts. It specifically diagnoses the tendency of models to prioritize text over audio (text dominance) under progressive levels of interference. Use when the user wants to benchmark on DEAF, or asks about evaluating thi...
Evaluates an agent's ability to iteratively collect clinical evidence and generate a ranked list of differential diagnoses for a simulated patient. It probes the system's diagnostic reasoning, evidence-gathering efficiency, and alignment with ground-truth pathologies. Use when the user wants to benchmark on DDXPlus, or asks about evaluating this task. Reports DDF1.
Probes large language models' ability to perform iterative medical diagnostic reasoning through a simulated patient-doctor dialogue. It evaluates whether models can gather clinical evidence, generate differential diagnoses, and converge on the correct ground-truth pathology within a limited number of interaction turns. Use when the user wants to benchmark on DDXPlus, or asks about evaluating this task. Reports diagnostic accuracy.
This benchmark evaluates the ability of deep learning models to detect AI-generated art (deeparts) versus conventional art (conarts) and identify their generative origins. It probes detector generalization across different state-of-the-art diffusion models and tests continual learning capabilities under evolving data streams with strict memory constraints. Use when the user wants to benchmark on DDDB, or asks about evaluating this task. Reports AA.
Evaluates the performance of machine learning and traditional algorithms for reconstructing particle tracks in drift chamber detectors. It probes hit-level matching accuracy, track-level reconstruction efficiency, charge identification correctness, and momentum resolution under realistic detector conditions. Use when the user wants to benchmark on DCTracks, or asks about evaluating this task. Reports track efficiency.
Evaluates the ability of multi-task learning models to predict post-click conversion rate (CVR) and click-through conversion rate (CTCVR) while mitigating selection bias in recommendation and search systems. It probes whether causal debiasing mechanisms improve ranking quality on both clicked and unclicked items across diverse e-commerce and industrial datasets. Use when the user wants to benchmark on Ali-CCP, Ali-Express (AE-ES), Ali-Express (AE-FR), Ali-Express (AE-NL), Ali-Express (AE-US),...
Evaluates the effectiveness of data curation strategies for language models by training base models on curated corpora and measuring performance on 53 downstream tasks. It isolates data quality effects from architectural and computational variables using fixed training recipes across multiple compute scales. Use when the user wants to benchmark on DCLM downstream tasks, or asks about evaluating this task. Reports MMLU 5-shot accuracy.
Evaluates the reconstruction fidelity of high-spatial-compression autoencoders and the generation quality and efficiency of latent diffusion models that utilize them. It benchmarks performance across multiple datasets and resolutions to assess trade-offs between compression ratio, image quality, and computational throughput. Use when the user wants to benchmark on ImageNet, FFHQ, MapillaryVistas, MJHQ, or asks about evaluating this task. Reports rFID, FID.
Evaluates whether large language models can discover genuinely new biological knowledge by generating correct scientific hypotheses or mechanisms from post-release literature, enforcing strict temporal separation to prevent data leakage. Use when the user wants to benchmark on DBench-Bio, or asks about evaluating this task. Reports Score.
Evaluates protein-ligand binding affinity prediction models on a modification-aware dataset, testing their ability to generalize across different train-test splits (new ligands, new proteins, modifications) and assessing robustness to wild-type overfitting and few-shot fine-tuning. Use when the user wants to benchmark on DAVIS-complete, or asks about evaluating this task. Reports Rp.
Evaluates large language models on their ability to parse, translate, and perform arithmetic reasoning with datetime information. It probes structured string formatting (ISO-8601) and multi-step temporal calculations across diverse linguistic contexts. Use when the user wants to benchmark on DATETIME, or asks about evaluating this task. Reports accuracy.
Evaluates zero-shot generalization of vision-language models across a broad suite of classification and retrieval tasks, with a focus on image-text alignment and retrieval accuracy. Use when the user wants to benchmark on DataComp Zero-Shot Suite, or asks about evaluating this task. Reports ImageNet accuracy.
Evaluates the zero-shot generalization capability of vision-language models across a diverse suite of image classification benchmarks. It measures how well pre-trained image-text alignment transfers to unseen downstream tasks without fine-tuning. Use when the user wants to benchmark on ImageNet, DataComp evaluation datasets, or asks about evaluating this task. Reports ImageNet.
Evaluates semi-supervised anomaly detection models for supply chain fraud under severe class imbalance and limited label availability. Probes the ability to leverage unsupervised pre-filtering and self-training to improve precision, recall, and F1-score while maintaining low false positive rates. Use when the user wants to benchmark on DataCo Smart Supply Chain Dataset, or asks about evaluating this task. Reports F1-Score.
Evaluates whether distributional or embedding similarity between a model's pretraining data and downstream tasks predicts few-shot or finetuned performance. It probes the 'similarity hypothesis' by measuring correlations between aggregate and example-level text similarities and model accuracy. Use when the user wants to benchmark on BIG-bench Lite, GLUE, or asks about evaluating this task. Reports correlation coefficient.
Evaluates a model's ability to retrieve relevant tables and text passages from a hybrid corpus to satisfy complex, multi-part analytical user requests (Data Product Requests). It probes multi-modal data integration and semantic clustering capabilities by requiring complete alignment between a request and its underlying data assets. Use when the user wants to benchmark on HybridQA, TAT-QA, ConvFinQA, or asks about evaluating this task. Reports Full Recall@100.
Evaluates the accuracy and computational efficiency of the DARF simulation framework for predicting speech recognition thresholds (SRT) in normal-hearing and hearing-impaired listeners across various acoustic maskers and hearing aid conditions. Use when the user wants to benchmark on Empirical SRT datasets (Hochmuth et al. 2015, Hülsmeier et al., Schädler et al. 2020a), or asks about evaluating this task. Reports SRT.
Evaluates speech enhancement models by measuring their ability to restore clean speech from noisy or degraded inputs. It probes perceptual quality, acoustic fidelity, and semantic preservation using human listening tests and objective feature-space distances. Use when the user wants to benchmark on DAPS, Noisy VCTK, or asks about evaluating this task. Reports MUSHRA.
Evaluates cross-domain patent retrieval systems by measuring how well they rank relevant patent documents or passages when queries and targets share or lack overlapping IPC3 classifications. Use when the user wants to benchmark on DAPFAM, or asks about evaluating this task. Reports NDCG@100.
Evaluates the quality of Chinese vision-language pre-training datasets by measuring downstream performance on cross-modal retrieval and multimodal reasoning benchmarks after continual pre-training. Use when the user wants to benchmark on Flickr30K-CN, MSCOCO-CN, MUGE, DCI-CN, DOCCI-CN, or asks about evaluating this task. Reports R@1/5/10.