Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 6,337–6,360 of 20,816 skills
Evaluates a vision transformer's ability to perform multi-label classification on chest X-ray images. It probes the model's capacity to detect multiple pathologies simultaneously and model inter-label dependencies using learnable label tokens. Use when the user wants to benchmark on NIH-CXR14, CheXpert-5, CheXpert-13, or asks about evaluating this task. Reports AUC (%).
Evaluates multivariate time series forecasting models on their ability to capture short-term local dependencies and long-term periodic patterns across diverse real-world datasets with varying temporal scales and frequencies. Use when the user wants to benchmark on Traffic, Solar-Energy, Electricity, Exchange-Rate, or asks about evaluating this task. Reports RSE.
Evaluates narrative understanding, commonsense reasoning, and topic coherence in speech-text models by selecting the most plausible continuation from multiple candidates. The benchmark tests both speech-to-speech and text-to-text modes to assess cross-modal alignment and reasoning capabilities under compute constraints. Use when the user wants to benchmark on HellaSwag (sHellaSWAG), StoryCloze, TopicStoryCloze, or asks about evaluating this task. Reports accuracy.
Evaluates supervised open information extraction (OIE) models on extracting schema-free predicate-argument tuples from sentences. Probes the model's ability to correctly identify predicates and arguments while maintaining syntactic head alignment and argument ordering. Use when the user wants to benchmark on LSOIE, or asks about evaluating this task. Reports F1.
Evaluates models' ability to detect and rank lexical semantic change in Spanish diachronic corpora. It probes both graded ranking of semantic shift magnitude and binary classification of sense gain/loss or change presence. Use when the user wants to benchmark on LSCDiscovery, or asks about evaluating this task. Reports Spearman rank correlation (SPR), F1 score.
Evaluates sequential recommendation models on next-item prediction tasks across various domains (movies, products, games) with varying sequence lengths and sparsity. Use when the user wants to benchmark on ML-1M, Amazon-Beauty, Amazon-Video, Amazon-Sports, Steam, XLong, or asks about evaluating this task. Reports Recall@10.
This evaluation probes the ability of a generative error correction model to refine audio-visual speech recognition transcripts under varying noise conditions. It measures how effectively multimodal cues (lip video and audio) combined with N-best hypotheses can reduce transcription errors compared to baseline systems. Use when the user wants to benchmark on LRS3, or asks about evaluating this task. Reports WER.
Evaluates large vision-language models on high-resolution remote sensing imagery by testing their ability to answer questions about color, count, position, and open-ended descriptions. It probes the model's perception and reasoning capabilities on complex, large-scale satellite/aerial images. Use when the user wants to benchmark on MME-RealWorld-RS, LRS-VQA, or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to generate accurate radiology reports from single-view X-ray images under varying quality conditions, from standard to severely degraded. It probes robustness to clinical acquisition artifacts and tests whether the model can extract quality-invariant diagnostic features without relying on historical patient data. Use when the user wants to benchmark on MIMIC-CXR LRRG Benchmarks, or asks about evaluating this task. Reports CheXbert F1.
Evaluates cross-lingual robustness of spoofing countermeasures by measuring spoof rejection rates on a multilingual synthetic-speech corpus at a fixed operating point calibrated on external benchmarks. Probes how language and synthesizer identity independently affect spoof detection performance. Use when the user wants to benchmark on Low-Resource Language Spoofing Corpus, or asks about evaluating this task. Reports spoof rejection rate (SRR).
Evaluates relation extraction models under low-resource conditions (8-shot, 10%, 100% training data) across diverse domains and languages. It probes few-shot learning capabilities, robustness to long-tailed class distributions, and the effectiveness of data augmentation and self-training strategies. Use when the user wants to benchmark on SemEval 2010 Task 8, TACREV, DialogRE, DuIE2.0, Wiki80, ChemProt, SciERC, CMeIE, or asks about evaluating this task. Reports Macro F1.
Evaluates cross-lingual transfer learning for Named Entity Recognition in low-resource Indian languages (Hindi and Marathi) by measuring how well models trained on combined or assisting-language datasets generalize to target language test sets compared to monolingual baselines. Use when the user wants to benchmark on IIT Bombay (Marathi), IJCNLP (Hindi), Wiki ANN (Hindi and Marathi), or asks about evaluating this task. Reports scores.
Evaluates automatic speech recognition (ASR) performance on low-resource and high-resource languages using synthetic audio generated from text augmentation. It probes the model's ability to generalize to unseen lexical and syntactic variations when trained on limited real speech data. Use when the user wants to benchmark on Vatlongos, Nashta, Kakabe, Shinekhen Buryat, LibriSpeech, or asks about evaluating this task. Reports WER.
Evaluates the effectiveness of a diffusion-based data augmentation method on mitigating long-tail bias and improving cross-domain generalization in remote-sensing semantic segmentation. It specifically probes whether synthetic label-image pairs can increase minority-class exposure while preserving domain realism and data distribution. Use when the user wants to benchmark on LoveDA, or asks about evaluating this task. Reports mIoU.
Evaluates monaural speech enhancement models by measuring how effectively they restore clean speech from noisy recordings. The benchmark probes the model's ability to handle diverse acoustic conditions and varying signal-to-noise ratios across two standard speech enhancement datasets. Use when the user wants to benchmark on VCTK+DEMAND, DNS Challenge 2020, or asks about evaluating this task. Reports PESQ.
Evaluates a model's ability to generate long-term (25s–50s) high-fidelity musical waveforms that are rhythmically synchronized with visual cues from diverse video scenarios like dancing and sports. It measures both the temporal alignment of generated beats with ground-truth audio and the overall subjective musical quality. Use when the user wants to benchmark on LORIS, or asks about evaluating this task. Reports F1.
This benchmark probes a model's ability to infer the exact number of training images used to fine-tune a Low-Rank Adaptation (LoRA) adapter solely from its learned weight matrices. It evaluates how well spectral and norm-based features of LoRA parameters correlate with and reveal the scale of the underlying training dataset. Use when the user wants to benchmark on LoRA-WiSE, or asks about evaluating this task. Reports MAE.
This evaluation probes the effectiveness of LoRA fine-tuning across 31 diverse NLP tasks by comparing base LLMs against their fine-tuned counterparts and proprietary models like GPT-4. It measures how much performance lift fine-tuning provides and whether smaller open-weight models can surpass larger closed-source models after adaptation. Use when the user wants to benchmark on magicoder, mmlu, glue_wnli, arc_combined, wikisql, boolq, customer_support, glue_cola, winogrande, glue_sst2, dbpedi...
Evaluates parameter-efficient fine-tuning (LoRA) for cross-domain few-shot object detection on aerial imagery. It probes the model's ability to generalize to new domains with limited labeled data while mitigating overfitting. Use when the user wants to benchmark on DOTA, DIOR, or asks about evaluating this task. Reports mAP@0.5.
Evaluates the effectiveness of transformer-specific dropout methods (e.g., HiddenKey, DropKey, HiddenCut) when combined with LoRA for parameter-efficient fine-tuning. It probes the model's ability to mitigate overfitting in LoRA settings across diverse natural language understanding and generation tasks. Use when the user wants to benchmark on GLUE, E2E, WebNLG, or asks about evaluating this task. Reports Accuracy, BLEU.
Evaluates automatic speech recognition (ASR) model performance across varying training data scales and model sizes, measuring generalization to in-domain and out-of-domain English speech benchmarks. Use when the user wants to benchmark on Loquacious Set, Librispeech, Voxpopuli, CommonVoice, or asks about evaluating this task. Reports WER.
Evaluates an adaptive dual-phase LLM inference acceleration system for multi-turn dialogues, probing its ability to maintain generation accuracy and reduce computational overhead across varying query positions in long-context conversations. It specifically tests whether the system can generalize beyond positional heuristics used by static KV cache compression methods. Use when the user wants to benchmark on MFQA-en, 2WikiMQA, Musique, HotpotQA, NrtvQA, Qasper, MultiNews, GovReport, QMSum, TRE...
Evaluates a model's ability to accurately regress one-loop scattering amplitudes across a high-dimensional kinematic phase space. It specifically probes precision in challenging regions and the reliability of uncertainty quantification inherent to Bayesian neural networks. Use when the user wants to benchmark on One-loop gg→γγg(g) amplitudes, or asks about evaluating this task. Reports Δ (relative amplitude error).
Probes long-context multi-document question answering by requiring models to synthesize evidence from all provided documents (10K–250K+ tokens) across financial reports and academic papers. It tests information extraction, comparison, clustering, and chain-of-reasoning capabilities in heterogeneous, document-level agentic retrieval settings. Use when the user wants to benchmark on Loong, or asks about evaluating this task. Reports Avg Score.