Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 8,833–8,856 of 21,231 skills
This protocol evaluates racial bias in image captioning models by measuring performance disparities between images containing lighter-skinned versus darker-skinned individuals. It probes whether models systematically generate lower-quality captions or exhibit different linguistic patterns for darker-skinned subjects compared to lighter-skinned ones, even when visual content is controlled. Use when the user wants to benchmark on COCO 2014 validation, or asks about evaluating this task. Reports...
This benchmark evaluates whether language models appropriately refuse or comply with contextually ambiguous, incomplete, unsupported, or safety-related requests. It probes a model's ability to distinguish between benign queries that should be answered and problematic queries that should be declined, while avoiding exaggerated over-refusal on safe prompts. Use when the user wants to benchmark on CoCoNot, or asks about evaluating this task. Reports compliance rate.
Evaluates the ability of vision-language models to re-rank candidate image captions based on their semantic alignment with extracted visual context. It probes how well models can leverage object-level visual information to improve caption relevance and accuracy. Use when the user wants to benchmark on COCO Captions (Karpathy test split), or asks about evaluating this task. Reports BERTScore (B-S).
Evaluates neural abstractive summarization models on their ability to generate factual, fluent, and relevant narrative summaries of randomized controlled trials (RCTs) from Cochrane systematic reviews. Probes the models' susceptibility to hallucination and their capacity to correctly infer the directionality of clinical findings. Use when the user wants to benchmark on Cochrane RCT Summaries, or asks about evaluating this task. Reports Manual Factuality.
Evaluates large language models' ability to generate correct, compilable COBOL code from natural language specifications, and to translate bidirectionally between COBOL and Java. It probes functional correctness, compilation reliability, and practical utility for legacy system modernization. Use when the user wants to benchmark on COBOLEval, COBOLCodeBench, COBOL-JavaTrans, or asks about evaluating this task. Reports Pass@1.
Evaluates the ability of debiasing frameworks to transform biased (MNAR) recommendation data into unbiased (MAR) representations, measuring how well debiased rankings align with ground-truth user preferences. It probes whether reweighting or perturbation mechanisms successfully mitigate selection and staleness biases without degrading predictive performance. Use when the user wants to benchmark on Coat, or asks about evaluating this task. Reports AUC.
Evaluates a model's ability to mitigate gender bias in downstream NLP tasks by measuring performance disparities across demographic groups. It probes whether models assign equal similarity scores to gender-swapped sentence pairs, maintain neutrality in natural language inference, and classify occupations without gender-based true positive rate gaps. Use when the user wants to benchmark on Bias-STS-B, Bias-NLI, Bias-in-Bios, or asks about evaluating this task. Reports average absolute differen...
Evaluates the ability of AI-generated image detectors to distinguish between real and synthetic images across diverse generative models, lossy compression formats, and real-world in-the-wild sources. It specifically probes out-of-distribution generalization and robustness to common post-processing transformations like JPEG compression, blurring, and noise. Use when the user wants to benchmark on Co-SpyBench, Co-SpyBench/in-the-wild, AIGCDetectBenchmark, GenImage, or asks about evaluating this...
Evaluates long-horizon agentic reasoning, tool-augmented decision making, and factual grounding under conflict-aware verification. It probes how well systems can audit divergent reasoning steps, maintain structured knowledge, and produce accurate answers across multi-hop and interdisciplinary tasks. Use when the user wants to benchmark on GAIA, HLE, Chinese-SimpleQA, or asks about evaluating this task. Reports accuracy.
Evaluates a joint deep learning architecture for simultaneous monocular depth estimation and semantic segmentation on aerial drone imagery, measuring prediction accuracy and inference speed against single-task and joint baselines. Use when the user wants to benchmark on MidAir, Aeroscapes, or asks about evaluating this task. Reports mIoU.
Evaluates the capability of models to detect section boundaries in clinical notes at the token level. It probes how well different architectures handle structured sentence-level segmentation versus unstructured freetext narrative variability in medical records. Use when the user wants to benchmark on MIMIC-IV Clinical Notes, or asks about evaluating this task. Reports Token-level F1.
Tests the model's ability to classify CNS-active versus CNS-inactive drugs and enrich active compounds from large virtual screening databases. It probes generalization on small-sample molecular datasets using external validation. Use when the user wants to benchmark on CNS Drug Dataset, or asks about evaluating this task. Reports AUC.
Evaluates abstractive summarization quality by scoring generated summaries against human-written references and expert rubric-based scores. It measures how well automatic metrics correlate with human judgments across different summarization systems. Use when the user wants to benchmark on CNNDM, or asks about evaluating this task. Reports COMET.
This evaluation benchmarks CNN-based models on automatic music tagging, measuring their ability to predict multiple genre, instrument, and mood labels from audio spectrograms. It assesses both standard classification performance and robustness to audio transformations like time-stretching and pitch shifting. Use when the user wants to benchmark on MagnaTagATune, Million Song Dataset, MTG-Jamendo, or asks about evaluating this task. Reports ROC-AUC.
Evaluates the trade-offs between predictive accuracy, model compression, and dynamic inference efficiency of CNN optimization techniques (pruning, quantization, early-exit) for edge deployment. It probes how different architectures handle static compression versus input-adaptive latency reduction under hardware-constrained conditions. Use when the user wants to benchmark on Unspecified classification dataset, or asks about evaluating this task. Reports accuracy (%).
Evaluates large language models' ability to generate valid, optimized molecular structures (SMILES) that satisfy multiple conflicting pharmacological and physicochemical property constraints. The benchmark probes the model's capacity for multi-objective reinforcement alignment, scaffold preservation, and strict adherence to property-wise improvement margins under both in-domain and out-of-distribution settings. Use when the user wants to benchmark on C-MuMOInstruct, or asks about evaluating t...
Evaluates the quality of generated responses in document-grounded conversations, specifically measuring how well models leverage external document context to produce engaging and fluent multi-turn dialogue. It assesses both automatic language modeling metrics and human-perceived response quality. Use when the user wants to benchmark on CMU.DoG, or asks about evaluating this task. Reports Perplexity.
Evaluates the zero-shot cross-modality transfer capability of open-vocabulary object detectors from RGB to X-ray imaging. It measures how well pre-trained RGB detectors can localize and classify objects in X-ray images without any fine-tuning or labeled X-ray data. Use when the user wants to benchmark on DET-COMPASS, PIXray, PIDray, CLCXray, DvXray, HiXray, or asks about evaluating this task. Reports AP.
Evaluates large language models' Chinese language understanding and multitask knowledge across 67 subjects spanning STEM, humanities, social sciences, and China-specific domains. It probes memorization, reasoning, and instruction-following capabilities in a multiple-choice question-answering format. Use when the user wants to benchmark on CMMLU, or asks about evaluating this task. Reports macro average accuracy.
Evaluates the quality of text-to-speech models trained on the CML-TTS dataset across seven low-resource languages. It probes speaker similarity preservation and text fidelity in synthesized audio under both seen and unseen speaker (zero-shot) conditions. Use when the user wants to benchmark on CML-TTS, or asks about evaluating this task. Reports SECS.
Evaluates music reward models on their ability to align with human aesthetic judgments and follow compositional multimodal instructions (text, lyrics, audio). It probes both absolute musicality scoring and relative pairwise preference ranking across diverse generation models. Use when the user wants to benchmark on PAM, MusicEval, Music Arena, CMI-Pref, or asks about evaluating this task. Reports Linear Correlation Coefficient (LCC), Spearman Rank Correlation (SRCC), Kendall-Tau (K-Tau), Pair...
Evaluates Chinese medical text embedding models across retrieval, reranking, and semantic textual similarity tasks. It probes the model's ability to capture domain-specific semantic alignment while measuring the trade-off between retrieval accuracy and inference efficiency. Use when the user wants to benchmark on CMedTEB, or asks about evaluating this task. Reports Avg.
Evaluates the transferability and quality of self-supervised multiview representations by measuring downstream performance on image classification, video action recognition, and semantic segmentation tasks. Use when the user wants to benchmark on ImageNet, UCF-101, HMDB-51, NYU-Depth-V2, STL-10, or asks about evaluating this task. Reports Top-1 classification accuracy (%).
Evaluates a contrastive multi-level graph neural network for session-based recommendation by measuring its ability to predict the next item in a user session using pairwise and high-order transition patterns. Use when the user wants to benchmark on Tmall, Diginetica, Nowplaying, or asks about evaluating this task. Reports Recall@K.