Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 8,593–8,616 of 21,228 skills
Measures social biases in masked language models by comparing the likelihood assigned to stereotypical versus anti-stereotypical sentence pairs. It quantifies how strongly models favor historically disadvantaged groups' stereotypes across nine demographic categories. Use when the user wants to benchmark on CrowS-Pairs, or asks about evaluating this task. Reports bias metric.
Evaluates conversational passage ranking by measuring how effectively a model ranks relevant documents across multi-turn search queries, balancing term similarity with contextual coherence. Use when the user wants to benchmark on TREC CAsT 2019, or asks about evaluating this task. Reports nDCG.
Evaluates algorithms for aggregating multiple noisy, crowdsourced transcriptions of the same audio recording into a single high-quality reference. It probes how well methods handle sequential textual noise, estimate worker reliability, and adapt across different audio quality domains. Use when the user wants to benchmark on CROWDSPEECH, VOXDIY, CROWDWSA2019, or asks about evaluating this task. Reports WER.
Evaluates the capability of decentralized federated learning (DFL) models to detect malware and classify benign states in IoT crowdsensing environments. It probes robustness under varying node counts, peer-to-peer network topologies, and data heterogeneity (IID vs. non-IID Dirichlet splits). Use when the user wants to benchmark on Crowdsensing Intrusion Detection Dataset, or asks about evaluating this task. Reports accuracy.
Evaluates object detectors' ability to identify humans in highly crowded and heavily occluded scenes. It covers three annotation levels (full body, visible body, head) and assesses cross-dataset generalization for pedestrian and head detection tasks. Use when the user wants to benchmark on CrowdHuman, or asks about evaluating this task. Reports mMR.
Evaluates dense optical flow estimation accuracy and long-term temporal consistency in complex crowd surveillance scenarios, specifically testing robustness to non-rigid, self-occluding motion and small object tracking. Use when the user wants to benchmark on CrowdFlow, or asks about evaluating this task. Reports EPE.
Evaluates the ability of pose estimation models to accurately predict 2D keypoints for humans and animals in crowded, occluded, and multi-instance scenarios. It probes robustness to detection ambiguity, overlapping instances, and the transferability of conditional pose inputs from bottom-up detectors to top-down refiners. Use when the user wants to benchmark on CrowdPose, OCHuman, COCO, Multi-Animal (SchoolingFish, Marmosets, Tri-Mouse), or asks about evaluating this task. Reports AP.
Evaluates the ability of generative dialogue state tracking models to accurately predict and maintain the complete set of user intents (domain-slot-value triples) across dialogue turns. It specifically probes cross-lingual and cross-ontology transfer capabilities by measuring how well models trained on one language or ontology generalize to another. Use when the user wants to benchmark on CrossWOZ-en, or asks about evaluating this task. Reports Joint Goal Accuracy.
Evaluates cross-lingual speech-to-speech translation (S2ST) systems on translation accuracy and prosody preservation. It measures how well a cascade-based S2ST pipeline preserves speaker identity and naturalness while translating speech across different language pairs. Use when the user wants to benchmark on CVSS-T, Indic-TTS, Fisher, MuST-C, VoxPopuli, or asks about evaluating this task. Reports BLEU.
Evaluates the quality of automatically induced cross-lingual summary alignments in the CrossSum dataset by measuring human agreement on whether two summaries correspond to the same source article. Use when the user wants to benchmark on CrossSum, or asks about evaluating this task. Reports alignment_accuracy.
Evaluates Vision-Language Models' ability to perform precise point-level geometric correspondence across multiple viewpoints. It probes fine-grained spatial grounding, visibility reasoning, cross-view correspondence judgment, and continuous 2D coordinate pointing. Use when the user wants to benchmark on CrossPoint-Bench, or asks about evaluating this task. Reports average accuracy.
Evaluates cross-lingual semantic similarity between news article pairs across four dimensions (Who, What, Where, When) in Ukrainian, Polish, Russian, and English. It probes a model's ability to align event-level information across languages while ignoring publication dates. Use when the user wants to benchmark on CrossNews-UA, or asks about evaluating this task. Reports macro-averaged F1-score.
Evaluates cross-domain named entity recognition by measuring how well models adapt from a source domain (CoNLL2003) to five specialized target domains. These domains feature unique, domain-specific entity types that test the model's ability to generalize beyond standard categories. Use when the user wants to benchmark on CrossNER, or asks about evaluating this task. Reports F1 score.
Evaluates multilingual image captioning models across 36 languages, probing their ability to generate stylistically coherent and culturally representative descriptions without relying on direct translation artifacts. It measures how well models generalize to low-resource and geographically diverse languages. Use when the user wants to benchmark on Crossmodal-3600, or asks about evaluating this task. Reports CIDEr.
Evaluates unsupervised cross-modality domain adaptation for medical image segmentation (Vestibular Schwannoma and Cochlea) and tumour grading (Koos classification) from ceT1 to T2 MRI. Use when the user wants to benchmark on crossMoDA, or asks about evaluating this task. Reports DSC.
Evaluates compositional generalization in medical vision-language models across a structured Modality–Anatomy–Task (MAT) schema. It probes zero-shot cross-task transfer, generalization to novel MAT combinations, and robustness under low-data regimes using a unified visual question answering interface. Use when the user wants to benchmark on CrossMed, or asks about evaluating this task. Reports top-1 classification accuracy.
Evaluates cross-lingual speech-to-text retrieval and intent detection capabilities across multiple datasets, testing how well speech queries can retrieve relevant text documents or classify intents without intermediate ASR or translation pipelines. Use when the user wants to benchmark on Kallaama-Retrieval-Eval, Fleurs-Retrieval-Eval, Urban Bus, WolBanking77, or asks about evaluating this task. Reports nDCG@5.
Evaluates zero-shot crosslingual generalization of multilingual LLMs after multitask finetuning. Probes language-agnostic task understanding, robustness to prompt translation, and scaling behavior across NLU, generative, and code tasks. Use when the user wants to benchmark on XNLI, XCOPA, XStoryCloze, XWinograd, HumanEval, or asks about evaluating this task. Reports accuracy.
Evaluates the robustness of multimodal LLMs against explicit and implicit jailbreak attacks while measuring their utility on benign queries. It probes whether a defense model can successfully refuse harmful image-text prompts without over-restricting safe inputs. Use when the user wants to benchmark on JailBreakV, VLGuard, FigStep, MM-SafetyBench, SIUO, MMBench, or asks about evaluating this task. Reports Attack Success Rate (ASR).
Evaluates a model's ability to generate novel, drug-like molecules with high binding affinity for unseen protein pockets in structure-based drug design. It probes the trade-offs between binding energy, molecular properties, and synthesis feasibility. Use when the user wants to benchmark on CrossDocked-100k, or asks about evaluating this task. Reports Vina Dock.
This evaluation probes the ability of multimodal foundation models to generate factual content without hallucination across text, image, and audio-visual modalities. It measures how well reference-free ranking methods correlate with human judgments or gold-standard references to rank model outputs by hallucination severity. Use when the user wants to benchmark on WikiBio, MHaluBench, AVHalluBench, or asks about evaluating this task. Reports System($ ho$).
Evaluates systems' ability to predict target-language pronoun class labels from source-language pronouns using lemmatized, POS-tagged translations and word alignments. It probes cross-lingual anaphora resolution and functional ambiguity handling in machine translation pipelines. Use when the user wants to benchmark on WMT 2016 Cross-lingual Pronoun Prediction Task, or asks about evaluating this task. Reports macro-averaged recall.
Evaluates the intelligibility, speaker similarity, and naturalness of synthesized speech in cross-lingual voice cloning and TTS scenarios. It also measures the accuracy of a language-agnostic speaking rate predictor for duration modeling across multiple languages. Use when the user wants to benchmark on Emilia, Seed-TTS-eval, LibriSpeech-PC test-clean, FLEURS, or asks about evaluating this task. Reports WER.
This evaluation probes a model's ability to perform cross-domain sequential recommendation by leveraging user interaction histories across two domains, even when user overlap is minimal or absent. It measures how well the model captures domain-specific and shared sequential patterns to rank candidate items accurately. Use when the user wants to benchmark on Micro Video, Amazon, or asks about evaluating this task. Reports AUC.