Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 3,337–3,360 of 22,874 skills
Evaluates an instance segmentation model's ability to accurately delineate and separate individual microstructural objects in high-resolution electron micrographs, particularly under varying instance densities. Use when the user wants to benchmark on UniEM-3M, or asks about evaluating this task. Reports mAP@0.5.
Evaluates retrieval and end-to-end generation performance of multimodal RAG systems on real-world PDF documents. It probes the ability of text-only, image-only, and multimodal (text-image fusion/joint) retrieval paradigms to locate relevant evidence and generate faithful, complete answers to cross-modality questions. Use when the user wants to benchmark on UNIDOC-BENCH, or asks about evaluating this task. Reports Precision@10, Recall@10, Faithfulness, Completeness.
Evaluates pixel-level biomedical image segmentation capability using convolutional networks. Probes the model's ability to precisely delineate cellular structures and membranes in electron and light microscopy images with limited training data. Use when the user wants to benchmark on EM segmentation challenge (ISBI 2012), PhC-U373, DIC-HeLa, or asks about evaluating this task. Reports warping error, IOU.
This protocol evaluates unbiased learning-to-rank models on their ability to correct position bias and propensity overestimation using implicit click feedback. It probes ranking quality under both dynamic online and static offline logging policies by comparing predicted rankings against ground truth relevance. Use when the user wants to benchmark on Yahoo! LETOR, Istella-S, or asks about evaluating this task. Reports NDCG@K.
Evaluates a robot's ability to infer human goals and execute household tasks from noisy, accented, or mispronounced spoken instructions. It probes robust speech perception, joint planning, and Theory of Mind in embodied human-robot collaboration under mixed-observability conditions. Use when the user wants to benchmark on UnclearInstruct, or asks about evaluating this task. Reports Accuracy.
Evaluates a CNN's ability to predict stellar atmospheric parameters and chemical abundances from low-resolution spectra, measuring both internal consistency across model runs and agreement with established spectroscopic pipeline measurements. Use when the user has predictions and gold and needs to compute Uncertainty.
Benchmarks the robustness of uncertainty estimation methods against label outliers and distribution shifts. It evaluates whether predicted prediction intervals and uncertainty quantifications maintain calibration and accuracy when training data is contaminated with noise or adversarial perturbations. Use when the user wants to benchmark on Synthetic 1D regression dataset, Real-world regression datasets, NYU-Depth-v2, or asks about evaluating this task. Reports Interval score.
Evaluates the robustness, calibration, and selective classification capability of uncertainty estimation methods (Deep Ensembles, MC Dropout, SVI, TTA) on histopathological whole slide images under domain shift and label noise. Use when the user wants to benchmark on Camelyon17, TCGA, or asks about evaluating this task. Reports AUARC.
Evaluates a model's ability to generate unbiased scene graphs by predicting pairwise relationships between objects in images. It specifically probes robustness to long-tailed predicate distributions by measuring per-class recall averaged across all predicate classes, rather than relying on global recall which favors head classes. Use when the user wants to benchmark on VG150, GQA200, or asks about evaluating this task. Reports mR@K.
Evaluates zero-shot and supervised machine translation quality across multiple language pairs using a multilingual encoder-decoder architecture. It probes the model's ability to translate between unseen language pairs (e.g., Spanish-French) using only monolingual data and reinforcement learning, without parallel training data for the target pair. Use when the user wants to benchmark on United Nations Parallel Corpus (UN corpus), or asks about evaluating this task. Reports BLEU.
Evaluates the ability of deep learning models to perform fine-grained named entity recognition for UMLS semantic types in biomedical text. It probes how well models handle class imbalance, contextual ambiguity, and domain shift between clinical notes and biomedical abstracts. Use when the user wants to benchmark on i2b2 2010, MedMentions(full), MedMentions(st21pv), or asks about evaluating this task. Reports F1.
Evaluates the cross-modality generalization and robustness of 3D medical segmentation foundation models by testing their ability to segment 13 whole-body organs in functional (PET) versus structural (CT/MRI) imaging using intrinsically paired intra-subject scans. Use when the user wants to benchmark on UMD Benchmark, or asks about evaluating this task. Reports Dice Similarity Coefficient (DSC).
This protocol evaluates the reliability of automatically generated relevance judgments (via the UMBRELA tool) compared to human assessments across different workflow conditions. It measures how well LLM-generated qrels align with human qrels in ranking retrieval systems using standard IR metrics and rank correlation. Use when the user wants to benchmark on TREC 2024 RAG Track, or asks about evaluating this task. Reports Kendall's τ.
Evaluates spoken dialogue models' ability to follow fine-grained speech style instructions (emotion, speed, volume, accent, language, composite) while maintaining general conversational competence. It measures both subjective audio quality/naturalness and objective content/emotion alignment against ground-truth style specifications. Use when the user wants to benchmark on UltraVoice Test Set, URO-Bench, or asks about evaluating this task. Reports MOS, IFR.
Evaluates multilingual LLMs on chat, math reasoning, and code generation across five languages (English, Chinese, Spanish, Russian, French) to measure the effectiveness of knowledge-enhanced supervised fine-tuning. Use when the user wants to benchmark on OMGEval, MGSM, Multilingual HumanEval, or asks about evaluating this task. Reports OMGEval score.
Evaluates text-to-image diffusion models on ultra-high-resolution generation, probing semantic alignment with prompts, fine-grained texture preservation, and overall perceptual quality at resolutions ≥4096px. Use when the user wants to benchmark on UltraHR-eval4K, Aesthetic-Eval@4096, or asks about evaluating this task. Reports FID.
Evaluates audio foundation models across understanding, generation, and codec capabilities. It probes semantic accuracy, timbre fidelity, acoustic quality, and multilingual speech comprehension using a unified taxonomy and standardized benchmarks. Use when the user wants to benchmark on SpeechCMMLU, SpeechHSK, LibriSpeech, AISHELL-1, or asks about evaluating this task. Reports WER.
Evaluates a chat model's ability to generate accurate, informative, and correct responses across diverse domains including commonsense, world knowledge, professional knowledge, mathematics, reasoning, and writing. It probes both factual correctness and response quality using automated LLM-based pairwise and independent scoring. Use when the user wants to benchmark on UltraChat Evaluation Set, or asks about evaluating this task. Reports ChatGPT scoring.
Evaluates the quality of pre-training datasets by measuring the downstream performance of models trained on them. It probes general language understanding, commonsense reasoning, and multilingual capabilities through standard zero-shot benchmarks. Use when the user wants to benchmark on MMLU, ARC-C, ARC-E, CommonSenseQA, HellaSwag, OpenbookQA, PIQA, SIQA, Winogrande, C-Eval, CMMLU, or asks about evaluating this task. Reports Average.
Evaluates cross-lingual transfer methods for Ukrainian text classification across toxicity, formality, and natural language inference tasks. It compares translation-based baselines, LLM prompting, and adapter/fine-tuning approaches on both machine-translated and semi-natural Ukrainian test sets. Use when the user wants to benchmark on Ukrainian Toxicity (Translated & Semi-natural), Ukrainian Formality (Translated & Semi-natural), Ukrainian NLI (Translated & Semi-natural), or asks about evalua...
Evaluates 3D medical image segmentation models on multi-modal MRI and CT scans across abdominal, brain, and whole-body anatomical domains. Probes the model's ability to produce accurate organ and tumor masks while measuring both volumetric overlap and boundary precision under zero-shot and fine-tuning settings. Use when the user wants to benchmark on AMOS, BTCV, BRATS, UKBOB, or asks about evaluating this task. Reports Dice Score.
Evaluates the ability of multimodal LLMs to predict binary disease risk from individual-specific clinical data, including tabular features and time-series spirograms. It tests how well serialized text and cross-modal embeddings integrate to produce accurate risk scores for conditions like asthma and diabetes. Use when the user wants to benchmark on UK Biobank, or asks about evaluating this task. Reports AUROC.
Evaluates the impact of data assimilation (DA) using the SPEnKF algorithm on a U-STN12 deep learning model for UK temperature forecasting. It probes the model's ability to integrate global atmospheric data (ERA5 T850) and surface observations (ASOS/ERA5 T2m) over a 120-hour lead time, measuring forecast accuracy degradation or improvement under varying noise levels and assimilation frequencies. Use when the user wants to benchmark on ERA5, ASOS, or asks about evaluating this task. Reports RMSE.
Evaluates Vietnamese language models' machine reading comprehension capabilities, specifically probing their ability to extract correct answer spans and correctly identify when a question cannot be answered from the given context. Use when the user wants to benchmark on UIT-ViQuAD 2.0, or asks about evaluating this task. Reports Exact Match (EM).