Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 3,937–3,960 of 23,482 skills
This benchmark probes an LLM agent's ability to perform spatial-temporal reasoning and generate executable code for Earth Observation tasks. It evaluates whether models can correctly answer yes/no questions derived from scientific articles by leveraging remote sensing data via Google Earth Engine. Use when the user wants to benchmark on UnivEARTH, or asks about evaluating this task. Reports accuracy.
Evaluates a video agent's capabilities across generation, understanding, editing, and segmentation tasks, while probing its agentic planning and memory mechanisms. It measures how well a unified agent architecture handles long-horizon, multi-step video workflows compared to monolithic baselines. Use when the user wants to benchmark on UniVA-Bench, or asks about evaluating this task. Reports MLLM Judge.
This evaluation measures the transcription accuracy of a fine-tuned automatic speech recognition model across four diverse speech benchmarks. It specifically probes the model's robustness to different speaking styles, accents, and linguistic contexts after applying a noise reduction step and a BART-based semantic correction pipeline. Use when the user wants to benchmark on LibriSpeech, Europarl-ASR, TED-LIUM, FLEURS, or asks about evaluating this task. Reports Word Error Rate (WER).
Evaluates text-to-SQL models on compositional generalization, out-of-domain robustness, and schema-question alignment across 18 diverse datasets and 12 domains. It probes the model's ability to handle long-form query decomposition and cross-domain SQL pattern diversity. Use when the user wants to benchmark on UNITE, or asks about evaluating this task. Reports accuracy.
Evaluates text summarization models across multiple dimensions including faithfulness, completeness, conciseness, domain stability, and abstractiveness. It tests how well summarizers handle diverse input contexts (domains, dialogue vs. non-dialogue, short vs. long texts) and the impact of PII redaction on hallucination. Use when the user wants to benchmark on UniSumEval, or asks about evaluating this task. Reports faithfulness.
Evaluates multimodal voice generation and conversion capabilities across face-driven, text-driven, and attribute-based tasks. Probes the model's ability to align facial, textual, and attribute descriptions with target speech while preserving speaker identity, content clarity, and naturalness. Use when the user wants to benchmark on LRS3, or asks about evaluating this task. Reports MOS-Match.
Evaluates a unified text-to-audio model's ability to generate speech, music, and sound effects from natural language instructions without reference audio. It probes instruction-following fidelity, acoustic quality, structural coherence, and the positive transfer effects of multi-modal joint training. Use when the user wants to benchmark on UniSonate Unified Corpus, Seed-TTS test set, SongEval benchmark, or asks about evaluating this task. Reports WER, SongEval.
Evaluates multimodal large language models on perceptual-level image understanding across three domains: Image Aesthetics & Art (IAA), Image Quality Assessment (IQA), and Image Structure & Texture Assessment (ISTA). It probes both continuous visual rating (VR) and discrete visual question answering (VQA) capabilities. Use when the user wants to benchmark on UniPercept-Bench, or asks about evaluating this task. Reports Acc..
Evaluates a unified discrete diffusion framework's capability to jointly generate and reason over image-text pairs. It probes unconditional and conditional generation quality, the effectiveness of classifier-free guidance, training and inference efficiency, and cross-modal retrieval and reasoning performance. Use when the user wants to benchmark on DataComp1B, CC12M, MS-COCO30k, Flickr, Winoground, or asks about evaluating this task. Reports FID.
Evaluates large language models' factual correctness by dynamically generating responses to factual questions and measuring how well hallucination detection and fact verification methods can predict the ground-truth factuality label of those responses. It probes a model's susceptibility to hallucination and the effectiveness of external evidence retrieval in verifying generated claims. Use when the user wants to benchmark on TriviaQA, NQ-Open, PopQA, 2WikiMultihopQA, HotpotQA, or asks about e...
Evaluates natural language generation models across multiple quality dimensions (e.g., coherence, fluency, consistency, relevance) by reframing assessment as a Boolean QA task. Measures how well automated scores align with human judgments using correlation metrics. Use when the user wants to benchmark on SummEval, Topical-Chat, SFRES, SFHOT, QAGS, or asks about evaluating this task. Reports Spearman correlation.
Evaluates an instance segmentation model's ability to accurately delineate and separate individual microstructural objects in high-resolution electron micrographs, particularly under varying instance densities. Use when the user wants to benchmark on UniEM-3M, or asks about evaluating this task. Reports mAP@0.5.
Evaluates retrieval and end-to-end generation performance of multimodal RAG systems on real-world PDF documents. It probes the ability of text-only, image-only, and multimodal (text-image fusion/joint) retrieval paradigms to locate relevant evidence and generate faithful, complete answers to cross-modality questions. Use when the user wants to benchmark on UNIDOC-BENCH, or asks about evaluating this task. Reports Precision@10, Recall@10, Faithfulness, Completeness.
Evaluates pixel-level biomedical image segmentation capability using convolutional networks. Probes the model's ability to precisely delineate cellular structures and membranes in electron and light microscopy images with limited training data. Use when the user wants to benchmark on EM segmentation challenge (ISBI 2012), PhC-U373, DIC-HeLa, or asks about evaluating this task. Reports warping error, IOU.
This protocol evaluates unbiased learning-to-rank models on their ability to correct position bias and propensity overestimation using implicit click feedback. It probes ranking quality under both dynamic online and static offline logging policies by comparing predicted rankings against ground truth relevance. Use when the user wants to benchmark on Yahoo! LETOR, Istella-S, or asks about evaluating this task. Reports NDCG@K.
Evaluates a robot's ability to infer human goals and execute household tasks from noisy, accented, or mispronounced spoken instructions. It probes robust speech perception, joint planning, and Theory of Mind in embodied human-robot collaboration under mixed-observability conditions. Use when the user wants to benchmark on UnclearInstruct, or asks about evaluating this task. Reports Accuracy.
Evaluates a CNN's ability to predict stellar atmospheric parameters and chemical abundances from low-resolution spectra, measuring both internal consistency across model runs and agreement with established spectroscopic pipeline measurements. Use when the user has predictions and gold and needs to compute Uncertainty.
Benchmarks the robustness of uncertainty estimation methods against label outliers and distribution shifts. It evaluates whether predicted prediction intervals and uncertainty quantifications maintain calibration and accuracy when training data is contaminated with noise or adversarial perturbations. Use when the user wants to benchmark on Synthetic 1D regression dataset, Real-world regression datasets, NYU-Depth-v2, or asks about evaluating this task. Reports Interval score.
Evaluates the robustness, calibration, and selective classification capability of uncertainty estimation methods (Deep Ensembles, MC Dropout, SVI, TTA) on histopathological whole slide images under domain shift and label noise. Use when the user wants to benchmark on Camelyon17, TCGA, or asks about evaluating this task. Reports AUARC.
Evaluates a model's ability to generate unbiased scene graphs by predicting pairwise relationships between objects in images. It specifically probes robustness to long-tailed predicate distributions by measuring per-class recall averaged across all predicate classes, rather than relying on global recall which favors head classes. Use when the user wants to benchmark on VG150, GQA200, or asks about evaluating this task. Reports mR@K.
Evaluates zero-shot and supervised machine translation quality across multiple language pairs using a multilingual encoder-decoder architecture. It probes the model's ability to translate between unseen language pairs (e.g., Spanish-French) using only monolingual data and reinforcement learning, without parallel training data for the target pair. Use when the user wants to benchmark on United Nations Parallel Corpus (UN corpus), or asks about evaluating this task. Reports BLEU.
Evaluates the ability of deep learning models to perform fine-grained named entity recognition for UMLS semantic types in biomedical text. It probes how well models handle class imbalance, contextual ambiguity, and domain shift between clinical notes and biomedical abstracts. Use when the user wants to benchmark on i2b2 2010, MedMentions(full), MedMentions(st21pv), or asks about evaluating this task. Reports F1.
Evaluates the cross-modality generalization and robustness of 3D medical segmentation foundation models by testing their ability to segment 13 whole-body organs in functional (PET) versus structural (CT/MRI) imaging using intrinsically paired intra-subject scans. Use when the user wants to benchmark on UMD Benchmark, or asks about evaluating this task. Reports Dice Similarity Coefficient (DSC).
This protocol evaluates the reliability of automatically generated relevance judgments (via the UMBRELA tool) compared to human assessments across different workflow conditions. It measures how well LLM-generated qrels align with human qrels in ranking retrieval systems using standard IR metrics and rank correlation. Use when the user wants to benchmark on TREC 2024 RAG Track, or asks about evaluating this task. Reports Kendall's τ.