Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 4,585–4,608 of 23,503 skills
This benchmark evaluates dynamic spatial reasoning and perception-memory integration in embodied environments. It probes models across three levels: static spatial perception, text-conditioned temporal memory, and visual-conditioned temporal memory, testing capabilities like object recognition, visual grounding, depth estimation, trajectory tracking, and long-horizon state reconstruction. Use when the user wants to benchmark on SpaMEM, or asks about evaluating this task. Reports mIoU.
Evaluates a model's ability to estimate the 6D pose (position and orientation) of non-cooperative spacecraft from monocular images. It probes generalization from photorealistic synthetic data to real space imagery and tests robustness to perceptual aliasing and orientation ambiguity. Use when the user wants to benchmark on URSO, SPEED, or asks about evaluating this task. Reports ESA Error.
Evaluates a multi-agent DRL framework (MAPPO) combined with whale optimization for maximizing satellite coverage and data rates in 6G sub-THz networks using reconfigurable intelligent surfaces (RIS). Use when the user wants to benchmark on Simulated LEO Satellite-RIS Network Environment, or asks about evaluating this task. Reports average data rate.
Evaluates named entity recognition performance across diverse domains and languages, specifically probing a model's ability to handle out-of-vocabulary words and morphological variation through hash-based embeddings versus traditional lookup embeddings. Use when the user wants to benchmark on CoNLL 2002, WNUT 2017, AnEM, Dutch Archaeology, OntoNotes 5.0, or asks about evaluating this task. Reports F1 score.
Evaluates multilingual encoder models on South Slavic languages (Croatian and Serbian) across three diverse NLP tasks: named entity recognition, parliamentary sentiment regression, and causal commonsense reasoning. Tests whether cost-efficient additional pretraining can match dedicated monolingual encoders without full from-scratch training. Use when the user wants to benchmark on hr500k, ReLDI-NormTagNER-hr, SETimes.SR, ReLDI-NormTagNER-sr, ParlaSent (HBS), COPA (Croatian & Serbian), or asks...
Evaluates biomedical named entity recognition (NER) capabilities on scientific literature. It probes a model's ability to identify and classify nine distinct bioentity types (e.g., genes, cell lines, diseases) within text extracted from published biological figures and captions. Use when the user wants to benchmark on SourceData-NLP, or asks about evaluating this task. Reports F1 score.
Evaluates a model's ability to identify which sentences in a source document contribute to an abstractive summary. It probes source sentence detection capability by ranking candidate sentences based on their inferred relevance to the summary. Use when the user wants to benchmark on SourceSum, or asks about evaluating this task. Reports NDCG.
This evaluation probes whether large language models exhibit source attribution bias, specifically penalizing arguments when the attributed source's expected ideological position conflicts with the argument's content (coherence bias). It measures how models adjust credibility ratings based on source-argument alignment and whether they explicitly reason about source credibility. Use when the user wants to benchmark on Source Attribution Bias Evaluation, or asks about evaluating this task. Repo...
This evaluation probes a model's ability to spatially localize sound sources in images or video frames given an accompanying audio clip. It measures how accurately the predicted bounding box overlaps with ground-truth annotations provided by multiple human annotators. Use when the user wants to benchmark on Flickr SoundNet Testset, VGG-Sound Source (VGG-SS), or asks about evaluating this task. Reports cIoU.
Evaluates text-queried sound separation models on natural and mixed audio. It measures how accurately a model isolates target sound sources from background noise or other sources, and how well it handles silence when the target is absent. Use when the user wants to benchmark on AudioSet, AudioCaps, ESC-50, or asks about evaluating this task. Reports SDRi.
Evaluates a model's ability to infer physical properties (air column length, container dimensions, flow rate, fill time, liquid weight) and classify container shapes solely from the acoustic characteristics of pouring liquids, without visual or tactile input. Use when the user wants to benchmark on Sound of Water 50, Wilson et al. [96] dataset, or asks about evaluating this task. Reports Mean Absolute Error (MAE).
Evaluates a model's ability to localize sound sources in audio-visual pairs by predicting spatial response maps or bounding boxes. It measures how effectively the model aligns audio signals with visual regions containing the corresponding sound, particularly testing robustness to semantically similar but mismatched cross-modal pairs. Use when the user wants to benchmark on VGGSound, SoundNet-Flickr, VGG-SS, SoundNet-Flickr-Test, or asks about evaluating this task. Reports cIoU.
Evaluates music source separation models on their ability to isolate individual instruments (vocals, bass, drums, other) from mixed audio tracks. It specifically probes robustness to label noise and bleeding artifacts in training data, as well as standard separation performance across different leaderboards. Use when the user wants to benchmark on SDXDB23_LabelNoise, SDXDB23_Bleeding, Standard (MDXDB21), or asks about evaluating this task. Reports SDR (Signal-to-Distortion Ratio).
This benchmark evaluates a model's ability to distinguish between real human-recorded songs and AI-generated synthetic songs. It specifically probes long-range temporal dependency modeling by testing performance on both short (5s) and long (120s) audio clips, while also measuring generalization to unseen generation algorithms and singers. Use when the user wants to benchmark on SONICS, or asks about evaluating this task. Reports F1 score.
Evaluates the effectiveness of an adversarial perturbation method (SongBsAb) designed to prevent illegal singing voice conversion. It probes the method's ability to disrupt singer identity and lyrical fidelity in converted audio while maintaining high audio quality and imperceptibility. Use when the user wants to benchmark on OpenSinger, NUS-48E, or asks about evaluating this task. Reports Lyric Word Error Rate (WER).
Probes the capability of joint entity and relation extraction for identifying software mentions and their attributes (URLs, versions, licenses) in scholarly articles. It specifically tests in-distribution performance and out-of-distribution generalization across two competition phases. Use when the user wants to benchmark on SOMD 2025, or asks about evaluating this task. Reports F1-score.
Evaluates the ability of token classification models to identify and categorize software mentions within academic sentences. It probes how well models handle class imbalance, subtoken segmentation, and syntactic complexity in scholarly text. Use when the user wants to benchmark on SOMD (Software Mention Detection in Scholarly Publications), or asks about evaluating this task. Reports F1-Score.
This benchmark evaluates a model's ability to perform open-vocabulary 3D instance segmentation on indoor point clouds. It probes the model's capacity to align 3D geometric features with free-form language instructions to generate accurate instance masks for both seen and unseen categories. Use when the user wants to benchmark on ScanNetv2, ScanNet200, Replica, or asks about evaluating this task. Reports AP.
Evaluates pretrained language models' in-distribution (ID) and out-of-distribution (OOD) classification accuracy under single-setting and multi-setting fine-tuning configurations. It probes how training dynamics and subset selection affect robustness and generalization across languages, sources, and tasks. Use when the user wants to benchmark on SST-2, IMDB, Yelp, Sentiment140, RTE, QQP, or asks about evaluating this task. Reports accuracy.
Evaluates a diffusion-based molecular language model's ability to generate chemically valid, drug-like molecules and optimize them for specific protein targets. It probes distribution matching, structural diversity, and target-aware binding affinity prediction. Use when the user wants to benchmark on ZINC-Curated, SMILES, SAFE, or asks about evaluating this task. Reports Novel Top-hit 5% Score.
Evaluates the reliability and discriminative power of automatic machine translation metrics by comparing their statistical significance against human MQM judgments. It measures how well a metric's pairwise system rankings align with human preferences using permutation-based p-values rather than hard binary decisions. Use when the user has predictions and gold and needs to compute Soft Pairwise Accuracy (SPA).
Evaluates neural models on three information extraction sub-tasks in materials science: detecting experiment-describing sentences, extracting and typing entity mentions (materials, values, devices), and filling experiment-specific slots (e.g., temperature, anode material). Use when the user wants to benchmark on SOFC-Exp Corpus, Synthesis Procedures Dataset, or asks about evaluating this task. Reports macro-average F1.
This benchmark evaluates a model's ability to recognize coordinated group activities in soccer matches by comparing two input modalities: raw video pixels and structured positional tracking data. It probes spatial-temporal reasoning, tactical formation understanding, and robustness to visual shifts by measuring how well models classify 10 distinct group actions from synchronized match footage. Use when the user wants to benchmark on SoccerNet-GAR, or asks about evaluating this task. Reports b...
Evaluates Vision-Language Models' ability to perform spatial, spatiotemporal, and social reasoning in dynamic, crowd-filled robot navigation scenarios. It probes scene understanding by asking VLMs to answer visual questions based on image sequences and Bird's Eye View (BEV) representations. Use when the user wants to benchmark on SocialNav-SUB, or asks about evaluating this task. Reports PA.