Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 7,873–7,896 of 20,853 skills
Evaluates one-shot singing voice conversion quality by measuring how naturally the converted audio sounds and how closely it matches the target speaker's voice, using only 20 seconds of target speech or singing data. Use when the user wants to benchmark on Database A, Database B, or asks about evaluating this task. Reports MOS naturalness.
Evaluates conversational recommendation systems across monolingual, multilingual, and cross-lingual settings. It probes a model's ability to generate relevant and fluent responses, select correct knowledge entities, maintain topic consistency, and successfully guide dialogues toward a recommendation target. Use when the user wants to benchmark on DuRecDial 2.0, or asks about evaluating this task. Reports F1.
Passage retrieval for web search queries, evaluating a model's ability to rank relevant documents from a large collection. It probes in-domain retrieval accuracy as well as out-of-domain and cross-lingual generalization, highlighting challenges like salient phrase mismatch, syntactic mismatch, and false negatives. Use when the user wants to benchmark on DuReader_retrieval, or asks about evaluating this task. Reports MRR@10.
Evaluates a dialogue system's ability to manage full-duplex speech interactions, specifically focusing on the timing and appropriateness of machine-to-user interruptions and user-to-machine interruptions, alongside system response latency. Use when the user wants to benchmark on duplex-dialogue-3k, or asks about evaluating this task. Reports FTED.
Evaluates reading comprehension and long-form text understanding by asking models to answer questions about movie plots. It specifically probes sensitivity to narrative length and semantic shifts between short and paraphrased long versions of the same story. Use when the user wants to benchmark on DuoRC, or asks about evaluating this task. Reports F1.
Evaluates semantic music tokenizers on their ability to preserve musical semantics for conditional generation, their efficiency for language modeling, and their audio reconstruction fidelity. It probes whether decoupled vocal-accompaniment tokenization yields better LM-friendliness and tagging performance without sacrificing perceptual quality. Use when the user wants to benchmark on MagnaTagATune, or asks about evaluating this task. Reports PPL@1024.
Evaluates the ability of human activity recognition models to classify dyadic kinesic functions and interactions across different subjects and physical locations. It probes robustness to viewpoint changes, background variations, occlusion, and modality-specific limitations (RGB vs. depth vs. 3D skeletons). Use when the user wants to benchmark on DUET, or asks about evaluating this task. Reports Cross-location accuracy (%), Cross-subject accuracy (%).
Ranks active compounds against decoys for a given protein target. It probes the model's ability to prioritize true binders in a large pool of inactive decoys and resist dataset biases. Use when the user wants to benchmark on DUD-E, AD, or asks about evaluating this task. Reports AUC.
This benchmark evaluates a model's ability to predict conversational turn-taking dynamics and agent actions from dual-channel speech audio. It probes the system's capacity to anticipate speech boundaries, detect backchannels, and classify continuous turn-taking states without relying on explicit silence timeouts or external labels. Use when the user wants to benchmark on otoSpeech, Switchboard, or asks about evaluating this task. Reports wF1.
Evaluates a unified vision tokenizer's capacity to decouple and jointly optimize low-level perceptual reconstruction and high-level semantic understanding. It probes zero-shot classification, cross-modal retrieval, image reconstruction fidelity, and downstream multimodal reasoning capabilities. Use when the user wants to benchmark on ImageNet-1K, Flickr8K, VQAv2, POPE, MME, SEED-IMG, MMBench, MM-Vet, or asks about evaluating this task. Reports Top-1 accuracy.
Evaluates a hybrid sequential and LLM-based framework for next-item movie recommendation. It probes the model's ability to capture temporal user preferences and semantic genre consistency to predict the next movie a user will watch. Use when the user wants to benchmark on MovieLens-1M, or asks about evaluating this task. Reports NDCG@5.
Evaluates a model's ability to learn sequentially from a stream of tasks without catastrophic forgetting, while adapting quickly to new tasks. It probes both task-aware (with task IDs) and task-free (without task IDs) continual learning settings, measuring final accuracy, forgetting, and knowledge transfer. Use when the user wants to benchmark on Split miniImageNet, CORE50, or asks about evaluating this task. Reports ACC.
Probes the joint functional correctness and security of LLM-generated code, while also evaluating an automated framework's ability to execute code in sandboxes and semantically judge test outcomes against human ground truth. Use when the user wants to benchmark on DualGauge-Bench, or asks about evaluating this task. Reports F1 Score.
This benchmark evaluates fine-grained real-time traffic light detection using synchronized dual-camera inputs. It probes a model's ability to accurately localize and classify multiple traffic light states across varying distances and sizes, while balancing detection speed and precision. Use when the user wants to benchmark on DualCam, or asks about evaluating this task. Reports F1-score.
Evaluates a model's ability to generate synchronized background audio and intelligible speech from video input, measuring audio quality, distribution matching, and audio-video temporal alignment. Use when the user wants to benchmark on DualBench, VGGSound, or asks about evaluating this task. Reports FAD↓.
Evaluates the ability of generative models to design dual-target ligands that simultaneously bind to two protein pockets with high affinity while maintaining favorable drug-like properties. It measures both binding strength and molecular quality across a large set of target pairs. Use when the user wants to benchmark on Dual-target drug design dataset, or asks about evaluating this task. Reports Dual High Affinity.
Evaluates the geometric reconstruction quality of Neural Radiance Fields by extracting 3D surfaces or edges using density gradients. It measures how accurately the predicted geometry aligns with ground truth point clouds across diverse real-world objects. Use when the user wants to benchmark on DTU benchmark dataset, or asks about evaluating this task. Reports completeness.
Evaluates a model's ability to classify drug-target interaction relations from biomedical text into one of ten specific interaction types. It probes multiclass relation extraction under conditions of severe class imbalance. Use when the user wants to benchmark on DrugProt, ChemProt, or asks about evaluating this task. Reports micro F1-score.
Evaluates a model's ability to predict continuous binding affinity for drug-target pairs across different cold-start and warm-start scenarios. It probes the model's generalization to unseen drugs, unseen targets, and fully seen interactions using regression metrics. Use when the user wants to benchmark on Davis, Metz, KIBA, or asks about evaluating this task. Reports RMSE.
This benchmark evaluates a model's ability to predict drug-target interactions by integrating molecular graphs and protein sequences into a heterogeneous interaction network. It probes the model's capacity to learn hierarchical graph representations and distinguish interacting from non-interacting drug-protein pairs. Use when the user wants to benchmark on DTI Benchmark, or asks about evaluating this task. Reports AUC.
Evaluates the ability of machine learning models to predict drug-target interactions in inductive settings where test drugs, targets, or both are unseen during training. It probes cold-start prediction capabilities and robustness to local class imbalance in sparse biological networks. Use when the user wants to benchmark on NR, GPCR, IC, E, DB, or asks about evaluating this task. Reports AUPR.
Evaluates a model's ability to predict continuous drug-target binding affinity and classify binary drug-target interactions. It probes geometry-aware representation learning, metric consistency, and generalization across diverse chemical-proteomic domains. Use when the user wants to benchmark on DTI-DG, BIOSNAP, BindingDB, DAVIS, or asks about evaluating this task. Reports PCC.
Evaluates the ability of molecular models to predict drug-target interactions (DTI) by classifying whether a given drug and protein target pair binds. It probes the model's capacity to integrate diverse molecular representations (sequences, graphs, structures) and interaction layers to distinguish positive binding pairs from negative ones. Use when the user wants to benchmark on Davis, BIOSNAP, or asks about evaluating this task. Reports ROC-AUC.
Evaluates drug-target affinity prediction models in cold-start settings (cold-drug and cold-target) to assess generalization to novel drugs or targets using transferred inter-molecular interaction knowledge. Use when the user wants to benchmark on Davis, Kiba, or asks about evaluating this task. Reports RMSE.