Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

21,231
skills in category
885
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 8,713–8,736 of 21,231 skills

Convdr EvalA

Evaluates conversational dense retrieval models on their ability to rank relevant documents given multi-turn conversational queries. It probes context capture, few-shot learning effectiveness, and robustness to noisy conversation history compared to query rewriting baselines. Use when the user wants to benchmark on TREC CAsT, OR-QuAC, or asks about evaluating this task. Reports NDCG@3, MRR@5.

researchpythongit
0
3
Convai2 EvalA

Evaluates open-domain chatbot capabilities in persona-driven multi-turn conversations, measuring response quality via automatic metrics and human judgments of engagement and persona consistency. Use when the user wants to benchmark on PERSONA-CHAT, or asks about evaluating this task. Reports Engagingness.

researchpythongo
0
3
Conv Search Rewriting EvalA

This evaluation probes a model's ability to rewrite conversational search queries to maximize retrieval effectiveness. It measures how well reformulated questions help both sparse and dense retrievers locate relevant passages across different dialogue contexts, including initial turns and topic shifts. Use when the user wants to benchmark on QReCC, TopiOCQA, or asks about evaluating this task. Reports MRR.

researchpythongo
0
3
Controllable Gen EvalA

Evaluates controllable image generation based on visual conditions (segmentation masks, edges, depth maps) by measuring how closely the generated image's extracted conditions match the input conditions. It tests spatial and structural controllability. Use when the user wants to benchmark on ControlNet++ dataset, or asks about evaluating this task. Reports mIoU (Seg. Mask).

researchpython
0
3
Contrastvae EvalA

Evaluates sequential recommendation models on Amazon review datasets, measuring ranking quality for next-item prediction. It specifically probes performance on long-tail items, sequence sparsity, and robustness to noisy inputs. Use when the user wants to benchmark on Amazon Beauty, Amazon Toys, Amazon Tools, Amazon Office, or asks about evaluating this task. Reports Recall@20.

researchpythonperformance
0
3
Contranerf EvalA

Evaluates generalizable Neural Radiance Field (NeRF) methods for novel view synthesis, specifically probing their ability to generalize from synthetic training data to real-world indoor and outdoor scenes. It measures rendering quality and geometric consistency across different domain gaps. Use when the user wants to benchmark on 3D-FRONT, ScanNet, DTU, LLFF, Google Scanned Object, or asks about evaluating this task. Reports PSNR.

researchpythongo
0
3
Continual Safety Alignment EvalA

Evaluates a model's ability to maintain safety alignment and task performance during sequential continual fine-tuning across multiple domains. It probes whether gradient-based sample selection prevents safety degradation (elastic reversion) and catastrophic forgetting while preserving general capabilities. Use when the user wants to benchmark on AdvBench, HarmBench, TruthfulQA, ARC-C, BoolQ, HellaSwag, Winogrande, GSM8K, MedMCQA, Squad_v2, or asks about evaluating this task. Reports ASR.

researchpythonperformance
0
3
Continual Ner EvalA

Evaluates a model's ability to perform continual learning in Named Entity Recognition (CL-NER) by incrementally learning new entity types while mitigating catastrophic forgetting of previously learned types. It specifically probes how well the model handles the 'Other-class' (miscellaneous/old entities) during incremental training and maintains performance across sequential learning steps. Use when the user wants to benchmark on OntoNotes5, i2b2, CoNLL2003, or asks about evaluating this task....

researchpythongo
0
3
Continual Multimodal EvalA

This benchmark evaluates a model's ability to sequentially learn a mix of visual understanding and generation tasks without catastrophically forgetting previously acquired knowledge. It specifically probes intra-modal retention (maintaining performance on earlier tasks) and inter-modal stability (preventing updates for one modality from degrading the other). Use when the user wants to benchmark on ScienceQA, TextVQA, GQA, VizWiz, ImageNet, CustomConcept101, or asks about evaluating this task....

researchpythongo
0
3
Continual Learning MetricsA

Evaluates a model's ability to retain knowledge from previously learned tasks while continuously training on new ones, and measures how past knowledge facilitates learning new tasks and improves performance on old ones. Use when the user has predictions and gold and needs to compute Average Performance (AP).

researchpythongo
0
3
Continual Learning Malware EvalA

Evaluates continual learning techniques for malware classification under domain, class, and task incremental settings. It measures how well models adapt to evolving malware distributions without catastrophic forgetting. The protocol compares complex CL methods against simple baselines like joint replay. Use when the user wants to benchmark on Drebin, EMBER, or asks about evaluating this task. Reports Mean accuracy.

researchpythongo
0
3
Continual Learning Accuracy EvalA

Evaluates a model's ability to learn sequentially across multiple tasks without catastrophic forgetting. It measures how well the model retains accuracy on previously learned tasks while adapting to new ones. Use when the user wants to benchmark on Split MNIST, Permuted MNIST, Split CIFAR-10/100, or asks about evaluating this task. Reports average classification accuracy.

researchpythongo
0
3
Continual Instruction Tuning EvalA

Evaluates the ability of large multimodal models to sequentially learn new instruction-following tasks without catastrophically forgetting previously acquired capabilities. It measures both retained performance on old tasks and the degree of forgetting across sequential training stages. Use when the user wants to benchmark on Flickr30k, TextCaps, VQA v2, OCR-VQA, GQA, VizWiz, TextVQA, or asks about evaluating this task. Reports Average performance ($A_t$).

researchpythonperformance
0
3
Contextual Sarcasm Detection EvalA

Evaluates a model's ability to detect sarcasm in contextual settings across Reddit comments, tweets, and multi-turn dialogues. It probes the capacity to capture sentiment incongruity and contextual cues rather than relying on surface-level lexical features. Use when the user wants to benchmark on SARC 2.0, Twitter, Sarcasm Corpus V2 Dialogues, or asks about evaluating this task. Reports F1-Score.

researchpythongo
0
3
Contextual Object Detection EvalA

Probes a multimodal large language model's ability to infer and localize objects within human-AI interaction contexts (e.g., cloze tests, captioning, QA) using open-vocabulary object names, rather than fixed class sets. Use when the user wants to benchmark on CODE, or asks about evaluating this task. Reports Acc@1.

researchpythongo
0
3
Contextual Ir EvalA

Evaluates information retrieval systems by measuring system-level performance metrics (dead links, response time, redundancy) and user-perceived relevance across different query topics and rank positions. Use when the user wants to benchmark on Custom IR Evaluation Corpus, or asks about evaluating this task. Reports Relevance Judgments.

researchpythonexpress
0
3
Contextual Earnings 22 EvalA

Evaluates speech-to-text systems on their ability to correctly recognize domain-specific custom vocabulary (e.g., company names, products) in real-world earnings call audio. It probes how well models leverage provided keyword contexts (local vs. global/noisy) to improve keyword recognition without introducing transcription artifacts. Use when the user wants to benchmark on Contextual Earnings-22, or asks about evaluating this task. Reports keyword F-score.

researchpythongo
0
3
Contextiq Retrieval EvalA

Evaluates zero-shot text-to-video retrieval for contextual advertising. It probes a model's ability to rank relevant video content based on natural language queries using multimodal signals (vision, audio, captions, metadata). Use when the user wants to benchmark on Val-1, Val-2, or asks about evaluating this task. Reports P@K.

researchpythonperformance
0
3
Contextformer EvalA

Evaluates whether integrating multimodal contextual metadata into pre-trained time series forecasting models improves prediction accuracy. It probes the model's ability to align external covariates with historical time series data to enhance forecast precision across multiple domains and horizons. Use when the user wants to benchmark on Synthetic ARMA(2,2), PEMS-SF, ETT (ETTm2), ECL, Beijing AQ, Store Sales, Monash (Bitcoin), Bitcoin + News, or asks about evaluating this task. Reports MSE.

researchpythongo
0
3
Contextdialog EvalA

Evaluates conversational context recall and utilization in voice interaction models, specifically measuring how well they remember and respond to past user and system utterances in multi-turn dialogues. It also probes the robustness of retrieval-augmented generation (RAG) when applied to speech-based models. Use when the user wants to benchmark on ContextDialog, or asks about evaluating this task. Reports GPT Score.

researchpythongit
0
3
Contextasr Bench EvalA

Evaluates how well ASR models and Large Audio Language Models (LALMs) leverage contextual world knowledge and linguistic reasoning to transcribe speech containing named entities. It tests performance across ten domains under three context conditions: no context, coarse-grained domain labels, and fine-grained technical terms. Use when the user wants to benchmark on ContextASR-Bench, or asks about evaluating this task. Reports Word Error Rate (WER).

researchpythongit
0
3
Context Tree Session Rec EvalA

This benchmark evaluates session-based recommendation models by predicting the immediate next item in a user's interaction sequence. It probes the model's ability to capture sequential dependencies and adapt to evolving user preferences and new items in both static and continuously updating environments. Use when the user wants to benchmark on MOOC, News, and RecSys Challenge Datasets, or asks about evaluating this task. Reports HR@k.

researchpythongo
0
3
Context Is Key EvalA

Evaluates a model's ability to integrate essential natural language context with numerical time series data to produce accurate forecasts. It probes multimodal reasoning, constraint satisfaction, and the capacity to leverage textual information for improving time series prediction. Use when the user wants to benchmark on CiK, or asks about evaluating this task. Reports RCRPS.

researchpythongo
0
3
Context Conflict Merge EvalA

Evaluates how language models merge conflicting generated and retrieved contexts in open-domain QA. It probes whether models exhibit a systematic bias toward generated contexts over retrieved ones when only one context contains the correct answer. Use when the user wants to benchmark on NQ-CC, TQA-CC, or asks about evaluating this task. Reports DiffGR.

researchpythongo
0
3