Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

22,846
skills in category
952
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 2,7372,760 of 22,846 skills

Zero Shot Transfer EvalA

This evaluation probes a vision-language model's ability to generalize to unseen image classification tasks without task-specific fine-tuning. It measures how well the model aligns visual features with natural language class descriptions to perform zero-shot classification across diverse domains. Use when the user wants to benchmark on ImageNet, CIFAR-10, Oxford-IIIT Pet, Food101, Stanford Cars, Kinetics700, EuroSAT, or asks about evaluating this task. Reports accuracy.

researchpythonperformance
0
3
Zero Shot Teacher Feedback EvalA

Evaluates zero-shot performance of LLMs in teacher coaching tasks, including scoring classroom transcripts against observation rubrics, identifying instructional highlights and missed opportunities, and generating actionable pedagogical suggestions. Use when the user wants to benchmark on CLASS & MQI Classroom Transcripts, or asks about evaluating this task. Reports Relevance.

researchpythongo
0
3
Zero Shot Retrieval Leakage EvalA

This evaluation probes the robustness of neural retrieval models to train-test data leakage by measuring how much performance on standard benchmarks artificially improves when training data contains near-duplicates or exact matches of test queries. It specifically assesses zero-shot transfer effectiveness under varying leakage conditions and training set sizes. Use when the user wants to benchmark on Robust04, TREC 2017 Common Core, TREC 2018 Common Core, or asks about evaluating this task. R...

researchpythongo
0
3
Zero Shot Human Classification EvalA

Evaluates the zero-shot transfer capability of a vision-language model on human-centric classification tasks, including activity recognition, age grouping, and emotion recognition, using pose-grounded text descriptions and subject-focused attention. Use when the user wants to benchmark on Stanford40, Emotic, LAGENDA-Body, LAGENDA-Face, UTKFace, FER+, or asks about evaluating this task. Reports top-k accuracy.

researchpythongo
0
3
Zero Shot Generalization EvalA

Evaluates a model's ability to generalize to unseen natural language tasks without task-specific fine-tuning or prompt tuning. It probes zero-shot performance across traditional NLP benchmarks and novel BIG-bench tasks using accuracy. Use when the user wants to benchmark on BIG-bench & Held-out NLP Tasks, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Zero Shot Cot Bias EvalA

Evaluates how zero-shot Chain-of-Thought (CoT) prompting affects social bias and toxicity in large language models compared to standard prompting. It measures performance degradation on stereotype benchmarks and the propensity to generate harmful outputs on harmful question tasks. Use when the user wants to benchmark on CrowS Pairs, StereoSet, BBQ, HarmfulQ, or asks about evaluating this task. Reports TD2 accuracy.

researchpythongo
0
3
Zero Shot Commonsense EvalA

Evaluates zero-shot commonsense reasoning capabilities of language models using multiple-choice questions. It specifically probes how prompt engineering and probability calibration strategies affect accuracy across different model sizes and architectures. Use when the user wants to benchmark on CommonsenseQA, COPA, OpenBookQA, PIQA, Social IQA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Zero Shot Asr EvalA

Evaluates the impact of perceptual audio enhancement (SAM-Audio) on zero-shot automatic speech recognition performance across Bengali and English noisy speech. It probes whether signal-level quality improvements translate to better machine transcription accuracy. Use when the user wants to benchmark on Bengali Noisy YouTube dataset, English noisy dataset, or asks about evaluating this task. Reports WER, CER.

researchpythonperformance
0
3
Zero Shot Adjustable Acceleration EvalA

This evaluation protocol assesses the capability of large language models to maintain task performance while dynamically pruning hidden activations during inference. It probes the model's robustness across natural language understanding, text generation, and instruction-tuning tasks under varying computational constraints and acceleration ratios. Use when the user wants to benchmark on IMDB, GLUE, WikiText-103, Penn Treebank (PTB), One Billion Word (1BW), LAMBADA, MMLU, or asks about evaluati...

researchpythongo
0
3
Yufeng Xguard EvalA

Evaluates the safety classification capabilities of guardrail models across multiple dimensions, including prompt/response safety detection, multilingual robustness, adversarial jailbreak resilience, and safe content completion. It also tests the model's ability to dynamically adapt to new moderation policies without retraining. Use when the user wants to benchmark on Aegis / Aegis2.0, WildGuard, StrongReject, SEval2.0, E-commerce Benchmark, Adaptive Policy Scope Benchmark, or asks about eval...

researchpythongo
0
3
Ytseg Segmentation EvalA

Evaluates a model's ability to detect topic boundaries in unstructured spoken transcriptions (text segmentation) and generate coherent chapter titles (smart chaptering). It probes hierarchical structuring, real-time/online processing constraints, and cross-domain generalization to meeting transcripts. Use when the user wants to benchmark on WIKI-727K, YTSEG, QMSUM, YTSEG[TITLES], or asks about evaluating this task. Reports F1.

researchpythongo
0
3
Youtube Implicit Rec EvalA

Evaluates recommender systems on implicit feedback datasets, testing their ability to rank relevant items for users. It probes model versatility across cold-start, offline, and instant recommendation scenarios using side information and sequential context features. Use when the user wants to benchmark on YouTube Implicit Feedback Subset, or asks about evaluating this task. Reports NDCG@100.

researchpythongo
0
3
Younger EvalA

This benchmark evaluates graph neural networks and generative models on AI-generated neural network architectures represented as directed acyclic graphs. It probes two capabilities: local component-level refinement (predicting data flows and operator types) and global end-to-end architecture generation. Use when the user wants to benchmark on Younger, or asks about evaluating this task. Reports F1.

researchpythonnode
0
3
Yoloe Lvis Coco EvalA

Evaluates open-vocabulary object detection and segmentation capabilities using text, visual, and prompt-free inputs on zero-shot and fine-tuned settings. Use when the user wants to benchmark on LVIS, COCO, or asks about evaluating this task. Reports Fixed AP.

researchpythongo
0
3
Yokai EvalA

Assesses large language models' knowledge of Japanese yokai (folklore creatures) through multiple-choice questions. It probes cultural and linguistic familiarity with Japanese folklore, revealing how training data exposure and language affect cross-cultural knowledge retention. Use when the user wants to benchmark on YokaiEval, or asks about evaluating this task. Reports correctness.

researchpythongo
0
3
Yodas Speech Recognition EvalA

Evaluates monolingual automatic speech recognition performance across multiple languages using a large-scale, YouTube-collected audio dataset. It probes the model's ability to accurately transcribe spoken language from noisy, real-world video subtitles after alignment filtering. Use when the user wants to benchmark on YODAS, or asks about evaluating this task. Reports CER.

researchpythontesting
0
3
Yelpchi Fraud Detection EvalA

This evaluation probes a graph neural network's ability to detect fraudulent or spam reviews within a multi-relational graph structure. It measures classification robustness against structural and semantic inconsistencies by training on varying fractions of labeled data and testing on the remainder. Use when the user wants to benchmark on YelpChi, or asks about evaluating this task. Reports F1-score.

researchpythonnode
0
3
Yelp13 Sentiment EvalA

This benchmark evaluates a model's ability to predict sentiment on long, complex documents by leveraging discourse structure. It probes whether incorporating hierarchical discourse trees improves sentiment classification and regression over standard sequential baselines, particularly for longer texts where sentiment is more subtle and diverse. Use when the user wants to benchmark on Yelp'13, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Yelp Rating Prediction EvalA

Predicts a 1-to-5 star rating for a restaurant review based on its text content. It probes a model's ability to capture sentiment, domain-specific linguistic patterns, and fine-grained textual features for multi-class classification. Use when the user wants to benchmark on Yelp Dataset, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Yahoo R6b Ctr EvalA

This evaluation probes a model's ability to optimize exploration-exploitation trade-offs in personalized click-through rate (CTR) prediction. It measures how effectively an exploration strategy improves cumulative user engagement and advertiser retention compared to standard ranking baselines. Use when the user wants to benchmark on Yahoo! R6B, or asks about evaluating this task. Reports CTR.

researchpythongo
0
3
Xwikis Summarisation EvalA

This benchmark evaluates the ability of multilingual and cross-lingual models to generate accurate English summaries from source documents in German, French, and Czech. It probes supervised, zero-shot, and few-shot cross-lingual transfer capabilities, as well as model robustness on out-of-domain news text. Use when the user wants to benchmark on XWikis, D_en→en, Voxeurop, or asks about evaluating this task. Reports ROUGE-L recall.

researchpythonperformance
0
3
Xtsc Bench EvalA

This benchmark evaluates the quality and stability of explanation methods for time series classification models. It probes four key properties: robustness to input perturbations, faithfulness to model predictions, computational complexity of the explanations, and reliability against known ground-truth informative features. Use when the user wants to benchmark on XTSC-Bench Synthetic Datasets, or asks about evaluating this task. Reports Faithfulness Correlation.

researchpythongo
0
3
Xtreme S EvalA

Evaluates cross-lingual speech representations across 102 languages and four task families: automatic speech recognition, speech translation, speech classification, and speech-text retrieval. It probes the ability of self-supervised speech models to generalize across languages, domains, and varying data regimes from high-resource to low-resource settings. Use when the user wants to benchmark on Fleurs, MLS, VoxPopuli, CoVoST-2, Minds-14, or asks about evaluating this task. Reports WER.

researchpythonperformance
0
3
Xtreme R EvalA

Evaluates zero-shot cross-lingual transfer by training models on English data and testing them on 50 typologically diverse languages across classification, QA, and retrieval tasks. It probes fine-grained diagnostic capabilities and cross-lingual alignment using structured performance breakdowns. Use when the user wants to benchmark on XQuAD, XCOPA, Mewsli-X, LAReQA, CheckList, or asks about evaluating this task. Reports Exact Match.

researchpythongo
0
3