Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

20,843
skills in category
869
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 7,633–7,656 of 20,843 skills

Evaluator Lm EvalA

This protocol evaluates an LLM's ability to act as a fine-grained text evaluator using custom score rubrics. It tests both absolute grading (assigning a 1–5 score and generating feedback based on a rubric and reference answer) and ranking grading (predicting human preference between two responses). Use when the user wants to benchmark on Feedback Bench, Vicuna Bench, MT Bench, FLASK Eval, MT Bench Human Judgments, HHH Alignment, or asks about evaluating this task. Reports Pearson correlation.

researchpythongo
0
3
Evalmuse 40k EvalA

Evaluates the fine-grained alignment and structural fidelity between generated images and their corresponding text prompts. It probes a model's ability to match specific visual elements (e.g., objects, colors, counts) and overall composition against human-annotated ground truth. Use when the user wants to benchmark on EvalMuse-40K, or asks about evaluating this task. Reports SRCC.

researchpythongo
0
3
Evalalign EvalA

Evaluates text-to-image generation models on two core capabilities: image faithfulness (consistency with real-world commonsense) and text-image alignment (adherence to the conditioning prompt). It probes whether generated images accurately reflect both visual realism and prompt instructions using a fine-grained, human-aligned framework. Use when the user wants to benchmark on EvalAlign, or asks about evaluating this task. Reports EvalAlign_f, EvalAlign_a.

researchpythonexpress
0
3
Eval4nlp 2023 Shared Task EvalA

Evaluates reference-free LLM prompting strategies as metrics for machine translation and summarization. It measures how well predicted quality scores correlate with human judgments (MQM for MT, human annotations for summarization). Use when the user wants to benchmark on Eval4NLP 2023 Shared Task (MT & Summarization), or asks about evaluating this task. Reports Kendall correlation.

researchpythongo
0
3
Eval Gauntlet And Lima Judge EvalA

This protocol evaluates instruction-tuned LLMs across two distinct paradigms: traditional closed-domain NLP benchmarks and open-ended generation quality. It probes whether performance on standard accuracy-based tasks aligns with preference judgments from a large language model judge, highlighting the tension between task-diverse versus style-aligned training data. Use when the user wants to benchmark on MosaicML Eval Gauntlet, LIMA test set, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Evade EvalA

Evaluates multimodal models' ability to detect evasive or deceptive content in e-commerce product listings. It probes fine-grained single-violation detection and long-context, rule-integrated reasoning across multiple overlapping policy categories. Use when the user wants to benchmark on EVADE, or asks about evaluating this task. Reports Full Accuracy.

researchpythongo
0
3
Eva Immunology Benchmark EvalA

Evaluates a multimodal foundation model's ability to predict drug efficacy, classify disease endotypes, and align cross-species transcriptomic and histological data for immunology and inflammation research. Use when the user wants to benchmark on I&I benchmark, IBDome dataset, or asks about evaluating this task. Reports AUROC.

researchpythongo
0
3
Eurospeech Asr EvalA

Assesses the utility of the EuroSpeech multilingual corpus for fine-tuning automatic speech recognition (ASR) models. It measures the reduction in word error rate achieved by training on this corpus compared to baseline models across under-resourced European languages. Use when the user wants to benchmark on EuroSpeech, or asks about evaluating this task. Reports Word Error Rate (WER).

researchpythongo
0
3
Eurosat EvalA

This benchmark evaluates a model's ability to classify land use and land cover types from multi-spectral satellite imagery. It probes the capability to distinguish between 10 distinct environmental classes using patch-based remote sensing inputs across different spectral band configurations. Use when the user wants to benchmark on EuroSAT, or asks about evaluating this task. Reports classification accuracy (%).

researchpythongo
0
3
Eureka Bench EvalA

Granular, capability-level analysis of large foundation models across multimodal reasoning, language understanding, safety, and stability. It dissects performance across fine-grained subcategories (e.g., geometric depth vs. height, single vs. multi-object detection) to reveal persistent failures and complementary strengths across models. Use when the user wants to benchmark on EUREKA-BENCH, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Eur Lex Sum EvalA

This benchmark evaluates long-form, multi- and cross-lingual summarization capabilities in the legal domain. It probes a model's ability to extract or generate concise summaries from lengthy, structurally complex EU legal documents across 24 official EU languages, including cross-lingual transfer scenarios. Use when the user wants to benchmark on EUR-Lex-Sum, or asks about evaluating this task. Reports ROUGE-1.

researchpythongit
0
3
Eulearn Genus Classification EvalA

This benchmark evaluates whether deep learning models can accurately classify 3D surfaces by their topological genus (Euler characteristic) using point clouds and mesh connectivity graphs. It specifically probes the models' ability to leverage structural and adjacency information rather than relying solely on raw geometric coordinates. Use when the user wants to benchmark on EuLearn, or asks about evaluating this task. Reports F1 score.

researchpythongo
0
3
Eu Taxonomy Kpi EvalA

Evaluates LLMs on extracting and predicting EU Taxonomy-compliant Key Performance Indicators from corporate sustainability reports. It probes two capabilities: multi-label classification of economic activities against regulatory definitions, and zero-shot regression of financial KPI percentages (Turnover, CapEx, OpEx) from unstructured text. Use when the user wants to benchmark on EU Taxonomy Sustainability Reports Dataset, or asks about evaluating this task. Reports multi-label text classifi...

researchpythongo
0
3
Etree Edge Ai EvalA

Evaluates decentralized model aggregation frameworks on edge devices under IID and Non-IID data distributions. It measures classification accuracy and convergence speed to compare hierarchical tree-based learning against centralized federated learning and fully decentralized gossip learning. Use when the user wants to benchmark on HAR Using Smartphones Dataset, Pendigits, or asks about evaluating this task. Reports classification accuracy.

researchpythongo
0
3
Etp Ecg Linear Zero Shot EvalA

This protocol evaluates the transferability and robustness of pre-trained ECG representations by measuring classification performance on downstream cardiac disease datasets. It probes both supervised linear probing capabilities and cross-modal zero-shot generalization to specific cardiac conditions without fine-tuning the encoder. Use when the user wants to benchmark on PTB-XL, CPSC2018, or asks about evaluating this task. Reports AUC.

researchpythongo
0
3
Etide EvalA

Evaluates the ability of event-based motion forecasting models to predict future binary event occurrence maps (ON/OFF channels) from a sequence of past frames. It probes structural preservation of sparse motion traces, temporal consistency, and downstream utility for segmentation and tracking under varying traffic and high-speed motion regimes. Use when the user wants to benchmark on ETram, E-3DTrack, or asks about evaluating this task. Reports aIoU.

researchpythongit
0
3
Ethos Hate Speech EvalA

Evaluates the ability of NLP models to detect and classify hate speech in social media comments. It probes both binary hate/non-hate classification and multi-label categorization across specific demographic/identity-based hate categories. The benchmark emphasizes handling overlapping labels and class imbalance typical of real-world user-generated text. Use when the user wants to benchmark on ETHOS, or asks about evaluating this task. Reports F1-score (macro).

researchpythongo
0
3
Ethiomt EvalA

Machine translation performance across multiple low-resource Ethiopian languages paired with English, evaluating both English-to-Ethiopian and Ethiopian-to-English directions. It probes how model initialization (training from scratch vs. fine-tuning a multilingual model) and available corpus size impact translation quality. Use when the user wants to benchmark on EthioMT, or asks about evaluating this task. Reports spBLEU.

researchpythonperformance
0
3
Ethioemo EvalA

Evaluates LLMs' ability to classify multiple emotions in low-resource Ethiopian languages (Amharic, Afan Oromo, Somali, Tigrinya) and English. It probes cross-lingual transfer capabilities, the effectiveness of zero-shot and few-shot prompting strategies, and the impact of fine-tuning on multi-label emotion understanding tasks. Use when the user wants to benchmark on EthioEmo, or asks about evaluating this task. Reports Weighted-averaged F1-score.

researchpythonexpress
0
3
Ethereum Fraud Detection EvalA

Evaluates a model's ability to detect fraudulent Ethereum accounts by analyzing transaction interaction graphs and behavioral patterns. It specifically probes the model's capacity to handle imbalanced label distributions and distinguish between Ponzi schemes and phishing scams using self-supervised feature learning. Use when the user wants to benchmark on Ponzi Scheme Dataset, Phish Scam Dataset, or asks about evaluating this task. Reports Binary-F1.

researchpythongit
0
3
Esv Intervention EvalA

Evaluates the psychological and behavioral impact of an AI-generated emotional self-voice intervention compared to text-only and control conditions on goal-related resilience, confidence, motivation, and emotional states. Use when the user wants to benchmark on Custom Human-Subject Intervention Dataset, or asks about evaluating this task. Reports Self-report questionnaire scores.

researchpythongo
0
3
Estonian Native Llm Benchmark EvalA

Evaluates large language models on native Estonian language capabilities across seven tasks covering factual recall, grammar, morphology, vocabulary, summarization, and structured information extraction. The benchmark emphasizes cultural and linguistic authenticity by using human-curated or native-source data without machine translation, assessing both general and domain-specific competencies. Use when the user wants to benchmark on Exams, Trivia, Declension, Words, Grammar, News, Speaker Nam...

researchpythongo
0
3
Essay Quality EvalA

Evaluates the quality and linguistic characteristics of argumentative essays generated by different AI models compared to human-written texts. It probes logical structure, vocabulary richness, syntactic complexity, and stylistic markers through expert human annotation. Use when the user wants to benchmark on Student Essay Dataset (90 topics), or asks about evaluating this task. Reports Mean Rating Score.

researchpythonexpress
0
3
Esperanto EvalA

Evaluates the robustness of AI-generated text detectors against back-translation manipulations. It probes whether detectors can maintain high true positive rates when AI-generated text is translated to intermediate languages and back-translated to English, preserving semantics while evading detection. Use when the user wants to benchmark on ESPERANTO, or asks about evaluating this task. Reports True Positive Rate (TPR).

researchpythongo
0
3