Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

20,820
skills in category
868
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 6,433–6,456 of 20,820 skills

Llama Vits Tts EvalA

Evaluates the naturalness, intelligibility, and emotional expressiveness of a non-autoregressive TTS model enhanced with LLM-derived semantic embeddings. Probes how well semantic tokens from Llama2 versus BERT improve acoustic quality and emotion similarity compared to baselines. Use when the user wants to benchmark on LJSpeech, 1-hour LJSpeech, EmoV_DB_bea_sem, or asks about evaluating this task. Reports ESMOS.

researchpythonexpress
0
3
Llama Energy Latency EvalA

This benchmark evaluates the inference latency and energy consumption of LLaMA models (7B-65B) across different GPU hardware (V100, A100) and sharding configurations. It probes the trade-offs between computational throughput, power usage, and hardware efficiency during text generation. Use when the user wants to benchmark on Alpaca, GSM8K, or asks about evaluating this task. Reports energy per second (Watts).

researchpythonperformance
0
3
Llama Berry EvalA

Evaluates LLMs on complex mathematical reasoning using search-based inference (SR-MCTS) rather than direct generation. It measures success rates across varying difficulty levels, from grade-school math to Olympiad-level problems, by testing both majority-vote and best-of-k strategies. Use when the user wants to benchmark on AIME24, AMC23, Math Odyssey, GPQA Diamond, OlympiadBench, College Math, MMLU STEM, GSM8K, GSMHard, MATH500, or asks about evaluating this task. Reports major@k.

researchpythongo
0
3
Livs T2i Alignment EvalA

Evaluates how well text-to-image models align with pluralistic, intersectional community preferences for urban public space design. Probes whether multi-criteria preference optimization (DPO) improves alignment over a baseline, and how prompt origin and annotator demographics influence preference consistency and rating distributions. Use when the user wants to benchmark on LIVS, or asks about evaluating this task. Reports preference_rate.

researchpython
0
3
Livexiv EvalA

This benchmark evaluates the multi-modal reasoning capabilities of Large Multimodal Models (LMMs) on scientific content scraped from ArXiv papers. It specifically probes visual question answering (VQA) on figures and table question answering (TQA) using multiple-choice formats derived from real-time academic publications. Use when the user wants to benchmark on LiveXiv, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Liveweb Ie EvalA

Evaluates web information extraction systems on live, dynamically evolving websites by testing their ability to identify target attributes and extract corresponding values from natural language queries. It probes the robustness of extraction pipelines against real-time web dynamics and complex layouts that break static HTML parsing. Use when the user wants to benchmark on LiveWeb-IE, or asks about evaluating this task. Reports F1 score.

researchpythontesting
0
3
Livefact EvalA

Evaluates LLMs on fake news detection under dynamic, time-evolving evidence streams. It probes both binary classification capability (Real/Fake) and reasoning capability (handling ambiguity when evidence is incomplete), while explicitly measuring benchmark data contamination and epistemic humility. Use when the user wants to benchmark on LiveFact November 2025 dataset, or asks about evaluating this task. Reports average_score.

researchpythongo
0
3
Liveclktb EvalA

Evaluates multilingual LLMs on their ability to transfer factual knowledge across languages using time-sensitive, real-world events that occur after the model's training cutoff. It measures both in-language factual recall and cross-lingual generalization performance across domains like music, movies, and sports. Use when the user wants to benchmark on LiveCLKTBench, or asks about evaluating this task. Reports Transfer Score.

researchpythongo
0
3
Liveclin EvalA

LiveClin probes real-world clinical reasoning and longitudinal case management by evaluating models on dynamic, multimodal patient scenarios. It tests the ability to maintain context across sequential diagnostic, treatment, and follow-up questions while resisting data contamination from static training corpora. Use when the user wants to benchmark on LiveClin, or asks about evaluating this task. Reports Case Accuracy.

researchpythongo
0
3
Livebench EvalA

Evaluates large language models across 18 tasks spanning math, coding, reasoning, language, instruction following, and data analysis. It uses dynamically updated, objectively scored questions from recent real-world sources to minimize test-set contamination and avoid LLM-judging biases. Use when the user wants to benchmark on LiveBench, or asks about evaluating this task. Reports LiveBench score.

researchpythongo
0
3
Liveaopsbench EvalB

Evaluates large language models' mathematical reasoning capabilities on Olympiad-level competition problems. It specifically probes whether models possess genuine problem-solving skills or merely rely on memorized pre-training data by using a continuously updated, timestamped benchmark to measure contamination-resistant accuracy. Use when the user wants to benchmark on AoPS24, Math, OlympiadBench, OmniMath, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Live Meta Mcg EvalA

Evaluates the ability of objective video quality assessment models to predict human-perceived quality of mobile cloud gaming videos distorted by compression and resizing artifacts. It benchmarks both general-purpose and gaming-specific no-reference models against human subjective ratings. Use when the user wants to benchmark on LIVE-Meta Mobile Cloud Gaming (LIVE-Meta MCG), or asks about evaluating this task. Reports SROCC.

researchpythongo
0
3
Litxbench EvalA

Evaluates an LLM's ability to extract structured material science experimental data from scientific literature text. It specifically probes the model's capacity to link extracted measurements to material processing lineages and adhere to a predefined schema. Use when the user wants to benchmark on LitXAlloy, or asks about evaluating this task. Reports F1 score.

researchpythongo
0
3
Litqa2 EvalA

Evaluates an agentic system's ability to retrieve relevant scientific literature, re-rank passages, and answer multiple-choice questions based on the retrieved text. It probes retrieval coverage, passage localization, and question-answering accuracy under information loss constraints. Use when the user wants to benchmark on LitQA2, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Liputan6 EvalA

Evaluates the ability of models to generate concise, accurate summaries of Indonesian news articles. It probes both extractive and abstractive summarization capabilities, highlighting performance gaps on more abstract summaries and the limitations of n-gram overlap metrics. Use when the user wants to benchmark on Liputan6, or asks about evaluating this task. Reports ROUGE F-1 (R1, R2, RL).

researchpythongit
0
3
Lip2wav EvalA

Evaluates a model's ability to synthesize natural, speaker-specific speech from unconstrained lip movements in large-vocabulary settings. It measures how well the model captures individual speaking styles and contextual cues from video frames. Use when the user wants to benchmark on Lip2Wav, GRID, TCD-TIMIT lip speaker corpus, or asks about evaluating this task. Reports mel reconstruction loss.

researchpythongo
0
3
Lip To Speech EvalA

Evaluates a model's ability to synthesize high-fidelity, intelligible speech directly from visual lip movements. It probes perceptual audio quality, content accuracy, and speaker identity preservation in a cross-dataset generalization setting. Use when the user wants to benchmark on LRS3-TED, LRS2-BBC, or asks about evaluating this task. Reports WER.

researchpython
0
3
Lip EvalA

Evaluates a model's ability to perform joint human semantic part segmentation and 16-keypoint pose estimation on diverse, unconstrained images with varying appearances, occlusions, and backgrounds. It probes the model's capacity to leverage structural body priors to resolve ambiguities in part boundaries and joint localization. Use when the user wants to benchmark on LIP, PASCAL-Person-Part, MPII Human Pose, ATR, or asks about evaluating this task. Reports mean IoU.

researchpython
0
3
Link Prediction EvalA

Evaluates a model's ability to predict missing or future edges in a graph based on its structural topology. It specifically probes whether the model learns meaningful graph patterns or merely exploits implicit degree biases inherent in the standard edge sampling procedure. Use when the user wants to benchmark on Empirical graphs (90 datasets), or asks about evaluating this task. Reports AUC-ROC.

researchpythonnode
0
3
Linguistic Shibboleth Hiring EvalA

Evaluates whether LLMs systematically penalize candidates for using hedging language in professional interview responses, despite identical substantive content. It probes the model's ability to decouple communication style from perceived technical competence and hiring suitability. Use when the user wants to benchmark on Linguistic Shibboleth Hiring Benchmark, or asks about evaluating this task. Reports average_score.

researchpythongo
0
3
Linguistic Probing EvalA

Evaluates how fine-tuning on downstream NLP tasks redistributes linguistic knowledge across transformer layers. It probes for part-of-speech tagging, syntactic chunking, and semantic tagging capabilities using linear classifiers on layer-wise hidden states. Use when the user wants to benchmark on Penn TreeBank, CoNLL 2000, Parallel Meaning Bank, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Linguistic DiversityA

Evaluates the lexical, semantic, and structural diversity of natural language instructions across datasets. It quantifies repetition, vocabulary richness, semantic coverage, and syntactic complexity to identify construction biases and limitations in dataset design. Use when the user has predictions and gold and needs to compute ROUGE-L.

researchpythongo
0
3
Linguasafe EvalA

Evaluates multilingual safety alignment of LLMs by measuring their ability to reject harmful prompts and accept benign ones across 12 languages and a hierarchical safety taxonomy. It probes both direct safety performance (vulnerability to harmful content) and indirect performance (oversensitivity to benign requests). Use when the user wants to benchmark on LinguaSafe, or asks about evaluating this task. Reports Vulnerability Score.

researchpythongo
0
3
Lingoqa EvalA

Evaluates vision-language models on autonomous driving video question answering, testing their ability to understand temporal visual context, describe scenes, predict actions, and justify answers based on driving scenarios. Use when the user wants to benchmark on LingoQA, or asks about evaluating this task. Reports Ling-Judge.

researchpythongo
0
3