Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

20,816
skills in category
868
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 6,361–6,384 of 20,816 skills

Loombench EvalA

Evaluates long-context language models across 22 benchmarks and 140 tasks, probing capabilities like long-form generation, information retrieval, and reasoning over extended contexts. It also assesses the efficiency of inference acceleration and RAG augmentation methods. Use when the user wants to benchmark on LOOMBench, or asks about evaluating this task. Reports task_accuracy.

researchpythongo
0
3
Loogle EvalA

Evaluates the ability of language models to comprehend and reason over long documents (up to 32k+ tokens) by testing short and long dependency tasks, including question answering, cloze completion, and summarization. Use when the user wants to benchmark on LooGLE, or asks about evaluating this task. Reports GPT4_score.

researchpythongo
0
3
Longvideo Bench EvalA

This benchmark evaluates long-context video-language understanding by testing a model's ability to retrieve specific moments from lengthy videos and reason over multimodal details. It distinguishes between single-moment visual perception and multi-moment relational reasoning across 17 fine-grained categories. Use when the user wants to benchmark on LongVideoBench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Longllmlingua EvalA

Evaluates the effectiveness of a question-aware prompt compression framework on long-context LLM tasks. It measures how well compressed prompts preserve key information and answer accuracy across multi-document QA, summarization, and code completion scenarios. Use when the user wants to benchmark on NaturalQuestions (Liu et al., 2023), LongBench (Bai et al., 2023), ZeroSCROLLS (Shaham et al., 2023), or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Longlamp EvalA

Evaluates a model's ability to generate personalized long-form text by integrating retrieved user profiles into a retrieval-augmented generation framework. It probes how well models can adapt their writing style and content to specific user attributes across different domains like emails, abstracts, reviews, and topic-based writing. Use when the user wants to benchmark on LongLaMP, or asks about evaluating this task. Reports METEOR.

researchpython
0
3
Longgenbench EvalA

Evaluates the ability of LLMs to maintain accuracy and logical consistency when generating long-text responses that answer multiple sequential questions from GSM8K or MMLU in a single pass. It specifically probes performance degradation as the number of generated questions increases. Use when the user wants to benchmark on LongGenBench-GSM8K, LongGenBench-MMLU, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Longform C EvalA

Evaluates instruction-following long-form text generation capabilities of LLMs across diverse in-domain and out-of-domain tasks, including news summarization, recipe and story generation, long-form QA, and multilingual generation, while also measuring general language understanding via MMLU. Use when the user wants to benchmark on LongForm-C, Writing Prompts, ELI5, Recipe Generation, MMLU, MLSUM, or asks about evaluating this task. Reports METEOR.

researchpython
0
3
Longcot EvalA

Probes a model's ability to maintain coherent state, plan, and execute multi-step reasoning over long, interdependent chains of thought spanning tens to hundreds of thousands of tokens across domains like chemistry, mathematics, and chess. Use when the user wants to benchmark on LongCoT, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Longbench Write EvalA

Evaluates an LLM's ability to generate ultra-long, coherent, and high-quality text (up to 20k words) while strictly adhering to explicit length constraints. It probes long-context generation capabilities, structural coherence over extended outputs, and instruction-following for length requirements. Use when the user wants to benchmark on LongBench-Write, or asks about evaluating this task. Reports Sq (Quality Score).

researchpythongo
0
3
Longbench V2 EvalA

Evaluates large language models' ability to comprehend and reason over realistic, extremely long contexts (up to 2M words) across six multitask domains. It probes deep understanding rather than shallow extraction by using challenging multiple-choice questions that require extended reasoning and careful reading. Use when the user wants to benchmark on LongBench v2, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Longbench Pro EvalA

Evaluates long-context understanding and reasoning capabilities of LLMs across bilingual (English/Chinese) tasks. It probes retrieval, ranking, ordering, multiple-choice, information extraction, and summarization under varying difficulty levels and context lengths. Use when the user wants to benchmark on LongBench Pro, or asks about evaluating this task. Reports LongBench Pro Score.

researchpythongo
0
3
Longbench EvalA

Evaluates large language models' ability to understand and process long contexts across bilingual (English and Chinese) multitask scenarios, including single/multi-document QA, summarization, few-shot learning, code completion, and synthetic tasks. Use when the user wants to benchmark on LongBench, or asks about evaluating this task. Reports F1.

researchpythongo
0
3
Longalign EvalA

Evaluates large language models' ability to follow instructions and retrieve information in long-context scenarios (up to 64k tokens), while also measuring their general capabilities and instruction-following performance in short-context settings. Use when the user wants to benchmark on LongBench-Chat, LongBench, MT-Bench, ARC, HellaSwag, TruthfulQA, MMLU, or asks about evaluating this task. Reports GPT-4 rating (1-10).

researchpythongo
0
3
Long Video Understanding EvalA

Evaluates a model's ability to perform temporal reasoning and question-answering on long-duration videos (several minutes to over an hour). It probes memory retention, attention allocation across extended sequences, and the capacity to filter irrelevant visual content while preserving critical frames. Use when the user wants to benchmark on LongVideoBench, MLVU, VideoMME (Long), LVBench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Long Video Qa EvalA

Evaluates long-form video understanding and multimodal reasoning capabilities across multiple-choice question answering tasks. It probes the model's ability to handle extended temporal dependencies, spatial-temporal reasoning, and tool-augmented retrieval in videos ranging from short clips to hour-long content. Use when the user wants to benchmark on LongVideoBench, VideoMME, LVBench, MLVU, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Long Term Vpr EvalA

Evaluates long-term visual place recognition (VPR) capabilities in dynamic underwater benthic environments. It probes a model's ability to geolocate camera views over multi-year intervals despite habitat changes, varying terrain ruggedness, and sub-decimeter registration errors. Use when the user wants to benchmark on Benthic Reference Sites Dataset, or asks about evaluating this task. Reports Recall@K.

researchpythondatabase
0
3
Long Term Motion EvalA

Evaluates long-term motion representations derived from dense point-tracking against image-based baselines across five perceptual tasks. It probes temporal generalization, motion representation efficiency, and the ability to capture spatio-temporal dynamics for classification and regression. Use when the user wants to benchmark on SSV2 (Temporal Dataset subset), Jester, VB100, RAVDESS, MITFabric, ADVIO, or asks about evaluating this task. Reports classification accuracy.

researchpythongo
0
3
Long Tail Session Rec EvalA

This evaluation probes a session-based recommendation model's ability to accurately predict the next item in a user's interaction sequence while mitigating popularity bias. It measures both standard ranking accuracy and the model's capacity to recommend long-tail items, ensuring recommendations align with user-specific item distribution preferences rather than just global popularity. Use when the user wants to benchmark on YOOCHOOSE, Last.fm, or asks about evaluating this task. Reports Recall...

researchpythongo
0
3
Long Sequence Modeling EvalA

Evaluates the capability of sequence modeling architectures to capture long-range dependencies across text, audio, and image modalities, as well as their computational efficiency and compatibility with standard Transformer and CNN backbones. Use when the user wants to benchmark on Long Range Arena (LRA), Speech Commands (SC), WikiText-103, GLUE, ImageNet-1k, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Long Range Arena EvalA

Evaluates efficient Transformer architectures on long-context sequence modeling tasks spanning text, images, and structured data. Probes capabilities in hierarchical reasoning, spatial navigation, and retrieval over sequences up to 16K tokens. Use when the user wants to benchmark on ListOps, Text Classification, Retrieval, Image Classification, Pathfinder / Path-X, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Long Ner EvalA

Evaluates named entity recognition capabilities, specifically probing a model's ability to handle class imbalance, out-of-vocabulary terms, and long or complex entity names across biomedical and general domain texts. Use when the user wants to benchmark on NCBI-disease, BC5CDR-disease, BC5CDR-chemical, BC4CHEMD, BC2GM, JNLPBA, LINNAEUS, Species-800, CoNLL-2003, WNUT-2017, or asks about evaluating this task. Reports F1.

researchpythongo
0
3
Long Form Scientific Summarization EvalA

Evaluates long-form scientific summarization models on their ability to generate relevant and faithful abstracts across clinical, chemical, and biomedical domains. It probes how calibration set construction and candidate selection strategies affect model performance on standard relevance and faithfulness metrics. Use when the user wants to benchmark on Scientific Summarization Datasets, or asks about evaluating this task. Reports Rouge-1 F1.

researchpythongo
0
3
Long Doc Rouge EvalA

Evaluates the quality of abstractive summaries for long scientific documents by measuring n-gram overlap between generated text and reference abstracts. Use when the user wants to benchmark on arXiv, PubMed, or asks about evaluating this task. Reports ROUGE-1.

researchpython
0
3
Long Cot Reasoning EvalA

This evaluation probes a model's ability to perform complex, multi-step reasoning across mathematics, coding, and scientific domains. It specifically measures the capacity to generate long chain-of-thought traces and produce correct final answers or executable code under strict generation constraints. Use when the user wants to benchmark on AIME24, AIME25, GPQA Diamond, LiveCodeBench v5, LiveCodeBench v6, or asks about evaluating this task. Reports average accuracy.

researchpythongo
0
3