Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

20,820
skills in category
868
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 6,385–6,408 of 20,820 skills

Long Cot Reasoning EvalA

This evaluation probes a model's ability to perform complex, multi-step reasoning across mathematics, coding, and scientific domains. It specifically measures the capacity to generate long chain-of-thought traces and produce correct final answers or executable code under strict generation constraints. Use when the user wants to benchmark on AIME24, AIME25, GPQA Diamond, LiveCodeBench v5, LiveCodeBench v6, or asks about evaluating this task. Reports average accuracy.

researchpythongo
0
3
Long Context Retrieval EvalA

Probes a model's ability to accurately retrieve hidden, specific information (needles) embedded within extremely long multimodal sequences (text, video, audio) and assesses its predictive stability over millions of tokens. Use when the user wants to benchmark on Paul Graham Essays (Synthetic), AlphaGo Documentary, VoxPopuli, or asks about evaluating this task. Reports recall.

researchpythongo
0
3
Long Context Reasoning EvalA

Evaluates a model's ability to perform multi-hop question answering and information retrieval over extremely long contexts (up to 128K tokens), while also measuring preservation of short-context reasoning and instruction-following capabilities. Use when the user wants to benchmark on LongBench v1, LongBench v2, MMLU, MATH-500, IFEval, Needle in a Haystack, RULER, or asks about evaluating this task. Reports pass@1 accuracy.

researchpythongo
0
3
Logiccat EvalA

This benchmark evaluates an LLM's ability to perform complex multi-step logical reasoning and chain-of-thought decomposition to generate correct SQL queries from natural language questions. It probes capabilities in mathematical deduction, physical knowledge integration, and cross-domain analytical querying. Use when the user wants to benchmark on LogicCat, or asks about evaluating this task. Reports Execution Accuracy (EX).

researchpythongo
0
3
Log Odds Mutation ScoringA

Evaluates a protein language model's zero-shot capability to score single-site amino acid mutations by comparing contextual likelihoods of wild-type versus mutant residues. It probes how well masked language modeling objectives capture evolutionary and structural constraints for mutant effect prediction. Use when the user has predictions and gold and needs to compute log-odds ratio.

researchpythongo
0
3
Loff Ta Image Classification EvalA

Evaluates the classification performance of a parameter-efficient model trained on cached foundation model features with tensor augmentations. It probes the model's ability to generalize across diverse image domains, object categories, and input resolutions using only lightweight classifier heads. Use when the user wants to benchmark on APTOS2019, DDSM, ISIC, AID, NABirds, Flowers102, StanfordCars, StanfordDogs, Oxford-III Pet, Caltech-101, SUN397, or asks about evaluating this task. Reports ...

researchpythongo
0
3
Lod Ood Detection EvalA

Evaluates a model's ability to distinguish in-distribution (ID) samples from out-of-distribution (OOD) samples using a threshold-free loss-difference clustering approach. It measures detection performance across standard benchmarks with diverse natural images and hard benchmarks where OOD classes share the same source dataset as ID. Use when the user wants to benchmark on CIFAR100, SVHN, Places, LSUN-Crop, LSUN-Resize, Textures, CIFAR10, TinyImageNet, or asks about evaluating this task. Repor...

researchpythongo
0
3
Locomo EvalA

This benchmark evaluates the long-term conversational memory of LLM agents by testing their ability to answer questions, summarize events, and generate multi-modal dialogues over very long, multi-session conversations. It probes how well models retain and reason over temporal and causal information across hundreds of turns and thousands of tokens. Use when the user wants to benchmark on LoCoMo, or asks about evaluating this task. Reports F1-score.

researchpythongo
0
3
Loco1 EvalA

Evaluates long-context retrieval capabilities on real-world documents where relevant information spans entire texts, such as legal contracts and medical notes. It specifically probes a model's ability to locate and rank relevant passages without relying on truncation or chunking strategies that often bias standard retrievers. Use when the user wants to benchmark on LoCoV1, or asks about evaluating this task. Reports nDCG@10.

researchpythonperformance
0
3
Localized Weather Prediction EvalA

Evaluates the ability of neural network architectures (KANs, TKANs, RNNs) to forecast localized weather variables (temperature, precipitation, pressure) one day ahead. Probes nonlinear time-series modeling and regression accuracy under varying data distributions, such as low precipitation versus high temperature variance. Use when the user wants to benchmark on Abidjan, Kigali, or asks about evaluating this task. Reports R².

researchpythongo
0
3
Local Prompt Ood EvalA

This evaluation probes a vision-language model's ability to distinguish in-distribution images from out-of-distribution samples using few-shot prompt tuning. It specifically tests fine-grained regional outlier detection by measuring how well the model separates known classes from diverse OOD datasets and semantically similar near-OOD subsets. Use when the user wants to benchmark on ImageNet-1K & OOD combination, or asks about evaluating this task. Reports FPR95.

researchpythongo
0
3
Local Planner Benchmarking EvalA

Compares local trajectory planners (DWB and TEB) for mobile manipulators by evaluating path smoothness, end-effector stability, trajectory deviation from a global path, and navigation accuracy/time in static and dynamic simulated environments. Use when the user wants to benchmark on Simulated Environments (Playground, Office, Warehouse), or asks about evaluating this task. Reports $\mathbf{p}_{e}$ (end-effector stability).

researchpythongo
0
3
Loasr Bench EvalA

Evaluates large speech language models on low-resource automatic speech recognition across 25 languages from 9 typologically diverse families. It probes cross-linguistic generalization, script bias (Latin vs. non-Latin), model scaling effects, and the impact of language-aware prompting on transcription accuracy. Use when the user wants to benchmark on LoASR-Bench, or asks about evaluating this task. Reports error rates.

researchpythonperformance
0
3
Load TimeA

Evaluates the data loading speed and rendering interactivity of the encube visual analytics framework on a tiled display system under varying data volumes and GPU memory constraints. Use when the user has predictions and gold and needs to compute Load time ($T_{\mathrm{Load}}$).

researchpythongo
0
3
Lo Hi EvalA

Evaluates machine learning models on two drug discovery tasks: Hit Identification (predicting activity for novel, structurally dissimilar molecules) and Lead Optimization (ranking minor molecular modifications to predict activity changes). It probes a model's ability to generalize to unseen chemical space and capture fine-grained structure-activity relationships. Use when the user wants to benchmark on DRD2-Hi, HIV-Hi, KDR-Hi, Sol-Hi, DRD2-Lo, KCNH2-Lo, KDR-Lo, or asks about evaluating this t...

researchpythongo
0
3
Lnmbench EvalA

Evaluates the robustness of label noise learning (LNL) methods on medical image classification tasks. It probes model performance under varying noise types (symmetric, instance-dependent, real-world), noise ratios, and class imbalance distributions across multiple imaging modalities. Use when the user wants to benchmark on PathMNIST, DermaMNIST, BloodMNIST, OrganCMNIST, DRTiD, Kaggle DR+, CheXpert, or asks about evaluating this task. Reports average classification accuracy.

researchpythongo
0
3
Lm Weather EvalA

Evaluates on-device meteorological variable forecasting and imputation capabilities using a federated learning framework with personalized adapters. It tests the model's ability to predict regional weather trends and handle missing data under data scarcity and heterogeneous distributions. Use when the user wants to benchmark on On-device Weather Series (ODW1/ODW2), or asks about evaluating this task. Reports MAE.

researchpythonperformance
0
3
Lm Loss And Benchmark EvalA

Evaluates the impact of multilingual data mixtures on language modeling capability and downstream task performance across multiple languages. It probes whether English dominance or high language count negatively interferes with multilingual model training. Use when the user wants to benchmark on mC4, FineWeb2, or asks about evaluating this task. Reports language modeling loss.

researchpythonperformance
0
3
Lm Eval Harness Benchmarks EvalA

Evaluates generative language models on a suite of multiple-choice and open-ended benchmarks covering reasoning, commonsense, multitask proficiency, and truthfulness. It measures accuracy across diverse domains to assess generalization and the impact of data combination strategies. Use when the user wants to benchmark on AI2 Reasoning Challenge (ARC), HellaSwag, MMLU, TruthfulQA, BigBench, HumanEval, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Llms4subjects EvalA

Evaluates LLMs' capability to perform multilingual subject tagging for technical library records by ranking relevant GND taxonomy subjects based on title and abstract. It probes the model's ability to handle large-scale taxonomies, bilingual semantic processing, and customizable top-k ranking for digital library classification. Use when the user wants to benchmark on all-subjects, tib-core, or asks about evaluating this task. Reports top-k ranked list.

researchpythongo
0
3
Llms4ol2024 EvalA

Evaluates LLMs on ontology learning tasks including term typing, taxonomy induction, and non-taxonomic relation extraction across multiple domains and few-shot/zero-shot settings. Use when the user wants to benchmark on LLMs4OL-2024, or asks about evaluating this task. Reports F1-score.

researchpythongo
0
3
Llmrec EvalA

Benchmarks off-the-shelf large language models on five recommendation tasks: rating prediction, sequential recommendation, direct recommendation, explanation generation, and review summarization. It probes both accuracy-driven prediction capabilities and natural language generation for explainability, revealing gaps between objective metric scores and human-perceived quality in recommendation contexts. Use when the user wants to benchmark on LLMRec Benchmark (includes Beauty dataset), or asks...

researchpythongo
0
3
Llmke Wikidata EvalA

Evaluates LLMs' ability to predict object entities given subject-relation pairs in Wikidata, testing knowledge retrieval, entity disambiguation, and domain-specific reasoning across 21 relations spanning 7 domains. Use when the user wants to benchmark on ISWC 2023 LM-KBC Challenge dataset, or asks about evaluating this task. Reports F1-score.

researchpythongo
0
3
Llmjudge EvalA

This benchmark evaluates the agreement and ranking consistency of LLM-generated relevance judgments against human assessments. It probes whether automated scoring methods can reliably replicate human relevance labels and maintain correct document ordering for information retrieval tasks. Use when the user wants to benchmark on LLMJudge test set, or asks about evaluating this task. Reports Cohen's \kappa.

researchpythonperformance
0
3