Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 6,385–6,408 of 20,820 skills
This evaluation probes a model's ability to perform complex, multi-step reasoning across mathematics, coding, and scientific domains. It specifically measures the capacity to generate long chain-of-thought traces and produce correct final answers or executable code under strict generation constraints. Use when the user wants to benchmark on AIME24, AIME25, GPQA Diamond, LiveCodeBench v5, LiveCodeBench v6, or asks about evaluating this task. Reports average accuracy.
Probes a model's ability to accurately retrieve hidden, specific information (needles) embedded within extremely long multimodal sequences (text, video, audio) and assesses its predictive stability over millions of tokens. Use when the user wants to benchmark on Paul Graham Essays (Synthetic), AlphaGo Documentary, VoxPopuli, or asks about evaluating this task. Reports recall.
Evaluates a model's ability to perform multi-hop question answering and information retrieval over extremely long contexts (up to 128K tokens), while also measuring preservation of short-context reasoning and instruction-following capabilities. Use when the user wants to benchmark on LongBench v1, LongBench v2, MMLU, MATH-500, IFEval, Needle in a Haystack, RULER, or asks about evaluating this task. Reports pass@1 accuracy.
This benchmark evaluates an LLM's ability to perform complex multi-step logical reasoning and chain-of-thought decomposition to generate correct SQL queries from natural language questions. It probes capabilities in mathematical deduction, physical knowledge integration, and cross-domain analytical querying. Use when the user wants to benchmark on LogicCat, or asks about evaluating this task. Reports Execution Accuracy (EX).
Evaluates a protein language model's zero-shot capability to score single-site amino acid mutations by comparing contextual likelihoods of wild-type versus mutant residues. It probes how well masked language modeling objectives capture evolutionary and structural constraints for mutant effect prediction. Use when the user has predictions and gold and needs to compute log-odds ratio.
Evaluates the classification performance of a parameter-efficient model trained on cached foundation model features with tensor augmentations. It probes the model's ability to generalize across diverse image domains, object categories, and input resolutions using only lightweight classifier heads. Use when the user wants to benchmark on APTOS2019, DDSM, ISIC, AID, NABirds, Flowers102, StanfordCars, StanfordDogs, Oxford-III Pet, Caltech-101, SUN397, or asks about evaluating this task. Reports ...
Evaluates a model's ability to distinguish in-distribution (ID) samples from out-of-distribution (OOD) samples using a threshold-free loss-difference clustering approach. It measures detection performance across standard benchmarks with diverse natural images and hard benchmarks where OOD classes share the same source dataset as ID. Use when the user wants to benchmark on CIFAR100, SVHN, Places, LSUN-Crop, LSUN-Resize, Textures, CIFAR10, TinyImageNet, or asks about evaluating this task. Repor...
This benchmark evaluates the long-term conversational memory of LLM agents by testing their ability to answer questions, summarize events, and generate multi-modal dialogues over very long, multi-session conversations. It probes how well models retain and reason over temporal and causal information across hundreds of turns and thousands of tokens. Use when the user wants to benchmark on LoCoMo, or asks about evaluating this task. Reports F1-score.
Evaluates long-context retrieval capabilities on real-world documents where relevant information spans entire texts, such as legal contracts and medical notes. It specifically probes a model's ability to locate and rank relevant passages without relying on truncation or chunking strategies that often bias standard retrievers. Use when the user wants to benchmark on LoCoV1, or asks about evaluating this task. Reports nDCG@10.
Evaluates the ability of neural network architectures (KANs, TKANs, RNNs) to forecast localized weather variables (temperature, precipitation, pressure) one day ahead. Probes nonlinear time-series modeling and regression accuracy under varying data distributions, such as low precipitation versus high temperature variance. Use when the user wants to benchmark on Abidjan, Kigali, or asks about evaluating this task. Reports R².
This evaluation probes a vision-language model's ability to distinguish in-distribution images from out-of-distribution samples using few-shot prompt tuning. It specifically tests fine-grained regional outlier detection by measuring how well the model separates known classes from diverse OOD datasets and semantically similar near-OOD subsets. Use when the user wants to benchmark on ImageNet-1K & OOD combination, or asks about evaluating this task. Reports FPR95.
Compares local trajectory planners (DWB and TEB) for mobile manipulators by evaluating path smoothness, end-effector stability, trajectory deviation from a global path, and navigation accuracy/time in static and dynamic simulated environments. Use when the user wants to benchmark on Simulated Environments (Playground, Office, Warehouse), or asks about evaluating this task. Reports $\mathbf{p}_{e}$ (end-effector stability).
Evaluates large speech language models on low-resource automatic speech recognition across 25 languages from 9 typologically diverse families. It probes cross-linguistic generalization, script bias (Latin vs. non-Latin), model scaling effects, and the impact of language-aware prompting on transcription accuracy. Use when the user wants to benchmark on LoASR-Bench, or asks about evaluating this task. Reports error rates.
Evaluates the data loading speed and rendering interactivity of the encube visual analytics framework on a tiled display system under varying data volumes and GPU memory constraints. Use when the user has predictions and gold and needs to compute Load time ($T_{\mathrm{Load}}$).
Evaluates machine learning models on two drug discovery tasks: Hit Identification (predicting activity for novel, structurally dissimilar molecules) and Lead Optimization (ranking minor molecular modifications to predict activity changes). It probes a model's ability to generalize to unseen chemical space and capture fine-grained structure-activity relationships. Use when the user wants to benchmark on DRD2-Hi, HIV-Hi, KDR-Hi, Sol-Hi, DRD2-Lo, KCNH2-Lo, KDR-Lo, or asks about evaluating this t...
Evaluates the robustness of label noise learning (LNL) methods on medical image classification tasks. It probes model performance under varying noise types (symmetric, instance-dependent, real-world), noise ratios, and class imbalance distributions across multiple imaging modalities. Use when the user wants to benchmark on PathMNIST, DermaMNIST, BloodMNIST, OrganCMNIST, DRTiD, Kaggle DR+, CheXpert, or asks about evaluating this task. Reports average classification accuracy.
Evaluates on-device meteorological variable forecasting and imputation capabilities using a federated learning framework with personalized adapters. It tests the model's ability to predict regional weather trends and handle missing data under data scarcity and heterogeneous distributions. Use when the user wants to benchmark on On-device Weather Series (ODW1/ODW2), or asks about evaluating this task. Reports MAE.
Evaluates the impact of multilingual data mixtures on language modeling capability and downstream task performance across multiple languages. It probes whether English dominance or high language count negatively interferes with multilingual model training. Use when the user wants to benchmark on mC4, FineWeb2, or asks about evaluating this task. Reports language modeling loss.
Evaluates generative language models on a suite of multiple-choice and open-ended benchmarks covering reasoning, commonsense, multitask proficiency, and truthfulness. It measures accuracy across diverse domains to assess generalization and the impact of data combination strategies. Use when the user wants to benchmark on AI2 Reasoning Challenge (ARC), HellaSwag, MMLU, TruthfulQA, BigBench, HumanEval, or asks about evaluating this task. Reports accuracy.
Evaluates LLMs' capability to perform multilingual subject tagging for technical library records by ranking relevant GND taxonomy subjects based on title and abstract. It probes the model's ability to handle large-scale taxonomies, bilingual semantic processing, and customizable top-k ranking for digital library classification. Use when the user wants to benchmark on all-subjects, tib-core, or asks about evaluating this task. Reports top-k ranked list.
Evaluates LLMs on ontology learning tasks including term typing, taxonomy induction, and non-taxonomic relation extraction across multiple domains and few-shot/zero-shot settings. Use when the user wants to benchmark on LLMs4OL-2024, or asks about evaluating this task. Reports F1-score.
Benchmarks off-the-shelf large language models on five recommendation tasks: rating prediction, sequential recommendation, direct recommendation, explanation generation, and review summarization. It probes both accuracy-driven prediction capabilities and natural language generation for explainability, revealing gaps between objective metric scores and human-perceived quality in recommendation contexts. Use when the user wants to benchmark on LLMRec Benchmark (includes Beauty dataset), or asks...
Evaluates LLMs' ability to predict object entities given subject-relation pairs in Wikidata, testing knowledge retrieval, entity disambiguation, and domain-specific reasoning across 21 relations spanning 7 domains. Use when the user wants to benchmark on ISWC 2023 LM-KBC Challenge dataset, or asks about evaluating this task. Reports F1-score.
This benchmark evaluates the agreement and ranking consistency of LLM-generated relevance judgments against human assessments. It probes whether automated scoring methods can reliably replicate human relevance labels and maintain correct document ordering for information retrieval tasks. Use when the user wants to benchmark on LLMJudge test set, or asks about evaluating this task. Reports Cohen's \kappa.