Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 6,577–6,600 of 20,836 skills
This benchmark evaluates large language models' capabilities in biology research, including literature retrieval, figure/table interpretation, database querying, protocol troubleshooting, and DNA/protein sequence manipulation. It probes whether models can perform multi-step, tool-dependent scientific reasoning or rely on memorization and heuristic guesswork. Use when the user wants to benchmark on LAB-Bench, or asks about evaluating this task. Reports accuracy.
Evaluates LLMs on multilingual proficiency across Spanish varieties and regional languages of Spain and Latin America (Basque, Catalan, Galician). It probes capabilities in natural language inference, reasoning, question answering, summarization, and linguistic acceptability using a resource-efficient few-shot configuration. Use when the user wants to benchmark on La Leaderboard (66 datasets), or asks about evaluating this task. Reports exact-match.
Evaluates abstractive text summarization models on their ability to generate concise, fluent, and coherent Marathi news summaries from longer source articles. It benchmarks performance against both a newly curated large-scale dataset (MahaSum) and an existing multilingual benchmark (XL-Sum Marathi subset). Use when the user wants to benchmark on XLsum, MahaSum, or asks about evaluating this task. Reports ROUGE.
Evaluates an agent's ability to safely manage power grid topology under unexpected line failures and fluctuating renewable energy generation. It probes robustness to sudden grid attacks and adaptability to changing energy mix proportions over a full year of seasonal scenarios. Use when the user wants to benchmark on Grid2Op (NeurIPS 2020 L2RPN), or asks about evaluating this task. Reports total_reward.
This evaluation probes the ability of an adaptive AI-to-AI deferral framework to selectively route clinical text classification tasks between domain-adapted BERT models and LLMs based on uncertainty signals. It measures whether intelligent routing improves classification accuracy while minimizing expensive LLM usage across binary and multi-class clinical NLP tasks. Use when the user wants to benchmark on ADE Corpus V2, MIMIC-IV Treatment Outcomes, or asks about evaluating this task. Reports F...
Probes the ability to detect phoneme-level mispronunciations in second-language (L2) speech by comparing acoustic features of learner speech against a voice-cloned native reference using dynamic time warping on MFCC envelopes. Use when the user wants to benchmark on L2-Arctic, or asks about evaluating this task. Reports accuracy.
Evaluates a streaming foreign accent conversion system's ability to neutralize non-native pronunciation while preserving speaker identity. The protocol uses self-reconstruction mode and compares synthesized outputs against offline-generated golden speaker utterances as a reference baseline. Use when the user wants to benchmark on L2-ARCTIC (Indian subset), or asks about evaluating this task. Reports non-native accent confidence.
Evaluates sentiment classification capability on Kyrgyz language text. It measures how effectively a model can distinguish between positive and negative sentiments using a manually annotated benchmark dataset. Use when the user wants to benchmark on kyrgyz-sst2, or asks about evaluating this task. Reports F1-score (Weighted).
This benchmark evaluates keyword spotting models on mobile devices by measuring classification accuracy, computational cost (FLOPs, parameters), and real-time inference latency on one-second audio utterances. It specifically tests whether temporal convolutions can replace 2D convolutions to reduce computational load while maintaining or improving accuracy. Use when the user wants to benchmark on Google Speech Commands Dataset, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates language models' ability to recognize formal game-theoretic structures (e.g., principal-agent conflict, signaling, strategic omission) in real-world knowledge work scenarios without explicit task hints. It measures the gap between a model's theoretical understanding of these concepts and its capacity for unprompted, practical problem framing. Use when the user wants to benchmark on KWBench, or asks about evaluating this task. Reports Pass Rate.
Evaluates knowledge-intensive visual grounding (KVG), requiring models to combine domain-specific reasoning with fine-grained visual perception to locate specific entities in images containing multiple similar objects. Use when the user wants to benchmark on KVG-Bench, or asks about evaluating this task. Reports accuracy.
Evaluates text-to-image models on knowledge-intensive generation across six high-school academic subjects and two languages. It probes scientific fidelity, logical reasoning, symbolic precision, and multilingual robustness using textbook-derived prompts and atomic checklist verification. Use when the user wants to benchmark on KVBench, or asks about evaluating this task. Reports performance_score.
Evaluates multimodal vision-language models on gastrointestinal endoscopy image understanding and clinical question answering. It probes factual recall, multi-step clinical reasoning across varying complexity levels, and robustness to realistic visual perturbations like motion blur and color shifts. Use when the user wants to benchmark on Kvasir-VQA-x1, or asks about evaluating this task. Reports BERT-F1.
Probes a model's ability to perform multimodal understanding and generation on gastrointestinal endoscopic images. It evaluates capabilities in descriptive captioning, answering clinical questions about visual findings, and synthesizing anatomically plausible medical images from text prompts. Use when the user wants to benchmark on Kvasir-VQA, or asks about evaluating this task. Reports BLEU.
Evaluates the quality of learned key-value (KV) cache eviction policies in preserving long-context reasoning and generation capabilities under strict memory constraints. It measures how well different compression strategies retain critical tokens without access to query-specific attention scores during the compression phase. Use when the user wants to benchmark on RULER-4k, OASST2-4k, BoolQ, ARC-Challenge, MMLU, HellaSwag, GovReport, or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to distinguish between questions it can answer confidently (known) and those it cannot (unknown). It probes the model's metacognitive uncertainty articulation and calibration under varying prompt conditions. Use when the user wants to benchmark on KUQ, or asks about evaluating this task. Reports F1-score.
Evaluates automatic speech recognition (ASR) models on spontaneous, radio-based Bambara speech containing real-world artifacts like code-switching, overlapping speakers, and background noise. It measures transcription accuracy under pragmatic normalization conditions. Use when the user wants to benchmark on Kunkado Test, Nyana-Eval, or asks about evaluating this task. Reports WER (%).
Evaluates the in-context learning capabilities of a relational foundation model on multi-table predictive tasks across diverse domains. It probes the model's ability to perform binary classification, multi-class classification, and regression directly on relational database structures without flattening or fine-tuning. Use when the user wants to benchmark on RelBenchV1, RelBenchV2, SALT, 4DBInfer, or asks about evaluating this task. Reports AUROC.
Evaluates recommendation models on live streaming data by testing their ability to rank relevant live rooms or streamers (top-K) and predict click-through probabilities (CTR), while accounting for real-time temporal dynamics and dynamic candidate pools. Use when the user wants to benchmark on KuaiLive, or asks about evaluating this task. Reports Recall@{5, 10, 20}.
Evaluates recommender system policies across three temporal granularities: request-level list-wise ranking, whole-session sequential recommendation under reinforcement learning, and cross-session user retention optimization. It measures how well simulated agents balance immediate engagement rewards with long-term user retention and list diversity. Use when the user wants to benchmark on KuaiRand, ML-1m, or asks about evaluating this task. Reports Average L-reward.
Evaluates session-based recommendation models on predicting the next item in a user's click sequence by integrating knowledge graph attributes and temporal dynamics between clicks. Use when the user wants to benchmark on Yoochoose, Diginetica, Last-fm, or asks about evaluating this task. Reports Recall@20.
Evaluates machine translation performance across 41 Creole languages, testing cross-lingual transfer and the impact of data cleaning and scale on translation quality. Use when the user wants to benchmark on Kreyol-MT, or asks about evaluating this task. Reports BLEU.
Evaluates multimodal large language models on a comprehensive suite of perception-language, nonverbal reasoning, OCR-free text understanding, and web page comprehension tasks. It measures zero-shot and few-shot cross-modal transfer, in-context learning, and the ability to align visual perception with language generation without external tools or fine-tuning. Use when the user wants to benchmark on MS COCO Caption, Flickr30k, VQAv2, VizWiz, Raven IQ Test, Rendered SST-2, HatefulMemes, WebSRC, ...
This benchmark evaluates vision-language models on multimodal medical reasoning using questions derived from the Korean Medical Licensing Examination. It probes the models' ability to integrate textual and visual evidence across diverse clinical imaging modalities, including cross-image reasoning when multiple scans are provided. Use when the user wants to benchmark on KorMedMCQA-V, or asks about evaluating this task. Reports accuracy.