Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 3,505–3,528 of 22,871 skills
Evaluates whether LLMs can correctly answer questions that require interacting with external tools, rather than relying on pre-trained knowledge. It probes tool selection, multi-step tool chaining, and reasoning over execution traces in an open-ended setting. Use when the user wants to benchmark on ToolQA, or asks about evaluating this task. Reports success rate.
Evaluates large language models' tool-use and function-calling capabilities, specifically probing multi-turn dialogues and agentic workflows such as search and memory retrieval. Use when the user wants to benchmark on BFCL-v4, τ-Bench, τ²-Bench, or asks about evaluating this task. Reports BFCL-v4 Overall.
Evaluates the safety and helpfulness of language model agents interacting with tools in a simulated environment. It measures how well an automated emulator and evaluator align with human judgments, and quantifies agent failure rates under standard and adversarial conditions. Use when the user wants to benchmark on ToolEmu Agent Trajectories, or asks about evaluating this task. Reports Cohen's κ (Quadratic-weighted).
This benchmark evaluates an LLM's ability to understand API documentation, select the correct tool, and generate valid arguments to fulfill a user's goal. It probes tool manipulation capabilities across a diverse set of 8 real-world applications, ranging from simple single-call tasks to complex multi-step reasoning. Use when the user wants to benchmark on ToolBench, or asks about evaluating this task. Reports success rate.
Evaluates a language model's ability to generalize tool-use capabilities to unseen tools and APIs through multi-turn interaction, parameter selection, and final response generation. It measures how well compact models trained on simulated data can adapt to real-world and out-of-dataset tool scenarios without task-specific fine-tuning. Use when the user wants to benchmark on ToolAlpaca Evaluation Set, GPT4Tools Test Set, or asks about evaluating this task. Reports Overall.
Evaluates an LLM's ability to autonomously invoke and coordinate multiple tools (search, code, calculator, etc.) for complex reasoning tasks. It probes both computational reasoning (math) and knowledge-intensive reasoning (open-domain QA) under a multi-tool collaborative setting. Use when the user wants to benchmark on AIME2024, AIME2025, MATH500, MATH, GSM8K, GAIA, HLE, WebWalker, HotpotQA, 2WikiMultihopQA, Musique, Bamboogle, or asks about evaluating this task. Reports LLM-based judging acc...
Evaluates foundation models' ability to decompose complex instructions, reason over subgoals, and dynamically select/call external APIs or tools to complete tasks across diverse domains like translation, mathematics, web search, and data processing. Use when the user wants to benchmark on MLQA, ASDiv, MathQA, RealTimeQA, HotpotQA, WebShop, ALFWorld, Curated (Map), Curated (Weather), Curated (Stock), Curated (Slides), Curated (Tables), Curated (KGs), Curated (Cooking), Curated (Movie), Curated...
Evaluates zero-shot and cross-dataset generalization of a tongue segmentation model adapted from SAM. It probes the model's ability to segment tongue regions in medical images without task-specific fine-tuning on the target datasets. Use when the user wants to benchmark on TongueSet1, BioHit (TongueSet2), Webset (TongueSet3), or asks about evaluating this task. Reports mIoU.
Evaluates deep learning-based network intrusion detection on IoT traffic, comparing hardware accelerators (Edge TPU vs ARM CPU) for classification accuracy, inference speed, and energy efficiency. Use when the user wants to benchmark on ToN-IoT, or asks about evaluating this task. Reports classification accuracy.
Evaluates large language models' Theory of Mind capabilities by testing their ability to infer mental states (beliefs, intentions, emotions) across multiple orders of reasoning using story-based narratives. The benchmark probes whether models can accurately track character perspectives and answer questions about what different agents know or believe in complex social scenarios. Use when the user wants to benchmark on TOMBENCH, or asks about evaluating this task. Reports accuracy.
Probes a model's ability to perform first- and second-order Theory of Mind reasoning by tracking agents' true and false beliefs about object locations. It specifically tests whether models can distinguish objective reality from subjective mental states while resisting heuristic shortcuts. Use when the user wants to benchmark on ToMi, or asks about evaluating this task. Reports aggregate accuracy.
This benchmark evaluates the robustness of language model tokenizers against real-world input perturbations, including orthographic errors, script variations, homoglyphs, diacritics, and stylistic changes across five languages. It isolates the impact of tokenizer design by testing identical model architectures with different tokenization strategies. Use when the user wants to benchmark on TokSuite, or asks about evaluating this task. Reports relative performance drop.
Evaluates how different tokenizer configurations (pre-tokenizer, fitting corpus, vocabulary size) affect downstream BERT performance on tasks requiring robustness vs. sensitivity to language variation. Use when the user wants to benchmark on AV (Authorship Verification), PAN, CORE, NUCLE, Dialect, GLUE, GLUE+typo, or asks about evaluating this task. Reports accuracy, F1.
Evaluates the privacy leakage of a token-level perturbation mechanism by measuring how easily an adversary can recover original tokens from their privatized embeddings. It probes the robustness of the privacy-preserving noise injection against nearest-neighbor-based inversion attacks. Use when the user has predictions and gold and needs to compute token embedding inversion accuracy.
Evaluates the ability of task-oriented dialogue systems to generate natural language responses while maintaining entity consistency and completing user goals across multiple domains. It probes end-to-end dialogue generation, dialogue state tracking, and response quality under both automated simulation and human evaluation. Use when the user wants to benchmark on DSTC8 Track 1 End-to-End Multi-Domain Dialogue Challenge, MultiWOZ 2.0 benchmark, or asks about evaluating this task. Reports Succes...
Evaluates a model's ability to perform multi-turn task-oriented dialogue by jointly tracking dialogue states and generating task-completing responses. It probes how well the system understands user intents, fills correct slots across multiple domains, and fulfills explicit user requests. Use when the user wants to benchmark on MultiWOZ2.0/2.1, In-Car, or asks about evaluating this task. Reports Comb.
Evaluates long-term vision-language tracking capability by measuring localization accuracy over extended video sequences while dynamically updating natural language descriptions to handle appearance changes and occlusions. Use when the user wants to benchmark on TNLLT, or asks about evaluating this task. Reports PR.
Evaluates natural language-based tracking on 2000 YouTube and surveillance videos, testing the model's ability to follow and adapt to language descriptions over time. Use when the user wants to benchmark on TNL2K, or asks about evaluating this task. Reports AUC.
Evaluates foundation models' multitask language understanding and reasoning capabilities in Traditional Chinese across diverse academic subjects including STEM, social sciences, humanities, and other domains. Use when the user wants to benchmark on TMMLU+, or asks about evaluating this task. Reports average accuracy (%).
Tests an LLM's ability to iteratively design transition metal complexes (TMCs) by maximizing specific properties (polarisability) or expanding multi-objective Pareto frontiers. Use when the user wants to benchmark on Pd(II) square planar complex space, or asks about evaluating this task. Reports Pareto frontier quality.
Evaluates Named Entity Recognition (NER) capabilities on Tagalog news text, specifically measuring performance across Person, Organization, and Location entities using supervised learning and zero-shot LLM prompting. Use when the user wants to benchmark on TLUNIFIED-NER, or asks about evaluating this task. Reports F1-score.
Evaluates large language models' proficiency in Tibetan across general knowledge comprehension and safety-critical domains. It probes the models' ability to handle low-resource language tasks, complex reasoning, and culturally sensitive alignment compared to English baselines. Use when the user wants to benchmark on Ti-MMLU, Ti-SafetyBench, or asks about evaluating this task. Reports Accuracy (ACC).
Evaluates large language models on Bangla language capabilities, specifically probing world knowledge, commonsense reasoning, physical reasoning, and reading comprehension. The benchmark uses multiple-choice and yes/no question formats to measure how well models understand and generate text in a low-resource language context. Use when the user wants to benchmark on Bangla MMLU, CommonsenseQA BN, OpenBookQA BN, PIQA BN, BoolQ BN, or asks about evaluating this task. Reports normalized accuracy.
Evaluates the ability of machine learning models to detect fraudulent financial transactions in real-time using aggregated transaction network features and basic attributes. It probes how well different feature engineering and classification approaches handle severe label imbalance and temporal data splits. Use when the user wants to benchmark on Ant Financial Transaction Dataset, or asks about evaluating this task. Reports F1 Score.