Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 4,225–4,248 of 23,477 skills
Evaluates conversational agents' ability to collaborate with a user simulator in a dual-control environment where both parties share tool access to a dynamic world. It probes coordination, communication under decentralized control, and adherence to domain-specific policies while resolving multi-step tasks. Use when the user wants to benchmark on $\tau^2$-Bench, or asks about evaluating this task. Reports pass^1.
This benchmark evaluates full-duplex voice agents on real-world conversational tasks across retail, airline, and telecom domains. It jointly measures task completion success and real-time interaction quality, including responsiveness, latency, interruption handling, and selectivity under varying acoustic conditions like noise, diverse accents, and natural turn-taking. Use when the user wants to benchmark on $\tau^2$-bench, or asks about evaluating this task. Reports pass@1.
Evaluates the task completion success rate of LLM agents performing long-horizon, tool-using agentic workflows across customer service and software engineering domains. Probes the agent's ability to execute environment-altering (mutating) actions safely and maintain goal alignment over extended trajectories. Use when the user wants to benchmark on $ au$-Bench Airline, $ au$-Bench Retail, $ au$-Bench-V Air, $ au$-Bench-V Ret, SWE-Bench Verified, or asks about evaluating this task. Reports score.
Evaluates cross-lingual sentence retrieval accuracy for low-resource languages by finding the most similar sentence in a target language corpus for each source sentence. It probes the alignment quality of vector spaces for languages with limited parallel data. Use when the user wants to benchmark on Tatoeba test setup (LASER), or asks about evaluating this task. Reports Accuracy.
Evaluates discrete reasoning capabilities over hybrid tabular and textual data, specifically focusing on financial question answering tasks that require arithmetic operations, counting, and span extraction. Use when the user wants to benchmark on FinQA, TAT-QA, TAT-DQA, or asks about evaluating this task. Reports EM.
Evaluates the quality of synthetically generated task automation instructions and tool invocation graphs. It probes whether the generated data is natural, appropriately complex, and correctly aligned with the underlying tool dependencies. Use when the user wants to benchmark on Hugging Face Tools, Multimedia Tools, Daily Life APIs, or asks about evaluating this task. Reports Alignment.
Evaluates a model's ability to perform 3D semantic segmentation on multi-parametric MRI scans of brain tumours. It probes the algorithm's capacity to distinguish and delineate multiple tumour sub-regions (edema, enhancing, non-enhancing) under high clinical variability in scanner hardware and acquisition protocols. Use when the user wants to benchmark on Task01_BrainTumour, or asks about evaluating this task. Reports semantic segmentation accuracy.
Evaluates the ability of neural dialogue systems to track user intent and belief states, generate contextually appropriate responses, and successfully complete task-oriented conversations (specifically restaurant search) by jointly modeling intent, belief, and database interaction. Use when the user wants to benchmark on Wizard-of-Oz restaurant dialogue corpus, or asks about evaluating this task. Reports Objective task success rate.
Evaluates the visual perceptual capabilities of large multimodal language models across object recognition, attribute recognition, spatial and temporal reasoning, and action recognition using programmatically generated image and video question-answering tasks. Use when the user wants to benchmark on Task-Me-Anything, or asks about evaluating this task. Reports accuracy.
Evaluates task graph scheduling algorithms by measuring their makespan on standard and adversarially modified network topologies. It probes how well algorithms minimize execution time across heterogeneous datasets and reveals performance reversals under minor structural changes. Use when the user wants to benchmark on Parallel Chains (and 15 other SAGA framework datasets), or asks about evaluating this task. Reports makespan.
This benchmark evaluates the parallel runtime performance and scalability of different programming systems by executing parameterized task graphs. It probes how efficiently systems handle varying degrees of parallelism, communication patterns, and computational versus memory-bound workloads. Use when the user wants to benchmark on Task Bench, or asks about evaluating this task. Reports minimum effective task granularity (METG).
This protocol evaluates a model's ability to maintain high accuracy on clean, unperturbed data while resisting adversarial attacks. It specifically probes the trade-off between clean accuracy and robustness under l_infinity perturbation constraints using both a custom simulated manifold dataset and the standard CIFAR-10 benchmark. Use when the user wants to benchmark on Transformed Hemisphere, CIFAR-10, or asks about evaluating this task. Reports Clean test accuracy.
Evaluates an algorithm's ability to identify minimal control input sets for steering complex networks to target states. It probes scalability, solution optimality, and the capacity to prioritize biologically relevant nodes (e.g., drug targets) across synthetic and biological interaction networks. Use when the user wants to benchmark on Breast DEF, Breast HCC1428, Ovarian DEF, Pancreatic AsPC-1, Social Interaction 1, Erdos-Renyi 1000, Scale Free 1000, Small World 1000, or asks about evaluating...
Evaluates whether world models can perform mapless path planning toward semantic targets in real-world environments. It probes spatio-temporal consistency, trajectory accuracy, and the ability to reason about explicit versus implicit goals without prior map information. Use when the user wants to benchmark on Target-Bench, or asks about evaluating this task. Reports WO.
Evaluates the ability of conditional generative models to produce chemically valid and structurally similar molecular graphs conditioned on specific target protein sequences. It probes both the generative quality (validity, uniqueness, novelty) and the chemical fidelity (similarity to known drugs) of the generated molecules. Use when the user wants to benchmark on Combined Drug-Target Dataset (BIOSNAP, BindingDB, DAVIS, DrugBank), or asks about evaluating this task. Reports Tanimoto similarity.
Evaluates a generative world model's ability to produce controllable road images, predict geographic coordinates from street-view imagery, and generate self-consistent navigation actions on held-out spatiotemporal data. Use when the user wants to benchmark on STRIDE, or asks about evaluating this task. Reports georeferencing_error_m.
Evaluates the performance, network traffic, and hardware overhead of the Tardis 2.0 cache coherence protocol under Total Store Order (TSO) consistency. It measures speedup and traffic reduction against a full-map MSI directory baseline and a baseline Tardis implementation across 20 multicore system benchmarks. Use when the user wants to benchmark on Splash2, PARSEC, HPCG, YCSB, TPCC, or asks about evaluating this task. Reports speedup.
Evaluates a model's ability to track arbitrary points in 3D space over time from monocular video, assessing spatio-temporal consistency, depth accuracy, and occlusion handling. It probes whether models can reconstruct scene geometry up to scale or only maintain relative depth consistency for local interactions. Use when the user wants to benchmark on TAPVid-3D, or asks about evaluating this task. Reports 3D Average Jaccard (3D-AJ).
Evaluates few-shot and zero-shot Russian language understanding across six tasks probing logical reasoning, multi-hop inference, commonsense knowledge, and ethical judgment. It also measures model robustness against linguistic adversarial perturbations like typos, deletions, and modality changes. Use when the user wants to benchmark on TAPE, or asks about evaluating this task. Reports accuracy.
Evaluates open-domain, multi-hop question answering by requiring models to aggregate data from multiple sources and generate structured answer tables. It probes capabilities in multi-hop reasoning, data normalization, unit conversion, and table construction. Use when the user wants to benchmark on TANQ, or asks about evaluating this task. Reports F1.
Evaluates novel view synthesis quality and 3D geometric reconstruction accuracy of NeRF variants on a single outdoor scene. Use when the user wants to benchmark on Tanks & Temples (Truck scene), or asks about evaluating this task. Reports CD.
Evaluates whether automatically composing domain-specific data augmentation transformations improves end-task classification performance. It probes the ability of a learned sequence model to generate effective augmentation pipelines compared to heuristic or random baselines across image and text domains. Use when the user wants to benchmark on MNIST, CIFAR-10, ACE (Employment relation extraction), DDSM (Mammography), or asks about evaluating this task. Reports test set accuracy.
Evaluates a model's ability to rank candidate answer sentences for a given question. It specifically probes the stability and robustness of pre-trained transformer models when transferred to a target domain using a two-stage fine-tuning protocol. Use when the user wants to benchmark on WikiQA, TREC-QA, or asks about evaluating this task. Reports MAP.
This benchmark evaluates the robustness of LLM safety safeguards against fine-tuning-based tampering attacks. It measures whether a model can maintain low accuracy on weaponized knowledge (forget) and high benign capabilities (retain) after undergoing various supervised fine-tuning attacks. Use when the user wants to benchmark on WMDP, MMLU, HarmBench, MT-Bench, or asks about evaluating this task. Reports Post-Attack Forget accuracy, Attack Success Rate (ASR).