Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

23,477
skills in category
979
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 4,225–4,248 of 23,477 skills

Tau2 Bench EvalA

Evaluates conversational agents' ability to collaborate with a user simulator in a dual-control environment where both parties share tool access to a dynamic world. It probes coordination, communication under decentralized control, and adherence to domain-specific policies while resolving multi-step tasks. Use when the user wants to benchmark on $\tau^2$-Bench, or asks about evaluating this task. Reports pass^1.

researchpythongit
0
3
Tau Voice EvalA

This benchmark evaluates full-duplex voice agents on real-world conversational tasks across retail, airline, and telecom domains. It jointly measures task completion success and real-time interaction quality, including responsiveness, latency, interruption handling, and selectivity under varying acoustic conditions like noise, diverse accents, and natural turn-taking. Use when the user wants to benchmark on $\tau^2$-bench, or asks about evaluating this task. Reports pass@1.

researchpythongo
0
3
Tau Bench EvalA

Evaluates the task completion success rate of LLM agents performing long-horizon, tool-using agentic workflows across customer service and software engineering domains. Probes the agent's ability to execute environment-altering (mutating) actions safely and maintain goal alignment over extended trajectories. Use when the user wants to benchmark on $ au$-Bench Airline, $ au$-Bench Retail, $ au$-Bench-V Air, $ au$-Bench-V Ret, SWE-Bench Verified, or asks about evaluating this task. Reports score.

researchpythongo
0
3
Tatoeba Similarity Search EvalA

Evaluates cross-lingual sentence retrieval accuracy for low-resource languages by finding the most similar sentence in a target language corpus for each source sentence. It probes the alignment quality of vector spaces for languages with limited parallel data. Use when the user wants to benchmark on Tatoeba test setup (LASER), or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Tat Llm EvalA

Evaluates discrete reasoning capabilities over hybrid tabular and textual data, specifically focusing on financial question answering tasks that require arithmetic operations, counting, and span extraction. Use when the user wants to benchmark on FinQA, TAT-QA, TAT-DQA, or asks about evaluating this task. Reports EM.

researchpythongo
0
3
Taskbench Dataset Quality EvalA

Evaluates the quality of synthetically generated task automation instructions and tool invocation graphs. It probes whether the generated data is natural, appropriately complex, and correctly aligned with the underlying tool dependencies. Use when the user wants to benchmark on Hugging Face Tools, Multimedia Tools, Daily Life APIs, or asks about evaluating this task. Reports Alignment.

researchpythonapi
0
3
Task01 Braintumour EvalA

Evaluates a model's ability to perform 3D semantic segmentation on multi-parametric MRI scans of brain tumours. It probes the algorithm's capacity to distinguish and delineate multiple tumour sub-regions (edema, enhancing, non-enhancing) under high clinical variability in scanner hardware and acquisition protocols. Use when the user wants to benchmark on Task01_BrainTumour, or asks about evaluating this task. Reports semantic segmentation accuracy.

researchpythongo
0
3
Task Oriented Dialogue EvalA

Evaluates the ability of neural dialogue systems to track user intent and belief states, generate contextually appropriate responses, and successfully complete task-oriented conversations (specifically restaurant search) by jointly modeling intent, belief, and database interaction. Use when the user wants to benchmark on Wizard-of-Oz restaurant dialogue corpus, or asks about evaluating this task. Reports Objective task success rate.

researchpythongo
0
3
Task Me Anything EvalA

Evaluates the visual perceptual capabilities of large multimodal language models across object recognition, attribute recognition, spatial and temporal reasoning, and action recognition using programmatically generated image and video question-answering tasks. Use when the user wants to benchmark on Task-Me-Anything, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Task Graph Scheduling EvalA

Evaluates task graph scheduling algorithms by measuring their makespan on standard and adversarially modified network topologies. It probes how well algorithms minimize execution time across heterogeneous datasets and reveals performance reversals under minor structural changes. Use when the user wants to benchmark on Parallel Chains (and 15 other SAGA framework datasets), or asks about evaluating this task. Reports makespan.

researchpythongo
0
3
Task Bench EvalA

This benchmark evaluates the parallel runtime performance and scalability of different programming systems by executing parameterized task graphs. It probes how efficiently systems handle varying degrees of parallelism, communication patterns, and computational versus memory-bound workloads. Use when the user wants to benchmark on Task Bench, or asks about evaluating this task. Reports minimum effective task granularity (METG).

researchpythonnode
0
3
Tart EvalA

This protocol evaluates a model's ability to maintain high accuracy on clean, unperturbed data while resisting adversarial attacks. It specifically probes the trade-off between clean accuracy and robustness under l_infinity perturbation constraints using both a custom simulated manifold dataset and the standard CIFAR-10 benchmark. Use when the user wants to benchmark on Transformed Hemisphere, CIFAR-10, or asks about evaluating this task. Reports Clean test accuracy.

researchpythongo
0
3
Target Controllability EvalA

Evaluates an algorithm's ability to identify minimal control input sets for steering complex networks to target states. It probes scalability, solution optimality, and the capacity to prioritize biologically relevant nodes (e.g., drug targets) across synthetic and biological interaction networks. Use when the user wants to benchmark on Breast DEF, Breast HCC1428, Ovarian DEF, Pancreatic AsPC-1, Social Interaction 1, Erdos-Renyi 1000, Scale Free 1000, Small World 1000, or asks about evaluating...

researchpythongo
0
3
Target Bench EvalA

Evaluates whether world models can perform mapless path planning toward semantic targets in real-world environments. It probes spatio-temporal consistency, trajectory accuracy, and the ability to reason about explicit versus implicit goals without prior map information. Use when the user wants to benchmark on Target-Bench, or asks about evaluating this task. Reports WO.

researchpythongo
0
3
Target Aware Molecular Generation EvalA

Evaluates the ability of conditional generative models to produce chemically valid and structurally similar molecular graphs conditioned on specific target protein sequences. It probes both the generative quality (validity, uniqueness, novelty) and the chemical fidelity (similarity to known drugs) of the generated molecules. Use when the user wants to benchmark on Combined Drug-Target Dataset (BIOSNAP, BindingDB, DAVIS, DrugBank), or asks about evaluating this task. Reports Tanimoto similarity.

researchpython
0
3
Tardis Stride EvalA

Evaluates a generative world model's ability to produce controllable road images, predict geographic coordinates from street-view imagery, and generate self-consistent navigation actions on held-out spatiotemporal data. Use when the user wants to benchmark on STRIDE, or asks about evaluating this task. Reports georeferencing_error_m.

researchpythonnode
0
3
Tardis 2.0 EvalA

Evaluates the performance, network traffic, and hardware overhead of the Tardis 2.0 cache coherence protocol under Total Store Order (TSO) consistency. It measures speedup and traffic reduction against a full-map MSI directory baseline and a baseline Tardis implementation across 20 multicore system benchmarks. Use when the user wants to benchmark on Splash2, PARSEC, HPCG, YCSB, TPCC, or asks about evaluating this task. Reports speedup.

researchpythondatabase
0
3
Tapvid 3d EvalA

Evaluates a model's ability to track arbitrary points in 3D space over time from monocular video, assessing spatio-temporal consistency, depth accuracy, and occlusion handling. It probes whether models can reconstruct scene geometry up to scale or only maintain relative depth consistency for local interactions. Use when the user wants to benchmark on TAPVid-3D, or asks about evaluating this task. Reports 3D Average Jaccard (3D-AJ).

researchpythongo
0
3
Tape EvalA

Evaluates few-shot and zero-shot Russian language understanding across six tasks probing logical reasoning, multi-hop inference, commonsense knowledge, and ethical judgment. It also measures model robustness against linguistic adversarial perturbations like typos, deletions, and modality changes. Use when the user wants to benchmark on TAPE, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Tanq EvalA

Evaluates open-domain, multi-hop question answering by requiring models to aggregate data from multiple sources and generate structured answer tables. It probes capabilities in multi-hop reasoning, data normalization, unit conversion, and table construction. Use when the user wants to benchmark on TANQ, or asks about evaluating this task. Reports F1.

researchpythongo
0
3
Tanks Temples Truck EvalA

Evaluates novel view synthesis quality and 3D geometric reconstruction accuracy of NeRF variants on a single outdoor scene. Use when the user wants to benchmark on Tanks & Temples (Truck scene), or asks about evaluating this task. Reports CD.

researchpythonangular
0
3
Tanda Augmentation EvalA

Evaluates whether automatically composing domain-specific data augmentation transformations improves end-task classification performance. It probes the ability of a learned sequence model to generate effective augmentation pipelines compared to heuristic or random baselines across image and text domains. Use when the user wants to benchmark on MNIST, CIFAR-10, ACE (Employment relation extraction), DDSM (Mammography), or asks about evaluating this task. Reports test set accuracy.

researchpythongo
0
3
Tanda As2 EvalA

Evaluates a model's ability to rank candidate answer sentences for a given question. It specifically probes the stability and robustness of pre-trained transformer models when transferred to a target domain using a two-stage fine-tuning protocol. Use when the user wants to benchmark on WikiQA, TREC-QA, or asks about evaluating this task. Reports MAP.

researchpythongo
0
3
Tamper Resistance EvalA

This benchmark evaluates the robustness of LLM safety safeguards against fine-tuning-based tampering attacks. It measures whether a model can maintain low accuracy on weaponized knowledge (forget) and high benign capabilities (retain) after undergoing various supervised fine-tuning attacks. Use when the user wants to benchmark on WMDP, MMLU, HarmBench, MT-Bench, or asks about evaluating this task. Reports Post-Attack Forget accuracy, Attack Success Rate (ASR).

researchpythongo
0
3