Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

22,954
skills in category
957
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 3,6013,624 of 22,954 skills

Token Embedding Inversion AccuracyA

Evaluates the privacy leakage of a token-level perturbation mechanism by measuring how easily an adversary can recover original tokens from their privatized embeddings. It probes the robustness of the privacy-preserving noise injection against nearest-neighbor-based inversion attacks. Use when the user has predictions and gold and needs to compute token embedding inversion accuracy.

researchpythongo
0
3
Tod Nlg EvalA

Evaluates the ability of task-oriented dialogue systems to generate natural language responses while maintaining entity consistency and completing user goals across multiple domains. It probes end-to-end dialogue generation, dialogue state tracking, and response quality under both automated simulation and human evaluation. Use when the user wants to benchmark on DSTC8 Track 1 End-to-End Multi-Domain Dialogue Challenge, MultiWOZ 2.0 benchmark, or asks about evaluating this task. Reports Succes...

researchpythongo
0
3
Tod EvalA

Evaluates a model's ability to perform multi-turn task-oriented dialogue by jointly tracking dialogue states and generating task-completing responses. It probes how well the system understands user intents, fills correct slots across multiple domains, and fulfills explicit user requests. Use when the user wants to benchmark on MultiWOZ2.0/2.1, In-Car, or asks about evaluating this task. Reports Comb.

researchpythongo
0
3
Tnllt EvalA

Evaluates long-term vision-language tracking capability by measuring localization accuracy over extended video sequences while dynamically updating natural language descriptions to handle appearance changes and occlusions. Use when the user wants to benchmark on TNLLT, or asks about evaluating this task. Reports PR.

researchpython
0
3
Tnl2k EvalA

Evaluates natural language-based tracking on 2000 YouTube and surveillance videos, testing the model's ability to follow and adapt to language descriptions over time. Use when the user wants to benchmark on TNL2K, or asks about evaluating this task. Reports AUC.

researchpythontesting
0
3
Tmmluplus EvalA

Evaluates foundation models' multitask language understanding and reasoning capabilities in Traditional Chinese across diverse academic subjects including STEM, social sciences, humanities, and other domains. Use when the user wants to benchmark on TMMLU+, or asks about evaluating this task. Reports average accuracy (%).

researchpythongo
0
3
Tmc Optimization EvalA

Tests an LLM's ability to iteratively design transition metal complexes (TMCs) by maximizing specific properties (polarisability) or expanding multi-objective Pareto frontiers. Use when the user wants to benchmark on Pd(II) square planar complex space, or asks about evaluating this task. Reports Pareto frontier quality.

researchpythongo
0
3
Tlunified Ner EvalA

Evaluates Named Entity Recognition (NER) capabilities on Tagalog news text, specifically measuring performance across Person, Organization, and Location entities using supervised learning and zero-shot LLM prompting. Use when the user wants to benchmark on TLUNIFIED-NER, or asks about evaluating this task. Reports F1-score.

researchpythongo
0
3
Tlue EvalA

Evaluates large language models' proficiency in Tibetan across general knowledge comprehension and safety-critical domains. It probes the models' ability to handle low-resource language tasks, complex reasoning, and culturally sensitive alignment compared to English baselines. Use when the user wants to benchmark on Ti-MMLU, Ti-SafetyBench, or asks about evaluating this task. Reports Accuracy (ACC).

researchpythongo
0
3
Titullms Bangla Benchmark EvalA

Evaluates large language models on Bangla language capabilities, specifically probing world knowledge, commonsense reasoning, physical reasoning, and reading comprehension. The benchmark uses multiple-choice and yes/no question formats to measure how well models understand and generate text in a low-resource language context. Use when the user wants to benchmark on Bangla MMLU, CommonsenseQA BN, OpenBookQA BN, PIQA BN, BoolQ BN, or asks about evaluating this task. Reports normalized accuracy.

researchpythongo
0
3
Titant Fraud Detection EvalA

Evaluates the ability of machine learning models to detect fraudulent financial transactions in real-time using aggregated transaction network features and basic attributes. It probes how well different feature engineering and classification approaches handle severe label imbalance and temporal data splits. Use when the user wants to benchmark on Ant Financial Transaction Dataset, or asks about evaluating this task. Reports F1 Score.

researchpythongo
0
3
Titan Sfdaod EvalA

Evaluates source-free domain adaptive object detection (SF-DAOD) and unsupervised domain adaptation (UDA) across natural and medical imaging domains. Probes a model's ability to align features and reduce pseudo-label noise when adapting to a target domain without access to target labels during training. Use when the user wants to benchmark on Cityscapes, Foggy Cityscapes, KITTI, SIM10k, BDD100k, RSNA-BSD1K, INBreast, DDSM, or asks about evaluating this task. Reports mAP.

researchpythonperformance
0
3
Tirauxcloud EvalA

Evaluates semantic segmentation models for day-and-night cloud detection using thermal infrared imagery, specifically testing how auxiliary environmental features improve segmentation accuracy and how well models transfer across different satellite sensors and resolutions. Use when the user wants to benchmark on Landsat Main, or asks about evaluating this task. Reports mIoU.

researchpythontesting
0
3
Tir Bench EvalA

Evaluates multimodal language models' ability to perform agentic reasoning with images, specifically requiring dynamic visual manipulation and tool-use to solve complex tasks like rotation, jigsaw assembly, and instrument reading. It probes whether models can iteratively process, crop, or transform visual inputs to extract information or solve spatial problems. Use when the user wants to benchmark on TIR-Bench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Tinymlperf EvalA

Evaluates the deployment efficiency and accuracy of neural networks on commodity microcontrollers (MCUs) under strict memory and latency constraints. It measures inference speed, memory footprint, and task-specific accuracy across vision, audio, and anomaly detection workloads. Use when the user wants to benchmark on TinyMLPerf (VWW, KWS, AD), or asks about evaluating this task. Reports Accuracy (%), Latency (ms).

researchpythongo
0
3
Tinybenchmarks Sampling EvalA

Evaluates the efficiency and accuracy of LLM benchmarking by testing how well a small, strategically selected subset of examples predicts overall model performance on standard evaluation scenarios. Use when the user wants to benchmark on HELM, MMLU, AlpacaEval 2.0, Open LLM Leaderboard, or asks about evaluating this task. Reports estimation error.

researchpythongo
0
3
Timit Tts EvalA

Evaluates deepfake detectors on distinguishing real from synthetic audio and video, and on attributing synthetic audio to specific TTS generators. It probes robustness to post-processing (DTW alignment, augmentation) and video compression. Use when the user wants to benchmark on TIMIT-TTS, or asks about evaluating this task. Reports AUC.

researchpythongo
0
3
Timid Robot Mistake Detection EvalA

This benchmark evaluates a model's ability to detect time-dependent and physical mistakes in robotic task executions from video. It specifically probes temporal reasoning, semantic task violation detection, and sim-to-real generalization by comparing frame-level anomaly predictions against ground-truth annotations. Use when the user wants to benchmark on BridgeData V2, Multi-robot dataset, or asks about evaluating this task. Reports F1.

researchpythongo
0
3
Timetom EvalA

Evaluates Large Language Models' Theory of Mind (ToM) reasoning capabilities across reading comprehension and interactive dialogue scenarios. It specifically probes the model's ability to track character beliefs over time, assess answerability, and determine information access, with a strong focus on first-order and higher-order (up to third-order) belief reasoning. Use when the user wants to benchmark on ToMI, BigToM, FanToM, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Timeseries Forecasting EvalA

Evaluates the forecasting accuracy of a time series foundation model across diverse real-world and synthetic datasets. It probes the model's ability to capture temporal dynamics, periodicity, and multi-scale patterns over a fixed context window to predict future values. Use when the user wants to benchmark on ETT1, ETT2, Exchange Rate, M1 Monthly, M1 Quarterly, M1 Yearly, M5, Monash M3, NN5, Traffic, Weather, M4 Monthly, Entsoe, Solar with Weather, UK Covid, Sensor Data, or asks about evaluat...

researchpython
0
3
Timerrecipe EvalA

Evaluates the effectiveness of individual architectural modules (e.g., normalization, decomposition, embedding, feedforward types) across diverse time-series forecasting scenarios to identify optimal configurations and predict performance without training. Use when the user wants to benchmark on PEMS03, ETT, Electricity, Social (Unemployment), or asks about evaluating this task. Reports MSE.

researchpythonperformance
0
3
Timer Time Series EvalA

Evaluates a large decoder-only Transformer model on standard time series benchmarks for forecasting, imputation, and anomaly detection. It probes the model's few-shot generalization, scalability, and robustness in data-scarce scenarios compared to encoder-only baselines. Use when the user wants to benchmark on ETT, ECL, Traffic, Weather, PEMS, UCR Anomaly Archive, or asks about evaluating this task. Reports MSE.

researchpython
0
3
Time Series Forecasting EvalA

Evaluates the capability of LLM-based models to forecast multivariate time series across multiple prediction horizons and real-world datasets. It probes pattern-aware temporal modeling and semantic alignment by measuring prediction accuracy under a channel-independent, rolling forecasting setup. Use when the user wants to benchmark on ETTh1, ETTh2, ETTm1, ETTm2, Weather, ECL, Traffic, or asks about evaluating this task. Reports MSE.

researchpython
0
3
Time Series Benchmark EvalA

Evaluates zero-shot forecasting performance of time series foundation models across diverse datasets and horizons. Probes model capability to capture structural temporal patterns (trend, seasonality, stationarity, complexity) and generalizes to unseen data without leakage. Use when the user wants to benchmark on TIME Benchmark, or asks about evaluating this task. Reports MASE, CRPS.

researchpythonperformance
0
3