Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

22,954
skills in category
957
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 3,6253,648 of 22,954 skills

Time Ra EvalA

Evaluates the ability of LLMs and MLLMs to diagnose anomalies in univariate and multivariate time series data. It probes the models' capacity for structured reasoning (generating a 'Thought') and precise action classification ('ActionID') based on raw or visualized temporal data. Use when the user wants to benchmark on RATs40K, or asks about evaluating this task. Reports Label Matching F1.

researchpythongo
0
3
Time Moe Forecasting EvalA

Evaluates long-term time series forecasting capabilities of foundation models in both zero-shot (unseen datasets) and in-distribution (fine-tuned) settings across multiple prediction horizons. Use when the user wants to benchmark on ETTh1, ETTh2, ETTm1, ETTm2, Weather, Global Temp, or asks about evaluating this task. Reports MSE.

researchpython
0
3
Tifa 100 EvalA

Evaluates how well an automatic multimodal transformer predicts fine-grained human feedback (quality scores, region-level misalignment heatmaps, and text misalignment annotations) on text-to-image generation tasks. Use when the user wants to benchmark on TIFA, or asks about evaluating this task. Reports correlation.

researchpython
0
3
Tiered Data Management EvalA

Evaluates the impact of tiered data management (L1–L3) on model performance across general knowledge, reasoning, math, and code domains. It compares models trained on different data quality tiers and different training strategies (mix vs. tiered) to validate data curation and scheduling efficacy. Use when the user wants to benchmark on OpenCompass Benchmarks, or asks about evaluating this task. Reports Average benchmark scores.

researchpythongo
0
3
Tidmad Denoising EvalA

Evaluates the ability of traditional and deep learning denoising algorithms to recover sinusoidal dark matter signals from ultra-long, noisy time series data collected by the ABRACADABRA experiment. Use when the user wants to benchmark on TIDMAD, or asks about evaluating this task. Reports mean square error.

researchpythongo
0
3
Tid 8 EvalA

Probes a model's ability to learn from inherently subjective or disagreed-upon annotations by treating each annotator's label as a separate example, rather than aggregating them into a single ground truth label. Use when the user wants to benchmark on TID-8, or asks about evaluating this task. Reports exact match accuracy.

researchpythongo
0
3
Tictactoe Manipulation EvalA

Evaluates a robot's ability to perform adversarial interaction and temporal reasoning through a structured pick-and-place board game. It tests the integration of vision-based contour segmentation, game-state reasoning via Minimax, and precise motion planning under hardware constraints. Use when the user wants to benchmark on Yale-CMU-Berkeley objects set, or asks about evaluating this task. Reports t_sub.

researchpythongo
0
3
Tib Bench EvalA

Evaluates vision-language models on their ability to generate accurate and linguistically coherent summaries of text-heavy multimodal presentations. It probes cross-modal alignment, long-context understanding, and the impact of different input modalities (raw video, slides, transcripts, interleaved pairs) and token budgets on summarization quality. Use when the user wants to benchmark on TIB-bench, or asks about evaluating this task. Reports $IbR_{overall}$.

researchpythonperformance
0
3
Tian Gong St Click EvalA

Evaluates click models on predicting user click sequences and estimating document relevance from search logs. It also measures how well models recover the underlying distribution of real click data (distributional coverage) and perform when document lists are poorly ranked. Use when the user wants to benchmark on TianGong-ST, or asks about evaluating this task. Reports LL.

researchpythongo
0
3
Thyme Scene Graph EvalA

Evaluates a model's ability to generate dynamic video scene graphs by predicting inter-object relationships and attributes across multiple temporal frames. It specifically probes temporal consistency, handling of occlusions, and modeling of long-range dependencies in both ground-level and aerial video footage. Use when the user wants to benchmark on ASPIRe, AeroEye-v1.0, or asks about evaluating this task. Reports R@20, mR@20.

researchpythongo
0
3
Thyme Multimodal EvalA

Evaluates multimodal large language models on image manipulation, visual perception, mathematical reasoning, and general vision-language tasks. It probes whether autonomous code generation and execution for image processing improves downstream accuracy and reduces hallucination. Use when the user wants to benchmark on MME-RealWorld, HR Bench, MathVista, Hallucination bench, MMStar, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Threat Intelligence EvalA

This benchmark evaluates an AI system's ability to extract actionable insights from threat intelligence reports and perform security reasoning. It probes multi-document comprehension, attack chain reconstruction, and MITRE ATT&CK framework mapping capabilities. Use when the user wants to benchmark on CyberSOCEval Threat Intelligence Reasoning, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Thiomi Baseline EvalA

Evaluates the quality and utility of a multimodal corpus for low-resource African languages. It does so by training and testing baseline models for automatic speech recognition, machine translation, and text-to-speech across multiple languages. Use when the user wants to benchmark on Thiomi Dataset, or asks about evaluating this task. Reports WER.

researchpythontesting
0
3
Thinkswitcher Math EvalA

Evaluates a model's ability to dynamically switch between short and long chain-of-thought reasoning modes based on task complexity, balancing mathematical problem-solving accuracy against computational efficiency. It probes whether a single reasoning model can adaptively select concise or elaborate reasoning paths without architectural changes or post-training. Use when the user wants to benchmark on GSM8K, MATH-500, AIME24, AIME25, LiveAoPS, Omni-MATH-500, OlympiadBench, or asks about evalua...

researchpythongo
0
3
Thinkjepa Ego Dex EvalA

Evaluates a model's ability to forecast future 3D hand/joint trajectories and latent video representations from egocentric video inputs. It probes long-horizon temporal consistency and physical plausibility in dexterous manipulation scenarios. Use when the user wants to benchmark on EgoDex, EgoExo4D, or asks about evaluating this task. Reports ADE.

researchpythongo
0
3
Theory Of Mind Qa EvalA

Evaluates a model's ability to track first-order and second-order false beliefs, distinguishing an agent's mental state from physical reality and memory. It probes whether systems can maintain consistent world-state representations when agents hold incorrect beliefs about object locations or events. Use when the user wants to benchmark on Sally-Anne & Icecream Van Tasks, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Theoremqa EvalA

Evaluates LLMs' ability to apply domain-specific theorems from mathematics, physics, computer science, and finance to solve complex scientific problems. It probes theorem-driven reasoning, numerical computation, and program generation capabilities. Use when the user wants to benchmark on TheoremQA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Themisio Io Sharing EvalA

Evaluates a policy-driven I/O sharing framework for burst buffers by measuring how effectively it allocates bandwidth, maintains fairness, and reduces interference across concurrent workloads. It probes the system's ability to enforce primitive and composite sharing policies under varying load conditions and compares performance against baseline schedulers. Use when the user wants to benchmark on ThemisIO Benchmark & Application Suite, or asks about evaluating this task. Reports sustained I/O...

researchpythonnode
0
3
Themis Coderewardbench EvalA

Evaluates code reward models on their ability to rank code pairs across five quality dimensions (correctness, efficiency, security, readability, maintainability) and eight programming languages. It probes whether models can generalize beyond functional correctness to assess non-functional code attributes and cross-lingual code preferences. Use when the user wants to benchmark on Themis-CodeRewardBench, or asks about evaluating this task. Reports preference accuracy.

researchpythongo
0
3
The Colosseum EvalA

Evaluates the generalization and robustness of robotic behavior cloning models under various environmental perturbations. It probes how well models trained on clean demonstrations can complete manipulation tasks when faced with changes in lighting, color, distractors, camera pose, and object properties. Use when the user wants to benchmark on The Colosseum, or asks about evaluating this task. Reports task-averaged success rate.

researchpythonperformance
0
3
Thaiocrbench EvalA

Evaluates vision-language models on Thai-language text-rich visual tasks, including document parsing, table/chart recognition, handwritten content extraction, and visual question answering. It probes fine-grained text recognition, structural layout understanding, and semantic reasoning in a low-resource, script-complex language setting. Use when the user wants to benchmark on ThaiOCRBench, or asks about evaluating this task. Reports BMFL.

researchpythongo
0
3
Thai Ser EvalA

Evaluates speech emotion recognition models on a culturally grounded Thai speech corpus, testing their ability to classify utterances into five emotion categories (neutral, angry, happy, sad, frustrated) across different recording environments and cross-corpus settings. Use when the user wants to benchmark on THAI-SER, or asks about evaluating this task. Reports weighted accuracy.

researchpythonrust
0
3
Tgbsseq EvalA

Evaluates temporal graph neural networks on future link prediction tasks, specifically probing their ability to generalize to unseen edges and capture complex sequential dynamics rather than memorizing repeated interactions. Use when the user wants to benchmark on ML-20M, Taobao, Yelp, GoogleLocal, Wikipedia, Reddit, Flickr, YouTube, Patent, WikiLink, or asks about evaluating this task. Reports MRR.

researchpythongo
0
3
Tg Redial EvalA

Evaluates a conversational recommender system's ability to naturally transition topics, recommend relevant items, and generate coherent responses within a dialogue. It probes the model's capacity to leverage historical interactions, user profiles, and topic sequences to maintain semantic flow and recommendation accuracy. Use when the user wants to benchmark on TG-ReDial, or asks about evaluating this task. Reports NDCG@k.

researchpythongo
0
3