Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

22,871
skills in category
953
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 3,5533,576 of 22,871 skills

Threat Intelligence EvalA

This benchmark evaluates an AI system's ability to extract actionable insights from threat intelligence reports and perform security reasoning. It probes multi-document comprehension, attack chain reconstruction, and MITRE ATT&CK framework mapping capabilities. Use when the user wants to benchmark on CyberSOCEval Threat Intelligence Reasoning, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Thiomi Baseline EvalA

Evaluates the quality and utility of a multimodal corpus for low-resource African languages. It does so by training and testing baseline models for automatic speech recognition, machine translation, and text-to-speech across multiple languages. Use when the user wants to benchmark on Thiomi Dataset, or asks about evaluating this task. Reports WER.

researchpythontesting
0
3
Thinkswitcher Math EvalA

Evaluates a model's ability to dynamically switch between short and long chain-of-thought reasoning modes based on task complexity, balancing mathematical problem-solving accuracy against computational efficiency. It probes whether a single reasoning model can adaptively select concise or elaborate reasoning paths without architectural changes or post-training. Use when the user wants to benchmark on GSM8K, MATH-500, AIME24, AIME25, LiveAoPS, Omni-MATH-500, OlympiadBench, or asks about evalua...

researchpythongo
0
3
Thinkjepa Ego Dex EvalA

Evaluates a model's ability to forecast future 3D hand/joint trajectories and latent video representations from egocentric video inputs. It probes long-horizon temporal consistency and physical plausibility in dexterous manipulation scenarios. Use when the user wants to benchmark on EgoDex, EgoExo4D, or asks about evaluating this task. Reports ADE.

researchpythongo
0
3
Theory Of Mind Qa EvalA

Evaluates a model's ability to track first-order and second-order false beliefs, distinguishing an agent's mental state from physical reality and memory. It probes whether systems can maintain consistent world-state representations when agents hold incorrect beliefs about object locations or events. Use when the user wants to benchmark on Sally-Anne & Icecream Van Tasks, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Theoremqa EvalA

Evaluates LLMs' ability to apply domain-specific theorems from mathematics, physics, computer science, and finance to solve complex scientific problems. It probes theorem-driven reasoning, numerical computation, and program generation capabilities. Use when the user wants to benchmark on TheoremQA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Themisio Io Sharing EvalA

Evaluates a policy-driven I/O sharing framework for burst buffers by measuring how effectively it allocates bandwidth, maintains fairness, and reduces interference across concurrent workloads. It probes the system's ability to enforce primitive and composite sharing policies under varying load conditions and compares performance against baseline schedulers. Use when the user wants to benchmark on ThemisIO Benchmark & Application Suite, or asks about evaluating this task. Reports sustained I/O...

researchpythonnode
0
3
Themis Coderewardbench EvalA

Evaluates code reward models on their ability to rank code pairs across five quality dimensions (correctness, efficiency, security, readability, maintainability) and eight programming languages. It probes whether models can generalize beyond functional correctness to assess non-functional code attributes and cross-lingual code preferences. Use when the user wants to benchmark on Themis-CodeRewardBench, or asks about evaluating this task. Reports preference accuracy.

researchpythongo
0
3
The Colosseum EvalA

Evaluates the generalization and robustness of robotic behavior cloning models under various environmental perturbations. It probes how well models trained on clean demonstrations can complete manipulation tasks when faced with changes in lighting, color, distractors, camera pose, and object properties. Use when the user wants to benchmark on The Colosseum, or asks about evaluating this task. Reports task-averaged success rate.

researchpythonperformance
0
3
Thaiocrbench EvalA

Evaluates vision-language models on Thai-language text-rich visual tasks, including document parsing, table/chart recognition, handwritten content extraction, and visual question answering. It probes fine-grained text recognition, structural layout understanding, and semantic reasoning in a low-resource, script-complex language setting. Use when the user wants to benchmark on ThaiOCRBench, or asks about evaluating this task. Reports BMFL.

researchpythongo
0
3
Thai Ser EvalA

Evaluates speech emotion recognition models on a culturally grounded Thai speech corpus, testing their ability to classify utterances into five emotion categories (neutral, angry, happy, sad, frustrated) across different recording environments and cross-corpus settings. Use when the user wants to benchmark on THAI-SER, or asks about evaluating this task. Reports weighted accuracy.

researchpythonrust
0
3
Tgbsseq EvalA

Evaluates temporal graph neural networks on future link prediction tasks, specifically probing their ability to generalize to unseen edges and capture complex sequential dynamics rather than memorizing repeated interactions. Use when the user wants to benchmark on ML-20M, Taobao, Yelp, GoogleLocal, Wikipedia, Reddit, Flickr, YouTube, Patent, WikiLink, or asks about evaluating this task. Reports MRR.

researchpythongo
0
3
Tg Redial EvalA

Evaluates a conversational recommender system's ability to naturally transition topics, recommend relevant items, and generate coherent responses within a dialogue. It probes the model's capacity to leverage historical interactions, user profiles, and topic sequences to maintain semantic flow and recommendation accuracy. Use when the user wants to benchmark on TG-ReDial, or asks about evaluating this task. Reports NDCG@k.

researchpythongo
0
3
Tfrb EvalA

Evaluates the causal reasoning and forecasting accuracy of LLMs and time-series models on multi-domain time-series data. It probes whether step-by-step reasoning and external event context improve numerical predictions or introduce narrative bias, particularly in stochastic versus pattern-rich domains. Use when the user wants to benchmark on TFRBench, or asks about evaluating this task. Reports MASE.

researchpythongo
0
3
Tfq Bench EvalA

Evaluates a model's ability to understand visual metaphors and image implications by verifying multiple factual and inferential propositions per image. It probes fine-grained visual perception, multi-hop reasoning, and theory of mind through structured true-false questioning. Use when the user wants to benchmark on TFQ-Bench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Tfbs Classification EvalA

Evaluates a model's ability to classify short DNA sequences as transcription factor binding sites or not, capturing its capacity to learn regulatory sequence patterns from genomic data. Use when the user wants to benchmark on TFBS classification, or asks about evaluating this task. Reports AUC.

researchpythonexpress
0
3
Textvr Retrieval EvalA

Evaluates cross-modal video retrieval models that must jointly process visual context and scene text (OCR tokens) to match sentence queries with relevant videos. Probes the model's ability to read, comprehend, and align fine-grained text semantics with visual frames in real-world scenarios. Use when the user wants to benchmark on TextVR, or asks about evaluating this task. Reports R@K (Recall@K).

researchpythongo
0
3
Texture Sam EvalA

Evaluates a model's ability to perform texture-aware segmentation by measuring how well it segments regions based on repeating texture patterns rather than semantic shape cues. It tests generalization on both synthetic texture-only images and natural images, while also checking for catastrophic forgetting on standard semantic benchmarks. Use when the user wants to benchmark on RWTD, STMD, ADE20K, or asks about evaluating this task. Reports mIoU.

researchpythonperformance
0
3
Texttabbench EvalA

Evaluates foundation models on tabular prediction tasks that require leveraging mixed categorical, numerical, and free-text features across diverse real-world domains. The benchmark tests whether models can maintain predictive performance when text features contain semantic ambiguity, synonym variation, or noise, while preserving structural tabular signals. Use when the user wants to benchmark on fraud, kick, osha, cards, complaints, spotify, airbnb, beer, houses, laptops, mercari, permits, w...

researchpythongo
0
3
Textatlas EvalA

Evaluates text-to-image generation models on their ability to render long, dense, and structurally complex text accurately within images. It probes both semantic alignment between the prompt and the generated image, and precise character/word-level OCR fidelity across diverse layouts, styles, and real-world scenes. Use when the user wants to benchmark on TextAtlasEval, or asks about evaluating this task. Reports OCR Accuracy (Acc.).

researchpython
0
3
Text Video Alignment EvalA

This evaluation protocol assesses how well text-to-video generation models align generated content with textual prompts across fine-grained attributes like object counts, colors, actions, and spatial relationships. It measures both semantic alignment and visual/motion quality to determine if refinement techniques successfully correct misalignments without degrading fidelity. Use when the user wants to benchmark on EvalCrafter, T2V-CompBench, or asks about evaluating this task. Reports Text-Vi...

researchpythongo
0
3
Text To Sql EvalA

Evaluates a model's ability to translate natural language questions into correct SQL queries for a complex, real-world industrial database. It also probes schema-linking precision by measuring how accurately the model identifies the required tables from the schema. Use when the user wants to benchmark on Industrial Energy Database Benchmark, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Text To Sql Annotation Error EvalA

Evaluates the reliability of text-to-SQL benchmarks by quantifying annotation error rates and measuring how these errors distort agent execution accuracy and leaderboard rankings. Use when the user wants to benchmark on BIRD, Spider 2.0-Snow, or asks about evaluating this task. Reports annotation error rate.

researchpythongo
0
3
Text Sanitization Reconstruction EvalA

Evaluates the vulnerability of differential privacy-based text sanitization methods by measuring how accurately an attacker can reconstruct original sensitive or personally identifiable information (PII) tokens from their sanitized counterparts. It probes the effectiveness of Bayesian inference-based reconstruction attacks against state-of-the-art sanitization defenses. Use when the user wants to benchmark on SST-2, AGNEWS, QNLI, Yelp, or asks about evaluating this task. Reports ASR.

researchpythongo
0
3