Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

23,477
skills in category
979
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 4,009–4,032 of 23,477 skills

Tsg Bench EvalA

Evaluates large language models' ability to understand and generate structured scene graphs from textual narratives. It probes spatial reasoning, action decomposition, and the capacity to map dynamic descriptions to discrete visual or structural elements. Use when the user wants to benchmark on TSG Bench, or asks about evaluating this task. Reports Exact Match (EM) / Accuracy, Precision, Recall, Macro F1.

researchpythongo
0
3
Tsfm Scaling EvalA

Evaluates how time series foundation models scale in forecasting accuracy and uncertainty calibration as model size, compute, and training data size increase. It probes both in-distribution generalization and out-of-distribution transfer capabilities across multiple standard time series forecasting benchmarks. Use when the user wants to benchmark on Monash subset, LSF subset, or asks about evaluating this task. Reports NLL.

researchpythonaws
0
3
Tsbow EvalA

Evaluates object detection models on traffic surveillance footage under diverse weather conditions and varying degrees of vehicle occlusion. It probes robustness to environmental degradation, scale variation, and dense urban traffic scenarios. Use when the user wants to benchmark on TSBOW, or asks about evaluating this task. Reports mAP50.

researchpythongit
0
3
Tsaqa EvalA

Evaluates large language models' ability to perform time series analysis and reasoning across six tasks (anomaly detection, classification, characterization, comparison, data transformation, and temporal relationship) using three question formats (true-or-false, multiple-choice, and puzzling). Use when the user wants to benchmark on TSAQA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Ts Insights EvalA

Evaluates the ability of large multimodal models to generate accurate, domain-agnostic natural language descriptions of time series trends. It probes cross-modal alignment between visual time series plots (or extracted features) and textual trend explanations. Use when the user wants to benchmark on TS-Insights, or asks about evaluating this task. Reports final_score.

researchpython
0
3
Ts Forecasting EvalA

Evaluates the forecasting accuracy and computational efficiency of deep learning models on multivariate time series data. It probes how architectural choices, preprocessing steps, and spatial-temporal processing configurations impact performance across varying forecasting horizons. Use when the user wants to benchmark on Weather, Solar-Energy, ECL, Traffic, or asks about evaluating this task. Reports MAE.

researchpythonperformance
0
3
Truthfulqa EvalA

Evaluates the factual accuracy and truthfulness of large language models by measuring their ability to select correct answers over common misconceptions. It probes the model's capacity to resist generating plausible but false statements across diverse categories like health, law, and politics. The benchmark specifically tests whether models can identify and output factually correct responses when presented with multiple candidate answers. Use when the user wants to benchmark on TruthfulQA, or...

researchpythongo
0
3
Truthfulqa Biogen Factuality EvalA

Evaluates an LLM's factual accuracy and hallucination mitigation across multiple-choice, short-form, and long-form generation tasks. It measures the trade-off between truthfulness and informativeness, and quantifies the exact number of supported versus unsupported facts in generated text. Use when the user wants to benchmark on TruthfulQA, BioGEN, or asks about evaluating this task. Reports Accuracy, True*Info, FActScore.

researchpythongo
0
3
Truemicl EvalA

This benchmark evaluates a model's ability to perform true multimodal in-context learning by requiring it to solve tasks that depend on both visual and textual information from provided demonstrations. It probes whether models can correctly attend to and utilize visual context in few-shot examples rather than relying on superficial textual patterns or prior knowledge. Use when the user wants to benchmark on TrueMICL, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Truck Driving Risk EvalA

Binary classification of truck driving risk on specific highway segments based on historical behavior, short-term trip dynamics, and real-time traffic conditions. It probes a model's ability to predict forward collision warning events using a small, highly imbalanced dataset of real-world trajectory data. Use when the user wants to benchmark on Truck Driving Risk Dataset, or asks about evaluating this task. Reports Accuracy.

researchpythontesting
0
3
Triviaqa EvalA

This benchmark evaluates reading comprehension on complex, compositional trivia questions that require multi-sentence reasoning and handling high lexical variability. It tests a model's ability to locate and extract precise answers from large, noisy evidence documents across different domains. Use when the user wants to benchmark on TriviaQA, or asks about evaluating this task. Reports exact match (EM).

researchpythongo
0
3
Triplesumm Video Summarization EvalA

Evaluates a model's ability to perform video summarization by predicting frame-level importance scores across visual, textual, and audio modalities. It probes the model's capacity for adaptive multimodal fusion and temporal dependency modeling to identify salient segments in long videos. Use when the user wants to benchmark on MoSu, Mr. HiSum, SumMe, TVSum, or asks about evaluating this task. Reports Kendall’s τ (kTau), Spearman’s ρ (sRho).

researchpythongo
0
3
Trilemma Of Truth EvalA

Evaluates large language models' ability to distinguish factually true statements from factually false and unverifiable ('neither') statements. It probes both prompt-based output probabilities and internal hidden activations to measure veracity classification accuracy and uncertainty quantification. Use when the user wants to benchmark on Trilemma of Truth Datasets, or asks about evaluating this task. Reports MCC.

researchpythongo
0
3
Trglue EvalA

This benchmark evaluates Turkish natural language understanding across multiple task types, including grammaticality judgment, sentiment analysis, paraphrase detection, and semantic textual similarity. It probes a model's ability to handle agglutinative morphology, culturally adapted expressions, and nuanced semantic equivalence in Turkish. Use when the user wants to benchmark on TrCoLA, TrSST-2, TrMRPC, TrSTS-B, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Trendfact EvalA

Evaluates fact-checking systems across three sub-tasks: retrieving relevant evidence, verifying claim truthfulness, and generating natural language explanations. It specifically probes a model's Hotspot Perception Ability (HPA) by measuring how effectively it allocates reasoning effort and computational resources based on the real-world influence of trending claims. Use when the user wants to benchmark on TrendFact, or asks about evaluating this task. Reports F1-macro.

researchpythongo
0
3
Trek 150 EvalA

Evaluates single-object visual tracking performance in first-person vision videos, specifically testing robustness to object manipulation, occlusions, and dynamic interactions under real-time execution constraints. Use when the user wants to benchmark on TREK-150, or asks about evaluating this task. Reports SS, NPS, GSR.

researchpythontesting
0
3
Treesatai Ts EvalA

Evaluates a model's ability to perform fine-grained tree species identification using multimodal Earth observation data, specifically leveraging temporal dynamics from optical and radar time series alongside high-resolution imagery. Use when the user wants to benchmark on TreeSatAI-TS, or asks about evaluating this task. Reports weighted F1.

researchpythongo
0
3
Treereview Peer Review EvalA

This benchmark evaluates an LLM's ability to perform deep, structured scientific peer review by generating comprehensive reviews and actionable feedback comments. It probes the model's capacity for hierarchical question decomposition, context-aware analysis of long documents, and alignment with human reviewer judgments across multiple quality dimensions. Use when the user wants to benchmark on TreeReview Benchmark (ICLR-2024, NeurIPS-2023, Nature Communications), or asks about evaluating this...

researchpythongo
0
3
Treeeval EvalA

TreeEval probes an LLM's ability to handle complex, adaptive reasoning through dynamically generated hierarchical questions. It evaluates how well a model's relative performance ranking aligns with established leaderboards like AlpacaEval2.0, while testing the framework's efficiency in distinguishing fine-grained capability differences without relying on static datasets. Use when the user wants to benchmark on TreeEval (Dynamic/Benchmark-Free), or asks about evaluating this task. Reports Spea...

researchpythongo
0
3
Treebench EvalA

Evaluates visual grounded reasoning by requiring models to precisely localize target objects in cluttered scenes and perform second-order reasoning about their interactions. It measures both the correctness of the final answer and the spatial accuracy of the traceable bounding box evidence. Use when the user wants to benchmark on TreeBench, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Tree Vs Sequence EvalA

This evaluation protocol compares recursive tree-based neural models against recurrent sequence-based models across multiple NLP tasks. It probes whether syntactic tree structures are necessary for learning representations, particularly for tasks requiring long-distance dependency modeling or hierarchical composition. Use when the user wants to benchmark on Stanford Sentiment Treebank, Pang Sentiment Dataset, UMD-QA, SemEval-2010 Task 8, Discourse Parsing, or asks about evaluating this task. ...

researchpythongo
0
3
Trec2025 Rag EvalA

Probes retrieval-augmented generation systems on complex, narrative-driven queries by decomposing information needs into sub-narratives. It evaluates document relevance based on sub-narrative coverage, measures response quality via strict vital recall of fully supported information nuggets, and assesses sentence-level factual grounding against cited documents. Use when the user wants to benchmark on MS MARCO V2.1, or asks about evaluating this task. Reports strict_vital_recall.

researchpython
0
3
Trec2024 Rag Nugget EvalA

Evaluates the factual accuracy and content grounding of RAG-generated answers by checking for the presence of key factual claims (nuggets) extracted from source documents. It also measures answer length to assess the trade-off between conciseness and completeness in system outputs. Use when the user wants to benchmark on TREC 2024 RAG Track, or asks about evaluating this task. Reports V_strict.

researchpython
0
3
Trec2022 Neuclir EvalA

Evaluates neural cross-language information retrieval systems on ad hoc, reranking, and monolingual tasks across Chinese, Persian, and Russian newswire collections. It measures how well models retrieve and rank relevant documents when queries are in English and documents are in other languages, or when queries are human-translated. Use when the user wants to benchmark on TREC 2022 NeuCLIR Collections, or asks about evaluating this task. Reports nDCG@10.

researchpythongo
0
3