Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

23,477
skills in category
979
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 4,201–4,224 of 23,477 skills

Tenspiler EvalA

Evaluates the correctness and performance of a verified lifting-based compiler that transpiles sequential C++/Python code into tensor operations across various DSLs and hardware accelerators. It measures synthesis efficiency, kernel execution speedup, and end-to-end performance including data transfer overhead. Use when the user wants to benchmark on TENSPILER benchmark suite, or asks about evaluating this task. Reports kernel_performance.

researchpythongo
0
3
Tenrec EvalA

Evaluates recommender systems across multiple tasks including click-through rate (CTR) prediction, sequential recommendation, and top-N item ranking. It probes cross-domain generalization, cold-start handling, and the sensitivity of ranking metrics to negative sampling strategies. Use when the user wants to benchmark on Tenrec, or asks about evaluating this task. Reports AUC.

researchpythongit
0
3
Tempreason EvalA

Evaluates large language models' ability to perform temporal reasoning across three complexity levels: time-time relations (L1), time-event relations (L2), and event-event relations (L3). It specifically probes models' robustness to historical and futuristic time periods, as well as their capacity for month-level intra-year reasoning. Use when the user wants to benchmark on TEMPREASON, or asks about evaluating this task. Reports EM.

researchpythongo
0
3
Tempqa Wd EvalA

Probes a system's ability to perform temporal question answering over knowledge bases by generating correct answers or SPARQL queries. It specifically evaluates generalization across different knowledge bases (Wikidata vs. Freebase) and interpretability through fine-grained intermediate annotations like entity/relation linking and λ-expressions. Use when the user wants to benchmark on TempQA-WD, or asks about evaluating this task. Reports F1.

researchpythongo
0
3
Temporalbench EvalA

TemporalBench probes LLM-based agents' ability to perform contextual and event-informed temporal reasoning across four distinct task families. It disentangles historical pattern interpretation, context-free forecasting, contextual alignment, and event-conditioned adaptation to reveal whether numerical prediction accuracy correlates with qualitative temporal judgment. Use when the user wants to benchmark on FreshRetailNet, PSML, Causal Chambers, MIMIC, or asks about evaluating this task. Repor...

researchpythongo
0
3
Temporal Graph Anomaly EvalA

Evaluates the ability of various data-driven models to detect emerging anomalies in temporal graphs derived from social media interactions. It probes how well different architectures generalize across different social platforms and remain robust to parameter variations and temporal/spatial shifts. Use when the user wants to benchmark on Twitter, Facebook, or asks about evaluating this task. Reports weighted F1 score.

researchpythonperformance
0
3
Temporal Domain Generalization EvalA

Evaluates a model's ability to generalize to future, unseen temporal domains without full retraining. It measures out-of-distribution accuracy on sequentially arriving target domains after training on historical source domains. Use when the user wants to benchmark on Yearbook, Rotated MNIST (RMNIST), FMoW, Huffpost, Arxiv, CLEAR-10/100, or asks about evaluating this task. Reports OOD_avg accuracy.

researchpythongo
0
3
Temporal Degradation EvalA

Evaluates how temporal misalignment between pretraining/fine-tuning data and evaluation data impacts model performance across classification and summarization benchmarks. Use when the user wants to benchmark on PubCLS, NewSum, TwiERC, AIC, PoliAff, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Tempobench EvalA

Evaluates large language models' temporal reasoning capabilities by decomposing performance into trace-based (TTE) and causal (TCE) components. It measures how well models handle structured logical specifications with varying complexity, isolating structural factors like horizon depth and information density. Use when the user wants to benchmark on TempoBench, or asks about evaluating this task. Reports exact-match accuracy.

researchpythongo
0
3
Tempo Sum EvalA

Evaluates text summarization models' temporal generalization by testing on datasets split by publication date, specifically probing how well models handle knowledge-conflicting future articles versus in-distribution past data. Use when the user wants to benchmark on BBC, CNN, or asks about evaluating this task. Reports FactCC.

researchpythontesting
0
3
Temmed Bench EvalA

Evaluates large vision-language models' ability to perform temporal reasoning on medical images by analyzing condition changes across multiple clinical visits. It probes capabilities in visual question answering, longitudinal clinical report generation, and selecting relevant image pairs based on temporal context. Use when the user wants to benchmark on TemMed-Bench, or asks about evaluating this task. Reports Avg..

researchpythongo
0
3
Tembed EvalA

Evaluates the quality and efficiency of tabular embedding models across four granularity levels (cell, row, column, table) and six downstream tasks including similarity search, triplet evaluation, prediction, and retrieval. It probes whether a single embedding approach can generalize universally across diverse structured data applications or if performance is highly task- and granularity-dependent. Use when the user wants to benchmark on TEmBed Benchmark Suite, or asks about evaluating this t...

researchpythongit
0
3
Teleqna EvalA

Evaluates large language models' domain-specific knowledge in telecommunications, covering general terminology, research concepts, and complex technical standards. It also benchmarks model performance against human telecom professionals under strict no-search conditions. Use when the user wants to benchmark on TeleQnA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Tedigan Text To Face EvalA

Evaluates a model's capability to synthesize high-resolution, diverse, and photorealistic face images conditioned on natural language prompts, and to perform text-guided editing of existing faces while preserving identity and irrelevant attributes. Use when the user wants to benchmark on Multi-Modal CelebA-HQ, or asks about evaluating this task. Reports FID.

researchpythongit
0
3
Tdc Admet EvalA

Evaluates molecular property prediction across 22 ADMET tasks. It probes the model's ability to generalize across diverse chemical properties using standardized benchmark splits for both regression and classification. Use when the user wants to benchmark on TDC ADMET Group, or asks about evaluating this task. Reports Classification AUROC.

researchpythonperformance
0
3
Tdc Adme Pk EvalA

Assesses pharmacological property prediction across ADME, PK, and toxicity tasks, including regression, classification, and correlation-based evaluation. Use when the user wants to benchmark on TDC Benchmark, or asks about evaluating this task. Reports AUROC / AUPRC.

researchpythongo
0
3
Tdbench EvalA

Evaluates vision-language models on top-down (aerial) image understanding by testing their ability to answer questions about rotated views. It measures rotational consistency to filter out hallucinations and decomposes performance into true knowledge versus lucky guessing via a probabilistic reliability framework. Use when the user wants to benchmark on TDBench, or asks about evaluating this task. Reports RotationalEval (RE).

researchpythongo
0
3
Tcm Best4sdt EvalA

This benchmark evaluates large language models' capabilities in Traditional Chinese Medicine (TCM) clinical reasoning, specifically focusing on syndrome differentiation and treatment decision-making. It probes the model's ability to accurately diagnose pathological patterns, formulate appropriate herbal prescriptions, and adhere to medical ethics and safety guidelines across 27 dimensions. Use when the user wants to benchmark on TCM-BEST4SDT, or asks about evaluating this task. Reports select...

researchpythongo
0
3
Tcga Histopathology EvalA

Evaluates the representation quality and generalization of self-supervised histopathology models across diverse patch-level diagnostic tasks and weakly supervised slide-level tasks using linear probing and fine-tuning on TCGA whole slide images. Use when the user wants to benchmark on TCGA Histopathology, or asks about evaluating this task. Reports average AUC.

researchpythongit
0
3
Tcab EvalA

Evaluates a model's ability to detect whether a given text instance has been adversarially perturbed (attack detection) and to identify the specific attack method used (attack labeling) across multiple text classification domains. Use when the user wants to benchmark on TCAB, or asks about evaluating this task. Reports balanced accuracy.

researchpythongo
0
3
Tbar EvalA

Evaluates the effectiveness of template-based automated program repair systems by applying fix patterns to buggy Java programs. It probes the system's ability to localize faults, generate syntactically valid patches, and pass test suites without breaking existing tests. Use when the user wants to benchmark on Defects4J, or asks about evaluating this task. Reports plausible_patch.

researchpythongo
0
3
Tb Classification EvalA

Evaluates deep learning models' ability to classify chest X-rays as tuberculosis or normal, comparing whole-image vs. lung-segmented inputs. Use when the user wants to benchmark on Kaggle CXR images and lung mask dataset, or asks about evaluating this task. Reports accuracy.

researchpythontesting
0
3
Tb Bench EvalA

This benchmark evaluates multi-modal large language models' ability to understand spatio-temporal traffic behaviors from ego-centric dashcam images and videos. It probes eight distinct perception tasks, including road detection, object-lane alignment, turning prediction, and ego-trajectory estimation, requiring both spatial reasoning and temporal tracking. Use when the user wants to benchmark on TB-Bench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Taxpraben EvalA

Evaluates LLMs on Chinese real-world tax practice tasks spanning classification, generation, structured prediction, and mixed matching. It probes capabilities across Bloom's taxonomy levels, from factual recall and understanding to complex tax strategy planning and risk prevention. Use when the user wants to benchmark on TaxPraBen, or asks about evaluating this task. Reports Overall Average.

researchpythongo
0
3