Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

22,871
skills in category
953
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 3,4813,504 of 22,871 skills

Tpu Workload EvalA

Evaluates the performance and energy efficiency of a Tensor Processing Unit (TPU) and alternative hardware designs across six specific neural network workloads. It probes how architectural parameters like memory bandwidth, clock rate, and matrix multiply unit size impact throughput and power consumption. Use when the user wants to benchmark on TPU Benchmark Workloads (MLP0, MLP1, LSTM0, LSTM1, CNN0, CNN1), or asks about evaluating this task. Reports Watt/die.

researchpythonperformance
0
3
Tpscalcbench EvalA

Evaluates large language models' ability to perform analytical calculations in hypersonic thermal protection system engineering using closed-form formulas and thermodynamic relations, without relying on external simulation tools. Use when the user wants to benchmark on TPS-CalcBench, or asks about evaluating this task. Reports relative_error.

researchpythonrust
0
3
Tpr@FprA

Evaluates the discriminative capability of a speaker verification model by measuring the true positive rate at fixed false positive rate thresholds. It probes how well the model's embedding space separates same-speaker pairs from different-speaker pairs under controlled error constraints. Use when the user has predictions and gold and needs to compute TPR@FPR.

researchpythongo
0
3
Tpcds Structural EvalA

Evaluates an LLM's ability to generate structurally complex SQL queries for real-world decision-making workloads. It probes the model's capacity to handle deep nesting, multiple joins, diverse column references, and complex filtering conditions compared to simpler benchmarks. Use when the user wants to benchmark on TPC-DS, or asks about evaluating this task. Reports structural_similarity.

researchpythongo
0
3
Toyadmos EvalA

Evaluates unsupervised anomalous sound detection systems on miniature machine operating sounds. It probes the ability of models to learn normal acoustic patterns and identify deviations caused by mechanical faults or environmental variations. Use when the user wants to benchmark on ToyADMOS, or asks about evaluating this task. Reports AUC-ROC.

researchpythongit
0
3
Toximol EvalA

Evaluates whether Multimodal Large Language Models (MLLMs) can generate structurally valid, low-toxicity alternative molecules from toxic inputs while adhering to drug-likeness, synthetic feasibility, and structural similarity constraints. It probes the model's ability to perform structure-aware molecular editing and cross-modal scientific reasoning. Use when the user wants to benchmark on ToxiMol, or asks about evaluating this task. Reports Toxicity Repair Success Rate.

researchpythongit
0
3
Toxicity Perspectives EvalA

Evaluates how well automated toxicity classifiers align with diverse human perceptions of harmful content, specifically measuring how demographic background and personal harassment experiences influence toxicity judgments. Use when the user wants to benchmark on Toxicity Perspectives Dataset, or asks about evaluating this task. Reports interrater agreement (Cohen's kappa).

researchpythongo
0
3
Toxicity EvalA

Evaluates the toxicity of text sequences (prompts and model continuations) by scoring them with a black-box API. It probes how well models generate non-toxic text and how sensitive toxicity metrics are to API updates and score drift over time. Use when the user wants to benchmark on REALTOXICITYPROMPTS, or asks about evaluating this task. Reports Toxic Fraction.

researchpythonapi
0
3
Toxicity Detection EvalA

Probes the ability of text generation models to produce non-toxic content by measuring average toxicity scores and comparing them via statistical significance testing. It specifically evaluates how accounting for classifier uncertainty affects the reliability of these comparisons. Use when the user wants to benchmark on BOLD, RealToxicityPrompts, or asks about evaluating this task. Reports Confidence Interval.

researchpythongo
0
3
Toxic Language Classification EvalA

Evaluates binary toxic language classification under extreme data scarcity and severe class imbalance. It probes how well classifiers can detect the minority 'threat' class when trained on a very small labeled dataset, and measures the effectiveness of various data augmentation techniques in improving recall and macro-F1. Use when the user wants to benchmark on Seed, or asks about evaluating this task. Reports macro-averaged F1-score.

researchpythongo
0
3
Toxic Comment EvalA

This benchmark evaluates the individual and group fairness of toxicity classifiers on online comments. It probes whether model predictions remain stable when sensitive identity tokens are swapped (individual fairness) and whether prediction accuracy is equitable across different demographic groups (group fairness). Use when the user wants to benchmark on Toxic Comment Classification Challenge, or asks about evaluating this task. Reports Balanced Accuracy (BA).

researchpythongit
0
3
Toxic Comment Classification EvalA

Evaluates transformer and RNN models on their ability to classify toxic comments while measuring classification accuracy and inference speed. It specifically probes identity-based bias by measuring how well models distinguish between toxic and normal comments across demographic subgroups. Use when the user wants to benchmark on Civil Comments, or asks about evaluating this task. Reports Macro AUROC.

researchpythongo
0
3
Tox21 Fsl EvalA

Evaluates graph neural networks for few-shot toxic molecule classification. It probes the model's ability to generalize from very limited labeled examples (shots) and adapt to new query sets using meta-learning and graph augmentation techniques. Use when the user wants to benchmark on Tox21, or asks about evaluating this task. Reports ROC-AUC Score.

researchpythonnode
0
3
Tox21 Challenge EvalA

Evaluates molecular toxicity prediction capabilities across diverse AI architectures (descriptor-based models, neural networks, tabular transformers, and zero-shot LLMs) on a standardized chemical safety benchmark. Use when the user wants to benchmark on Tox21 Challenge dataset, or asks about evaluating this task. Reports performance.

researchpythongo
0
3
Towervision Multilingual Vl EvalA

This evaluation probes the multilingual vision-language capabilities of models across text recognition, cultural understanding, multimodal translation, and video reasoning. It specifically tests cross-lingual generalization and cultural grounding in both image and video domains across high- and low-resource languages. Use when the user wants to benchmark on ALM-Bench, OCRBench, cc-OCR, TextVQA, CoMMuTE, Multi30K, ViMUL-Bench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Tovo Consensus EvalA

This evaluation probes a model's ability to classify text content according to a user-defined toxicity taxonomy. It measures how closely the model's predictions align with gold labels generated through a multi-model voting process, and tests generalization to out-of-domain categories. Use when the user wants to benchmark on ToVo, or asks about evaluating this task. Reports consensus rate.

researchpythongo
0
3
Touchdown EvalA

Evaluates an agent's capacity to navigate through street-view environments using natural language instructions and resolve complex spatial descriptions to locate a hidden target object within a panoramic image. Use when the user wants to benchmark on Touchdown, or asks about evaluating this task. Reports pixel distance.

researchpythongo
0
3
Toto Ts Forecasting EvalA

Evaluates zero-shot and fine-tuned time series forecasting capabilities on real-world observability telemetry and general-purpose benchmarks. Probes model robustness to high-dimensional, nonstationary, multivariate series with skewed distributions and varying temporal intervals. Use when the user wants to benchmark on Boom, Boomlet, GIFT-Eval, LSF, or asks about evaluating this task. Reports CRPS.

researchpython
0
3
Topofair Fairness EvalA

Evaluates fairness-aware link prediction models on synthetic graphs with controlled topological biases. It probes how structural properties like assortativity, heterogeneity, and class imbalance impact fairness metrics (SP, EO) and predictive accuracy (Hit@10, AUC). Use when the user wants to benchmark on Opinion use case, Friendship use case, Collab use case, Real datasets (Collab, Polblogs, Facebook), or asks about evaluating this task. Reports Statistical Parity (SP), Equalized Odds (EO).

researchpythonnode
0
3
Topoc Cancer Diagnosis EvalA

Evaluates histopathology image classification for ovarian and breast cancer diagnosis using topological deep learning features combined with CNNs. Probes the model's ability to differentiate cancer subtypes and benign/malignant cases from microscopic tissue tiles. Use when the user wants to benchmark on UBC-OCEAN, BREAKHIS, or asks about evaluating this task. Reports Balanced Accuracy.

researchpythonperformance
0
3
Topiocqa EvalA

Evaluates open-domain conversational question answering with topic switching, requiring models to maintain context across multiple turns and dynamically retrieve relevant documents to answer evolving questions. Use when the user wants to benchmark on TOPIOCQA, or asks about evaluating this task. Reports F1.

researchpythongo
0
3
Topic Trend Detection EvalA

Measures the change in sentiment trend towards a specific topic over time or across datasets, requiring temporal or comparative analysis. Use when the user wants to benchmark on Twitter2015-test, or asks about evaluating this task. Reports avgDiff.

researchpython
0
3
Topic Level Polarity EvalA

Predicts the sentiment polarity associated with a specific topic within a tweet, requiring topic-aware sentiment classification and contextual disambiguation. Use when the user wants to benchmark on Twitter2015-test, or asks about evaluating this task. Reports macro-averaged F1.

researchpythonperformance
0
3
Toolsandbox EvalA

Evaluates LLM tool-use capabilities in a stateful, conversational, and interactive setting. It probes the model's ability to handle implicit state dependencies, canonicalize arguments, handle insufficient information, and maintain efficiency across single/multiple tool calls and user turns. Use when the user wants to benchmark on ToolSandbox, or asks about evaluating this task. Reports average similarity score.

researchpythongo
0
3