Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

23,477
skills in category
979
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 3,961–3,984 of 23,477 skills

Ultrachat EvalA

Evaluates a chat model's ability to generate accurate, informative, and correct responses across diverse domains including commonsense, world knowledge, professional knowledge, mathematics, reasoning, and writing. It probes both factual correctness and response quality using automated LLM-based pairwise and independent scoring. Use when the user wants to benchmark on UltraChat Evaluation Set, or asks about evaluating this task. Reports ChatGPT scoring.

researchpythongit
0
3
Ultra Fineweb EvalA

Evaluates the quality of pre-training datasets by measuring the downstream performance of models trained on them. It probes general language understanding, commonsense reasoning, and multilingual capabilities through standard zero-shot benchmarks. Use when the user wants to benchmark on MMLU, ARC-C, ARC-E, CommonSenseQA, HellaSwag, OpenbookQA, PIQA, SIQA, Winogrande, C-Eval, CMMLU, or asks about evaluating this task. Reports Average.

researchpythonperformance
0
3
Ukrainian Text Classification EvalA

Evaluates cross-lingual transfer methods for Ukrainian text classification across toxicity, formality, and natural language inference tasks. It compares translation-based baselines, LLM prompting, and adapter/fine-tuning approaches on both machine-translated and semi-natural Ukrainian test sets. Use when the user wants to benchmark on Ukrainian Toxicity (Translated & Semi-natural), Ukrainian Formality (Translated & Semi-natural), Ukrainian NLI (Translated & Semi-natural), or asks about evalua...

researchpythongo
0
3
Ukbob Medical Segmentation EvalA

Evaluates 3D medical image segmentation models on multi-modal MRI and CT scans across abdominal, brain, and whole-body anatomical domains. Probes the model's ability to produce accurate organ and tumor masks while measuring both volumetric overlap and boundary precision under zero-shot and fine-tuning settings. Use when the user wants to benchmark on AMOS, BTCV, BRATS, UKBOB, or asks about evaluating this task. Reports Dice Score.

researchpythonperformance
0
3
Ukb Disease Risk EvalA

Evaluates the ability of multimodal LLMs to predict binary disease risk from individual-specific clinical data, including tabular features and time-series spirograms. It tests how well serialized text and cross-modal embeddings integrate to produce accurate risk scores for conditions like asthma and diabetes. Use when the user wants to benchmark on UK Biobank, or asks about evaluating this task. Reports AUROC.

researchpythongit
0
3
Uk Weather Da EvalA

Evaluates the impact of data assimilation (DA) using the SPEnKF algorithm on a U-STN12 deep learning model for UK temperature forecasting. It probes the model's ability to integrate global atmospheric data (ERA5 T850) and surface observations (ASOS/ERA5 T2m) over a 120-hour lead time, measuring forecast accuracy degradation or improvement under varying noise levels and assimilation frequencies. Use when the user wants to benchmark on ERA5, ASOS, or asks about evaluating this task. Reports RMSE.

researchpythongo
0
3
Uit Viquad 2.0 EvalA

Evaluates Vietnamese language models' machine reading comprehension capabilities, specifically probing their ability to extract correct answer spans and correctly identify when a question cannot be answered from the given context. Use when the user wants to benchmark on UIT-ViQuAD 2.0, or asks about evaluating this task. Reports Exact Match (EM).

researchpythongo
0
3
Ugr16 EvalA

Evaluates how data preprocessing choices—such as observation selection, flow directionality, and feature engineering—affect the performance of unsupervised anomaly detection models on network traffic. Use when the user wants to benchmark on UGR'16, or asks about evaluating this task. Reports AUC.

researchpythonperformance
0
3
Ueof EvalA

Evaluates the accuracy of event-based optical flow estimation models in underwater environments. It probes how well algorithms handle low-texture, turbid, and refractive conditions compared to terrestrial benchmarks. Use when the user wants to benchmark on UEOF, or asks about evaluating this task. Reports AEE.

researchpythongo
0
3
Ucr Augmented EvalA

Evaluates time series classifiers' reliance on temporal structure by measuring accuracy degradation when temporal alignment is disrupted via padding. It contrasts performance on the original UCR benchmark against a perturbed version to isolate the contribution of temporal correlations versus tabular features. Use when the user wants to benchmark on UCR Augmented, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Ucla Asv EvalA

Evaluates automatic speaker verification (ASV) robustness to speaking-style mismatches between enrollment and test utterances. It measures how well data augmentation techniques can compensate for style variability without requiring multi-style training data. Use when the user wants to benchmark on UCLA database, or asks about evaluating this task. Reports EER.

researchpythondatabase
0
3
Uci Gp Benchmark EvalA

Evaluates the time-accuracy trade-offs of approximate Gaussian Process regression methods against exact baselines and simple models across multiple UCI regression datasets. It measures how quickly approximations converge to near-exact performance while tracking predictive quality over time. Use when the user wants to benchmark on UCI regression datasets, or asks about evaluating this task. Reports NLPD.

researchpythonperformance
0
3
Ucfe EvalA

Evaluates large language models' ability to handle dynamic, multi-turn financial dialogues across diverse user personas and task types, measuring their adaptability to shifting user needs and financial expertise. Use when the user wants to benchmark on UCFE, or asks about evaluating this task. Reports Elo score.

researchpythongo
0
3
Uav Wildlife Tracking EvalA

This evaluation probes an autonomous UAV navigation model's ability to track wildlife by predicting flight commands that match expert pilot behavior. It measures how well the model maintains optimal camera framing and altitude for behavioral video collection. Use when the user wants to benchmark on KABR, or asks about evaluating this task. Reports % of actions matching original flight.

researchpythongo
0
3
UXEB Decay FittingA

Evaluates the capability to characterize and quantify correlated noise in near-term quantum processors by measuring the exponential decay of average fidelity under random quantum circuits. Use when the user has predictions and gold and needs to compute uXEB decay fitting (ENR extraction).

researchpythongo
0
3
UMLIP High Temp Mof EvalA

Evaluates the accuracy of universal machine-learned interatomic potentials (uMLIPs) in predicting energy, forces, and stress tensors during high-temperature molecular dynamics simulations of metal-organic frameworks (MOFs), including their stability and thermal decomposition behavior. Use when the user wants to benchmark on High-Temperature MOF AIMD Benchmark, or asks about evaluating this task. Reports energy MAE.

researchpythonperformance
0
3
U2flow EvalA

Evaluates the accuracy of dense optical flow estimation and the reliability of per-pixel uncertainty quantification in an unsupervised setting. It probes the model's ability to handle occlusions, textureless regions, and domain shifts without ground-truth flow supervision. Use when the user wants to benchmark on KITTI, Sintel, or asks about evaluating this task. Reports EPE.

researchpythongo
0
3
U Math EvalA

Evaluates LLMs' ability to solve university-level mathematical problems, both text-based and multimodal. It also includes a meta-evaluation component to assess how well models can judge the correctness of free-form mathematical solutions. Use when the user wants to benchmark on U-MATH, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Twnertc EvalA

Evaluates the annotation quality of the TWNERTC corpus for Turkish named entity recognition (NER) and text categorization (TC) by comparing automated labels against human-verified ground truths. It measures how well coarse-grained and fine-grained entity types, as well as domain categories, align with human judgment across different noise-reduction post-processing variants. Use when the user wants to benchmark on TWNERTC, or asks about evaluating this task. Reports F1-Score.

researchpythongo
0
3
Twinviews Bias EvalA

Evaluates whether reward models exhibit political bias by measuring the average reward scores assigned to politically left-leaning versus right-leaning statements on the same topics. The protocol compares mean reward differences across model sizes and training runs to detect systematic left-leaning skew. Use when the user wants to benchmark on TwinViews-13k, or asks about evaluating this task. Reports average_reward.

researchpythongit
0
3
Twin 2k 500 EvalA

Evaluates the ability of LLMs to simulate individual-level human behavior across demographic, psychological, cognitive, economic, and behavioral economics domains. It measures test-retest accuracy and replication of known behavioral biases using a large-scale dataset of 2,058 U.S. individuals. Use when the user wants to benchmark on Twin-2K-500, or asks about evaluating this task. Reports test-retest accuracy.

researchpythongit
0
3
Twiff Bench EvalA

Evaluates a model's ability to perform dynamic visual reasoning by generating temporally grounded, physically plausible future frames and textual explanations. It probes both the quality of the step-by-step reasoning process and the correctness of the final answer in open-ended video scenarios. Use when the user wants to benchmark on TwiFF-Bench, Seed-Bench-R1, or asks about evaluating this task. Reports Answer score.

researchpythongit
0
3
Tweettopic EvalA

Evaluates language models on classifying tweets into predefined topics, testing both single-label and multi-label classification capabilities. It probes robustness to social media noise, short-form content, and topic overlap in real-world settings. Use when the user wants to benchmark on TweetTopic, or asks about evaluating this task. Reports Macro F1.

researchpythongo
0
3
Tweeteval EvalA

Evaluates the ability of language models to classify short social media posts across seven distinct Twitter-specific tasks, including sentiment, emotion, hate speech, and irony detection. It probes domain adaptation by comparing models pre-trained on generic text versus those further trained on large-scale Twitter corpora. Use when the user wants to benchmark on TweetEval, or asks about evaluating this task. Reports M-F1.

researchpythongo
0
3