Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

20,838
skills in category
869
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 7,249–7,272 of 20,838 skills

Geobench EvalA

Evaluates the transferability and semantic grounding of remote sensing foundation models on image-level classification and pixel-level semantic segmentation tasks across multiple geospatial benchmarks. Use when the user wants to benchmark on GeoBench, or asks about evaluating this task. Reports F1.

researchpythongo
0
3
Geo880 Atis EvalA

Evaluates a neural semantic parser's ability to map natural language utterances directly to executable SQL queries. It probes compositional generalization and schema grounding by measuring whether predicted queries return the exact same results as gold queries on a target database. Use when the user wants to benchmark on GEO880, ATIS, or asks about evaluating this task. Reports denotation accuracy.

researchpythongo
0
3
Genspace Alignment EvalA

Evaluates how well automated metrics and VLMs align with human judgments on spatially-aware image generation tasks across nine sub-domains. Use when the user wants to benchmark on GenSpace Human Alignment Test Set, or asks about evaluating this task. Reports agreement.

researchpythongo
0
3
Genimage EvalA

This benchmark probes a model's ability to distinguish real from AI-generated images across multiple diffusion and GAN generators, and to identify specific synthetic flaws in generated images. It evaluates standard detection accuracy, cross-generator generalization, and robustness to common image perturbations like blur, rotation, and brightness shifts. Use when the user wants to benchmark on GenImage, GENHARD, GENEXPLAIN, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Geneval T2iA

Evaluates text-to-image generation by measuring how accurately models follow complex prompts with multiple objects, attributes, and spatial constraints. Use when the user wants to benchmark on GenEval, or asks about evaluating this task. Reports GenEval Overall.

researchpython
0
3
Geneval EvalA

Evaluates text-to-image alignment by testing whether models correctly render specified objects, attributes, and spatial relations across six categories: single object, two objects, counting, colors, position, and attribute binding. Use when the user wants to benchmark on GenEval, or asks about evaluating this task. Reports Overall.

researchpythongo
0
3
Generative Unfolding EvalA

Evaluates a generative ML model's ability to correct detector effects (unfolding) for highly boosted hadronic top-quark decays. It probes the model's capacity to reconstruct high-dimensional kinematic phase space while mitigating simulation-induced model bias and accurately extracting the top_mass_measurement. Use when the user wants to benchmark on CMS benchmark top-pair simulation, or asks about evaluating this task. Reports top_mass_measurement.

researchpythonperformance
0
3
Generalad EvalA

Evaluates cross-domain anomaly detection capability across semantic, near-distribution, and industrial benchmarks. Probes the model's ability to detect pixel-level defects and semantic novelties using a self-supervised transformer discriminator that attends to distorted features without task-specific tuning. Use when the user wants to benchmark on CIFAR-10, CIFAR-100, Fashion-MNIST, View, Aircraft-FGVC, Stanford Cars, MVTec-AD, MVTec-LOCO, VisA, MPDD, or asks about evaluating this task. Repor...

researchpythonperformance
0
3
General Nlu EvalA

Evaluates whether integrating external knowledge (textual descriptions or embeddings) improves performance on general natural language understanding tasks compared to baseline pre-trained language models. It probes the model's ability to leverage external semantic information to enhance representation learning and decision-making across classification, regression, and sequence labeling benchmarks. Use when the user wants to benchmark on GLUE, Penn Treebank, CoNLL-2003, or asks about evaluatin...

researchpythongo
0
3
Geneol EvalA

Evaluates training-free sentence embedding quality by aggregating LLM-generated semantic variations. Probes semantic similarity preservation and cross-task robustness without model fine-tuning. Use when the user wants to benchmark on STS benchmark, MTEB, or asks about evaluating this task. Reports Spearman rank correlation (cosine similarity).

researchpythongo
0
3
Gene Bench EvalA

This benchmark evaluates how well various deep learning models encode biological knowledge by predicting 312 ground-truth gene properties. It probes capabilities across five domains: genomic, regulatory, localization, biological processes, and protein features, using binary, multi-label, and multi-class classification tasks. Use when the user wants to benchmark on GeneBench, or asks about evaluating this task. Reports AUC.

researchpythongit
0
3
Genderpair EvalA

Evaluates gender bias in large language models by measuring the model's preference or generation likelihood across stereotypical versus counterfactual gendered prompts. It specifically probes inclusivity and diversity by including marginalized gender identities such as transgender and non-binary groups. Use when the user wants to benchmark on GenderPair, or asks about evaluating this task. Reports Bias-Pair Ratio.

researchpythongo
0
3
Genderbias Vl EvalA

This benchmark probes the gender bias of Large Vision-Language Models (LVLMs) in occupation inference tasks. It uses counterfactual visual question pairs to measure how model predictions change when the perceived gender of a subject is swapped, evaluating both cognitive accuracy and fairness under individual and causal fairness frameworks. Use when the user wants to benchmark on GenderBias-VL, or asks about evaluating this task. Reports Idealized Score (Ipss).

researchpythongo
0
3
Gendeg EvalA

Evaluates the out-of-distribution (OoD) generalization and within-distribution performance of All-In-One Image Restoration (AIOR) models across six degradation types (haze, rain, snow, motion blur, raindrop, low-light) when trained with synthetic degradation data. Use when the user wants to benchmark on O-HAZE, LHP, RainDS, RSVD, GoPro, or asks about evaluating this task. Reports LPIPS.

researchpythongo
0
3
Gen Nerf EvalA

Evaluates the rendering quality and computational efficiency of a generalizable Neural Radiance Field (NeRF) model for novel view synthesis. It measures how accurately the model reconstructs unseen scenes from a few source views, balancing image fidelity against computational cost and hardware throughput. Use when the user wants to benchmark on NeRF Synthetic, LLFF, DeepVoxels, or asks about evaluating this task. Reports PSNR.

researchpythongo
0
3
Gemini Robotics 15 EvalA

Evaluates a robot's ability to execute short-horizon and multi-step manipulation tasks across diverse embodiments, environments, and visual/instructional variations. It specifically probes zero-shot cross-embodiment skill transfer and the impact of explicit 'thinking' traces on task progress and success. Use when the user wants to benchmark on Gemini Robotics 1.5 Benchmark, or asks about evaluating this task. Reports progress score.

researchpythongo
0
3
Gemex EvalA

Evaluates large vision-language models on chest X-ray diagnosis by testing their ability to answer medical questions, provide textual reasoning, and ground answers to specific visual regions in radiographs. Use when the user wants to benchmark on GEMeX, or asks about evaluating this task. Reports AR-score.

researchpythongo
0
3
Gem EvalA

Evaluates natural language generation models across diverse tasks including content planning, surface realization, and communicative goals. It probes lexical similarity, semantic equivalence, faithfulness, and output diversity using both reference-based and reference-free automated metrics. Use when the user wants to benchmark on CommonGen, Czech Restaurant, DART, E2E clean, MLSum, Schema-Guided, ToTTo, XSum, WebNLG, Turk, ASSET, WikiLingua, or asks about evaluating this task. Reports BLEU.

researchpythongo
0
3
Gem Cot Mixed Task EvalA

Evaluates the ability of LLMs to perform zero-shot and few-shot reasoning across a heterogeneous mix of unseen and known task types without manual task-specific prompting. It probes dynamic demonstration routing, clustering-based generalization, and streaming adaptation in mixed-task scenarios. Use when the user wants to benchmark on AQUA-RAT, MultiArith, AddSub, GSM8K, SingleEq, SVAMP, Last Letter Concatenation, Coin Flip, StrategyQA, CSQA, BIG-Bench Hard (BBH), or asks about evaluating this...

researchpythongo
0
3
GecoA

Evaluates geometric consistency in text-to-video generation by measuring structural and motion coherence across camera trajectories, detecting deformation and occlusion artifacts in static scenes. Use when the user has predictions and gold and needs to compute Fused.

researchpythongo
0
3
Gebench EvalA

Evaluates image generation models' ability to function as dynamic GUI environments, probing temporal coherence, multi-step interaction logic, spatial grounding, and visual fidelity across sequential state transitions. Use when the user wants to benchmark on GEBench, or asks about evaluating this task. Reports GE-Score.

researchpythongo
0
3
Gdro Tabular Imbalance EvalA

Assesses deep learning models' ability to classify highly imbalanced binary tabular data by comparing standard empirical risk minimization against group distributionally robust optimization. Use when the user wants to benchmark on Multiple benchmark imbalanced tabular datasets, or asks about evaluating this task. Reports g-mean.

researchpythonperformance
0
3
Gdibench EvalA

Evaluates document intelligence by decoupling visual and reasoning complexity into graded difficulty levels (V0–V2, R0–R2). It probes a model’s ability to extract, reason over, and generalize across diverse document types while mitigating catastrophic forgetting during fine-tuning. Use when the user wants to benchmark on GDI-Bench, or asks about evaluating this task. Reports Accuracy / normalized edit distance.

researchpythongo
0
3
Gcn Node Classification EvalA

Semi-supervised node classification on citation and knowledge graphs. It probes the model's ability to learn graph-structured representations and classify nodes using only a small fraction of labeled examples. Use when the user wants to benchmark on Citeseer, Cora, Pubmed, NELL, or asks about evaluating this task. Reports prediction accuracy.

researchpythongo
0
3