Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

20,842
skills in category
869
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 7,273–7,296 of 20,842 skills

Gebench EvalA

Evaluates image generation models' ability to function as dynamic GUI environments, probing temporal coherence, multi-step interaction logic, spatial grounding, and visual fidelity across sequential state transitions. Use when the user wants to benchmark on GEBench, or asks about evaluating this task. Reports GE-Score.

researchpythongo
0
3
Gdro Tabular Imbalance EvalA

Assesses deep learning models' ability to classify highly imbalanced binary tabular data by comparing standard empirical risk minimization against group distributionally robust optimization. Use when the user wants to benchmark on Multiple benchmark imbalanced tabular datasets, or asks about evaluating this task. Reports g-mean.

researchpythonperformance
0
3
Gdibench EvalA

Evaluates document intelligence by decoupling visual and reasoning complexity into graded difficulty levels (V0–V2, R0–R2). It probes a model’s ability to extract, reason over, and generalize across diverse document types while mitigating catastrophic forgetting during fine-tuning. Use when the user wants to benchmark on GDI-Bench, or asks about evaluating this task. Reports Accuracy / normalized edit distance.

researchpythongo
0
3
Gcn Node Classification EvalA

Semi-supervised node classification on citation and knowledge graphs. It probes the model's ability to learn graph-structured representations and classify nodes using only a small fraction of labeled examples. Use when the user wants to benchmark on Citeseer, Cora, Pubmed, NELL, or asks about evaluating this task. Reports prediction accuracy.

researchpythongo
0
3
Gcai Constitution EvalA

Evaluates the moral grounding, coherence, fairness, and real-world applicability of AI alignment constitutions through human surveys, alongside the downstream safety alignment and general capabilities of fine-tuned language models. Use when the user wants to benchmark on BABELSCAPE/ALERT, MMLU, Social Bias BBQ, or asks about evaluating this task. Reports 5-point Likert rating.

researchpythongo
0
3
Gca Tool Use EvalA

Evaluates an LLM's ability to use region-specific climate tools in a multi-step agentic pipeline. It probes structured tool invocation, argument schema adherence, step-wise reasoning, and end-to-end answer accuracy on Gulf-focused climate queries. Use when the user wants to benchmark on GCA-DS, or asks about evaluating this task. Reports AnsAcc.

researchpythongo
0
3
Gazeta Russian Summarization EvalA

This benchmark evaluates the quality of abstractive and extractive text summarization in Russian. It measures how well generated summaries capture the key information and stylistic qualities of the original news articles compared to human-written references. Use when the user wants to benchmark on Gazeta, or asks about evaluating this task. Reports ROUGE.

researchpythongit
0
3
Gazebo Simulation EvalA

Evaluates an MPC-based autonomous driving controller's ability to perform collision avoidance and lane maneuvers (overtaking, merging, following) in dynamic environments using a physics-based simulator. It tests the controller's real-time feasibility and trajectory smoothness under varying traffic densities. Use when the user wants to benchmark on Gazebo Simulation, or asks about evaluating this task. Reports computation time.

researchpythonangular
0
3
Gaussianvlm EvalA

Evaluates a 3D vision-language model's ability to perform object-centric and scene-centric reasoning tasks, including captioning, question answering, embodied planning, and dialogue. It probes spatial grounding, semantic abstraction, and robust generalization to out-of-domain real-world scene representations. Use when the user wants to benchmark on ScanRefer, ScanQA, Nr3D, SQA3D, 3D-LLM ScanNet subset, ScanNet++ (OOD object counting), or asks about evaluating this task. Reports Exact-match ac...

researchpython
0
3
Gass T2i Diversity EvalA

Evaluates text-to-image generation models on their ability to produce diverse, high-quality, and semantically aligned images under fixed prompts. It specifically probes disentangled diversity by measuring prompt-dependent semantic variation versus prompt-independent background/style variation. Use when the user wants to benchmark on ImageNet-1K, DrawBench, or asks about evaluating this task. Reports VS.

researchpythongo
0
3
Gas Saturation EvalA

Evaluates a neural operator's ability to predict long-term multiphase flow dynamics (gas saturation and pressure buildup) in porous media using sparse time snapshots. It probes data efficiency, generalization to unseen time steps, and computational resource usage compared to baseline spectral methods. Use when the user wants to benchmark on Synthetic multiphase flow dataset (gas saturation & pressure buildup), or asks about evaluating this task. Reports R^2.

researchpythongit
0
3
Garments2look EvalA

Probes the ability of virtual try-on and image editing models to synthesize high-fidelity, multi-reference outfit images. It evaluates whether models can preserve fine-grained garment details, maintain correct layering orders, and adhere to specific styling techniques while keeping the target person's pose consistent. Use when the user wants to benchmark on Garments2Look, DressCode-MR, or asks about evaluating this task. Reports FID↓.

researchpythongit
0
3
Garment Folding EvalA

Evaluates a robot's ability to manipulate deformable fabric garments using a single arm. It probes two capabilities: flattening a crumpled T-shirt to maximize coverage, and folding a flattened T-shirt to match a goal configuration while minimizing wrinkles. Use when the user wants to benchmark on Google Reach T-shirt folding environment, or asks about evaluating this task. Reports max_coverage_pct.

researchpythongo
0
3
Gar Bench EvalA

Evaluates a multimodal LLM's ability to perform precise, context-aware visual understanding at the region level. It probes fine-grained perception, compositional reasoning across multiple visual prompts, and detailed localized captioning for both images and videos. Use when the user wants to benchmark on GAR-Bench-VQA, GAR-Bench-Cap, DLC-Bench, Ferret-Bench, MDVP-Bench, LVIS, PACO, VideoRefer-Bench, or asks about evaluating this task. Reports Overall score.

researchpythongo
0
3
Gaps EvalA

Evaluates AI clinicians on clinical reasoning depth, answer completeness, robustness to input perturbations, and safety risk mitigation. It probes how well models retrieve, synthesize, and apply evidence-based guidelines under varying cognitive loads and adversarial conditions. Use when the user wants to benchmark on GAPS-NCCN-NSCLC-preview, or asks about evaluating this task. Reports GAPS score (normalized rubric-based).

researchpythongo
0
3
Gap Overlap Kg EvalA

Evaluates a knowledge graph's ability to perform gap and overlap analysis on life insurance contracts by answering scenario-based competency questions. It probes the system's capacity for structured, evidence-grounded reasoning to determine claim coverage, denial, or non-applicability across heterogeneous contract types. Use when the user wants to benchmark on Insurance Contract KG Benchmark, or asks about evaluating this task. Reports accuracy.

researchpythongit
0
3
Gap EvalA

Evaluates a model's ability to resolve gendered ambiguous pronouns to their correct antecedent names in natural text. It specifically probes for gender bias and the reliance on syntactic or contextual cues over surface-level heuristics. Use when the user wants to benchmark on GAP, or asks about evaluating this task. Reports F1 score.

researchpythongo
0
3
Gaoyao EvalA

Evaluates the multilingual and multicultural capabilities of large language models across 26 languages and 51 cultures. It probes cognitive abilities (e.g., reasoning, reading, translation) and cultural understanding (monocultural and cross-cultural contexts) to identify geographical performance disparities and benchmark saturation. Use when the user wants to benchmark on GaoYao, or asks about evaluating this task. Reports accuracy, win_rate.

researchpythongo
0
3
Gamma Glaucoma EvalA

Evaluates multi-modal medical image analysis models for glaucoma staging by jointly processing 2D fundus images and 3D OCT volumes. It probes the model's ability to fuse cross-modality features and correctly classify patients into normal, early, or progressive glaucoma stages. Use when the user wants to benchmark on GAMMA Challenge, or asks about evaluating this task. Reports kappa.

researchpythongo
0
3
Gameplayqa EvalA

GameplayQA evaluates multi-modal large language models' ability to understand decision-dense, first-person synchronized multi-video environments. It probes capabilities in agent-state tracking, temporal reasoning, and cross-video event alignment across three cognitive difficulty levels. Use when the user wants to benchmark on GameplayQA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Gamephysics Video Search EvalA

Evaluates zero-shot video retrieval capability using natural language queries to locate specific objects, compound descriptions, and game physics bugs in unstructured gameplay footage. It probes the model's ability to generalize across diverse open-world game genres and visual styles without fine-tuning. Use when the user wants to benchmark on GamePhysics, or asks about evaluating this task. Reports top-k accuracy.

researchpython
0
3
Gamefactory EvalA

This evaluation protocol assesses a video generation model's ability to follow discrete and continuous action inputs while maintaining semantic alignment with text prompts and preserving the original model's visual domain. It measures action-following accuracy, camera pose consistency, text-video semantic relevance, and overall video generation quality across in-domain and open-domain scenes. Use when the user wants to benchmark on GF-Minecraft, VPT (Find Cave), or asks about evaluating this ...

researchpython
0
3
Game Of 24 EvalA

Evaluates an LLM's ability to perform algorithmic search and recursive reasoning within a single generation window, without external tree search or iterative prompting. It probes systematic exploration, pruning, and backtracking capabilities in a mathematical constraint satisfaction task. Use when the user wants to benchmark on Game of 24, or asks about evaluating this task. Reports Success rate.

researchpythongo
0
3
Game EvalA

Evaluates a vision-action model's ability to predict actions and scale ratios from video game footage, testing in-distribution and out-of-distribution generalization across 2D and 3D games. Use when the user wants to benchmark on Video Games (In-Distribution & OOD), or asks about evaluating this task. Reports Pearson correlation.

researchpythongo
0
3