Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

20,839
skills in category
869
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 7,105–7,128 of 20,839 skills

Guardrail Robustness EvalA

Evaluates the robustness and generalization of LLM safety guardrails against adversarial jailbreak prompts, measuring their ability to correctly classify harmful vs. benign inputs under both known benchmark distributions and novel, contextually framed attacks. Use when the user wants to benchmark on Adversarial Guardrail Benchmark, or asks about evaluating this task. Reports Overall Accuracy.

researchpythongo
0
3
Guacamol Molecule Generation EvalA

This evaluation probes a molecular generative model's ability to produce chemically valid, diverse, and structurally realistic molecules. It measures how well the generated molecules match the physicochemical property distributions of real compounds while maintaining high novelty and uniqueness rates. Use when the user wants to benchmark on GuacaMol benchmark suite, or asks about evaluating this task. Reports KL divergence.

researchpythongo
0
3
Gtsinger EvalA

Evaluates singing voice synthesis models on technique-controllable generation, singer similarity, and audio quality across multiple languages and vocal techniques. It probes the model's ability to accurately control specific singing techniques (e.g., vibrato, mixed voice) while maintaining naturalness and timbre fidelity. Use when the user wants to benchmark on GTSinger, or asks about evaluating this task. Reports MOS-Q.

researchpythongo
0
3
Gtpbd Mm EvalA

Evaluates multimodal terraced parcel extraction by measuring pixel-level segmentation accuracy, edge-level boundary recovery, and object-level structural consistency across image-only, image+text, and image+text+DEM input settings. Use when the user wants to benchmark on GTPBD-MM, or asks about evaluating this task. Reports mIoU.

researchpythongo
0
3
Gtpbd EvalA

Evaluates fine-grained agricultural parcel delineation, boundary detection, and cross-domain generalization on high-resolution remote sensing imagery of terraced terrain. It benchmarks semantic segmentation, edge extraction, and parcel extraction models across multiple geographic domains. Use when the user wants to benchmark on GTPBD, or asks about evaluating this task. Reports IoU.

researchpythongo
0
3
Gtb Dti EvalA

Evaluates structure-based drug-target interaction (DTI) prediction models on regression (binding affinity) and classification (binding status) tasks across six standard bioinformatics datasets. Use when the user wants to benchmark on DAVIS, KIBA, BindingDB, Human, Cycles, Drugbank, or asks about evaluating this task. Reports PCC, ROC-AUC.

researchpythonperformance
0
3
Gt23d Bench EvalA

Evaluates the quality and alignment of generated 3D assets against text prompts across multiple dimensions, including textual alignment, texture fidelity, geometry correctness, and multi-view consistency. It measures how well automated metrics correlate with human preferences to provide a reliable assessment of general text-to-3D generation methods. Use when the user wants to benchmark on GT23D-Bench, or asks about evaluating this task. Reports Texture Fidelity.

researchpython
0
3
Gsr Bench EvalA

Evaluates multimodal LLMs' ability to understand and disambiguate spatial relations (e.g., on, under, left of, right of, in front of, behind) between objects in images. It isolates spatial reasoning from object grounding by providing depth maps, bounding boxes, and segmentation masks alongside images. Use when the user wants to benchmark on GSR-Bench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Gsm8k V EvalA

This benchmark evaluates vision-language models' ability to perform multi-step mathematical reasoning using purely visual, comic-style narratives instead of text. It specifically probes challenges in inter-image semantic understanding, object grounding, and extracting numerical relationships from multi-panel visual contexts. Use when the user wants to benchmark on GSM8K-V, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Gsm8k EvalA

Evaluates a model's ability to perform multi-step arithmetic reasoning by generating natural language solutions to grade school math word problems and verifying their correctness. Use when the user wants to benchmark on GSM8K, or asks about evaluating this task. Reports solve rate.

researchpythongo
0
3
Gseval Pixel Grounding EvalA

Evaluates a model's ability to perform open-vocabulary, fine-grained pixel grounding by generating accurate segmentation masks from complex, long-form referring expressions across multiple granularities (stuff, part, multi-object, single-object). Use when the user wants to benchmark on GSEval, gRefCOCO, RefCOCOm, RefCOCO, RefCOCOg, or asks about evaluating this task. Reports cIoU / gIoU.

researchpythonexpress
0
3
Gscan EvalA

Evaluates systematic generalization in grounded language understanding by testing whether models can interpret natural language commands within dynamic grid-world environments. It probes compositional generalization across novel object properties, navigation directions, contextual size references, action-argument bindings, and adverbial modifiers. Use when the user wants to benchmark on gSCAN, or asks about evaluating this task. Reports exact match accuracy.

researchpythongo
0
3
Gsc Speech Commands EvalA

Evaluates keyword spotting models trained on real versus synthetic speech data, measuring how ASR-based filtering of hallucinated synthetic commands affects classification accuracy on the Google Speech Commands dataset. Use when the user wants to benchmark on Google Speech Commands (GSC), or asks about evaluating this task. Reports Accuracy (%).

researchpythongo
0
3
Gru D EvalA

This evaluation probes a model's ability to handle multivariate time series with missing values by jointly learning temporal dependencies and informative missing patterns. It tests classification performance on clinical and synthetic datasets, measuring how well the model exploits masking and time-interval information for early prediction and multi-task diagnosis. Use when the user wants to benchmark on Gesture, PhysioNet Challenge 2012, MIMIC-III, or asks about evaluating this task. Reports ...

researchpythonperformance
0
3
Group Fairness Reward EvalA

Evaluates whether reward models assign equal average scores to high-quality responses across different demographic/occupational groups. It probes for systematic bias in how models rank expert-written abstracts based on the author's discipline. Use when the user wants to benchmark on arXiv Metadata (Curated), or asks about evaluating this task. Reports Normalized Maximum Group Difference.

researchpythonexpress
0
3
Groundnext EvalA

Evaluates vision-language models on UI element localization and grounding across desktop, mobile, and web interfaces. It measures how accurately a model can identify and locate specific UI components based on text instructions, and assesses their effectiveness in multi-step agentic tasks. Use when the user wants to benchmark on SSPro, OSW-G, MMB-GUI, SSv2, UI-V, OSWorld-Verified, or asks about evaluating this task. Reports average performance.

researchpythongo
0
3
Groundlie360 EvalA

This benchmark evaluates a model's ability to detect and localize multimodal misinformation across text, speech, and video. It probes fine-grained cross-modal reasoning by requiring binary veracity classification, sub-type categorization, and precise grounding of fake content at the token, frame, and bounding-box levels. Use when the user wants to benchmark on GroundLie360, or asks about evaluating this task. Reports F1 score.

researchpythongo
0
3
Groundingme EvalA

Evaluates multimodal large language models' (MLLMs) visual grounding capabilities across four dimensions: discriminative object distinction, spatial relational understanding, handling occlusion/size constraints, and the ability to reject ungroundable queries. It measures how well models can localize objects in images and whether they hallucinate or correctly refuse impossible requests. Use when the user wants to benchmark on GroundingME, or asks about evaluating this task. Reports Accuracy@0.5.

researchpythongo
0
3
Grounding Video Reasoning EvalA

Evaluates video understanding models on physical event reasoning across six domains (gravity, fluids, collisions, deformation, friction, state changes). It probes spatio-temporal grounding by requiring models to predict what happens, when it happens, and where it happens, while measuring robustness to input perturbations like shuffling, ablation, and frame masking. Use when the user wants to benchmark on Physical Video Reasoning Benchmark, or asks about evaluating this task. Reports LGM.

researchpythongo
0
3
Grounder EvalA

This benchmark evaluates a model's ability to localize arbitrary natural language phrases within images. It probes phrase grounding capabilities by requiring the model to attend to relevant image regions and select a bounding box that matches the textual description, without relying on explicit bounding box supervision during training. Use when the user wants to benchmark on Flickr 30k Entities, ReferItGame, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Grounded Ecg Understanding EvalA

Evaluates a multimodal LLM's ability to interpret 12-lead ECG signals and images, providing clinically grounded diagnoses, detailed feature annotations, and evidence-based reasoning. It also tests cardiac abnormality detection and automated report generation across multiple public ECG datasets. Use when the user wants to benchmark on MIMIC-IV-ECG, ECG-Bench (PTB-XL, CPSC2018, G12EC, CODE-15%, CSN), PTB-XL Report, ECG-QA, or asks about evaluating this task. Reports DiagnosisAccuracy.

researchpythonexpress
0
3
Groundcocoa EvalA

Evaluates compositional and conditional reasoning in LLMs by requiring them to match complex, logically constrained user preferences to specific flight booking options. It probes the model's ability to handle interdependent requirements and atypical constraints without external reasoning engines. Use when the user wants to benchmark on GroundCocoa, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Gromov Wasserstein SimilarityA

Evaluates how well Gromov-Wasserstein distance captures functional similarity between neural network layer representations, enabling the identification of structural transitions and latent sub-networks across varying dimensionalities without task-specific supervision. Use when the user has predictions and gold and needs to compute Gromov-Wasserstein distance.

researchpythongo
0
3
Grl Perturbation Sensitivity EvalA

Evaluates graph neural network robustness and feature/structure reliance by measuring performance degradation under 13 structured perturbations to node features and graph topology. It classifies datasets based on their sensitivity profiles to structural vs. feature information. Use when the user wants to benchmark on GRL Benchmark Collection (49 datasets), or asks about evaluating this task. Reports sensitivity_profile.

researchpythonnode
0
3