Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

20,839
skills in category
869
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 7,153–7,176 of 20,839 skills

Graph2vec EvalA

Evaluates the ability of graph representation learning methods to capture structural equivalence for downstream graph classification and clustering tasks. It probes whether learned embeddings can effectively distinguish between different graph classes or group structurally similar graphs without explicit supervision. Use when the user wants to benchmark on Benchmark Graph Classification Datasets (MUTAG, PTC, PROTEINS, NCI1, NCI109), Android Malware Detection Dataset, AMD Malware Clustering Da...

researchpythongo
0
3
Graph To Vision EvalA

This benchmark probes a vision-language model's ability to jointly interpret and reason across multiple related graph images. It specifically tests cross-modal integration, structural understanding, and instruction-following accuracy when processing homogeneous and heterogeneous graph groupings. Use when the user wants to benchmark on Graph-to-Vision Benchmark, or asks about evaluating this task. Reports instruction-following accuracy.

researchpythonnode
0
3
Graph Gen Benchmark EvalA

This benchmark evaluates how effectively graph generative models can produce synthetic graphs that serve as reliable proxies for benchmarking Graph Neural Networks. It measures the fidelity of generated graphs by comparing GNN performance metrics trained on synthetic data against those trained on the original real-world graphs. Use when the user wants to benchmark on Cora, Citeseer, Pubmed, AmazonC, AmazonP, MS CS, MS Physic, or asks about evaluating this task. Reports MSE.

researchpythonnode
0
3
Graph Counterfactual Fairness EvalA

Evaluates graph neural networks for node classification fairness by measuring prediction accuracy alongside statistical fairness metrics (demographic parity, equal opportunity) and a novel graph counterfactual fairness metric that quantifies how much node predictions change when sensitive attributes of the node and its neighbors are perturbed. Use when the user wants to benchmark on Synthetic, Bail, Credit, or asks about evaluating this task. Reports δ_CF.

researchpythonnode
0
3
Graph Classification Tfgw EvalA

Evaluates the ability of graph neural networks and optimal transport-based methods to classify graphs by learning discriminative representations that capture both structural and feature dissimilarities. It probes expressiveness beyond the Weisfeiler-Lehman test and generalization on heterogeneous real-world graph structures. Use when the user wants to benchmark on 4-CYCLES, SKIP-CIRCLES, MUTAG, PTC, ENZYMES, PROTEIN, NCI1, IMDB-B, IMDB-M, COLLAB, or asks about evaluating this task. Reports ac...

researchpythongo
0
3
Graph Classification Accuracy EvalA

Evaluates the ability of graph representation models to correctly classify entire graphs based on their structural topology and, optionally, node or edge attributes. It probes whether local structural summaries or complex neural architectures can capture discriminative patterns for tasks like social network or chemical compound categorization. Use when the user wants to benchmark on IMDB BINARY, IMDB MULTI, COLLAB, REDDIT BINARY, REDDIT 5K, REDDIT 12K, ENZYMES, PROTEINS, D&D, MUTAG, PTC, NCI1...

researchpythongo
0
3
Graph Alignment EvalA

Evaluates a model's ability to perform structural graph alignment by predicting a node-to-node correspondence between two graphs that maximizes shared edges. It probes the model's capacity for equivariant representation learning and combinatorial optimization on graph structures. Use when the user wants to benchmark on Graph Alignment Benchmark (Synthetic & Real-world), or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Granular Change AccuracyA

Evaluates Dialogue State Tracking (DST) models by measuring performance based on per-turn belief state changes rather than raw slot or turn-level accuracy. It aims to provide partial credit for partially correct predictions and reduce bias from error timing and distribution across dialogue turns. Use when the user has predictions and gold and needs to compute Granular Change Accuracy.

researchpythongo
0
3
Granger Causal Inference EvalA

Probes a model's ability to identify causal genomic regulatory relationships (ATAC-seq peaks to RNA-seq genes) using temporal causal inference on single-cell multimodal data. It evaluates how well predicted peak-gene associations align with independent biological proxies like eQTLs and chromatin interactions, testing robustness to high-dimensional sparsity and partial temporal orderings. Use when the user wants to benchmark on sci-CAR, SNARE-seq, SHARE-seq, or asks about evaluating this task....

researchpythongo
0
3
Gram Dti EvalA

Evaluates multimodal drug-target interaction prediction and zero-shot retrieval capabilities across multiple benchmark datasets and cold-start scenarios. Use when the user wants to benchmark on Activation, Yamanishi_08, Hetionet, Inhibition, or asks about evaluating this task. Reports AUPR.

researchpythonperformance
0
3
Gradsafe Jailbreak Detection EvalA

Evaluates the ability to detect jailbreak or unsafe prompts in LLM inputs using gradient-based analysis. It probes zero-shot and adapted detection capabilities against established moderation APIs and LLM-based detectors. Use when the user wants to benchmark on ToxicChat, XSTest, or asks about evaluating this task. Reports AUPRC.

researchpythongo
0
3
Grad Tts EvalA

Evaluates text-to-speech synthesis quality, inference efficiency, and probabilistic modeling accuracy of a diffusion-based model. It probes the trade-off between synthesis fidelity and computational cost by varying reverse diffusion steps, and measures human-perceived audio quality against strong baselines. Use when the user wants to benchmark on LJSpeech, or asks about evaluating this task. Reports MOS.

researchpythongo
0
3
Gracore EvalA

Evaluates large language models' ability to comprehend and reason over graph structures presented as textual descriptions. It probes capabilities ranging from basic graph understanding and semantic reasoning to complex graph theory reasoning across pure and heterogeneous graphs. Use when the user wants to benchmark on GraCoRe, or asks about evaluating this task. Reports score.

researchpythongo
0
3
Grables EvalA

Evaluates whether tabular models can capture extension-sensitive inter-row dependencies (e.g., global counts, overlaps) compared to graph-based message-passing models, and tests if hybrid approaches combining tabular features with graph-derived representations improve performance. Use when the user wants to benchmark on Synthetic transactions dataset, Retail transaction dataset, relbench-trial, or asks about evaluating this task. Reports ROC-AUC.

researchpythongo
0
3
Gqa EvalA

Evaluates visual reasoning and compositional question answering on real-world images. It probes a model's ability to understand scene relationships, answer multi-step questions, and maintain logical consistency across related queries. Use when the user wants to benchmark on GQA, or asks about evaluating this task. Reports Accuracy.

researchpythonrust
0
3
Gpu Power Cap EvalA

Evaluates the performance and power efficiency trade-offs of NVIDIA H100 and H200 GPUs under varying power caps, isolating compute-bound (DGEMM) and memory-bound (STriad) workloads to analyze architectural scaling and frequency throttling dynamics. Use when the user wants to benchmark on cuBLAS DGEMM, TheBandwidthBenchmark (STriad kernel), or asks about evaluating this task. Reports Throughput (TFlop/s or TB/s).

researchpythonnode
0
3
Gpu Memory Co Optimization EvalA

Evaluates the trade-off between system memory footprint and task latency when co-executing multiple workloads under different integrated CPU/GPU memory management policies on embedded platforms. It measures how strategically assigning Device, Managed, or Host-Pinned memory policies affects peak memory consumption, average GPU execution time, and overall GPU utilization during multitasking. Use when the user wants to benchmark on Rodinia Benchmark Suite (subset), DJI Drone Object Detection, Au...

researchpythonperformance
0
3
Gpu Inference Benchmark EvalA

Evaluates GPU inference performance across different neural network models, numerical precision modes, and batch sizes. It measures how architectural differences and execution parallelism impact throughput, latency, and memory utilization under production-like conditions. Use when the user wants to benchmark on ResNet models (ResNet-18, ResNet-50, ResNet-101) with synthetic inputs, or asks about evaluating this task. Reports throughput (images/sec).

researchpythongit
0
3
Gptscore EvalA

Evaluates the correlation between automated scoring functions (GPTScore variants) and human judgments across multiple text generation tasks. It probes the ability of instruction-based LLMs to serve as training-free, customizable evaluators that align with human preference. Use when the user wants to benchmark on SummEval, RealSumm, NEWSROOM, QXSUM, MQM-2020, BAGEL, SFRES, FED, or asks about evaluating this task. Reports Spearman correlation.

researchpython
0
3
Gptaraeval EvalA

Evaluates large language models on Arabic natural language understanding and generation across 44 tasks and over 60 datasets, covering both Modern Standard Arabic and dialectal varieties. Use when the user wants to benchmark on GPTAraEval Benchmark Suite, or asks about evaluating this task. Reports macro-F1.

researchpythongo
0
3
Gpt4v Ocr EvalA

Evaluates the optical character recognition and document understanding capabilities of GPT-4V across multiple tasks including scene text recognition, handwritten text recognition, mathematical expression recognition, and table structure recognition. It probes the model's robustness to different languages, handwriting styles, complex layouts, and input resolutions. Use when the user wants to benchmark on CUTE80, SCUT-CTW1500, Total-Text, WordArt, ReCTS, MLT19, IAM, CASIA-HWDB, CROHME2014, HME1...

researchpythongo
0
3
Gpt4 Based Exact MatchA

Evaluates whether a language model's final numerical answer to a grade-school math word problem matches the ground truth. It uses an external LLM to extract the final answer from the model's generated solution and compares it against the gold answer. Use when the user has predictions and gold and needs to compute GPT4-based-Exact-Match.

researchpythongo
0
3
Gpt3 Few Shot EvalA

Evaluates the few-shot, one-shot, and zero-shot learning capabilities of large autoregressive language models across diverse NLP tasks including language modeling, cloze completion, question answering, translation, and commonsense reasoning. Use when the user wants to benchmark on Penn Tree Bank (PTB), LAMBADA, HellaSwag, StoryCloze 2016, Natural Questions, WebQuestions, TriviaQA, WMT14/WMT16 Translation, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Gpro EvalA

Evaluates large vision-language models on complex mathematical and visual reasoning tasks, measuring both correctness and computational efficiency. It specifically probes the model's ability to avoid excessive chain-of-thought generation (overthinking) by dynamically routing computation between fast perception, slow perception, and slow reasoning paths. Use when the user wants to benchmark on MathVision, MathVerse, MathVista, DynaMath, MM-Vet, or asks about evaluating this task. Reports accur...

researchpythongo
0
3