All authors
qhjqhj00 avatar

Claude Skills by qhjqhj00

github.com/qhjqhj00
7,636 skillsA× 7,623B× 11C× 1D× 10 installs2,408 views
Gptscore EvalA

Evaluates the correlation between automated scoring functions (GPTScore variants) and human judgments across multiple text generation tasks. It probes the ability of instruction-based LLMs to serve as training-free, customizable evaluators that align with human preference. Use when the user wants to benchmark on SummEval, RealSumm, NEWSROOM, QXSUM, MQM-2020, BAGEL, SFRES, FED, or asks about evaluating this task. Reports Spearman correlation.

researchpython
0
3
Gpu Inference Benchmark EvalA

Evaluates GPU inference performance across different neural network models, numerical precision modes, and batch sizes. It measures how architectural differences and execution parallelism impact throughput, latency, and memory utilization under production-like conditions. Use when the user wants to benchmark on ResNet models (ResNet-18, ResNet-50, ResNet-101) with synthetic inputs, or asks about evaluating this task. Reports throughput (images/sec).

researchpythongit
0
3
Gpu Memory Co Optimization EvalA

Evaluates the trade-off between system memory footprint and task latency when co-executing multiple workloads under different integrated CPU/GPU memory management policies on embedded platforms. It measures how strategically assigning Device, Managed, or Host-Pinned memory policies affects peak memory consumption, average GPU execution time, and overall GPU utilization during multitasking. Use when the user wants to benchmark on Rodinia Benchmark Suite (subset), DJI Drone Object Detection, Au...

researchpythonperformance
0
3
Gpu Power Cap EvalA

Evaluates the performance and power efficiency trade-offs of NVIDIA H100 and H200 GPUs under varying power caps, isolating compute-bound (DGEMM) and memory-bound (STriad) workloads to analyze architectural scaling and frequency throttling dynamics. Use when the user wants to benchmark on cuBLAS DGEMM, TheBandwidthBenchmark (STriad kernel), or asks about evaluating this task. Reports Throughput (TFlop/s or TB/s).

researchpythonnode
0
3
Gqa EvalA

Evaluates visual reasoning and compositional question answering on real-world images. It probes a model's ability to understand scene relationships, answer multi-step questions, and maintain logical consistency across related queries. Use when the user wants to benchmark on GQA, or asks about evaluating this task. Reports Accuracy.

researchpythonrust
0
3
Grables EvalA

Evaluates whether tabular models can capture extension-sensitive inter-row dependencies (e.g., global counts, overlaps) compared to graph-based message-passing models, and tests if hybrid approaches combining tabular features with graph-derived representations improve performance. Use when the user wants to benchmark on Synthetic transactions dataset, Retail transaction dataset, relbench-trial, or asks about evaluating this task. Reports ROC-AUC.

researchpythongo
0
3
Gracore EvalA

Evaluates large language models' ability to comprehend and reason over graph structures presented as textual descriptions. It probes capabilities ranging from basic graph understanding and semantic reasoning to complex graph theory reasoning across pure and heterogeneous graphs. Use when the user wants to benchmark on GraCoRe, or asks about evaluating this task. Reports score.

researchpythongo
0
3
Grad Tts EvalA

Evaluates text-to-speech synthesis quality, inference efficiency, and probabilistic modeling accuracy of a diffusion-based model. It probes the trade-off between synthesis fidelity and computational cost by varying reverse diffusion steps, and measures human-perceived audio quality against strong baselines. Use when the user wants to benchmark on LJSpeech, or asks about evaluating this task. Reports MOS.

researchpythongo
0
3
Gradsafe Jailbreak Detection EvalA

Evaluates the ability to detect jailbreak or unsafe prompts in LLM inputs using gradient-based analysis. It probes zero-shot and adapted detection capabilities against established moderation APIs and LLM-based detectors. Use when the user wants to benchmark on ToxicChat, XSTest, or asks about evaluating this task. Reports AUPRC.

researchpythongo
0
3
Gram Dti EvalA

Evaluates multimodal drug-target interaction prediction and zero-shot retrieval capabilities across multiple benchmark datasets and cold-start scenarios. Use when the user wants to benchmark on Activation, Yamanishi_08, Hetionet, Inhibition, or asks about evaluating this task. Reports AUPR.

researchpythonperformance
0
3
Granger Causal Inference EvalA

Probes a model's ability to identify causal genomic regulatory relationships (ATAC-seq peaks to RNA-seq genes) using temporal causal inference on single-cell multimodal data. It evaluates how well predicted peak-gene associations align with independent biological proxies like eQTLs and chromatin interactions, testing robustness to high-dimensional sparsity and partial temporal orderings. Use when the user wants to benchmark on sci-CAR, SNARE-seq, SHARE-seq, or asks about evaluating this task....

researchpythongo
0
3
Granite R2 Retrieval EvalA

This evaluation protocol assesses the retrieval and reranking capabilities of encoder-based embedding models across diverse domains including general text, code, long documents, tables, and multi-turn conversations. It also measures encoding speed to evaluate efficiency in large-scale document ingestion pipelines. Use when the user wants to benchmark on MTEB-v2, BEIR, COIR, MLDR, LongEmbed, Table IR, MT-RAG, IBM Documentation, Miracl, or asks about evaluating this task. Reports NDCG@10.

ai-agentspythongo
0
3
Granular Change AccuracyA

Evaluates Dialogue State Tracking (DST) models by measuring performance based on per-turn belief state changes rather than raw slot or turn-level accuracy. It aims to provide partial credit for partially correct predictions and reduce bias from error timing and distribution across dialogue turns. Use when the user has predictions and gold and needs to compute Granular Change Accuracy.

researchpythongo
0
3
Graph Alignment EvalA

Evaluates a model's ability to perform structural graph alignment by predicting a node-to-node correspondence between two graphs that maximizes shared edges. It probes the model's capacity for equivariant representation learning and combinatorial optimization on graph structures. Use when the user wants to benchmark on Graph Alignment Benchmark (Synthetic & Real-world), or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Graph Classification Accuracy EvalA

Evaluates the ability of graph representation models to correctly classify entire graphs based on their structural topology and, optionally, node or edge attributes. It probes whether local structural summaries or complex neural architectures can capture discriminative patterns for tasks like social network or chemical compound categorization. Use when the user wants to benchmark on IMDB BINARY, IMDB MULTI, COLLAB, REDDIT BINARY, REDDIT 5K, REDDIT 12K, ENZYMES, PROTEINS, D&D, MUTAG, PTC, NCI1...

researchpythongo
0
3
Graph Classification Tfgw EvalA

Evaluates the ability of graph neural networks and optimal transport-based methods to classify graphs by learning discriminative representations that capture both structural and feature dissimilarities. It probes expressiveness beyond the Weisfeiler-Lehman test and generalization on heterogeneous real-world graph structures. Use when the user wants to benchmark on 4-CYCLES, SKIP-CIRCLES, MUTAG, PTC, ENZYMES, PROTEIN, NCI1, IMDB-B, IMDB-M, COLLAB, or asks about evaluating this task. Reports ac...

researchpythongo
0
3
Graph Counterfactual Fairness EvalA

Evaluates graph neural networks for node classification fairness by measuring prediction accuracy alongside statistical fairness metrics (demographic parity, equal opportunity) and a novel graph counterfactual fairness metric that quantifies how much node predictions change when sensitive attributes of the node and its neighbors are perturbed. Use when the user wants to benchmark on Synthetic, Bail, Credit, or asks about evaluating this task. Reports δ_CF.

researchpythonnode
0
3
Graph Gen Benchmark EvalA

This benchmark evaluates how effectively graph generative models can produce synthetic graphs that serve as reliable proxies for benchmarking Graph Neural Networks. It measures the fidelity of generated graphs by comparing GNN performance metrics trained on synthetic data against those trained on the original real-world graphs. Use when the user wants to benchmark on Cora, Citeseer, Pubmed, AmazonC, AmazonP, MS CS, MS Physic, or asks about evaluating this task. Reports MSE.

researchpythonnode
0
3
Graph To Vision EvalA

This benchmark probes a vision-language model's ability to jointly interpret and reason across multiple related graph images. It specifically tests cross-modal integration, structural understanding, and instruction-following accuracy when processing homogeneous and heterogeneous graph groupings. Use when the user wants to benchmark on Graph-to-Vision Benchmark, or asks about evaluating this task. Reports instruction-following accuracy.

researchpythonnode
0
3
Graph2vec EvalA

Evaluates the ability of graph representation learning methods to capture structural equivalence for downstream graph classification and clustering tasks. It probes whether learned embeddings can effectively distinguish between different graph classes or group structurally similar graphs without explicit supervision. Use when the user wants to benchmark on Benchmark Graph Classification Datasets (MUTAG, PTC, PROTEINS, NCI1, NCI109), Android Malware Detection Dataset, AMD Malware Clustering Da...

researchpythongo
0
3
Graphcheck Factcheck EvalA

This evaluation probes a model's ability to perform multihop fact-checking over long-form documents and open-domain QA contexts. It measures how well the system identifies factual inconsistencies or supports claims by reasoning over complex, lengthy grounding texts across general and medical domains. Use when the user wants to benchmark on AggreFact-CNN, AggreFact-Xsum, Summeval, ExpertQA, COVID-Fact, SCIFact, PubHealth, or asks about evaluating this task. Reports balanced accuracy.

researchpythongo
0
3
Graphfusionsbr EvalA

Evaluates session-based recommendation systems by predicting the next item in a user's interaction sequence. It probes the model's ability to capture high-order item relationships and leverage external knowledge graphs for accurate, context-aware ranking. Use when the user wants to benchmark on Tmall, RetailRocket, KKBox, or asks about evaluating this task. Reports P@10.

researchpythongo
0
3
Graphgen EvalA

Evaluates the ability of LLMs to answer knowledge-intensive questions across atomic, aggregated, and multi-hop reasoning scenarios in agricultural, medical, and general domains. It measures how well supervised fine-tuning with synthetic knowledge-graph data improves closed-book QA performance. Use when the user wants to benchmark on SeedEval, PQArefEval, HotpotEval, or asks about evaluating this task. Reports ROUGE-F.

researchpythonaws
0
3
Graphic Design Bench EvalA

Evaluates AI systems' ability to perceive, reason about, and generate professional graphic design artifacts across layout, typography, vector graphics, template semantics, and animation. It probes multi-constraint satisfaction through design-native metrics measuring spatial accuracy, perceptual quality, and semantic alignment. Use when the user wants to benchmark on LICA layered-composition dataset, or asks about evaluating this task. Reports mIoU.

designpythongo
0
3
Graphlog EvalA

Evaluates the ability of Graph Neural Networks to induce, compose, and generalize logical rules across synthetic knowledge graphs. It probes relational reasoning, multi-task learning capacity, and catastrophic forgetting in continual learning settings. Use when the user wants to benchmark on GraphLog, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Graphood Drugood EvalA

Evaluates graph neural networks' out-of-distribution (OOD) generalization across synthetic, image, molecular, and text graph datasets. It measures how well models maintain performance when tested on domain-shifted splits (e.g., different graph sizes or molecular scaffolds) compared to in-distribution data. Use when the user wants to benchmark on GraphOOD & DrugOOD, or asks about evaluating this task. Reports ROC-AUC, Accuracy.

researchpythongo
0
3
Graphpb Mos EvalA

Evaluates the naturalness and prosody quality of synthesized Chinese speech by measuring how closely the generated audio matches human-like pausing and rhythm. It probes the model's ability to capture hierarchical syntactic-semantic dependencies for prosody boundary prediction in text-to-speech systems. Use when the user wants to benchmark on Databaker dataset, or asks about evaluating this task. Reports MOS.

researchpythonphp
0
3
Graphrag Bench EvalA

Evaluates Graph Retrieval-Augmented Generation (GraphRAG) frameworks against vanilla RAG across fact retrieval, complex reasoning, contextual summarization, and creative generation tasks. It measures generation quality, retrieval effectiveness, graph structural complexity, and computational efficiency to determine when graph-based retrieval provides measurable benefits over dense vector retrieval. Use when the user wants to benchmark on Novel Dataset, Medical Dataset, or asks about evaluating...

ai-agentspythongo
0
3
Graphrec EvalA

Evaluates the predictive accuracy of social recommendation models by forecasting user-item ratings. It jointly leverages user-item interaction graphs and user-user social graphs to learn co-embeddings, testing the model's ability to integrate heterogeneous social tie strengths and opinion signals into rating prediction. Use when the user wants to benchmark on Ciao, Epinions, or asks about evaluating this task. Reports RMSE.

researchpythongo
0
3
Graphsum Mds EvalA

Evaluates multi-document summarization performance and model explainability by comparing sentence vs. paragraph inputs and analyzing how attention weights correlate with reference summary similarity to reveal positional bias. Use when the user wants to benchmark on MultiNews, WikiSum, or asks about evaluating this task. Reports ROUGE-F (1/2/L).

datapythonperformance
0
3
Graphwalker EvalA

Evaluates a graph-guided in-context learning framework for clinical reasoning on electronic health records. It probes the model's ability to select non-redundant, interacting demonstrations based on patient semantic graphs and information gain signals to improve prediction accuracy on mortality, length-of-stay, readmission, and clinical QA tasks. Use when the user wants to benchmark on MIMIC-III, MIMIC-IV, CMB, MedQA, CMB-clin, or asks about evaluating this task. Reports AUROC, AUPRC.

ai-agentspythongo
0
3
Grasp EvalA

Evaluates multimodal language models' ability to understand language grounding and intuitive physics principles through video-based question answering. It probes capabilities like object detection, feature recognition, and physical plausibility reasoning using simulated Unity environments. Use when the user wants to benchmark on GRASP, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Grasp Pruning EvalA

Evaluates the test accuracy of single-shot pruning methods at initialization on image classification tasks. It measures how well a pruned sub-network can be trained and generalizes compared to baselines like SNIP and random pruning. Use when the user wants to benchmark on CIFAR-10, CIFAR-100, Tiny-ImageNet, ImageNet, or asks about evaluating this task. Reports test accuracy.

researchpythongo
0
3
Grasp Sparql EvalA

This evaluation probes an LLM's ability to generate correct SPARQL queries from natural language questions across diverse knowledge graphs. It measures how well the model can navigate graph structures, handle complex queries, and produce executable results that match ground-truth answers. Use when the user wants to benchmark on WebQuestionsSP (WQSP), ComplexWebQuestions (CWQ), QALD-7, QALD-10, SPINACH, WikiWebQuestions (WWQ), or asks about evaluating this task. Reports F1-score.

researchpythonperformance
0
3
Grasp Success EvalA

This benchmark evaluates the functional impact of 6D object pose estimation and 3D mesh reconstruction methods on robotic grasping performance. It measures how geometric inaccuracies and spatial pose errors propagate to affect the success rate of physics-based grasping attempts in simulation. Use when the user wants to benchmark on YCB-Video (YCB-V), or asks about evaluating this task. Reports grasping success.

researchpythonperformance
0
3
Graspclutter6d EvalA

Evaluates robotic perception and manipulation capabilities in highly cluttered, real-world environments. It benchmarks instance segmentation, 6D object pose estimation, and 6-DoF grasp detection under varying levels of occlusion and scene complexity. Use when the user wants to benchmark on GraspClutter6D, or asks about evaluating this task. Reports Grasp Success Rate (GSR).

researchpythongo
0
3
Gravity Inversion EvalA

Evaluates a 3D gravity inversion framework's ability to recover subsurface density structures from gravity data, measuring both computational efficiency (scaling and speedup) and geological accuracy (recovery of synthetic anomalies and fit to field observations). Use when the user wants to benchmark on Synthetic and Field Gravity Inversion Benchmarks, or asks about evaluating this task. Reports speedup.

researchpythonangular
0
3
Grecx EvalA

Evaluates the recommendation accuracy and inference efficiency of GNN-based and matrix factorization models on real-world interaction datasets. It specifically probes how well models perform under standardized evaluation conditions that account for varying negative sampling strategies and representation dimensions. Use when the user wants to benchmark on yelp2018, gowalla, amazon-book, or asks about evaluating this task. Reports NDCG@20.

researchpythongo
0
3
Greek Llm Benchmark EvalA

Evaluates open-source (Llama-70b) and closed-source (GPT-4o mini) LLMs across seven distinct NLP tasks in Modern Greek. It probes capabilities in classification, sequence labeling, text generation, and machine translation to assess model performance in a lesser-resourced language setting. Use when the user wants to benchmark on SemEval-2020 Task 12 (OffensEval-2020 Greek), Greek Native Corpus (GNC), Global Voices Greek MT Corpus, Areios Pagos Legal Summarization Corpus, University Help Desk I...

ai-agentspythongo
0
3
Greekbarbench EvalA

This benchmark probes free-text legal reasoning and multi-hop statutory citation using Greek Bar exam questions. It requires models to analyze case facts, cite relevant Greek legal articles, and produce open-ended legal analysis. Performance is measured across three dimensions: factual accuracy, correct statutory citation, and quality of legal reasoning. Use when the user wants to benchmark on GreekBarBench, or asks about evaluating this task. Reports Mean.

researchpythongo
0
3
Greekmmlu EvalA

This benchmark evaluates large language models' ability to answer multiple-choice questions across 45 diverse academic, professional, and governmental subjects in Greek. It specifically probes native language fluency, cultural grounding, and domain-specific knowledge retention under zero-shot and few-shot prompting conditions. Use when the user wants to benchmark on GreekMMLU, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Greenphase EvalA

Evaluates the capability of seismic models to detect earthquake events and precisely pick P- and S-wave arrival times from continuous three-component waveform data. It measures both detection accuracy and temporal picking precision under a fixed time-tolerance constraint. Use when the user wants to benchmark on STEAD (Stanford Earthquake Dataset), or asks about evaluating this task. Reports F1.

researchpythongo
0
3
Grefcoco EvalA

Tests a model's ability to ground natural language expressions that may refer to zero, one, or multiple objects in an image. The model must output a corresponding set of bounding boxes rather than a single box, evaluating its capacity for multi-target and no-target referring expression comprehension. Use when the user wants to benchmark on gRefCOCO, or asks about evaluating this task. Reports set-matching accuracy.

researchpythonexpress
0
3
Gres EvalA

Evaluates a model's ability to segment arbitrary numbers of target objects (including zero) in an image based on a natural language expression. It probes multi-target localization, no-target rejection, and robustness to complex linguistic structures like counting and compound relations. Use when the user wants to benchmark on gRefCOCO, or asks about evaluating this task. Reports generalized IoU (gIoU).

researchpythongo
0
3
Gridnethd EvalA

Evaluates 3D semantic segmentation capabilities for power line infrastructure using multi-modal LiDAR and image data. It probes a model's ability to accurately classify geometric and visual features into 11 distinct classes, including critical assets like pylons, cables, and insulators. Use when the user wants to benchmark on GridNet-HD, or asks about evaluating this task. Reports mIoU.

researchpythongo
0
3
Gridtopix EvalA

Evaluates the ability of embodied agents to learn long-horizon planning and navigation tasks using only terminal rewards, and tests the effectiveness of distilling policies from simplified gridworld experts into visual agents via imitation learning. Use when the user wants to benchmark on PointGoal Navigation, Furniture Moving, 3 vs. 1 Football, or asks about evaluating this task. Reports SPL.

researchpythongo
0
3
Griffin EvalA

Evaluates aerial-ground cooperative 3D object detection and multi-object tracking in simulated urban environments. Probes cross-view feature alignment, occlusion handling, and communication efficiency under dynamic drone altitudes. Use when the user wants to benchmark on Griffin, or asks about evaluating this task. Reports AP, AMOTA.

researchpythonangular
0
3
Grin Drive EvalA

Evaluates a model's ability to generate polygon segmentation masks for generalized referring navigable regions in autonomous driving scenes based on natural language navigation instructions. It specifically probes handling of single-target, multi-target, and no-target scenarios without biasing towards trivial existence predictions. Use when the user wants to benchmark on GRiN-Drive, or asks about evaluating this task. Reports msIoU.

researchpythongo
0
3
Grinsztajn45 EvalA

Evaluates the predictive performance of deep neural networks on tabular data across classification and regression tasks, comparing them against tree-based models and other DNNs. It probes how well architectures handle numerical-only versus heterogeneous (numerical + categorical) features at different dataset scales. Use when the user wants to benchmark on Grinsztajn45, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Grl Perturbation Sensitivity EvalA

Evaluates graph neural network robustness and feature/structure reliance by measuring performance degradation under 13 structured perturbations to node features and graph topology. It classifies datasets based on their sensitivity profiles to structural vs. feature information. Use when the user wants to benchmark on GRL Benchmark Collection (49 datasets), or asks about evaluating this task. Reports sensitivity_profile.

researchpythonnode
0
3