All authors
qhjqhj00 avatar

Claude Skills by qhjqhj00

github.com/qhjqhj00
7,636 skillsA× 7,623B× 11C× 1D× 12 installs2,582 views
Futurex EvalA

Evaluates LLM agents' ability to forecast real-world future events under uncertainty. It probes reasoning depth, tool-use/search capability, and temporal validity by requiring models to answer dynamic, live-updated questions before event resolution. Use when the user wants to benchmark on FutureX, or asks about evaluating this task. Reports overall_score.

researchpythongo
0
3
Fuxi 2 Weather Forecast EvalA

Evaluates the accuracy and temporal consistency of 1-hourly global weather forecasts up to 90 hours lead time. It probes the model's ability to capture both large-scale atmospheric patterns and fine-scale variability across meteorological, energy, aviation, and marine variables. Use when the user wants to benchmark on ERA5 (2018 testing data), or asks about evaluating this task. Reports RMSE.

datapythontesting
0
3
Fuximt Xxzh EvalA

Evaluates multilingual machine translation capability for Chinese-involved pairs (xx-zh), measuring translation quality across varying levels of parallel data availability. It probes how well models leverage cross-lingual knowledge transfer and handle data scarcity in low-resource settings. Use when the user wants to benchmark on xx-zh translation pairs, or asks about evaluating this task. Reports BLEU.

researchpythongo
0
3
Fwi EvalA

Evaluates the ability of deep learning models to perform full waveform inversion (FWI) by predicting subsurface velocity maps from seismic data under varying source frequencies and locations. It probes generalization across different source configurations and robustness to noise and missing traces. Use when the user wants to benchmark on FWI-F, FWI-L, FWI-FL, or asks about evaluating this task. Reports L2 relative error.

researchpythonperformance
0
3
Fysics EvalA

Evaluates multimodal large language models' ability to perceive, reason about, and generate physical attributes and laws from images, videos, and audio. It probes causal physical reasoning, material property mapping, and cross-modal consistency rather than superficial pattern matching. Use when the user wants to benchmark on FysicsEval, or asks about evaluating this task. Reports average score.

researchpythongo
0
3
G4satbench EvalA

This benchmark evaluates the capability of Graph Neural Networks to solve Boolean satisfiability (SAT) problems. It probes whether GNNs can accurately predict formula satisfiability, generate satisfying variable assignments, and identify unsatisfiable cores, while assessing their ability to learn search heuristics from graph-structured logical representations. Use when the user wants to benchmark on G4SATBench, or asks about evaluating this task. Reports classification accuracy.

researchpythongo
0
3
Gabeorlanski Bc EvalA

Compute gabeorlanski/bc_eval via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of gabeorlanski/bc_eval.

developmentpython
0
3
Gadbench EvalA

Evaluates supervised graph anomaly detection capabilities on static attributed graphs. It benchmarks models across transductive and inductive settings, homogeneous and heterogeneous graph structures, and compares traditional tree ensembles with neighbor aggregation against standard and specialized GNNs. Use when the user wants to benchmark on Reddit, Weibo, Amazon, Yelp, T-Fin, Ellip, Tolo, Quest, DGraph, T-Social, or asks about evaluating this task. Reports AUPRC.

researchpythongo
0
3
Gaeleval EvalA

Evaluates LLMs' morphosyntactic competence, machine translation quality, and culturally grounded question-answering abilities in Scottish Gaelic. It probes how well models handle minority language grammar, idiomatic usage, and domain-specific cultural knowledge without relying on English-centric prompting. Use when the user wants to benchmark on GaelEval, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Gai Nerf EvalA

Evaluates the accuracy and generalization of wireless channel prediction models across diverse indoor environments and frequency bands. It probes the model's ability to predict received signal strength (RSSI) and channel state information (CSI) given spatial coordinates and environmental geometry, while testing robustness to physical scene changes and cross-frequency translation. Use when the user wants to benchmark on Our Own Datasets, Argos channel dataset, NewRF simulated data, or asks abo...

researchpythongo
0
3
Gaia EvalA

GAIA probes the ability of AI assistants to perform real-world, conceptually simple tasks that require multi-step reasoning, tool use, and multi-modal processing. It measures robustness in practical everyday reasoning and factual validation on questions explicitly designed to be outside the model's training data. Use when the user wants to benchmark on GAIA, or asks about evaluating this task. Reports score.

researchpythongo
0
3
Gamayun EvalA

Evaluates multilingual LLM capabilities across general knowledge, reasoning, mathematics, and cultural understanding in English, Russian, and other languages. It probes zero-shot and few-shot performance on standardized benchmarks and custom cultural knowledge tests. Use when the user wants to benchmark on MMLU, GSM8K, MERA, RuBIN, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Game EvalA

Evaluates a vision-action model's ability to predict actions and scale ratios from video game footage, testing in-distribution and out-of-distribution generalization across 2D and 3D games. Use when the user wants to benchmark on Video Games (In-Distribution & OOD), or asks about evaluating this task. Reports Pearson correlation.

researchpythongo
0
3
Game Of 24 EvalA

Evaluates an LLM's ability to perform algorithmic search and recursive reasoning within a single generation window, without external tree search or iterative prompting. It probes systematic exploration, pruning, and backtracking capabilities in a mathematical constraint satisfaction task. Use when the user wants to benchmark on Game of 24, or asks about evaluating this task. Reports Success rate.

researchpythongo
0
3
Gamefactory EvalA

This evaluation protocol assesses a video generation model's ability to follow discrete and continuous action inputs while maintaining semantic alignment with text prompts and preserving the original model's visual domain. It measures action-following accuracy, camera pose consistency, text-video semantic relevance, and overall video generation quality across in-domain and open-domain scenes. Use when the user wants to benchmark on GF-Minecraft, VPT (Find Cave), or asks about evaluating this ...

researchpython
0
3
Gamephysics Video Search EvalA

Evaluates zero-shot video retrieval capability using natural language queries to locate specific objects, compound descriptions, and game physics bugs in unstructured gameplay footage. It probes the model's ability to generalize across diverse open-world game genres and visual styles without fine-tuning. Use when the user wants to benchmark on GamePhysics, or asks about evaluating this task. Reports top-k accuracy.

researchpython
0
3
Gameplayqa EvalA

GameplayQA evaluates multi-modal large language models' ability to understand decision-dense, first-person synchronized multi-video environments. It probes capabilities in agent-state tracking, temporal reasoning, and cross-video event alignment across three cognitive difficulty levels. Use when the user wants to benchmark on GameplayQA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Gamma Glaucoma EvalA

Evaluates multi-modal medical image analysis models for glaucoma staging by jointly processing 2D fundus images and 3D OCT volumes. It probes the model's ability to fuse cross-modality features and correctly classify patients into normal, early, or progressive glaucoma stages. Use when the user wants to benchmark on GAMMA Challenge, or asks about evaluating this task. Reports kappa.

researchpythongo
0
3
Gaoyao EvalA

Evaluates the multilingual and multicultural capabilities of large language models across 26 languages and 51 cultures. It probes cognitive abilities (e.g., reasoning, reading, translation) and cultural understanding (monocultural and cross-cultural contexts) to identify geographical performance disparities and benchmark saturation. Use when the user wants to benchmark on GaoYao, or asks about evaluating this task. Reports accuracy, win_rate.

researchpythongo
0
3
Gap EvalA

Evaluates a model's ability to resolve gendered ambiguous pronouns to their correct antecedent names in natural text. It specifically probes for gender bias and the reliance on syntactic or contextual cues over surface-level heuristics. Use when the user wants to benchmark on GAP, or asks about evaluating this task. Reports F1 score.

researchpythongo
0
3
Gap Overlap Kg EvalA

Evaluates a knowledge graph's ability to perform gap and overlap analysis on life insurance contracts by answering scenario-based competency questions. It probes the system's capacity for structured, evidence-grounded reasoning to determine claim coverage, denial, or non-applicability across heterogeneous contract types. Use when the user wants to benchmark on Insurance Contract KG Benchmark, or asks about evaluating this task. Reports accuracy.

researchpythongit
0
3
Gaps EvalA

Evaluates AI clinicians on clinical reasoning depth, answer completeness, robustness to input perturbations, and safety risk mitigation. It probes how well models retrieve, synthesize, and apply evidence-based guidelines under varying cognitive loads and adversarial conditions. Use when the user wants to benchmark on GAPS-NCCN-NSCLC-preview, or asks about evaluating this task. Reports GAPS score (normalized rubric-based).

researchpythongo
0
3
Gar Bench EvalA

Evaluates a multimodal LLM's ability to perform precise, context-aware visual understanding at the region level. It probes fine-grained perception, compositional reasoning across multiple visual prompts, and detailed localized captioning for both images and videos. Use when the user wants to benchmark on GAR-Bench-VQA, GAR-Bench-Cap, DLC-Bench, Ferret-Bench, MDVP-Bench, LVIS, PACO, VideoRefer-Bench, or asks about evaluating this task. Reports Overall score.

researchpythongo
0
3
Garment Folding EvalA

Evaluates a robot's ability to manipulate deformable fabric garments using a single arm. It probes two capabilities: flattening a crumpled T-shirt to maximize coverage, and folding a flattened T-shirt to match a goal configuration while minimizing wrinkles. Use when the user wants to benchmark on Google Reach T-shirt folding environment, or asks about evaluating this task. Reports max_coverage_pct.

researchpythongo
0
3
Garments2look EvalA

Probes the ability of virtual try-on and image editing models to synthesize high-fidelity, multi-reference outfit images. It evaluates whether models can preserve fine-grained garment details, maintain correct layering orders, and adhere to specific styling techniques while keeping the target person's pose consistent. Use when the user wants to benchmark on Garments2Look, DressCode-MR, or asks about evaluating this task. Reports FID↓.

researchpythongit
0
3
Gas Saturation EvalA

Evaluates a neural operator's ability to predict long-term multiphase flow dynamics (gas saturation and pressure buildup) in porous media using sparse time snapshots. It probes data efficiency, generalization to unseen time steps, and computational resource usage compared to baseline spectral methods. Use when the user wants to benchmark on Synthetic multiphase flow dataset (gas saturation & pressure buildup), or asks about evaluating this task. Reports R^2.

researchpythongit
0
3
Gass T2i Diversity EvalA

Evaluates text-to-image generation models on their ability to produce diverse, high-quality, and semantically aligned images under fixed prompts. It specifically probes disentangled diversity by measuring prompt-dependent semantic variation versus prompt-independent background/style variation. Use when the user wants to benchmark on ImageNet-1K, DrawBench, or asks about evaluating this task. Reports VS.

researchpythongo
0
3
Gaussianvlm EvalA

Evaluates a 3D vision-language model's ability to perform object-centric and scene-centric reasoning tasks, including captioning, question answering, embodied planning, and dialogue. It probes spatial grounding, semantic abstraction, and robust generalization to out-of-domain real-world scene representations. Use when the user wants to benchmark on ScanRefer, ScanQA, Nr3D, SQA3D, 3D-LLM ScanNet subset, ScanNet++ (OOD object counting), or asks about evaluating this task. Reports Exact-match ac...

researchpython
0
3
Gazebo Simulation EvalA

Evaluates an MPC-based autonomous driving controller's ability to perform collision avoidance and lane maneuvers (overtaking, merging, following) in dynamic environments using a physics-based simulator. It tests the controller's real-time feasibility and trajectory smoothness under varying traffic densities. Use when the user wants to benchmark on Gazebo Simulation, or asks about evaluating this task. Reports computation time.

researchpythonangular
0
3
Gazeta Russian Summarization EvalA

This benchmark evaluates the quality of abstractive and extractive text summarization in Russian. It measures how well generated summaries capture the key information and stylistic qualities of the original news articles compared to human-written references. Use when the user wants to benchmark on Gazeta, or asks about evaluating this task. Reports ROUGE.

researchpythongit
0
3
Gbc EvalA

Evaluates the cross-modal alignment and generalization of CLIP models trained on various captioning formats. It probes zero-shot classification, bidirectional image-text retrieval, compositional reasoning, dense semantic segmentation, and fine-grained text-to-image generation control. Use when the user wants to benchmark on ImageNet-1k, Flickr30k, MS-COCO, SugarCrepe, ShareGPT4V-cap100k, ADE20K, DCI, or asks about evaluating this task. Reports SugarCrepe.

ai-agentspythongo
0
3
Gca Tool Use EvalA

Evaluates an LLM's ability to use region-specific climate tools in a multi-step agentic pipeline. It probes structured tool invocation, argument schema adherence, step-wise reasoning, and end-to-end answer accuracy on Gulf-focused climate queries. Use when the user wants to benchmark on GCA-DS, or asks about evaluating this task. Reports AnsAcc.

researchpythongo
0
3
Gcai Constitution EvalA

Evaluates the moral grounding, coherence, fairness, and real-world applicability of AI alignment constitutions through human surveys, alongside the downstream safety alignment and general capabilities of fine-tuned language models. Use when the user wants to benchmark on BABELSCAPE/ALERT, MMLU, Social Bias BBQ, or asks about evaluating this task. Reports 5-point Likert rating.

researchpythongo
0
3
Gcn Node Classification EvalA

Semi-supervised node classification on citation and knowledge graphs. It probes the model's ability to learn graph-structured representations and classify nodes using only a small fraction of labeled examples. Use when the user wants to benchmark on Citeseer, Cora, Pubmed, NELL, or asks about evaluating this task. Reports prediction accuracy.

researchpythongo
0
3
Gdibench EvalA

Evaluates document intelligence by decoupling visual and reasoning complexity into graded difficulty levels (V0–V2, R0–R2). It probes a model’s ability to extract, reason over, and generalize across diverse document types while mitigating catastrophic forgetting during fine-tuning. Use when the user wants to benchmark on GDI-Bench, or asks about evaluating this task. Reports Accuracy / normalized edit distance.

researchpythongo
0
3
Gdro Tabular Imbalance EvalA

Assesses deep learning models' ability to classify highly imbalanced binary tabular data by comparing standard empirical risk minimization against group distributionally robust optimization. Use when the user wants to benchmark on Multiple benchmark imbalanced tabular datasets, or asks about evaluating this task. Reports g-mean.

researchpythonperformance
0
3
Gebench EvalA

Evaluates image generation models' ability to function as dynamic GUI environments, probing temporal coherence, multi-step interaction logic, spatial grounding, and visual fidelity across sequential state transitions. Use when the user wants to benchmark on GEBench, or asks about evaluating this task. Reports GE-Score.

researchpythongo
0
3
GecoA

Evaluates geometric consistency in text-to-video generation by measuring structural and motion coherence across camera trajectories, detecting deformation and occlusion artifacts in static scenes. Use when the user has predictions and gold and needs to compute Fused.

researchpythongo
0
3
Gedit Imgedit Bench EvalA

Evaluates the quality of AI-generated image edits by measuring semantic consistency, perceptual quality, and overall fidelity against source images and edit instructions. Use when the user wants to benchmark on GEdit-Bench, ImgEdit-Bench, or asks about evaluating this task. Reports Overall (GEdit-Bench).

ai-agentspythongo
0
3
Gem Cot Mixed Task EvalA

Evaluates the ability of LLMs to perform zero-shot and few-shot reasoning across a heterogeneous mix of unseen and known task types without manual task-specific prompting. It probes dynamic demonstration routing, clustering-based generalization, and streaming adaptation in mixed-task scenarios. Use when the user wants to benchmark on AQUA-RAT, MultiArith, AddSub, GSM8K, SingleEq, SVAMP, Last Letter Concatenation, Coin Flip, StrategyQA, CSQA, BIG-Bench Hard (BBH), or asks about evaluating this...

researchpythongo
0
3
Gem EvalA

Evaluates natural language generation models across diverse tasks including content planning, surface realization, and communicative goals. It probes lexical similarity, semantic equivalence, faithfulness, and output diversity using both reference-based and reference-free automated metrics. Use when the user wants to benchmark on CommonGen, Czech Restaurant, DART, E2E clean, MLSum, Schema-Guided, ToTTo, XSum, WebNLG, Turk, ASSET, WikiLingua, or asks about evaluating this task. Reports BLEU.

researchpythongo
0
3
Gemex EvalA

Evaluates large vision-language models on chest X-ray diagnosis by testing their ability to answer medical questions, provide textual reasoning, and ground answers to specific visual regions in radiographs. Use when the user wants to benchmark on GEMeX, or asks about evaluating this task. Reports AR-score.

researchpythongo
0
3
Gemini Embedding EvalA

Evaluates the quality of multilingual and code-aware text embeddings across diverse tasks including retrieval, classification, clustering, and cross-lingual matching. It probes the model's ability to generate generalizable representations that perform well across 250+ languages, English, and programming languages. Use when the user wants to benchmark on MMTEB, XTREME-UP, XOR-Retrieve, or asks about evaluating this task. Reports Task Mean.

ai-agentspythongo
0
3
Gemini Robotics 15 EvalA

Evaluates a robot's ability to execute short-horizon and multi-step manipulation tasks across diverse embodiments, environments, and visual/instructional variations. It specifically probes zero-shot cross-embodiment skill transfer and the impact of explicit 'thinking' traces on task progress and success. Use when the user wants to benchmark on Gemini Robotics 1.5 Benchmark, or asks about evaluating this task. Reports progress score.

researchpythongo
0
3
Gen Nerf EvalA

Evaluates the rendering quality and computational efficiency of a generalizable Neural Radiance Field (NeRF) model for novel view synthesis. It measures how accurately the model reconstructs unseen scenes from a few source views, balancing image fidelity against computational cost and hardware throughput. Use when the user wants to benchmark on NeRF Synthetic, LLFF, DeepVoxels, or asks about evaluating this task. Reports PSNR.

researchpythongo
0
3
Gendeg EvalA

Evaluates the out-of-distribution (OoD) generalization and within-distribution performance of All-In-One Image Restoration (AIOR) models across six degradation types (haze, rain, snow, motion blur, raindrop, low-light) when trained with synthetic degradation data. Use when the user wants to benchmark on O-HAZE, LHP, RainDS, RSVD, GoPro, or asks about evaluating this task. Reports LPIPS.

researchpythongo
0
3
Genderbias Vl EvalA

This benchmark probes the gender bias of Large Vision-Language Models (LVLMs) in occupation inference tasks. It uses counterfactual visual question pairs to measure how model predictions change when the perceived gender of a subject is swapped, evaluating both cognitive accuracy and fairness under individual and causal fairness frameworks. Use when the user wants to benchmark on GenderBias-VL, or asks about evaluating this task. Reports Idealized Score (Ipss).

researchpythongo
0
3
Genderpair EvalA

Evaluates gender bias in large language models by measuring the model's preference or generation likelihood across stereotypical versus counterfactual gendered prompts. It specifically probes inclusivity and diversity by including marginalized gender identities such as transgender and non-binary groups. Use when the user wants to benchmark on GenderPair, or asks about evaluating this task. Reports Bias-Pair Ratio.

researchpythongo
0
3
Gene Bench EvalA

This benchmark evaluates how well various deep learning models encode biological knowledge by predicting 312 ground-truth gene properties. It probes capabilities across five domains: genomic, regulatory, localization, biological processes, and protein features, using binary, multi-label, and multi-class classification tasks. Use when the user wants to benchmark on GeneBench, or asks about evaluating this task. Reports AUC.

researchpythongit
0
3
Geneol EvalA

Evaluates training-free sentence embedding quality by aggregating LLM-generated semantic variations. Probes semantic similarity preservation and cross-task robustness without model fine-tuning. Use when the user wants to benchmark on STS benchmark, MTEB, or asks about evaluating this task. Reports Spearman rank correlation (cosine similarity).

researchpythongo
0
3