All authors
qhjqhj00 avatar

Claude Skills by qhjqhj00

github.com/qhjqhj00
7,636 skillsA× 7,623B× 11C× 1D× 10 installs2,408 views
Gromov Wasserstein SimilarityA

Evaluates how well Gromov-Wasserstein distance captures functional similarity between neural network layer representations, enabling the identification of structural transitions and latent sub-networks across varying dimensionalities without task-specific supervision. Use when the user has predictions and gold and needs to compute Gromov-Wasserstein distance.

researchpythongo
0
3
Ground Motion Synthesis EvalA

Evaluates the ability of a generative model to synthesize realistic 3-component broadband ground motion acceleration time histories conditioned on seismic parameters. It probes the model's capacity to match empirical spectral intensities (PSA, FAS, EAS) and capture aleatory variability across different frequency bands and tectonic settings. Use when the user wants to benchmark on BBP dataset, Kik-net dataset, or asks about evaluating this task. Reports Normalized model residual (epsilon).

datapythongo
0
3
Groundcocoa EvalA

Evaluates compositional and conditional reasoning in LLMs by requiring them to match complex, logically constrained user preferences to specific flight booking options. It probes the model's ability to handle interdependent requirements and atypical constraints without external reasoning engines. Use when the user wants to benchmark on GroundCocoa, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Grounded Ecg Understanding EvalA

Evaluates a multimodal LLM's ability to interpret 12-lead ECG signals and images, providing clinically grounded diagnoses, detailed feature annotations, and evidence-based reasoning. It also tests cardiac abnormality detection and automated report generation across multiple public ECG datasets. Use when the user wants to benchmark on MIMIC-IV-ECG, ECG-Bench (PTB-XL, CPSC2018, G12EC, CODE-15%, CSN), PTB-XL Report, ECG-QA, or asks about evaluating this task. Reports DiagnosisAccuracy.

researchpythonexpress
0
3
Grounder EvalA

This benchmark evaluates a model's ability to localize arbitrary natural language phrases within images. It probes phrase grounding capabilities by requiring the model to attend to relevant image regions and select a bounding box that matches the textual description, without relying on explicit bounding box supervision during training. Use when the user wants to benchmark on Flickr 30k Entities, ReferItGame, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Grounding Video Reasoning EvalA

Evaluates video understanding models on physical event reasoning across six domains (gravity, fluids, collisions, deformation, friction, state changes). It probes spatio-temporal grounding by requiring models to predict what happens, when it happens, and where it happens, while measuring robustness to input perturbations like shuffling, ablation, and frame masking. Use when the user wants to benchmark on Physical Video Reasoning Benchmark, or asks about evaluating this task. Reports LGM.

researchpythongo
0
3
Groundingme EvalA

Evaluates multimodal large language models' (MLLMs) visual grounding capabilities across four dimensions: discriminative object distinction, spatial relational understanding, handling occlusion/size constraints, and the ability to reject ungroundable queries. It measures how well models can localize objects in images and whether they hallucinate or correctly refuse impossible requests. Use when the user wants to benchmark on GroundingME, or asks about evaluating this task. Reports Accuracy@0.5.

researchpythongo
0
3
Groundlie360 EvalA

This benchmark evaluates a model's ability to detect and localize multimodal misinformation across text, speech, and video. It probes fine-grained cross-modal reasoning by requiring binary veracity classification, sub-type categorization, and precise grounding of fake content at the token, frame, and bounding-box levels. Use when the user wants to benchmark on GroundLie360, or asks about evaluating this task. Reports F1 score.

researchpythongo
0
3
Groundnext EvalA

Evaluates vision-language models on UI element localization and grounding across desktop, mobile, and web interfaces. It measures how accurately a model can identify and locate specific UI components based on text instructions, and assesses their effectiveness in multi-step agentic tasks. Use when the user wants to benchmark on SSPro, OSW-G, MMB-GUI, SSv2, UI-V, OSWorld-Verified, or asks about evaluating this task. Reports average performance.

researchpythongo
0
3
Groundset EvalA

Probes zero-shot spatial understanding and grounding capabilities of multimodal LLMs on high-resolution remote sensing imagery. Evaluates generalization across captioning, classification, detection, segmentation, and VQA tasks using verified cadastral vector annotations. Use when the user wants to benchmark on GroundSet, or asks about evaluating this task. Reports F1@0.5.

ai-agentspythonexpress
0
3
Group Fairness Reward EvalA

Evaluates whether reward models assign equal average scores to high-quality responses across different demographic/occupational groups. It probes for systematic bias in how models rank expert-written abstracts based on the author's discipline. Use when the user wants to benchmark on arXiv Metadata (Curated), or asks about evaluating this task. Reports Normalized Maximum Group Difference.

researchpythonexpress
0
3
Gru D EvalA

This evaluation probes a model's ability to handle multivariate time series with missing values by jointly learning temporal dependencies and informative missing patterns. It tests classification performance on clinical and synthetic datasets, measuring how well the model exploits masking and time-interval information for early prediction and multi-task diagnosis. Use when the user wants to benchmark on Gesture, PhysioNet Challenge 2012, MIMIC-III, or asks about evaluating this task. Reports ...

researchpythonperformance
0
3
Gsc Speech Commands EvalA

Evaluates keyword spotting models trained on real versus synthetic speech data, measuring how ASR-based filtering of hallucinated synthetic commands affects classification accuracy on the Google Speech Commands dataset. Use when the user wants to benchmark on Google Speech Commands (GSC), or asks about evaluating this task. Reports Accuracy (%).

researchpythongo
0
3
Gscan EvalA

Evaluates systematic generalization in grounded language understanding by testing whether models can interpret natural language commands within dynamic grid-world environments. It probes compositional generalization across novel object properties, navigation directions, contextual size references, action-argument bindings, and adverbial modifiers. Use when the user wants to benchmark on gSCAN, or asks about evaluating this task. Reports exact match accuracy.

researchpythongo
0
3
Gseval Pixel Grounding EvalA

Evaluates a model's ability to perform open-vocabulary, fine-grained pixel grounding by generating accurate segmentation masks from complex, long-form referring expressions across multiple granularities (stuff, part, multi-object, single-object). Use when the user wants to benchmark on GSEval, gRefCOCO, RefCOCOm, RefCOCO, RefCOCOg, or asks about evaluating this task. Reports cIoU / gIoU.

researchpythonexpress
0
3
Gsm8k EvalA

Evaluates a model's ability to perform multi-step arithmetic reasoning by generating natural language solutions to grade school math word problems and verifying their correctness. Use when the user wants to benchmark on GSM8K, or asks about evaluating this task. Reports solve rate.

researchpythongo
0
3
Gsm8k V EvalA

This benchmark evaluates vision-language models' ability to perform multi-step mathematical reasoning using purely visual, comic-style narratives instead of text. It specifically probes challenges in inter-image semantic understanding, object grounding, and extracting numerical relationships from multi-panel visual contexts. Use when the user wants to benchmark on GSM8K-V, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Gsr Bench EvalA

Evaluates multimodal LLMs' ability to understand and disambiguate spatial relations (e.g., on, under, left of, right of, in front of, behind) between objects in images. It isolates spatial reasoning from object grounding by providing depth maps, bounding boxes, and segmentation masks alongside images. Use when the user wants to benchmark on GSR-Bench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Gt23d Bench EvalA

Evaluates the quality and alignment of generated 3D assets against text prompts across multiple dimensions, including textual alignment, texture fidelity, geometry correctness, and multi-view consistency. It measures how well automated metrics correlate with human preferences to provide a reliable assessment of general text-to-3D generation methods. Use when the user wants to benchmark on GT23D-Bench, or asks about evaluating this task. Reports Texture Fidelity.

researchpython
0
3
Gtb Dti EvalA

Evaluates structure-based drug-target interaction (DTI) prediction models on regression (binding affinity) and classification (binding status) tasks across six standard bioinformatics datasets. Use when the user wants to benchmark on DAVIS, KIBA, BindingDB, Human, Cycles, Drugbank, or asks about evaluating this task. Reports PCC, ROC-AUC.

researchpythonperformance
0
3
Gte EvalA

Evaluates the cross-task generalization and retrieval quality of a general-purpose text embedding model across classification, retrieval, clustering, reranking, semantic similarity, summarization, and code search tasks. It measures how well zero-shot and unsupervised embeddings transfer to diverse downstream benchmarks without task-specific fine-tuning. Use when the user wants to benchmark on SST-2, BEIR, MTEB (English subset), CodeSearchNet, or asks about evaluating this task. Reports nDCG@10.

ai-agentspythongo
0
3
Gtpbd EvalA

Evaluates fine-grained agricultural parcel delineation, boundary detection, and cross-domain generalization on high-resolution remote sensing imagery of terraced terrain. It benchmarks semantic segmentation, edge extraction, and parcel extraction models across multiple geographic domains. Use when the user wants to benchmark on GTPBD, or asks about evaluating this task. Reports IoU.

researchpythongo
0
3
Gtpbd Mm EvalA

Evaluates multimodal terraced parcel extraction by measuring pixel-level segmentation accuracy, edge-level boundary recovery, and object-level structural consistency across image-only, image+text, and image+text+DEM input settings. Use when the user wants to benchmark on GTPBD-MM, or asks about evaluating this task. Reports mIoU.

researchpythongo
0
3
Gtsinger EvalA

Evaluates singing voice synthesis models on technique-controllable generation, singer similarity, and audio quality across multiple languages and vocal techniques. It probes the model's ability to accurately control specific singing techniques (e.g., vibrato, mixed voice) while maintaining naturalness and timbre fidelity. Use when the user wants to benchmark on GTSinger, or asks about evaluating this task. Reports MOS-Q.

researchpythongo
0
3
Guacamol Molecule Generation EvalA

This evaluation probes a molecular generative model's ability to produce chemically valid, diverse, and structurally realistic molecules. It measures how well the generated molecules match the physicochemical property distributions of real compounds while maintaining high novelty and uniqueness rates. Use when the user wants to benchmark on GuacaMol benchmark suite, or asks about evaluating this task. Reports KL divergence.

researchpythongo
0
3
Guardrail Robustness EvalA

Evaluates the robustness and generalization of LLM safety guardrails against adversarial jailbreak prompts, measuring their ability to correctly classify harmful vs. benign inputs under both known benchmark distributions and novel, contextually framed attacks. Use when the user wants to benchmark on Adversarial Guardrail Benchmark, or asks about evaluating this task. Reports Overall Accuracy.

researchpythongo
0
3
Guardreasoner Vl EvalA

Evaluates the ability of vision-language models to detect harmful content in user prompts and AI responses across text, image, and multimodal inputs. It probes safety alignment and reasoning capabilities by measuring classification accuracy on diverse safety benchmarks. Use when the user wants to benchmark on ToxicChat, HarmBench, OpenAIModeration, AegisSafetyTest, SimpleSafetyTests, WildGuardTest, HarmImageTest, SPA-VL-Eval, SafeRLHF, BeaverTails, XSTestResponse, or asks about evaluating thi...

researchpythongo
0
3
Gui Agent Halluc EvalA

This protocol evaluates GUI agents on visual grounding, action execution, and hallucination rates across mobile, desktop, and web interfaces. It measures how well models localize UI elements, execute multi-step tasks under varying instruction granularities, and avoid perception or reasoning errors. Use when the user wants to benchmark on ScreenSpot-V2, ScreenSpot-Pro, AndroidControl, GUI-Odyssey, or asks about evaluating this task. Reports Action Type Accuracy (Type), Grounding Accuracy (GR),...

researchpythongo
0
3
Gui Agent Kv EvalA

Evaluates the accuracy and efficiency of GUI agents under varying KV cache compression budgets across visual grounding, offline action prediction, and online task completion benchmarks. Use when the user wants to benchmark on ScreenSpotV2, ScreenSpot-Pro, AndroidControl, Multimodal-Mind2Web, AgentNetBench, OSWorld-Verified, or asks about evaluating this task. Reports step accuracy.

researchpythongo
0
3
Gui Ceval EvalA

Evaluates multimodal large language models and agents on Chinese mobile GUI interaction tasks. It probes atomic capabilities like visual perception, grounding, and planning, as well as end-to-end execution reliability in both offline simulation and real-device online environments. Use when the user wants to benchmark on GUI-CEval, or asks about evaluating this task. Reports Online Agent success rate.

researchpython
0
3
Gui Grounding Agent EvalA

Evaluates a GUI agent's ability to precisely locate UI elements (grounding) and execute multi-step tasks in real-world desktop/web environments. It probes spatial reasoning, text/icon matching, and long-horizon planning robustness without lookahead. Use when the user wants to benchmark on ScreenSpot-Pro, ScreenSpot-V2, OSWorld-G, OSWorld, WindowsAgentArena, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Gui Grounding EvalA

Evaluates a model's ability to locate specific UI elements on screenshots based on natural language instructions. It measures both the recall of candidate generation and the precision of visual discrimination to select the correct bounding box. Use when the user wants to benchmark on MMBench-GUI, ScreenSpot-Pro, UI-Vision, ScreenSpot-v2, UI-I2E-Bench, OSWorld-G, or asks about evaluating this task. Reports Top-1 Accuracy.

researchpythongo
0
3
Gui Grounding Navigation EvalA

This evaluation probes a GUI agent's ability to localize UI elements via grounding and execute multi-step navigation tasks across mobile, web, and desktop platforms. It measures spatial perception, action planning consistency, and cross-platform generalization under both offline and online interaction settings. Use when the user wants to benchmark on ScreenSpot-V2, ScreenSpot-Pro, AndroidControl, AndroidWorld, ChiM-Nav, Ubu-Nav, or asks about evaluating this task. Reports success rate, Step S...

researchpythongo
0
3
Guicourse Gui Nav EvalA

Evaluates vision-language models' ability to navigate graphical user interfaces by predicting correct action types and precise screen coordinates. It probes OCR, pixel-level grounding, and multi-step task planning across web and mobile environments. Use when the user wants to benchmark on GUIAct, Mind2Web, AITW, or asks about evaluating this task. Reports StepSR.

researchpythongo
0
3
Guide Research Idea EvalA

This evaluation probes a system's ability to act as a scientific advisor by predicting whether research hypotheses will be accepted at a top-tier AI conference. It measures alignment with expert peer-review decisions using ranking-based precision and recall metrics on a held-out set of conference submissions. Use when the user wants to benchmark on ICLR 2025 Submissions Test Set, or asks about evaluating this task. Reports Top-30% Precision.

researchpythongo
0
3
Guiodyssey EvalA

Evaluates multimodal agents' ability to perform cross-app GUI navigation on mobile devices by predicting correct UI actions based on screen states and task instructions. It probes spatial reasoning, action planning, and the model's capacity to leverage historical context across multiple applications. Use when the user wants to benchmark on GUIOdyssey, or asks about evaluating this task. Reports Action Matching Score (AMS).

researchpythongo
0
3
Guydav Restrictedpython Code EvalA

Compute guydav/restrictedpython_code_eval via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of guydav/restrictedpython_code_eval.

developmentpython
0
3
Gw Glitch Mitigation EvalA

Evaluates the ability of a joint signal-glitch model to accurately recover compact binary coalescence parameters and reconstruct glitch waveforms when they overlap in real gravitational-wave detector data. It compares a baseline model ignoring glitches against a full model that jointly fits both components. Use when the user wants to benchmark on LIGO O3 data segments, or asks about evaluating this task. Reports mismatch.

researchpythongo
0
3
Gw Reconstruction EvalA

Evaluates the fidelity of model-agnostic Bayesian waveform reconstruction for unbound binary black hole fly-bys across different detector noise environments and frame function parameterizations. It also quantifies the astrophysical detection sensitivity and expected event rates for current and next-generation interferometers. Use when the user wants to benchmark on Simulated Hyperbolic BBH Encounters, or asks about evaluating this task. Reports overlap.

researchpythongo
0
3
Gwlans EvalA

Predicts target words in computer-aided translation based on source sentences, translation context (prefix, suffix, zero, bidirectional), and human-typed characters. It probes the model's ability to handle discontinuous context and weak positional information in real-world CAT scenarios. Use when the user wants to benchmark on GWLAN Benchmark, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Gwrfrf Spatial Spectrum EvalA

This benchmark evaluates a model's ability to synthesize accurate spatial radio-frequency spectra at target transmitter locations using neighboring spectra and scene geometry. It probes both single-scene prediction accuracy and cross-scene generalization capabilities in wireless propagation environments. Use when the user wants to benchmark on RFID Dataset, MATLAB Dataset, or asks about evaluating this task. Reports MSE.

researchpython
0
3
Gwtc2 Bbh Mass Distribution EvalA

Evaluates the ability of semi-parametric and parametric models to recover the astrophysical primary mass distribution of binary black holes from gravitational wave observations, specifically testing for features like the ~35 M⊙ peak and low-mass structure. Use when the user wants to benchmark on GWTC-2 catalog, or asks about evaluating this task. Reports marginal likelihood.

researchpythongo
0
3
Gwtc3 Spin Population EvalA

Evaluates the ability of population inference models to explain the observed spin distributions of binary black holes in the GWTC-3 catalog. It probes how well different astrophysical formation scenarios (e.g., field vs. dynamical assembly, zero-spin subpopulations) fit the gravitational-wave data. Use when the user wants to benchmark on GWTC-3, or asks about evaluating this task. Reports Bayes factor ($\mathcal{B}$).

researchpythongit
0
3
Gym V EvalA

Evaluates agentic vision models on zero-shot generalization across 179 procedurally generated environments spanning 10 domains. It measures task completion via answer correctness for single-turn interactions and cumulative performance via normalized episodic return for multi-turn interactions. Use when the user wants to benchmark on Gym-V, or asks about evaluating this task. Reports normalized episodic return.

researchpythongo
0
3
H Sinn Turbulence EvalA

Evaluates a Convolutional Autoencoder's ability to compress and reconstruct geophysical turbulence fields while preserving high-order statistical moments. It specifically probes the model's capacity to capture non-Gaussian, intermittent structures like extreme vertical drafts without degrading point-wise accuracy. Use when the user wants to benchmark on Stratified turbulence simulation data, or asks about evaluating this task. Reports MAPE on kurtosis ($K_w$).

datapythontesting
0
3
H2seqrec EvalA

Evaluates sequential recommendation models on predicting the next item a user will interact with, capturing temporal dynamics and handling sparse user-item interactions. Use when the user wants to benchmark on AMT, Goodreads, or asks about evaluating this task. Reports HR@K, NDCG@K.

researchpythongo
0
3
H2vu Benchmark EvalA

Evaluates multimodal large language models on hierarchical and holistic video understanding, specifically probing temporal reasoning, countercommonsense comprehension, trajectory state tracking, and first-person streaming video analysis. Use when the user wants to benchmark on H²VU, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
H3wb EvalA

Evaluates 3D whole-body human pose estimation and lifting capabilities. It probes a model's ability to reconstruct 133-keypoint 3D skeletons from complete 2D poses, occluded/incomplete 2D poses, or monocular RGB images, with specific focus on body, face, and hand regions. Use when the user wants to benchmark on H3WB, or asks about evaluating this task. Reports MPJPE.

researchpythongit
0
3
Habibi Tts EvalA

Evaluates zero-shot and unified-dialectal text-to-speech synthesis across multiple Arabic dialects. It measures transcription accuracy, speaker similarity, and audio naturalness to assess how well a model preserves dialectal features and voice identity without dialect-specific fine-tuning. Use when the user wants to benchmark on Habibi Benchmark, or asks about evaluating this task. Reports WER-O.

researchpython
0
3
Habitat Benchmark EvalA

Evaluates a robot's ability to perform long-horizon mobile manipulation and object rearrangement tasks in simulated environments. It probes hierarchical planning, whole-body continuous control, and robust recovery from failures across multi-step subtask sequences. Use when the user wants to benchmark on Habitat Benchmark, or asks about evaluating this task. Reports completion rate.

researchpythongo
0
3