All authors
qhjqhj00 avatar

Claude Skills by qhjqhj00

github.com/qhjqhj00
7,636 skillsA× 7,623B× 11C× 1D× 10 installs2,093 views
Unpie EvalA

Assesses multimodal language models' ability to resolve lexical ambiguity in puns using visual context. It probes visual-textual alignment, multimodal literacy, and the capacity to disambiguate or reconstruct ambiguous text when provided with explanatory or disambiguating images. Use when the user wants to benchmark on UNPIE, or asks about evaluating this task. Reports exact-match accuracy.

researchpythongo
0
3
Unrealzoo EvalA

Evaluates embodied AI agents' capabilities in complex, photo-realistic 3D open-world environments. Specifically probes visual navigation on unstructured terrain, active visual tracking across diverse scenes, and social tracking under dynamic distractions, varying morphologies, and different control frequencies. Use when the user wants to benchmark on UnrealZoo, or asks about evaluating this task. Reports Success Rate (SR).

researchpythongo
0
3
Unsafebench EvalA

Evaluates the effectiveness of image safety classifiers in detecting various unsafe content categories across real-world and AI-generated images. It also probes classifier robustness to distribution shifts caused by artistic representations and grid layouts in AI-generated content. Use when the user wants to benchmark on UnsafeBench, or asks about evaluating this task. Reports F1-Score.

researchpythongo
0
3
Unseen Object 6d Pose EvalA

Evaluates a model's ability to estimate the 6D pose (rotation and translation) of novel, unseen 3D objects in real-world scenes without retraining, using only their mesh models and partial RGBD inputs. It specifically probes robustness to pose ambiguity, partial observability, and real-world noise. Use when the user wants to benchmark on GraspNet-1Billion, YCB-Video, or asks about evaluating this task. Reports IADD.

researchpythongo
0
3
Unseen Speaker Ser EvalA

Evaluates a model's ability to recognize emotions in speech from speakers it has never encountered during training. It probes cross-speaker generalization and robustness to acoustic variability across multiple languages and recording conditions. Use when the user wants to benchmark on CREMA-D, IEMOCAP, RAVDESS, EmoDB, CaFE, BhavVani, or asks about evaluating this task. Reports WF1.

researchpythongo
0
3
Unstereo EvalA

Evaluates whether language models exhibit gender bias when processing sentence pairs that have been filtered to remove explicit gendered language and stereotypical co-occurrences. It measures the model's ability to generate gender-neutral completions and checks for systematic preference toward male or female pronouns in stereotype-free contexts. Use when the user wants to benchmark on USE-5, USE-10, USE-20, WB (Winobias), WG (Winogender), or asks about evaluating this task. Reports US fairnes...

researchpythongo
0
3
Unsupervised Lexical Semantic Change EvalA

This benchmark evaluates a model's ability to detect unsupervised lexical semantic change across diachronic corpus pairs. It probes two capabilities: binary classification of whether a word's sense has been gained or lost, and ranking the intensity of semantic change relative to a gold standard. Use when the user wants to benchmark on SemEval-2020 Task 1, or asks about evaluating this task. Reports accuracy, Spearman’s rank-order correlation coefficient.

researchpythongo
0
3
Unsupervised Near Duplicate EvalA

Evaluates the ability of image descriptors to distinguish near-duplicate image pairs from non-duplicates under extreme specificity constraints, simulating large-scale forensic or fraud detection scenarios. Use when the user wants to benchmark on MFND (Mir-Flickr Near-Duplicate), CLAIMS, Holidays, California-ND, or asks about evaluating this task. Reports sensitivity at false positive rate (FPR).

researchpythongo
0
3
Unsupervised Relation Extraction EvalA

Evaluates the ability of language models to perform unsupervised relation extraction by predicting relation labels or tokens from contextual text. It probes factual grounding and context-constrained generation capabilities across varying relation types and corpus sources. Use when the user wants to benchmark on T-REx, Google-RE, ZSRE, TACRED, or asks about evaluating this task. Reports F1.

researchpythongo
0
3
Unsw Nb15 EvalA

Evaluates network intrusion detection capability by classifying network traffic flows as benign or malicious (or specific attack types) using graph-structured representations of network connections. It probes the model's ability to learn from adaptive graph construction and contrastive learning under resource-constrained conditions. Use when the user wants to benchmark on UNSW-NB15, or asks about evaluating this task. Reports accuracy.

researchpythonnode
0
3
Unsw Nb15 Nids EvalA

Evaluates the classification accuracy and computational efficiency of machine learning models for network intrusion detection on a realistic dataset of contemporary traffic and synthetic attacks. It also assesses the privacy preservation and data utility of a Pearson Correlation Coefficient (PCC) feature selection and Least Squares Method (LSM) data distortion pipeline. Use when the user wants to benchmark on UNSW-NB15, or asks about evaluating this task. Reports Accuracy.

datapythongo
0
3
Up5 Fairness EvalA

Evaluates recommendation accuracy and counterfactual fairness of LLM-based recommendation models. It measures ranking performance using Hit@k metrics and assesses bias by calculating the AUC for predicting sensitive user attributes from recommendations. Use when the user wants to benchmark on MovieLens-1M, Insurance, or asks about evaluating this task. Reports Hit@1.

researchpythontesting
0
3
Uquad1.0 EvalA

This benchmark evaluates Machine Reading Comprehension (MRC) capabilities in Urdu by testing a model's ability to extract correct answer spans from context paragraphs in response to questions. It probes span prediction accuracy, handling of multiple valid answers, and performance across different question types and named entities. Use when the user wants to benchmark on UQuAD1.0, or asks about evaluating this task. Reports F1.

researchpythongo
0
3
Urban Pathfinding EvalA

Evaluates real-time urban pathfinding algorithms under dynamic traffic and weather conditions. It measures how well traditional graph search methods and deep learning models predict optimal routes and minimize travel time in a simulated Berlin city environment. Use when the user wants to benchmark on Berlin Urban Simulation, or asks about evaluating this task. Reports Average Travel Time (s).

researchpythongo
0
3
Urban Syn Uda EvalA

Evaluates the utility of a semi-procedurally generated synthetic driving dataset (UrbanSyn) for unsupervised domain adaptation (UDA) in semantic segmentation. It probes whether combining multiple synthetic sources reduces the domain gap and improves pixel-level classification accuracy on real-world urban driving benchmarks. Use when the user wants to benchmark on UrbanSyn, GTAV, Synscapes, Cityscapes, BDD100K, Mapillary Vistas, or asks about evaluating this task. Reports self-labeling accuracy.

researchpythongo
0
3
Urban Tree Detection EvalA

Evaluates the capability of segmentation models to detect and delineate individual tree crowns and canopy coverage from high-resolution aerial imagery across diverse urban and tropical environments. Use when the user wants to benchmark on Zurich Municipal Tree Inventory & Swisstopo Imagery, WeRobotics Open AI Challenge (Tonga), or asks about evaluating this task. Reports Recall.

researchpythongo
0
3
Urbanverse EvalA

Evaluates the fidelity of automated urban scene reconstruction from city-tour videos and the generalization capability of reinforcement learning navigation policies trained in these simulations. It measures how well generated scenes match real-world semantics and layouts, and how effectively policies transfer to unseen simulated and real-world environments. Use when the user wants to benchmark on KITTI-360, CraftBench, AutoBench, or asks about evaluating this task. Reports success_rate.

researchpythongo
0
3
Urdu Mner EvalA

Evaluates the ability of models to recognize named entities (Person, Location, Organization, Miscellaneous) in Urdu social media posts by jointly processing textual and visual inputs. It probes cross-modal alignment, handling of low-resource language morphological complexity, and ambiguity resolution using visual context. Use when the user wants to benchmark on Twitter2015-Urdu, or asks about evaluating this task. Reports F1 score.

content-marketingpythongo
0
3
Url Benchmark EvalA

Evaluates how well uncertainty estimates for latent representations predict the correctness of the representation itself, specifically whether the nearest neighbor in the embedding space belongs to the same class. It tests the transferability and scalability of uncertainty quantification methods across different backbones and unseen datasets. Use when the user wants to benchmark on ImageNet-1k, or asks about evaluating this task. Reports R-AUROC.

researchpython
0
3
Uro Bench EvalA

Evaluates end-to-end speech-to-speech dialogue models across three core dimensions: understanding, reasoning, and oral conversation. It probes multilingual proficiency, multi-turn dialogue handling, and the ability to generate paralinguistic and emotional cues in audio responses. Use when the user wants to benchmark on URO-Bench, or asks about evaluating this task. Reports Task Accomplish Score.

researchpythongo
0
3
Ursa Gan EvalA

Evaluates cross-domain speech recognition and enhancement robustness by training downstream models on generatively simulated target-domain data. Probes the ability of ASR and SE systems to generalize to unseen acoustic conditions, channel mismatches, and compound noise-channel distortions. Use when the user wants to benchmark on Hakka Across Taiwan (HAT), Taiwanese Across Taiwan (TAT), VoiceBank-DEMAND (VBD), HAT-ESC, or asks about evaluating this task. Reports CER.

researchpythongo
0
3
Us Grid Forecasting EvalA

Evaluates the ability of deep learning architectures (SSMs, Transformers, RNNs) to forecast hourly electricity load across major US power grids. It probes how well models capture temporal patterns, handle varying prediction horizons, and integrate exogenous weather covariates for accurate grid-scale forecasting. Use when the user wants to benchmark on US ISO Hourly Load Data (EIA-930), or asks about evaluating this task. Reports MSE (%).

researchpythonexpress
0
3
Usamo Proof EvalA

This benchmark probes the ability of large language models to generate rigorous, step-by-step mathematical proofs for high-school olympiad-level problems. It evaluates logical coherence, justification of assumptions, and adherence to formal proof standards rather than just numerical correctness. Use when the user wants to benchmark on 2025 USA Math Olympiad, or asks about evaluating this task. Reports proof_points.

researchpythongo
0
3
Usb Summarization EvalA

Evaluates multiple text summarization capabilities including extractive/abstractive generation, factuality verification, factual error correction, topic-constrained generation, sentence compression, evidence extraction, and unsupported span detection across diverse domains. Use when the user wants to benchmark on Extractive Summarization (EXT), Abstractive Summarization (ABS), Factuality Classification (FAC), Fixing Factuality (FIX), Topic-based Summarization (TOPIC), Multi-sentence Compressi...

researchpythongo
0
3
Usc Asr EvalA

Evaluates automatic speech recognition (ASR) systems on Uzbek language audio by measuring character and word error rates against manually transcribed ground truth. It probes the model's ability to accurately transcribe low-resource speech data without relying on external linguistic resources or pronunciation dictionaries. Use when the user wants to benchmark on USC, or asks about evaluating this task. Reports WER.

researchpythonexpress
0
3
User Claim Distribution EvalA

This evaluation probes the semantic, epistemic, and verifiability characteristics of real-world user-submitted fact-checking requests. It measures how public demand for verification aligns with or diverges from synthetic benchmark distributions, highlighting gaps in current misinformation evaluation corpora. Use when the user wants to benchmark on User Fact-Checking Claims Dataset, or asks about evaluating this task. Reports Veracity Score.

datapythongo
0
3
Usmankiani256 MeteorA

Compute Usmankiani256/meteor via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of Usmankiani256/meteor.

developmentpython
0
3
Utd Mhad EvalA

Evaluates a model's ability to predict future human joint positions over a 15-frame horizon using past observations, while testing continual learning capabilities across different subjects and curriculum-based fine-tuning. Use when the user wants to benchmark on UTD-MHAD, or asks about evaluating this task. Reports MSE.

researchpythontesting
0
3
Utility Aware Data Pricing EvalA

Evaluates whether token-level quality signals and empirical training gain metrics can accurately predict real data utility for LLM fine-tuning, outperforming traditional row- or token-count baselines. It probes the framework's predictive alignment, ranking fidelity, and robustness to adversarial or low-value data across multiple domains. Use when the user wants to benchmark on Alpaca, GSM8K, CodeXGLUE-Python, or asks about evaluating this task. Reports Spearman rank correlation.

researchpythonperformance
0
3
Utkface Age Estimation EvalA

This evaluation protocol assesses the accuracy of a lightweight neural network for predicting a person's age from a single facial image. It focuses on regression-based age estimation to determine how well compact models generalize to held-out test data while maintaining deployment efficiency. Use when the user wants to benchmark on UTKFace, or asks about evaluating this task. Reports MAE.

researchpythongit
0
3
Uvh 26 EvalA

Evaluates object detection models on a domain-specific Indian traffic dataset, probing their ability to localize and classify 14 heterogeneous vehicle types under surveillance viewpoints with varying occlusion and scale. Use when the user wants to benchmark on UVH-26, or asks about evaluating this task. Reports mAP(50:95).

researchpythongo
0
3
Uvrb EvalA

Evaluates zero-shot generalization of video embedding models across 16 diverse retrieval tasks and domains. It probes capabilities like spatial/temporal reasoning, compositional understanding, and partially relevant matching, revealing how well models generalize beyond standard benchmarks. Use when the user wants to benchmark on UVRB (Universal Video Retrieval Benchmark), or asks about evaluating this task. Reports Recall@1 (R@1).

ai-agentspython
0
3
V Measure ScoreA

Compute the v_measure_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute v_measure_score, or asks how to score with v_measure_score.

documentationpython
0
3
V Triune EvalA

Evaluates vision-language models on a unified suite of visual reasoning and perception tasks, measuring generalization across real-world benchmarks, mathematical reasoning, and object detection/grounding capabilities. Use when the user wants to benchmark on MEGA-Bench Core, MMMU, MathVista, COCO, OVDEval, CountBench, OCRBench, ScreenSpot-Pro, or asks about evaluating this task. Reports MEGA-Bench Core weighted average.

researchpythongo
0
3
V2c Animation EvalA

Evaluates visually-driven voice cloning by measuring speech quality, temporal alignment, length consistency, speaker identity preservation, and emotion transfer accuracy against ground-truth audio. Use when the user wants to benchmark on V2C-Animation, or asks about evaluating this task. Reports MCD-DTW-SL.

researchpythongo
0
3
V2c Chem Dubbing EvalA

This evaluation probes a model's ability to generate synchronized, emotionally faithful, and speaker-identifiable speech for movie dubbing tasks. It measures audio-visual alignment, spectral similarity, and the preservation of speaker identity and emotional tone against ground-truth recordings. Use when the user wants to benchmark on V2C, Chem, or asks about evaluating this task. Reports LSE-D.

researchpythongo
0
3
V2v Got Planning EvalA

Evaluates cooperative autonomous driving planning capabilities using a multimodal LLM with graph-of-thoughts reasoning. It measures trajectory prediction accuracy and collision avoidance under occlusion-aware perception and planning-aware prediction scenarios. Use when the user wants to benchmark on V2V-GoT-QA, or asks about evaluating this task. Reports L2 error.

researchpythongo
0
3
V2v Qa EvalA

Evaluates a multi-modal LLM's ability to fuse 3D perception features from multiple connected vehicles to answer safety-critical driving queries. It probes spatial grounding, notable object identification near planned waypoints, and collision-avoidance trajectory planning in cooperative autonomous driving scenarios. Use when the user wants to benchmark on V2V-QA, or asks about evaluating this task. Reports F1.

researchpythongo
0
3
V2x Radar EvalA

Evaluates 3D object detection capabilities for autonomous driving using multi-modal sensors (LiDAR, camera, 4D radar) in single-agent (roadside and vehicle-mounted) and cooperative perception setups. It probes robustness to adverse weather conditions and communication delays in cooperative scenarios. Use when the user wants to benchmark on V2X-Radar, or asks about evaluating this task. Reports AP@IoU.

researchpythongo
0
3
V3det EvalA

Probes an object detector's ability to localize and classify instances across a vast, hierarchical vocabulary of over 13,000 categories. It evaluates both closed-set detection performance and open-vocabulary generalization to novel, unseen categories. Use when the user wants to benchmark on V3Det, or asks about evaluating this task. Reports AP.

researchpythongo
0
3
Vaani Asr Lid EvalA

This evaluation protocol assesses the utility of the Vaani dataset for fine-tuning automatic speech recognition (ASR) and spoken language identification (LID) models across diverse Indian languages and regions. It measures performance gains from fine-tuning on Vaani's transcribed audio and images against established benchmarks, highlighting regional dialectal variations and low-resource language capabilities. Use when the user wants to benchmark on Vaani, FLEURS, Kathbath, or asks about evalu...

researchpythongit
0
3
Vabench EvalA

Evaluates the quality, cross-modal consistency, synchronization, and spatial audio rendering of text-to-audio-video and image-to-audio-video generation models. It probes physical plausibility, emotional expressiveness, and stereo separation across seven real-world sound categories. Use when the user wants to benchmark on VABench, or asks about evaluating this task. Reports Audio-Visual Align.

researchpythongo
0
3
Vad Anticipation EvalA

Evaluates a model's ability to detect anomalous events in surveillance videos and anticipate their occurrence in future frames. It specifically probes scene-dependent anomaly recognition and multi-step temporal anticipation. Use when the user wants to benchmark on ShanghaiTech, CUHK Avenue, IITB Corridor, NWPU Campus, ShanghaiTech-sd, or asks about evaluating this task. Reports AUC (%).

researchpythonperformance
0
3
Vad Prediction EvalA

Evaluates a model's ability to predict continuous emotional dimensions (Valence, Arousal, Dominance) from text. Specifically probes the model's capacity to capture affective polarization signals in parliamentary discourse. Use when the user wants to benchmark on Knesset VAD Annotation, or asks about evaluating this task. Reports Pearson correlation.

researchpythongo
0
3
Vae Malware Detection EvalA

Evaluates the effectiveness of Variational Autoencoder (VAE)-derived latent space features for malware classification using traditional machine learning models. It probes robustness to data partitioning, random seed initialization, and computational efficiency without hyperparameter tuning. Use when the user wants to benchmark on EMBER, BODMAS, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Vaexbench EvalA

Evaluates multimodal large language models' ability to perform extractive and abstractive spatiotemporal reasoning on egocentric videos. It probes long-horizon memory, object tracking, spatial orientation, and metric distance estimation under both multiple-choice and free-form generation settings. Use when the user wants to benchmark on VAEX-Bench, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Valerie22 EvalA

This protocol evaluates the perceptual fidelity and cross-domain generalization capability of the VALERIE22 synthetic urban dataset by training a semantic segmentation model on it and testing on real-world automotive datasets. It specifically probes how dataset diversity (unique 3D assets) and training scale affect downstream perception performance. Use when the user wants to benchmark on VALERIE22, Cityscapes, A2D2, BDD100K, India Driving Dataset, Mapillary Vistas, or asks about evaluating t...

researchpythontesting
0
3
ValiditysoftA

Evaluates the faithfulness and semantic plausibility of model-agnostic XAI techniques by generating soft counterfactuals via token-level perturbations. It measures whether perturbations actually change model predictions and whether the generated explanations align with the true causal impact of those changes. Use when the user has predictions and gold and needs to compute Validitysoft, Csoft.

researchpythongo
0
3
Vallp TerA

Compute Vallp/ter via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of Vallp/ter.

developmentpython
0
3
Valueground EvalA

Evaluates whether multimodal large language models (MLLMs) can maintain consistent culture-conditioned value judgments when response options are replaced with minimally contrastive visual proxies. It probes cross-modal prediction stability and the ability to ground textual value tendencies in subtle visual contrasts. Use when the user wants to benchmark on ValueGround, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3