All authors
qhjqhj00 avatar

Claude Skills by qhjqhj00

github.com/qhjqhj00
7,636 skillsA× 7,623B× 11C× 1D× 10 installs2,103 views
Dreamlip Zero Shot EvalA

Evaluates the zero-shot transfer capability of language-image pre-trained models across image-text retrieval, semantic segmentation, image classification, and vision-language reasoning tasks. Use when the user wants to benchmark on ImageNet, MSCOCO, Flickr30K, ADE20K-150, VOC-20, or asks about evaluating this task. Reports R@K, Top-1 accuracy.

ai-agentspythongo
0
3
Dreamomni2 EvalA

Evaluates a model's ability to perform multimodal instruction-based image editing and generation, specifically testing adherence to text instructions while manipulating concrete objects and abstract attributes (e.g., texture, style) using multiple reference images. Use when the user wants to benchmark on DreamOmni2 benchmark, or asks about evaluating this task. Reports success editing ratio.

researchpythongo
0
3
Dressipi Sbr EvalA

Evaluates a session-based recommendation model's ability to predict the next item a user will purchase based on their recent browsing history. It specifically probes how well the model handles cold-start scenarios and varying data availability by measuring ranking quality and hit rates on short retail sessions. Use when the user wants to benchmark on Dressipi, or asks about evaluating this task. Reports Recall@20.

researchpythongo
0
3
Driveact EvalA

Evaluates fine-grained driver action recognition in constrained in-cabin environments using multimodal video inputs (RGB, IR, Depth). It probes the model's ability to classify 34 specific driver activities under variable illumination and occlusion by measuring both overall and per-class recognition accuracy. Use when the user wants to benchmark on Drive&Act, or asks about evaluating this task. Reports Top-1 accuracy.

researchpythonperformance
0
3
Drivebench EvalA

Evaluates the reliability, visual grounding, and corruption resilience of vision-language models in autonomous driving. It probes whether models genuinely interpret degraded visual inputs or rely on textual priors and hallucinated reasoning when visual cues are missing or corrupted. Use when the user wants to benchmark on DriveBench, or asks about evaluating this task. Reports GPT score.

researchpythongo
0
3
Drivecritic EvalA

Evaluates a model's ability to judge autonomous driving trajectory pairs based on context-aware reasoning, safety, and human preferences, rather than relying on rigid rule-based thresholds. It probes whether the model can integrate visual and symbolic context to reason about nuanced traffic situations like lateral buffer maintenance or stop sign compliance. Use when the user wants to benchmark on DriveCritic, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Drivelmmo1 EvalA

Evaluates step-by-step visual reasoning capabilities of multimodal models in autonomous driving scenarios, covering perception, prediction, and planning. It assesses both the logical coherence of intermediate reasoning steps and the accuracy of final answers. Use when the user wants to benchmark on DriveLMM-o1, or asks about evaluating this task. Reports final reasoning score.

ai-agentspythongit
0
3
Driver Dojo EvalA

Evaluates the generalization capability of reinforcement learning policies for autonomous driving across procedurally generated traffic scenarios. It probes how well agents trained on a fixed set of road layouts and traffic dynamics perform when transferred to unseen environments with varying vehicle interactions and partial observability. Use when the user wants to benchmark on Driver Dojo, or asks about evaluating this task. Reports Interquartile Mean (IQM) reward.

researchpythongo
0
3
Drivinggen EvalA

Evaluates generative video world models for autonomous driving by jointly assessing visual realism, trajectory plausibility, temporal and agent-level consistency, and ego-conditioned motion controllability over a 100-frame prediction horizon. It benchmarks both general-purpose and driving-specific models to reveal trade-offs between photorealism and physical motion fidelity. Use when the user wants to benchmark on DrivingGen, or asks about evaluating this task. Reports Avg. Rank.

researchpythongo
0
3
Drivingvqa EvalA

Evaluates a vision-language model's ability to perform multi-label multiple-choice question answering on real-world driving scenarios, requiring precise visual grounding and spatial reasoning to select all correct answers from a set of options. Use when the user wants to benchmark on DrivingVQA, or asks about evaluating this task. Reports exam score.

researchpythongo
0
3
Droid Robot Manipulation EvalA

Evaluates the robustness and generalization of robot manipulation policies when co-trained with diverse, in-the-wild data. It probes the model's ability to handle distractors, novel objects, and scene variations across short, medium, and long-horizon tasks. Use when the user wants to benchmark on DROID Evaluation Tasks, or asks about evaluating this task. Reports success rate.

researchpython
0
3
Droidspan EvalA

Evaluates the sustainability and robustness of a dynamic behavioral profiling approach (DroidSpan) for Android malware detection over time and against code obfuscation, compared to a static baseline (MamaDroid). Use when the user wants to benchmark on all-data, oldBen+oldMal, MalObf, or asks about evaluating this task. Reports F1-measure.

researchpythongo
0
3
Dronevehicle EvalA

Evaluates aerial vehicle detection capability using aligned RGB and infrared image pairs. It specifically probes a model's ability to fuse cross-modal features and handle uncertainty in low-light or complex urban backgrounds. Use when the user wants to benchmark on DroneVehicle, or asks about evaluating this task. Reports mAP.

researchpythongo
0
3
Droughted EvalA

Evaluates time-series forecasting models on predicting U.S. drought severity across 1 to 6 week horizons using meteorological and static features. It probes both regression accuracy and multi-class classification performance for drought monitoring levels. Use when the user wants to benchmark on DroughtED, or asks about evaluating this task. Reports MAE.

researchpythongo
0
3
Droughtset EvalA

Evaluates spatiotemporal forecasting models on predicting three drought indices (soil moisture, evaporative stress index, and solar-induced chlorophyll fluorescence) across the U.S. CONUS using weekly climate and vegetation data. It also assesses the models' ability to classify drought events based on soil moisture percentiles. Use when the user wants to benchmark on DroughtSet, or asks about evaluating this task. Reports MAE.

datapythontesting
0
3
Drsm Certified Robustness EvalA

Evaluates the standard classification accuracy and certified robustness of a malware detector against adversarial byte perturbations. It measures how well the model maintains correct predictions under a bounded perturbation budget using a de-randomized smoothing defense with window ablation. Use when the user wants to benchmark on PACE, or asks about evaluating this task. Reports Standard Accuracy.

researchpythongit
0
3
Drug Discovery Benchmarks EvalA

Evaluates a multi-modal foundation model's capability across classification, regression, and generation tasks in drug discovery. It probes the model's ability to predict cell types, assess drug efficacy and safety, design antibody CDR regions, and estimate binding affinities for proteins and small molecules. Use when the user wants to benchmark on Zheng68k, MoleculeNet (BBBP/ClinTox), GDSC (Cancer-Drug Response 1-3), SAbDab, Weber TCR Benchmark, SKEMPI S1131, DTI Benchmark, or asks about eval...

researchpythongo
0
3
Drug Pair Scoring EvalA

This evaluation benchmarks deep learning architectures on predicting drug-drug interactions, polypharmacy side effects, and drug synergy. It measures how well models encode molecular graphs and combine them to score pairwise biological outcomes across multiple pharmacological domains. Use when the user wants to benchmark on TWOSIDES, Drugbank DDI, DrugComb, DrugCombDB, OncolyPharm, or asks about evaluating this task. Reports AUROC.

researchpythonperformance
0
3
Drug Target Interaction EvalA

Evaluates computational models on predicting binary drug-target interactions using standardized bioactivity data. It probes the model's ability to learn molecular and protein representations and generalize across different data splits (lenient, cold-ligand, cold-target). Use when the user wants to benchmark on Curated DTI dataset, or asks about evaluating this task. Reports AUROC.

researchpythongo
0
3
Drugbank Hetionet EvalA

Evaluates the ability of matrix completion algorithms to predict missing biological interactions (drug-target or compound-disease) using sparse association matrices and side information. Use when the user wants to benchmark on DrugBank, Hetionet (Drug Repurposing), or asks about evaluating this task. Reports AUPR.

researchpythongo
0
3
Drugcareqa EvalA

Evaluates an AI system's ability to perform integrated clinical decision-making by simulating real-world online medical consultations. It probes the model's capacity to reason through patient symptoms, generate accurate diagnoses, and recommend appropriate medications within a unified workflow. Use when the user wants to benchmark on DrugCareQA, or asks about evaluating this task. Reports diagnostic and medication recommendation accuracy.

researchpythongo
0
3
Drugood EvalA

Evaluates the out-of-distribution (OOD) generalization and robustness of graph neural networks and sequence models on molecular binding affinity prediction tasks under various domain shifts and annotation noise levels. Use when the user wants to benchmark on DrugOOD, or asks about evaluating this task. Reports AUROC.

researchpythongo
0
3
Drugpc EvalA

Probes multi-step therapeutic reasoning and tool-use for drug-related questions, including interactions, contraindications, and patient-specific treatment strategies. It tests the model's ability to dynamically select biomedical tools, retrieve verified knowledge, and generate evidence-grounded answers. Use when the user wants to benchmark on DrugPC, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Drugplayground EvalA

Evaluates LLMs' ability to generate accurate, chemically plausible drug property descriptions and to produce meaningful text embeddings for drug discovery. It probes descriptive accuracy, lexical/structural alignment with ground truth, and embedding similarity for downstream representation tasks. Use when the user wants to benchmark on MolTextNet, or asks about evaluating this task. Reports Normalized Total score.

researchpythongit
0
3
Drunper Metrica TesiA

Compute Drunper/metrica_tesi via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of Drunper/metrica_tesi.

developmentpython
0
3
Drv Code Security EvalA

Evaluates the effectiveness of a Detect-Repair-Verify (DRV) workflow for fixing security vulnerabilities in LLM-generated code across different programming languages and granularity scopes (project, requirement, file). It measures how well iterative repair converges to a state that is both functionally correct and secure. Use when the user wants to benchmark on Custom LLM-generated code artifacts (JS, PHP, Python), or asks about evaluating this task. Reports S\C Yield Rate.

securitypythonphp
0
3
Drvoice EvalA

Evaluates a speech-text voice conversation model's capabilities in speech-to-text understanding, speech-to-speech generation, and overall speech quality. It probes modality alignment, reasoning, open-ended QA, and instruction following across multiple audio benchmarks. Use when the user wants to benchmark on OpenAudioBench, VoiceBench, UltraEval-Audio, Big Bench Audio, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Ds 1000 EvalA

This benchmark evaluates a model's ability to generate correct, executable Python code for data science tasks, specifically focusing on NumPy operations. It probes functional correctness under natural language descriptions and tests robustness against surface-form and semantic perturbations of original StackOverflow problems. Use when the user wants to benchmark on numpy-100, or asks about evaluating this task. Reports pass@1.

researchpythonapi
0
3
Dsbench EvalA

Evaluates Vision-Language Models' ability to perceive and reason about safety-critical scenarios in autonomous driving, covering both external environmental hazards (e.g., traffic rules, obstacles, weather) and in-cabin driver states (e.g., fatigue, distraction, emotion). It probes fine-grained hazard recognition, regulatory compliance, and multi-step safety reasoning under diverse, high-risk conditions. Use when the user wants to benchmark on DSBench, or asks about evaluating this task. Repo...

researchpythongo
0
3
Dsd Scene Analysis EvalA

Evaluates the ability of vision-language models to generate detailed, technically accurate scene descriptions from images, leveraging high-fidelity human annotations and peer-ranked photography data. Use when the user wants to benchmark on DataSeeds.AI Sample Dataset (DSD), or asks about evaluating this task. Reports BLEU-4.

researchpythonaws
0
3
Dsf Gan Utility EvalA

Evaluates the predictive utility of synthetic tabular data generated by a GAN. It measures how well a downstream classifier or regressor trained on the synthetic samples performs when evaluated on a strictly held-out real validation set. Use when the user wants to benchmark on Two distinct tabular datasets (names in Appendix A), or asks about evaluating this task. Reports model performance.

researchpythonperformance
0
3
Dsin Ctr EvalA

This benchmark evaluates a model's ability to predict click-through rates (CTR) by leveraging session-aware user behavior sequences. It probes how well a system can decompose historical interactions into time-separated sessions, model cross-session interest evolution, and adaptively weight session interests relative to a target item. Use when the user wants to benchmark on Advertising Dataset, Recommender Dataset, or asks about evaluating this task. Reports AUC.

researchpythonperformance
0
3
Dsp EvalA

Evaluates retrieval-augmented language models on open-domain, multi-hop, and conversational question answering by testing their ability to dynamically search for evidence, bootstrap in-context demonstrations, and generate accurate answers without fine-tuning. Use when the user wants to benchmark on Open-SQuAD, HotPotQA, QReCC, or asks about evaluating this task. Reports EM.

researchpythongo
0
3
Dsp Toxicity Prediction EvalA

Evaluates machine learning models' ability to predict diarrhetic shellfish poisoning (DSP) toxicity events in mussels using long-term environmental and phytoplankton monitoring data. It probes the model's capacity to integrate biological indicators (toxic species abundance) with abiotic drivers (salinity, river flow, temperature) for binary hazard forecasting. Use when the user wants to benchmark on Gulf of Trieste HAB monitoring dataset, or asks about evaluating this task. Reports F1 score.

datapythongo
0
3
Dst EvalA

Evaluates cross-lingual and zero-shot dialogue state tracking by measuring a model's ability to predict correct slot-value pairs in target languages using limited or translated training data. Use when the user wants to benchmark on Parallel MultiWoZ, Multilingual WoZ, or asks about evaluating this task. Reports Joint Goal Accuracy.

researchpythongo
0
3
Dst Jga EvalA

Evaluates a model's ability to track dialogue state in task-oriented conversations by predicting slot-value pairs across multiple domains. It measures how well the model maintains accurate belief states over multi-turn interactions, both with its own previous predictions and with ground-truth history. Use when the user wants to benchmark on RiSAWOZ, MultiWOZ, CrossWOZ, or asks about evaluating this task. Reports Joint Goal Accuracy (JGA).

researchpythongo
0
3
Dstc10 Spoken EvalA

Evaluates task-oriented dialogue systems on spoken conversations to measure robustness against ASR errors and disfluencies. It probes multi-domain dialogue state tracking, knowledge-seeking turn detection, knowledge selection, and response generation capabilities under realistic speech conditions. Use when the user wants to benchmark on DSTC10, DSTC9, MultiWOZ 2.1, or asks about evaluating this task. Reports Joint Goal Accuracy.

researchpythongo
0
3
Dstc11 Track2 Intent Induction EvalA

Evaluates a model's ability to automatically induce conversation intents by clustering utterances from task-oriented dialogues without prior intent labels. It measures how well the induced clusters align with ground-truth intent categories using supervised clustering metrics. Use when the user wants to benchmark on DSTC11 Track 2, or asks about evaluating this task. Reports ACC.

researchpythongo
0
3
Dstc11 Track3 EvalA

Evaluates a system's ability to track dialogue state in spoken conversations, specifically measuring robustness to ASR errors, disfluencies, and proper noun mismatches. Use when the user wants to benchmark on DSTC11 Track 3, or asks about evaluating this task. Reports JGA.

researchpythongo
0
3
Dt Pens EvalA

This benchmark evaluates a model's ability to generate personalized news headlines by accurately capturing user interests from implicit feedback (clicks and dwell times) while filtering out noise. It probes the system's capacity to align generated text with both lexical patterns and semantic meaning relative to ground-truth headlines tailored to specific user preferences. Use when the user wants to benchmark on DT-PENS, or asks about evaluating this task. Reports ROUGE-1.

researchpythontesting
0
3
Dta Affinity Prediction EvalA

Evaluates a model's ability to predict the binding affinity between small molecule drugs and protein targets. It probes regression accuracy, ranking consistency, and correlation strength on standardized drug-target interaction datasets. Use when the user wants to benchmark on Davis, KIBA, or asks about evaluating this task. Reports MSE.

researchpythongo
0
3
Dta Coldstart EvalA

Evaluates drug-target affinity prediction models in cold-start settings (cold-drug and cold-target) to assess generalization to novel drugs or targets using transferred inter-molecular interaction knowledge. Use when the user wants to benchmark on Davis, Kiba, or asks about evaluating this task. Reports RMSE.

researchpythonperformance
0
3
Dti Benchmark EvalA

Evaluates the ability of molecular models to predict drug-target interactions (DTI) by classifying whether a given drug and protein target pair binds. It probes the model's capacity to integrate diverse molecular representations (sequences, graphs, structures) and interaction layers to distinguish positive binding pairs from negative ones. Use when the user wants to benchmark on Davis, BIOSNAP, or asks about evaluating this task. Reports ROC-AUC.

researchpythontesting
0
3
Dti Binding Affinity EvalA

Evaluates a model's ability to predict continuous drug-target binding affinity and classify binary drug-target interactions. It probes geometry-aware representation learning, metric consistency, and generalization across diverse chemical-proteomic domains. Use when the user wants to benchmark on DTI-DG, BIOSNAP, BindingDB, DAVIS, or asks about evaluating this task. Reports PCC.

researchpythongo
0
3
Dti Inductive Prediction EvalA

Evaluates the ability of machine learning models to predict drug-target interactions in inductive settings where test drugs, targets, or both are unseen during training. It probes cold-start prediction capabilities and robustness to local class imbalance in sparse biological networks. Use when the user wants to benchmark on NR, GPCR, IC, E, DB, or asks about evaluating this task. Reports AUPR.

researchpythongo
0
3
Dti Prediction EvalA

This benchmark evaluates a model's ability to predict drug-target interactions by integrating molecular graphs and protein sequences into a heterogeneous interaction network. It probes the model's capacity to learn hierarchical graph representations and distinguish interacting from non-interacting drug-protein pairs. Use when the user wants to benchmark on DTI Benchmark, or asks about evaluating this task. Reports AUC.

researchpythonnode
0
3
Dti Regression EvalA

Evaluates a model's ability to predict continuous binding affinity for drug-target pairs across different cold-start and warm-start scenarios. It probes the model's generalization to unseen drugs, unseen targets, and fully seen interactions using regression metrics. Use when the user wants to benchmark on Davis, Metz, KIBA, or asks about evaluating this task. Reports RMSE.

researchpython
0
3
Dti Relation Extraction EvalA

Evaluates a model's ability to classify drug-target interaction relations from biomedical text into one of ten specific interaction types. It probes multiclass relation extraction under conditions of severe class imbalance. Use when the user wants to benchmark on DrugProt, ChemProt, or asks about evaluating this task. Reports micro F1-score.

researchpythongo
0
3
Dtu Nerf Edge Detection EvalA

Evaluates the geometric reconstruction quality of Neural Radiance Fields by extracting 3D surfaces or edges using density gradients. It measures how accurately the predicted geometry aligns with ground truth point clouds across diverse real-world objects. Use when the user wants to benchmark on DTU benchmark dataset, or asks about evaluating this task. Reports completeness.

researchpython
0
3
Dual Target Drug Design EvalA

Evaluates the ability of generative models to design dual-target ligands that simultaneously bind to two protein pockets with high affinity while maintaining favorable drug-like properties. It measures both binding strength and molecular quality across a large set of target pairs. Use when the user wants to benchmark on Dual-target drug design dataset, or asks about evaluating this task. Reports Dual High Affinity.

researchpythongo
0
3