All authors
qhjqhj00 avatar

Claude Skills by qhjqhj00

github.com/qhjqhj00
7,636 skillsA× 7,623B× 11C× 1D× 10 installs2,098 views
Densepose Coco EvalA

Evaluates a model's ability to perform dense human pose estimation by predicting per-pixel body part labels and UV coordinates on a 3D surface model. It measures how well the model handles real-world variations in scale, pose, occlusion, and background clutter. Use when the user wants to benchmark on COCO-DensePose, or asks about evaluating this task. Reports AP.

researchpythongo
0
3
Dental Triagebench EvalA

Evaluates multimodal clinical reasoning by requiring models to integrate radiographic images (OPGs) and patient complaints to predict hierarchical dental triage labels. It probes the model's ability to perform precise, multi-label treatment referrals and broad specialty-level routing in a zero-shot clinical setting. Use when the user wants to benchmark on Dental-TriageBench, or asks about evaluating this task. Reports Macro-F1.

researchpythongo
0
3
Depression Diagnosis Chat EvalA

Evaluates a model's ability to conduct depression-diagnosis-oriented dialogues by tracking psychological states, generating appropriate responses, summarizing patient symptoms, and classifying depression/suicide severity. It also assesses conversational qualities like fluency, empathy, and doctor-likeness through human evaluation. Use when the user wants to benchmark on MedDialog, or asks about evaluating this task. Reports BLEU-2, Average weighted F1.

researchpythongo
0
3
Depth Anything Ac EvalA

Evaluates zero-shot monocular relative depth estimation robustness under complex environmental conditions such as low light, adverse weather (rain, fog, snow), and synthetic noise. It probes the model's ability to recover fine-grained spatial relationships and object boundaries from degraded inputs without fine-tuning. Use when the user wants to benchmark on DA-2K (multi-condition), NuScenes-night, Robotcar-night, Driving-Stereo, KITTI-C, KITTI, NYU-D, Sintel, ETH3D, DIODE, or asks about eval...

researchpython
0
3
Depth EvalA

Evaluates the effectiveness of a hierarchically pre-trained encoder-decoder model (DEPTH) against a standard T5 baseline on discourse understanding, natural language inference, sentiment analysis, grammar checking, and instruction following. Use when the user wants to benchmark on MNLI, SST2, CoLA, DiscoEval, Natural Instructions, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Depthcues EvalA

Evaluates whether large vision models inherently understand human monocular depth cues (e.g., occlusion, perspective, texture gradient) through classification tasks, and measures their downstream monocular depth estimation performance on standard datasets. Use when the user wants to benchmark on DepthCues, NYUv2, DIW, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Dermabench EvalA

Evaluates vision-language models on dermatological visual question answering and clinical reasoning. It probes the model's ability to understand skin lesions across diverse Fitzpatrick skin types, answer structured diagnostic questions, and reason about morphology and distribution. Use when the user wants to benchmark on DermaBench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Dermx Benchmark EvalA

Evaluates the diagnostic accuracy and explainability of various Convolutional Neural Network architectures on dermatological image classification. It specifically probes how well different architectures localize clinically relevant skin characteristics using Grad-CAM heatmaps compared to human dermatologists. Use when the user wants to benchmark on DermXDB, or asks about evaluating this task. Reports image-level Grad-CAM F1 score.

researchpythongo
0
3
Descrip3d 3d Scene EvalA

Evaluates large language models' ability to perform 3D scene understanding tasks, including single and multi-object visual grounding, 3D scene captioning, and contextual question answering, using object-level text descriptions for relational reasoning. Use when the user wants to benchmark on ScanRefer, Multi3DRefer, Scan2Cap, ScanQA, SQA3D, or asks about evaluating this task. Reports Acc@0.25 / Acc@0.5, F1@0.25 / F1@0.5, CIDEr@0.5 / CIDEr, EM / EM-R.

researchpythongo
0
3
Designbench EvalA

Evaluates multimodal large language models (MLLMs) on front-end web development tasks, including code generation, editing, and repair across multiple frameworks (React, Vue, Angular, HTML/CSS). It probes capabilities in visual-to-code translation, framework-specific syntax handling, code localization, and component reuse. Use when the user wants to benchmark on DesignBench, or asks about evaluating this task. Reports Compilation Success Rate (CSR).

developmentjavascriptpython
0
3
Detailed Localized Captioning EvalA

Evaluates a model's ability to generate detailed, region-specific descriptions for images and videos, ranging from keywords to multi-sentence captions. It probes fine-grained visual grounding, attribute recognition, and hallucination resistance by comparing generated text against reference captions or using attribute-level positive/negative judgments. Use when the user wants to benchmark on DLC-Bench, LVIS, PACO, Flickr30k Entities, Ref-L4, HC-STVG, VideoRefer-Bench-D, or asks about evaluatin...

researchpythongo
0
3
Detailmaster EvalA

Evaluates a text-to-image model's ability to faithfully render long, descriptive prompts. It probes fine-grained semantic alignment across character presence, attributes, spatial relationships, and scene composition, as well as overall aesthetic and alignment quality using preference models. Use when the user wants to benchmark on DetailMaster, or asks about evaluating this task. Reports CharacterPresence.

researchpythongo
0
3
Detailverifybench EvalA

Evaluates multimodal large language models' ability to pinpoint erroneous content at the token level within long-form image captions. It probes whether models can distinguish between visually grounded facts and hallucinated details by localizing specific tokens that contradict the input image. Use when the user wants to benchmark on DetailVerifyBench, or asks about evaluating this task. Reports token-level F1.

researchpythongo
0
3
Detecting Llm Peer Reviews EvalA

Evaluates the efficacy of covert watermarking techniques embedded in manuscript PDFs to force LLM-generated peer reviews to contain specific hidden markers. It also tests the robustness of these watermarks against common reviewer defenses like paraphrasing, detection prompts, and page cropping, as well as the performance of cryptic prompt injection via gradient-based optimization. Use when the user wants to benchmark on ICLR 2024 submissions, ICLR 2021 submissions, ICLR 2024 submissions (cont...

researchpythongo
0
3
Devnet Anomaly Detection EvalA

Evaluates the capability of anomaly detection models to identify rare or deviant data points using only a small set of labeled anomalies as prior knowledge. It probes data efficiency, robustness to varying anomaly contamination levels in unlabeled training data, and the ability to rank anomalies effectively under severe class imbalance. Use when the user wants to benchmark on donors, census, fraud, celeba, backdoor, URL, campaign, news20, thyroid, or asks about evaluating this task. Reports A...

researchpythonperformance
0
3
Dexcanvas Success EvalA

This evaluation probes a robot policy's ability to successfully reproduce human-demonstrated dexterous manipulation trajectories in physics simulation. It measures robustness by testing performance under nominal conditions and under controlled initial pose perturbations. Use when the user wants to benchmark on DexCanvas, or asks about evaluating this task. Reports success rate.

researchpythonexpress
0
3
Dexycb EvalA

Evaluates joint perception and manipulation capabilities for hand-object interactions, specifically 2D detection, 6D object pose estimation, and 3D hand pose estimation on real-world RGB-D sequences. Use when the user wants to benchmark on DexYCB, or asks about evaluating this task. Reports precision-coverage.

researchpythonperformance
0
3
Df3dv 1k EvalA

Evaluates the robustness of distractor-free novel view synthesis methods against large-scale, diverse distractor scenarios. It measures how well radiance field and 3D Gaussian Splatting models can reconstruct clean 3D scenes from cluttered or dynamically changing inputs without degrading static background quality. Use when the user wants to benchmark on DF3DV-1K, DF3DV-41, or asks about evaluating this task. Reports PSNR.

researchpythongo
0
3
Dfjsp Qa EvalA

Evaluates the ability of a quantum annealer (D-Wave) to solve distributed flexible job shop scheduling problems (DFJSP) compared to classical simulated annealing. It probes solver performance in terms of solution quality (energy, makespan, constraint satisfaction) and computational efficiency (runtime scaling) across varying problem sizes. Use when the user wants to benchmark on Custom DFJSP instances (wool textile industry), or asks about evaluating this task. Reports System energy.

researchpythongo
0
3
Dfm Dialogue EvalA

Evaluates a unified dialogue foundation model across representation, knowledge distillation, and generation capabilities on diverse dialogue-oriented tasks. It probes the model's ability to perform intent detection, slot filling, semantic parsing, dialogue state tracking, text-to-SQL, and end-to-end task-oriented dialogue generation. Use when the user wants to benchmark on DialoGLUE, MULTIWOZ2.0, MULTIWOZ2.2, Spider, CoSQL, CLINC150, BANKING77, HWU64, RESTAURANT8K, DSTC8, TOP, PERSONALCHAT, C...

researchpythongo
0
3
Dfme EvalA

Evaluates automatic dynamic facial micro-expression recognition (MER) models on a large-scale spontaneous micro-expression dataset. It probes the model's ability to classify subtle, high-frame-rate facial movements across seven emotion categories while handling class imbalance and variable video lengths. Use when the user wants to benchmark on DFME, or asks about evaluating this task. Reports Accuracy (ACC).

researchpythongo
0
3
Dgfnet Av Sep EvalA

Evaluates audio-visual models on their ability to separate target musical instrument sounds from mixed audio using synchronized video cues. It probes cross-modal feature alignment and dynamic fusion of audio and visual signals for source separation in complex environments. Use when the user wants to benchmark on MUSIC, MUSIC-21, or asks about evaluating this task. Reports SDR.

researchpythongo
0
3
Dgser EvalA

Evaluates a model's ability to perform sequential next-item recommendation by modeling dynamic collaborative signals and temporal user preferences. It tests how well the system captures high-order item transitions and time-annotated graph structures to predict the next interaction in a user's history. Use when the user wants to benchmark on Amazon-CDs, Amazon-Games, Amazon-Beauty, or asks about evaluating this task. Reports NDCG@10.

researchpython
0
3
Dharmaocr Benchmark EvalA

Evaluates structured OCR extraction fidelity and text degeneration rates on printed, handwritten, and legal documents. Measures how well models adhere to JSON schemas while minimizing pathological generation loops. Use when the user wants to benchmark on DharmaOCR-Benchmark, or asks about evaluating this task. Reports Score.

researchpythongo
0
3
Dhen Ctr EvalA

Evaluates the effectiveness of a deep hierarchical ensemble network for large-scale click-through rate (CTR) prediction. It probes the model's ability to capture complex, non-overlapping feature interactions across multiple layers and scale efficiently on industrial-scale data. Use when the user wants to benchmark on Industrial in-house dataset, or asks about evaluating this task. Reports Normalized Entropy (NE) loss.

researchpythongo
0
3
Dhoroni EvalA

Evaluates a model's ability to perform multi-dimensional discourse analysis on Bengali climate news articles. It probes capabilities in stance detection, authenticity verification, political influence identification, and various information extraction tasks related to environmental reporting. Use when the user wants to benchmark on Dhoroni, or asks about evaluating this task. Reports F1 Score.

researchpythonperformance
0
3
Dia Safety EvalA

Evaluates the safety of conversational AI models by measuring their tendency to generate unsafe responses at both the utterance level and within conversational context. It specifically probes context-sensitive unsafety, where responses appear safe in isolation but become harmful when conditioned on prior dialogue history. Use when the user wants to benchmark on DiaSafety, or asks about evaluating this task. Reports proportion.

researchpythongo
0
3
Diabetes Note Classification EvalA

This benchmark evaluates the ability of machine learning models to perform binary classification on free-text electronic health record (EHR) progress notes related to diabetes. It probes how well different architectures (CNNs, RNNs, SVMs, hybrids) capture local linguistic patterns and generalize across different hospital datasets. Use when the user wants to benchmark on BWH/UTP Clinical Notes, or asks about evaluating this task. Reports AUC.

researchpythongo
0
3
Diabetica EvalA

Evaluates large language models on diabetes care and management tasks, probing their ability to recall foundational medical knowledge, make clinical decisions in case studies, generate precise text, and reason through open-ended patient queries. Use when the user wants to benchmark on Diabetica, or asks about evaluating this task. Reports accuracy.

ai-agentspythongo
0
3
Diacr Ita EvalA

Binary classification of lexical semantic change for target words across two diachronic time periods. It probes whether models can reliably detect meaning shifts in Italian using corpus pairs from newspapers and books. Use when the user wants to benchmark on DIACR-Ita, or asks about evaluating this task. Reports classification accuracy.

researchpythongo
0
3
Diagnosisarena EvalA

Clinical diagnostic reasoning capability of LLMs, requiring them to generate plausible diagnoses from patient case descriptions and imaging/symptom details. It probes the model's ability to perform complex, multi-step medical deduction and generalize across 28 clinical specialties. Use when the user wants to benchmark on DiagnosisArena, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Dial Turning Rl EvalA

Evaluates a robot's ability to perform contact-based manipulation and learn a continuous control policy via reinforcement learning to match a target angle on a potentiometer. Use when the user wants to benchmark on DeltaZ Dial Turning Task, or asks about evaluating this task. Reports reward.

researchpythongit
0
3
Dialectal Sentiment Classification EvalA

Evaluates sentiment classification performance across different English dialects (en-US, en-AU, en-UK, en-IN) and tests how label proximity, review length, and sentiment density affect model generalization. Use when the user wants to benchmark on Google Place Reviews (Dialectal Sentiment), or asks about evaluating this task. Reports F1-Score.

researchpythongo
0
3
Dialectalarabicmmlu EvalA

Evaluates large language models' ability to understand and reason across multiple Arabic dialects and standard Arabic across diverse academic and professional domains. It measures dialectal generalization and sensitivity to linguistic context by comparing performance under default, dialect-conditioned, and dialect-identification prompts. Use when the user wants to benchmark on DialectalArabicMMLU, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Dialogstudio Response EvalA

Evaluates a model's ability to generate task-oriented and knowledge-grounded dialogue responses in zero-shot and few-shot settings. It measures lexical overlap and unigram F1 against ground-truth responses to assess generalization across open-domain and multi-domain conversational tasks. Use when the user wants to benchmark on CoQA, MultiWOZ 2.2, or asks about evaluating this task. Reports ROUGE-L.

researchpython
0
3
Dialogue Safety Robustness EvalA

Evaluates the robustness of offensive language detection models against adversarial human attacks in single-turn and multi-turn dialogue contexts. It measures classifier resilience when exposed to iterative, context-aware attacks designed to evade safety filters. Use when the user wants to benchmark on Wikipedia Toxic Comments, or asks about evaluating this task. Reports Weighted-F1.

researchpythongo
0
3
Dialogue Summarization EvalA

Evaluates the quality of unsupervised abstractive dialogue summarization across multiple domains by comparing generated summaries against human references using standard n-gram and LCS overlap metrics. It tests the model's ability to compress and rephrase conversational transcripts into coherent summaries without training data. Use when the user wants to benchmark on AMI, ICSI, DialogSum, SAMSum, MediaSum, SummScreen, ADS, or asks about evaluating this task. Reports ROUGE-1.

researchpython
0
3
Dialseg711 Seg EvalA

Evaluates dialogue segmentation on a benchmark constructed by joining disparate task-oriented dialogues. It probes the model's ability to detect abrupt, artificial context shifts and identify segment boundaries in synthetic multi-intent conversations. Use when the user wants to benchmark on DialSeg711, or asks about evaluating this task. Reports Pk.

researchpythongo
0
3
Diamonds EvalA

Evaluates Theory of Mind and participant-centric reasoning in multi-party dialogues by testing a model's ability to track dynamic numerical variables, filter distractors, and reason from specific character perspectives (including false beliefs) rather than using omniscient context. Use when the user wants to benchmark on DIAMONDs, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Dianjin R1 EvalA

Evaluates large language models' financial reasoning capabilities and general problem-solving skills across multiple benchmarks. It measures how well models can answer domain-specific financial questions and general math/science reasoning tasks, while also assessing compliance rule adherence in Chinese financial contexts. Use when the user wants to benchmark on CFLUE, FinQA, CCC, MATH-500, GPQA-Diamond, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Dice CoefficientA

Evaluates a disease-weighted attention refinement framework for medical image analysis. It probes the model's ability to align cross-attention maps with radiologist annotations (grounding) while maintaining diagnostic accuracy on chest X-ray classification tasks. Use when the user has predictions and gold and needs to compute Dice Coefficient.

researchpythongo
0
3
Dien Ctr EvalA

Evaluates a model's ability to predict click-through rates (CTR) by modeling sequential user behavior and dynamically evolving latent interests relative to a target item. Use when the user wants to benchmark on Amazon Books, Amazon Electronics, Industrial (Taobao), or asks about evaluating this task. Reports AUC.

researchpythontesting
0
3
Difair EvalA

This benchmark evaluates a language model's ability to disentangle factual gender knowledge from gender bias in masked language modeling. It measures whether a model can correctly predict gendered tokens in gender-specific contexts while remaining gender-neutral in gender-neutral contexts, revealing the trade-off between fairness and factual performance. Use when the user wants to benchmark on DIFAIR, or asks about evaluating this task. Reports GIS.

researchpythongo
0
3
Diffaware Ctxtaware EvalA

Evaluates whether LLMs recognize meaningful demographic group differences (Difference Awareness) and understand when differential treatment is contextually appropriate (Contextual Awareness), challenging the standard 'color-blind' fairness paradigm. Use when the user wants to benchmark on DiffAware and CtxtAware Benchmark Suite, or asks about evaluating this task. Reports win rate.

researchpythongo
0
3
Differential Auditing EvalA

This evaluation probes an adversarial auditing framework where a blue team must identify a compromised model among a pair of nearly identical models. It tests the ability to detect hidden backdoors, misaligned behaviors, or injected instructions using various probing strategies under varying levels of prior knowledge. Use when the user wants to benchmark on CIFAR-10, Truthful QA, HHH, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Diffseg30k EvalA

Pixel-level localization of diffusion-based AI edits in images, shifting from whole-image classification to semantic segmentation to identify precisely which regions have been altered by generative models. Use when the user wants to benchmark on DiffSeg30k, or asks about evaluating this task. Reports localization accuracy.

researchpython
0
3
Diffurec EvalA

Evaluates a model's ability to predict the next item in a user's interaction sequence by capturing dynamic preferences and multi-aspect item representations. It probes sequential recommendation performance under varying sequence lengths and item popularities. Use when the user wants to benchmark on Amazon Beauty, Amazon Toys, Movielens-1M, Steam, or asks about evaluating this task. Reports HR@K.

researchpythontesting
0
3
Diffusion Instruction Tuning EvalA

This protocol evaluates the vision-language alignment and zero-shot generalization capabilities of fine-tuned VLMs across diverse multimodal tasks. It measures how effectively aligning VLM cross-attention with diffusion model attention maps improves performance on document understanding, reasoning, real-world visual comprehension, and hallucination detection benchmarks. Use when the user wants to benchmark on AI2D, ChartQA, OCRBench, DocVQA, InfoVQA, MME, MMBench, ScienceQA, MMStar, MMMU, Rea...

researchpythongo
0
3
Diffusion Policy EvalA

Evaluates visuomotor policy learning in robotics by measuring how well a model generates sequential actions to complete manipulation tasks under both state and image observations. It probes the policy's ability to handle multimodal action distributions, long-horizon dependencies, and latency robustness across rigid and fluid object manipulation. Use when the user wants to benchmark on Robomimic, Push-T, Block Push, Franka Kitchen, or asks about evaluating this task. Reports success_rate.

researchpythongo
0
3
Diffusion Rep EvalA

Evaluates whether conditional diffusion models learn semantically meaningful and factorized representations by measuring generation accuracy against ground truth latent coordinates and the predictive power of internal model embeddings over those coordinates. Use when the user wants to benchmark on Synthetic 2D Gaussian Bump Dataset, or asks about evaluating this task. Reports predicted label accuracy.

datapythongo
0
3