All authors
qhjqhj00 avatar

Claude Skills by qhjqhj00

github.com/qhjqhj00
7,636 skillsA× 7,623B× 11C× 1D× 10 installs2,233 views
Dualbench EvalA

Evaluates a model's ability to generate synchronized background audio and intelligible speech from video input, measuring audio quality, distribution matching, and audio-video temporal alignment. Use when the user wants to benchmark on DualBench, VGGSound, or asks about evaluating this task. Reports FAD↓.

researchpython
0
3
Dualcam EvalA

This benchmark evaluates fine-grained real-time traffic light detection using synchronized dual-camera inputs. It probes a model's ability to accurately localize and classify multiple traffic light states across varying distances and sizes, while balancing detection speed and precision. Use when the user wants to benchmark on DualCam, or asks about evaluating this task. Reports F1-score.

researchpythongit
0
3
Dualgauge EvalA

Probes the joint functional correctness and security of LLM-generated code, while also evaluating an automated framework's ability to execute code in sandboxes and semantically judge test outcomes against human ground truth. Use when the user wants to benchmark on DualGauge-Bench, or asks about evaluating this task. Reports F1 Score.

researchpythonsecurity
0
3
Dualnet Cl EvalA

Evaluates a model's ability to learn sequentially from a stream of tasks without catastrophic forgetting, while adapting quickly to new tasks. It probes both task-aware (with task IDs) and task-free (without task IDs) continual learning settings, measuring final accuracy, forgetting, and knowledge transfer. Use when the user wants to benchmark on Split miniImageNet, CORE50, or asks about evaluating this task. Reports ACC.

researchpythonperformance
0
3
Dualrec Movielens EvalA

Evaluates a hybrid sequential and LLM-based framework for next-item movie recommendation. It probes the model's ability to capture temporal user preferences and semantic genre consistency to predict the next movie a user will watch. Use when the user wants to benchmark on MovieLens-1M, or asks about evaluating this task. Reports NDCG@5.

researchpythonperformance
0
3
Dualtoken EvalA

Evaluates a unified vision tokenizer's capacity to decouple and jointly optimize low-level perceptual reconstruction and high-level semantic understanding. It probes zero-shot classification, cross-modal retrieval, image reconstruction fidelity, and downstream multimodal reasoning capabilities. Use when the user wants to benchmark on ImageNet-1K, Flickr8K, VQAv2, POPE, MME, SEED-IMG, MMBench, MM-Vet, or asks about evaluating this task. Reports Top-1 accuracy.

researchpythongo
0
3
Dualturn Turn Taking EvalA

This benchmark evaluates a model's ability to predict conversational turn-taking dynamics and agent actions from dual-channel speech audio. It probes the system's capacity to anticipate speech boundaries, detect backchannels, and classify continuous turn-taking states without relying on explicit silence timeouts or external labels. Use when the user wants to benchmark on otoSpeech, Switchboard, or asks about evaluating this task. Reports wF1.

researchpythongo
0
3
Duccio Nas EvalA

Evaluates hardware-aware neural architecture search (NAS) methods on edge IoT tasks, measuring classification accuracy alongside hardware constraints like memory footprint, latency, and computational complexity (OPs) on a RISC-V IoT SoC. It benchmarks both mask-based and path-based differentiable NAS approaches across image classification, visual wake words, keyword spotting, and anomaly detection tasks. Use when the user wants to benchmark on CIFAR-10, MSCOCO 2014, Speech Commands v2, DCASE2...

businesspythongo
0
3
Dud E Virtual Screening EvalA

Ranks active compounds against decoys for a given protein target. It probes the model's ability to prioritize true binders in a large pool of inactive decoys and resist dataset biases. Use when the user wants to benchmark on DUD-E, AD, or asks about evaluating this task. Reports AUC.

researchpythonperformance
0
3
Duet Dyadic Har EvalA

Evaluates the ability of human activity recognition models to classify dyadic kinesic functions and interactions across different subjects and physical locations. It probes robustness to viewpoint changes, background variations, occlusion, and modality-specific limitations (RGB vs. depth vs. 3D skeletons). Use when the user wants to benchmark on DUET, or asks about evaluating this task. Reports Cross-location accuracy (%), Cross-subject accuracy (%).

researchpythongo
0
3
DunnindexA

Compute the DunnIndex metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute DunnIndex, or asks how to score with DunnIndex.

documentationpython
0
3
Duo Tok EvalA

Evaluates semantic music tokenizers on their ability to preserve musical semantics for conditional generation, their efficiency for language modeling, and their audio reconstruction fidelity. It probes whether decoupled vocal-accompaniment tokenization yields better LM-friendliness and tagging performance without sacrificing perceptual quality. Use when the user wants to benchmark on MagnaTagATune, or asks about evaluating this task. Reports PPL@1024.

researchpythonperformance
0
3
Duorc EvalA

Evaluates reading comprehension and long-form text understanding by asking models to answer questions about movie plots. It specifically probes sensitivity to narrative length and semantic shifts between short and paraphrased long versions of the same story. Use when the user wants to benchmark on DuoRC, or asks about evaluating this task. Reports F1.

researchpythongo
0
3
Duplex Dialogue EvalA

Evaluates a dialogue system's ability to manage full-duplex speech interactions, specifically focusing on the timing and appropriateness of machine-to-user interruptions and user-to-machine interruptions, alongside system response latency. Use when the user wants to benchmark on duplex-dialogue-3k, or asks about evaluating this task. Reports FTED.

researchpythongo
0
3
Duq Weather Forecasting EvalA

Evaluates a deep learning model's capability to perform spatio-temporal weather forecasting and quantify predictive uncertainty. It tests the model's ability to fuse historical observations with numerical weather prediction (NWP) data to generate accurate point forecasts and reliable 90% prediction intervals over a 37-hour horizon. Use when the user wants to benchmark on Beijing weather dataset, or asks about evaluating this task. Reports SS_avg.

datapython
0
3
Dureader Retrieval EvalA

Passage retrieval for web search queries, evaluating a model's ability to rank relevant documents from a large collection. It probes in-domain retrieval accuracy as well as out-of-domain and cross-lingual generalization, highlighting challenges like salient phrase mismatch, syntactic mismatch, and false negatives. Use when the user wants to benchmark on DuReader_retrieval, or asks about evaluating this task. Reports MRR@10.

researchpythongo
0
3
Durecdial 2.0 EvalA

Evaluates conversational recommendation systems across monolingual, multilingual, and cross-lingual settings. It probes a model's ability to generate relevant and fluent responses, select correct knowledge entities, maintain topic consistency, and successfully guide dialogues toward a recommendation target. Use when the user wants to benchmark on DuRecDial 2.0, or asks about evaluating this task. Reports F1.

researchpythongo
0
3
Duriansc Svc EvalA

Evaluates one-shot singing voice conversion quality by measuring how naturally the converted audio sounds and how closely it matches the target speaker's voice, using only 20 seconds of target speech or singing data. Use when the user wants to benchmark on Database A, Database B, or asks about evaluating this task. Reports MOS naturalness.

researchpythongo
0
3
Dutch Ade Corpus EvalA

This benchmark evaluates transformer and Bi-LSTM models for detecting adverse drug events (ADEs) in Dutch clinical free text. It probes named entity recognition for drugs and disorders, relation classification for ADE and prescribing indication pairs, and document-level ADE detection. The protocol emphasizes handling class imbalance and evaluating performance across strict/lenient entity matching and single vs. grouped ADE relations. Use when the user wants to benchmark on Dutch ADE corpus, I...

researchpythonperformance
0
3
Dutch Book Review Sentiment EvalA

Evaluates the effectiveness of Universal Language Model Fine-tuning (ULMFiT) versus traditional SVM classifiers on low-resource, domain-specific sentiment classification tasks using small training sets. Use when the user wants to benchmark on Dutch book reviews, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Dutch Financial Benchmark EvalA

Evaluates LLMs on domain-specific financial tasks in Dutch, including sentiment analysis, named entity recognition, relation extraction, query answering, and headline classification. It also tests cross-lingual adaptability by benchmarking the Dutch model on English financial data. Use when the user wants to benchmark on Dutch Financial Benchmark, English Financial Benchmark, or asks about evaluating this task. Reports zero-shot performance.

researchpythongo
0
3
Dutch Llm Bench EvalA

Evaluates Dutch LLMs on reasoning, sentiment analysis, linguistic acceptability, world knowledge, and word sense disambiguation using zero-shot multiple-choice and binary classification tasks. Use when the user wants to benchmark on ARC (Dutch), DBRD, Dutch CoLA, Global MMLU (Dutch), XLWIC-NL, or asks about evaluating this task. Reports accuracy.

datapythongo
0
3
Dutch Medical Dialogue EvalA

Evaluates the quality of synthetically generated Dutch medical dialogues across structural, lexical, and qualitative dimensions to assess conversational naturalness and domain-specific accuracy. Use when the user wants to benchmark on Synthetic Dutch Medical Dialogues, or asks about evaluating this task. Reports MSTTR.

datapythongo
0
3
Dvbench EvalA

Evaluates Vision Large Language Models' ability to understand safety-critical driving videos across a hierarchical taxonomy of 25 abilities, including perception, temporal-spatial reasoning, and risk assessment. Use when the user wants to benchmark on DVBench, or asks about evaluating this task. Reports Top-1 accuracy.

researchpythongo
0
3
Dvcs Cff Extraction EvalA

This benchmark evaluates machine learning models' ability to extract Compton Form Factors (CFFs) from deeply virtual Compton scattering (DVCS) cross-section data. It specifically probes how well models adhere to quantum chromodynamics (QCD) constraints, generalize across kinematic regions, and accurately quantify both aleatoric and epistemic uncertainties during the extraction process. Use when the user wants to benchmark on DVCS unpolarized proton target data, or asks about evaluating this t...

datapythongo
0
3
Dvd Dst EvalA

Evaluates a model's ability to track visual objects and their attributes across turns in video-grounded dialogues. It probes long-term cross-modal dependency resolution and precise state decoding under controlled, bias-free synthetic dialogue conditions. Use when the user wants to benchmark on DVD-DST, or asks about evaluating this task. Reports Joint Acc.

researchpythongo
0
3
Dvfs Latency Energy EvalA

Evaluates the accuracy of a data-driven DVFS-aware latency model for DNN inference on GPUs against a traditional FLOPs-based benchmark. It probes the model's ability to predict real-world inference time and energy consumption under varying frequency settings, deadlines, and cooperative offloading scenarios. Use when the user wants to benchmark on CIFAR10, or asks about evaluating this task. Reports inference time (ms).

researchpythonperformance
0
3
Dvitel CodebleuA

Compute dvitel/codebleu via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of dvitel/codebleu.

developmentpython
0
3
Dvqa EvalA

This benchmark evaluates a model's ability to perform visual reasoning and information extraction on bar chart data visualizations. It specifically probes whether systems can accurately read chart-specific labels, handle out-of-vocabulary terms, and answer natural language questions about quantitative relationships and chart structure. Use when the user wants to benchmark on DVQA, or asks about evaluating this task. Reports exact-match accuracy.

researchpythongo
0
3
Dw Bench EvalA

Evaluates LLMs' ability to reason about data warehouse graph topologies, specifically focusing on foreign key path enumeration, data lineage impact analysis, and multi-hop graph traversal. It probes whether models can perform structural graph reasoning versus relying on lexical cues, using heterogeneous schema graphs with foreign key and lineage edges. Use when the user wants to benchmark on DW-Bench, or asks about evaluating this task. Reports Micro-EM.

researchpythongo
0
3
Dwrf Weather Forecast EvalA

Evaluates the accuracy, uncertainty quantification, and physical consistency of high-resolution ensemble weather forecasts for renewable energy applications. It probes a model's ability to downscale coarse atmospheric data to 1 km resolution while preserving multi-scale turbulence, thermodynamic constraints, and extreme event probabilities. Use when the user wants to benchmark on Northwestern Gobi Desert Wind Farm & ERA5 Reanalysis, or asks about evaluating this task. Reports RMSE, CRPS.

datapythongo
0
3
Dy Meter EvalA

Evaluates online anomaly detection models under concept drift by testing their ability to adapt to evolving data distributions without retraining. It probes instance-level sensitivity to context-dependent anomalies across continuous and discrete streaming scenarios. Use when the user wants to benchmark on Ionosphere, Pima, Satellite, Mammography, BGL, NSL-KDD, KDD99, Activity Recognition, Internal Bleeding, NASA, GaitPhase, EPG, ECG, Machine temperature, CPU utilization, INSECTS-Abr, INSECTS-...

researchpythongo
0
3
Dyabd Segmentation EvalA

This benchmark evaluates the segmentation capabilities of deep learning models on dynamic abdominal MRI scans. It specifically probes how well models handle extreme anatomical variability caused by real-time muscle motion during breathing and Valsalva maneuvers, across few-shot, prompt-based, and fully automatic inference settings. Use when the user wants to benchmark on DyABD, or asks about evaluating this task. Reports Dice.

researchpythongit
0
3
Dyad Arch EvalA

Evaluates a block-sparse linear layer approximation (DYAD) against dense baselines across standard NLP and vision benchmarks, measuring accuracy preservation and computational efficiency. Use when the user wants to benchmark on BLIMP, OPENLLM, GLUE+, MNIST, or asks about evaluating this task. Reports BLIMP accuracy.

researchpythongo
0
3
Dynamath EvalA

Evaluates the robustness of Vision-Language Models in mathematical reasoning by measuring performance across dynamically generated variants of seed questions. It probes how well models handle numerical, geometric, and contextual perturbations while maintaining consistent logical deduction. Use when the user wants to benchmark on DynaMath, or asks about evaluating this task. Reports average-case accuracy.

researchpythongo
0
3
Dynamic Audio Visual Nav EvalA

Evaluates an embodied agent's ability to navigate towards and catch a moving, previously unheard sound source in unmapped 3D environments using only audio and visual observations. It probes spatial reasoning, temporal memory, and robustness to noisy or distractor audio scenarios. Use when the user wants to benchmark on Replica, Matterport3D, or asks about evaluating this task. Reports DSPL.

researchpythongo
0
3
Dynamic House Simulator EvalA

Evaluates an agent's ability to perform temporal link prediction and object search in partially observable, dynamic environments by predicting object locations, ranking location likelihoods, and navigating to objects sequentially. Use when the user wants to benchmark on Dynamic House Simulator, or asks about evaluating this task. Reports NDCG.

researchpythongo
0
3
Dynamic Nerf Soccer EvalA

Evaluates the ability of dynamic NeRF models to perform photorealistic novel view synthesis in large-scale, dynamic sports environments. It probes spatiotemporal modeling capabilities, specifically how well models handle fast-moving small objects (like a soccer ball) and scale variations across different camera configurations. Use when the user wants to benchmark on Synthetic Soccer Scenes (Multi-Camera), or asks about evaluating this task. Reports PSNR.

researchpython
0
3
Dynamic Superb EvalA

Evaluates instruction-tuned speech models on their ability to perform diverse speech and audio tasks using natural language instructions. It probes zero-shot generalization by testing performance on seen versus unseen tasks and instructions across six dimensions: content, speaker, semantics, degradation, paralinguistics, and audio. Use when the user wants to benchmark on Dynamic-SUPERB, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Dynamic Superb Phase2 EvalA

Evaluates instruction-based universal speech and audio models across 180 tasks spanning speech, music, and environmental audio. It probes capabilities like automatic speech recognition, emotion recognition, speaker verification, and audio classification using a unified instruction-following framework. Use when the user wants to benchmark on Dynamic-SUPERB Phase-2, or asks about evaluating this task. Reports relative_score.

researchpythongo
0
3
Dynamic Topic Quality EvalA

Evaluates dynamic topic models by measuring topic coherence and diversity across chronological time slices, and assesses the utility of learned document-topic distributions via downstream text classification and clustering tasks. Use when the user wants to benchmark on NeurIPS, ACL, UN, NYT, WHO, or asks about evaluating this task. Reports Topic Coherence (TC).

researchpythongo
0
3
Dynamic Unlearning EvalA

Evaluates the effectiveness and robustness of LLM unlearning methods by measuring residual knowledge retrieval across dynamically generated single-hop, multi-hop, and alias-based queries, alongside the retention of adjacent and general knowledge. Use when the user wants to benchmark on RWKU, TOFU, or asks about evaluating this task. Reports Multi-hop Forgetting Criterion.

researchpythongo
0
3
Dynamic Urban View Synthesis EvalA

Evaluates novel view synthesis and image reconstruction quality in dynamic urban environments containing fast-moving objects and varying environmental conditions. It measures how well a model can render unseen viewpoints and reconstruct training views while handling dynamic geometry and pose drift. Use when the user wants to benchmark on Argoverse 2, KITTI, VKITTI2, or asks about evaluating this task. Reports PSNR.

researchpythongo
0
3
Dynamicare Medical Diagnosis EvalA

Evaluates a dynamic multi-agent framework's ability to perform interactive, open-ended medical diagnosis by querying patients and ranking potential diagnoses. It also assesses the system's performance on interactive multiple-choice medical QA and the quality of simulated patient responses. Use when the user wants to benchmark on MIMIC-Patient, MEDIQ, or asks about evaluating this task. Reports Hit@K.

ai-agentspythongo
0
3
Dysarthric Asr EvalA

Evaluates the ability of ASR and LLM-enhanced decoding models to accurately transcribe dysarthric speech across varying severity levels and domains. It probes robustness to phonetic distortions, grammatical consistency, and cross-dataset generalization. Use when the user wants to benchmark on TORGO, UASpeech, or asks about evaluating this task. Reports WER.

researchpythongo
0
3
Dysql Bench EvalA

Evaluates a model's ability to perform dynamic, multi-turn Text-to-SQL interactions that support full CRUD operations. It probes contextual reasoning, adaptability to evolving user requests, and error recovery within stateful database dialogues. Use when the user wants to benchmark on DySQL-Bench, or asks about evaluating this task. Reports state-equivalence accuracy.

researchpythongo
0
3
Dzen EvalA

Evaluates foundation models' ability to answer multiple-choice academic questions in both English and Dzongkha across varying grade levels and scientific subjects. It specifically probes factual recall, procedural application, and multi-step reasoning capabilities in a low-resource multilingual setting. Use when the user wants to benchmark on DZEN, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
E Care EvalA

Evaluates a model's ability to perform commonsense causal reasoning by predicting the reasonableness of causal facts, and to generate conceptually grounded natural language explanations for those causal relationships. Use when the user wants to benchmark on e-CARE, or asks about evaluating this task. Reports Accuracy (%).

researchpythongo
0
3
E Commerce Related Search EvalA

Evaluates the effectiveness of AI-generated related search queries in an e-commerce setting by measuring their ability to drive user engagement and purchases compared to a production baseline. Use when the user wants to benchmark on eBay user interaction logs, or asks about evaluating this task. Reports click-through rate (CTR).

researchpythonperformance
0
3
E2e Gmner EvalA

Evaluates a model's ability to perform end-to-end grounded multimodal named entity recognition, jointly identifying entity spans in text, predicting their semantic types, and grounding them to corresponding bounding boxes in an associated image. It probes the model's capacity for multimodal alignment, structured generation, and robustness to annotation noise via chain-of-thought reasoning. Use when the user wants to benchmark on Twitter-GMNER, Twitter-FMNERG, or asks about evaluating this tas...

researchpythongo
0
3