All authors
qhjqhj00 avatar

Claude Skills by qhjqhj00

github.com/qhjqhj00
7,636 skillsA× 7,623B× 11C× 1D× 10 installs2,098 views
Covid Net Cxr 2 EvalA

Binary classification of chest X-ray images to detect SARS-CoV-2 infection. It probes a model's ability to distinguish COVID-19 positive cases from negative cases (including no pneumonia and non-SARS-CoV-2 pneumonia) using a large, multinational dataset. Use when the user wants to benchmark on COVID-Net CXR-2 benchmark dataset, or asks about evaluating this task. Reports Sensitivity.

researchpythongo
0
3
Covid19 Cxr Ct Detection EvalA

Evaluates a lightweight CNN's ability to classify chest radiology images (CXR and CT scans) as positive or negative for COVID-19, and in a three-class setting. It also tests cross-modality generalization by training on CT and testing on CXR data. Use when the user wants to benchmark on CXR/CT Chest Radiology Dataset, or asks about evaluating this task. Reports accuracy.

researchpythontesting
0
3
Covid19 Cxr Detection EvalA

Evaluates a deep CNN's ability to classify chest X-ray images into COVID-19 positive and negative/healthy categories using region-edge and channel-boosted features. Use when the user wants to benchmark on Three datasets (names not specified in section), or asks about evaluating this task. Reports 95% Confidence Interval (CI).

researchpythongo
0
3
Covid19 Xray Classification EvalA

This evaluation probes a model's ability to classify chest X-rays as COVID-19 positive or negative using a cross-modal distillation setup where CT images are only used during training. It specifically tests the robustness of transfer learning under extremely small, patient-level paired cohorts and prevalence-heavy validation splits. Use when the user wants to benchmark on COVID-19 Image Data Collection, or asks about evaluating this task. Reports Accuracy.

researchpythonperformance
0
3
Covidx Classification EvalA

Evaluates deep learning models' ability to classify chest X-ray images into three diagnostic categories: normal, non-COVID-19 pneumonia, and COVID-19. It probes medical image classification performance under realistic class imbalance and tests whether models rely on clinically relevant lung regions or artifacts. Use when the user wants to benchmark on COVIDx, or asks about evaluating this task. Reports F-score.

researchpythongo
0
3
CovochevalA

Evaluates zero-shot conversational voice cloning systems on their ability to generate natural, expressive speech that matches a target speaker's timbre and spontaneous style without prior training on the target speaker. It measures pronunciation accuracy, speaker similarity, and subjective qualities like naturalness, quality, and spontaneous style. Use when the user wants to benchmark on HQ-Conversations / CoVoC Test Prompts, or asks about evaluating this task. Reports FS.

researchpythongo
0
3
Covost2 St Mt EvalA

Evaluates multilingual speech-to-text translation and automatic speech recognition across 22 languages. It probes the model's ability to transcribe spoken audio and translate it into English (or from English) under monolingual, bilingual, and multilingual training regimes. Use when the user wants to benchmark on CoVoST 2, or asks about evaluating this task. Reports BLEU.

researchpythongo
0
3
Cow Bench EvalA

Evaluates modal, spatial, and temporal consistency in general world models through 18 sub-tasks across six task categories. It uses human-designed checklists to verify fine-grained physical laws, causal reasoning, and cross-modal alignment, moving beyond perceptual metrics to hard verification. Use when the user wants to benchmark on CoW-Bench, or asks about evaluating this task. Reports checklist_score.

researchpythongo
0
3
Cp Bench EvalA

This benchmark evaluates speech-LLMs on contextual and paralinguistic reasoning tasks. It probes the models' ability to integrate linguistic content with emotional, prosodic, and social cues from in-the-wild speech data to answer specific question types. Use when the user wants to benchmark on CP-Bench, or asks about evaluating this task. Reports LLaMA-3-70B judge score.

researchpythongo
0
3
Cpllab SyntaxgymA

Compute cpllab/syntaxgym via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of cpllab/syntaxgym.

developmentpython
0
3
Cppe 5 EvalA

Evaluates object detection models on fine-grained medical personal protective equipment (PPE) in complex, real-world scenes. It probes a model's ability to localize and classify coveralls, face shields, gloves, masks, and goggles from non-canonical perspectives, measuring detection accuracy across multiple IoU thresholds and object scales. Use when the user wants to benchmark on CPPE-5, or asks about evaluating this task. Reports AP (mean Average Precision).

researchpythongo
0
3
Cps3d Seg EvalA

Evaluates 3D point cloud segmentation models for detecting surface defects on integrated circuit package substrates. It probes the model's ability to accurately classify high-density point clouds into normal and defect categories under industrial inspection conditions. Use when the user wants to benchmark on CPS3D-Seg, or asks about evaluating this task. Reports mIoU.

researchpythongo
0
3
Cpsc2018 Ecg Classification EvalA

Evaluates deep learning models on multi-lead ECG signal classification for arrhythmia detection under artificially balanced conditions. It probes the model's ability to extract discriminative temporal-spatial features from raw 12-lead cardiac signals and maintain robustness against various types of physiological noise. Use when the user wants to benchmark on CPSC2018, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Cpsdbench EvalA

Evaluates LLMs on Chinese public security domain tasks including text classification, information extraction, question answering, and text generation. Probes domain-specific accuracy, reliability, and contextual understanding in high-stakes law enforcement scenarios. Use when the user wants to benchmark on Weibo Sentiment Analysis, Rumor Detection, Telecommunication Fraud Detection, Drug-Related Case Reports, Public Security Case Reading Comprehension, Public Security Case Summary, or asks ab...

datapythongo
0
3
Cpt Merging EvalA

Evaluates the effectiveness of merging Continual Pretraining (CPT) models to build domain-specialized financial LLMs. It probes the model's ability to recover lost general knowledge, exhibit cross-domain complementarity, and demonstrate emergent reasoning capabilities through weight-space integration. Use when the user wants to benchmark on Financial Benchmark, or asks about evaluating this task. Reports Macro-Gain.

researchpythonperformance
0
3
Cqa EvalA

Evaluates a model's ability to answer multiple-choice commonsense reasoning questions. It probes whether providing natural language explanations (human or model-generated) alongside questions improves reasoning performance compared to a baseline without explanations. Use when the user wants to benchmark on CQA, or asks about evaluating this task. Reports Accuracy (%).

researchpythongo
0
3
Cracknex EvalA

Evaluates few-shot crack segmentation performance under low-light conditions using illumination-invariant features. It tests the model's ability to generalize from well-illuminated support images to unseen low-light query images in both synthetic and real-world scenarios. Use when the user wants to benchmark on ll_CrackSeg9k, LCSD, or asks about evaluating this task. Reports mIOU.

researchpythonperformance
0
3
Craft EvalA

Evaluates the ability of instruction-tuned LLMs to perform domain-specific multiple-choice question answering and text generation tasks. It measures how well models fine-tuned on synthetic, corpus-retrieved data generalize to held-out human-annotated benchmarks in biology, medicine, commonsense, recipe generation, and summarization. Use when the user wants to benchmark on ScienceQA (BioQA), MedMCQA (MedQA), CommonsenseQA 2.0 (CSQA), RecipeNLG (RecipeGen), CNN-DailyMail (Summarization), or ask...

researchpythongo
0
3
Crag EvalA

Evaluates the factual reliability and hallucination resistance of Retrieval-Augmented Generation (RAG) systems on realistic, dynamic, and long-tail questions. It measures how well models avoid generating incorrect information and appropriately abstain when knowledge is missing. Use when the user wants to benchmark on CRAG, or asks about evaluating this task. Reports truthfulness.

researchpythongo
0
3
Craibench EvalA

CrAIBench probes the robustness of Web3 AI agents against context manipulation attacks, specifically memory injection and prompt injection. It evaluates whether agents can maintain user intent and resist adversarial goals when malicious instructions are embedded in historical memory or active prompts. Use when the user wants to benchmark on CrAIBench, or asks about evaluating this task. Reports Targeted Attack Success Rate (ASR).

researchpythongo
0
3
CramersvA

Compute the CramersV metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute CramersV, or asks how to score with CramersV.

documentationpythongo
0
3
CramervonmisesA

Compute the cramervonmises metric — provided by scipy.stats. Use when the user has predictions and ground-truth and needs to compute cramervonmises, or asks how to score with cramervonmises.

documentationpythongo
0
3
Craw4llm EvalA

Evaluates the efficiency and data quality of web crawling strategies for LLM pretraining by measuring downstream model performance after training on crawled or selected documents. It compares graph-connectivity-based, random, and pretraining-influence-based URL scoring methods against an oracle baseline. Use when the user wants to benchmark on ClueWeb22-A (English subset), or asks about evaluating this task. Reports Average performance on 22 core tasks (DCLM evaluation recipe).

researchpythongo
0
3
Crbench EvalA

Evaluates whether multimodal large language models can perform genuine visual reasoning on charts by inferring values from axes and scales, rather than relying on OCR or pre-existing annotations. It probes the model's ability to interpret complex visual structures and perform multi-step estimation on both synthetic and real-world charts. Use when the user wants to benchmark on CRBench, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Creatidesign EvalA

This benchmark evaluates a diffusion model's ability to generate graphic designs that precisely adhere to multiple heterogeneous conditions, including primary visual subjects, secondary layout elements, and textual prompts. It probes fine-grained multi-subject preservation, semantic layout alignment, and overall compositional harmony. Use when the user wants to benchmark on CreatiDesign Validation Set, or asks about evaluating this task. Reports Avg. Score.

designpythongo
0
3
Creation Mmbench EvalA

Evaluates context-aware creative intelligence in multimodal and text-only models by assessing their ability to generate creative, contextually relevant content while maintaining visual factuality across diverse functional and creative writing tasks. Use when the user wants to benchmark on Creation-MMBench, or asks about evaluating this task. Reports VFS, Reward.

researchpythongit
0
3
Creative Fatigue Detection EvalA

Evaluates change point detection algorithms for identifying the onset of creative fatigue in digital advertising campaigns. It probes early warning capabilities, precision-recall trade-offs, and detection latency against ground-truth performance degradation events. Use when the user wants to benchmark on Synthetic Gradual Decline, Synthetic Sharp Decline, or asks about evaluating this task. Reports Delay (days).

datapythongo
0
3
Credit Card Fraud Detection EvalA

Evaluates a model's ability to detect fraudulent credit card transactions in a streaming context by learning topological and sequential patterns from transaction graphs without manual feature engineering. Use when the user wants to benchmark on Credit Card Transaction Dataset (Feb-Sep), or asks about evaluating this task. Reports AP.

researchpythonnode
0
3
Creditprint EvalA

Evaluates a model's ability to predict user creditworthiness based on geographic mobility footprints. It probes whether spatiotemporal visitation patterns and region-level credit signals can reliably distinguish users who pay their mobile phone bills from those who do not. Use when the user wants to benchmark on Hangzhou user mobility dataset, or asks about evaluating this task. Reports AUC.

researchpython
0
3
Crew Wildfire EvalA

Probes LLM-based multi-agent coordination in dynamic, partially observable wildfire disaster response scenarios. It evaluates capabilities such as spatial reasoning, task designation, plan adaptation, and heterogeneous team collaboration under stochastic dynamics and long-horizon objectives. Use when the user wants to benchmark on CREW-Wildfire, or asks about evaluating this task. Reports task success.

researchpythongo
0
3
Crews Ews Budget EvalA

Evaluates an AI system's ability to extract and classify budget allocations for Early Warning System (EWS) investments from heterogeneous financial PDF reports. It probes multi-label classification, numerical budget extraction with tolerance, and evidence retrieval/mapping in climate finance contexts. Use when the user wants to benchmark on MDB Evidence Set, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Crisismmd Classification EvalA

Evaluates multimodal deep learning models for disaster response by classifying social media posts into informativeness categories (informative vs. not-informative) and humanitarian content categories (e.g., affected individuals, rescue efforts, infrastructure damage). It tests the model's ability to jointly learn from text and image modalities to improve classification performance over unimodal baselines. Use when the user wants to benchmark on CrisisMMD, or asks about evaluating this task. R...

content-marketingpythongo
0
3
Criteo Attribution Bidding EvalA

Evaluates the efficiency of display advertising bidding strategies by measuring how well they align advertiser payments with actual conversion attribution. It probes the model's ability to predict conversion probability and adjust bids dynamically to avoid overbidding after early clicks, ultimately maximizing advertiser utility under budget constraints. Use when the user wants to benchmark on Criteo Attribution Dataset, or asks about evaluating this task. Reports U_A.

businesspythongo
0
3
Criteo Ctr EvalA

Evaluates the predictive quality and system efficiency of deep learning recommendation models on click-through rate prediction. It measures how well parameter-sharing compression techniques maintain model accuracy while reducing memory footprint and improving training and inference latency. Use when the user wants to benchmark on criteo-kaggle, criteo-tb, or asks about evaluating this task. Reports AUC.

researchpythongo
0
3
Criteo Dlrm EvalA

Evaluates the predictive accuracy and inference efficiency of deep learning recommendation models (DLRM) with compressed embedding tables on large-scale advertising click-through rate datasets. It measures Area Under the ROC Curve (AUC) to assess model quality and samples per second to quantify inference throughput under memory-constrained conditions. Use when the user wants to benchmark on CriteoTB, Criteo Kaggle, or asks about evaluating this task. Reports AUC.

researchpythongo
0
3
Critical Icu Prediction EvalA

Evaluates traditional machine learning and deep learning models on a large-scale, multi-institutional OMOP CDM dataset for ICU clinical prediction. It probes the ability of models to forecast patient outcomes (mortality, length of stay, readmission, sepsis) using early admission temporal features. Use when the user wants to benchmark on CRITICAL, or asks about evaluating this task. Reports AUROC.

datapythontesting
0
3
Critical Point Uncertainty EvalA

Evaluates the correctness and computational efficiency of closed-form algorithms for computing critical point probabilities in 2D scalar fields under various parametric and nonparametric noise models. The protocol compares these analytical solutions against Monte Carlo sampling baselines across synthetic and real-world scientific datasets to validate accuracy and speed. Use when the user wants to benchmark on Ackley function (synthetic), Gaussian mixture model (synthetic), E3SM climate data, ...

researchpythongo
0
3
Criticality Detection AccuracyA

Evaluates the ability to detect the onset of systemic instability (criticality) in simulated AI systems by monitoring performance variance across multiple benchmarks. It probes whether a derivative-based threshold can reliably flag phase transitions before functional collapse. Use when the user has predictions and gold and needs to compute percentage of correct classifications.

researchpythongo
0
3
CriticalsuccessindexA

Compute the CriticalSuccessIndex metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute CriticalSuccessIndex, or asks how to score with CriticalSuccessIndex.

documentationpythonperformance
0
3
Crl Biometric EvalA

Evaluates a model's ability to learn generalizable biometric feature representations in a continual learning setting, specifically measuring generalization to unseen identities across sequential learning steps rather than retaining knowledge of previously seen classes. Use when the user wants to benchmark on CRL-face, CRL-person, LFW, Megaface, or asks about evaluating this task. Reports Top 1 accuracy.

researchpythongit
0
3
Crohme Hme100k EvalA

Evaluates handwritten mathematical expression recognition (HMER) by measuring exact LaTeX sequence matching, tolerant symbol-level error rates, and structural tree prediction accuracy on complex handwritten formulas. Use when the user wants to benchmark on CROHME, HME100K, or asks about evaluating this task. Reports ExpRate.

researchpythongo
0
3
Croma EvalA

Evaluates self-supervised remote sensing representations across classification and segmentation tasks using optical and radar-optical inputs. Probes representation quality via finetuning, linear/nonlinear probing, kNN, and clustering. Use when the user wants to benchmark on BigEarthNet, fMoW-Sentinel, EuroSAT, Canadian Cropland, DFC2020, DW-Expert, MARIDA, or asks about evaluating this task. Reports mAP, Top 1 Acc., mIoU.

researchpythongo
0
3
Crop And Zoom Tool Use EvalA

Evaluates vision-language models' ability to use a crop-and-zoom tool for high-resolution visual question answering, disentangling intrinsic capability improvements from tool-induced gains and harms across multiple benchmarks. Use when the user wants to benchmark on VStar, HR-Bench 4k/8k, VisualProbe Easy/Medium/Harm, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Cross Continual Rl EvalA

Evaluates continual reinforcement learning capabilities in robotic simulation, specifically measuring how well agents retain performance on previously learned tasks while learning new sequential tasks. It probes catastrophic forgetting, transfer effects, and intrinsic task difficulty across line-following, object-pushing, and reaching benchmarks. Use when the user wants to benchmark on CRoSS, or asks about evaluating this task. Reports average cumulated score.

researchpythongo
0
3
Cross Domain Ctr EvalA

Evaluates cross-domain knowledge transfer for click-through rate (CTR) prediction by measuring how well a model trained on a source domain generalizes to a target domain with non-overlapping features. It probes context-aware feature translation and explicit knowledge augmentation in recommendation systems. Use when the user wants to benchmark on Amazon, Taobao, Alibaba Production, or asks about evaluating this task. Reports AUC.

researchpythongo
0
3
Cross Domain Meta Dl EvalA

Probes few-shot image classification generalization across diverse domains and highly variable task regimes (2–20 ways, 1–20 shots) without relying on pre-trained backbones. Use when the user wants to benchmark on Meta-Album, or asks about evaluating this task. Reports accuracy.

researchpythontesting
0
3
Cross Domain Object Detection EvalA

Evaluates an object detection model's ability to generalize across domain shifts (e.g., real-to-artistic, clear-to-foggy, synthetic-to-real) using only labeled source data and unlabeled target data during training. It measures how well the model mitigates domain bias and adapts to unseen target distributions without target annotations. Use when the user wants to benchmark on PASCAL VOC 2007+2012, Clipart1k, Watercolor2k, Cityscapes, Foggy Cityscapes, SIM10K, or asks about evaluating this task...

researchpythongo
0
3
Cross Domain Sequential Rec EvalA

This evaluation probes a model's ability to perform cross-domain sequential recommendation by leveraging user interaction histories across two domains, even when user overlap is minimal or absent. It measures how well the model captures domain-specific and shared sequential patterns to rank candidate items accurately. Use when the user wants to benchmark on Micro Video, Amazon, or asks about evaluating this task. Reports AUC.

researchpythongo
0
3
Cross Domain Text To Sql EvalA

Evaluates the reliability of cross-domain text-to-SQL benchmarks by exposing flaws in automated metrics like execution accuracy and exact set match, and by introducing human-in-the-loop validation to handle schema ambiguity and query equivalence. Use when the user wants to benchmark on Spider, Spider-DK, BIRD, or asks about evaluating this task. Reports Execution Accuracy.

databasespythongo
0
3
Cross Lingual F5 Tts EvalA

Evaluates the intelligibility, speaker similarity, and naturalness of synthesized speech in cross-lingual voice cloning and TTS scenarios. It also measures the accuracy of a language-agnostic speaking rate predictor for duration modeling across multiple languages. Use when the user wants to benchmark on Emilia, Seed-TTS-eval, LibriSpeech-PC test-clean, FLEURS, or asks about evaluating this task. Reports WER.

researchpythongit
0
3