All authors
qhjqhj00 avatar

Claude Skills by qhjqhj00

github.com/qhjqhj00
7,636 skillsA× 7,623B× 11C× 1D× 10 installs2,053 views
Chartqa EvalA

Evaluates multimodal models' ability to answer questions about charts by extracting visual and textual information. It probes robustness to missing labels and geometric perturbations, distinguishing between simple pattern matching and true structural reasoning. Use when the user wants to benchmark on ChartQA, Charixv, or asks about evaluating this task. Reports Relaxed Accuracy (RA).

researchpythongo
0
3
Chartsumm EvalA

Evaluates automatic chart-to-text summarization models on their ability to generate accurate, fluent, and informative summaries from chart metadata and data tables. It probes factual correctness, trend capture, and hallucination resistance across short and long summary formats. Use when the user wants to benchmark on ChartSumm, or asks about evaluating this task. Reports BLEU.

researchpythongit
0
3
Chartverse EvalA

Evaluates visual language models on complex chart understanding and reasoning tasks, including data extraction, comparison, and trend analysis across diverse chart types and domains. Use when the user wants to benchmark on ChartQA-Pro, CharXiv, ChartMuseum, ChartX, ChartBench, EvoChart, or asks about evaluating this task. Reports average score.

researchpythongo
0
3
Charxiv EvalA

Evaluates multimodal large language models' ability to understand real-world charts through descriptive and reasoning tasks. It probes capabilities like information extraction, pattern recognition, counting, compositional reasoning, and robustness to chart complexity (e.g., multiple subplots) and unanswerable queries. Use when the user wants to benchmark on CharXiv, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Chatbot Arena EvalA

Evaluates large language models by collecting human preference votes on pairwise responses to real-world prompts, then ranks them using Bradley-Terry models to measure alignment and real-world utility. Use when the user wants to benchmark on Chatbot Arena, or asks about evaluating this task. Reports BT coefficients.

researchpythongo
0
3
Chatgpt Multitask EvalA

Evaluates ChatGPT's zero-shot multitask capabilities across summarization, machine translation, sentiment analysis, question answering, dialogue, and misinformation detection. It probes the model's generalization, reasoning, multilingual understanding, and task-specific performance without fine-tuning. Use when the user wants to benchmark on CNN/DM, SAMSum, FLoRes-200, NusaX, bAbI, EntailmentBank, CLUTRR, StepGame, Pep-3k, COVID-Social, COVID-Scientific, MultiWOZ2.2, OpenDialKG, or asks about...

researchpythongo
0
3
Chatgpt Nlp EvalA

Evaluates ChatGPT's few-shot generation and reasoning capabilities across multiple NLP tasks, including question answering, commonsense reasoning, natural language inference, and sentiment analysis. It probes the model's ability to follow task-specific formalizations, leverage retrieved demonstrations, and mitigate hallucination through self-verification. Use when the user wants to benchmark on SQuADv2, TQA, MRQA-OOD, CSQA, StrategyQA, RTE, CommitmentBank, SST-2, IMDB, Yelp, or asks about eva...

researchpythongo
0
3
Chatqa Conversational Qa EvalA

Evaluates multi-turn conversational question answering over both long and short documents, including text-only and tabular contexts. It probes a model's ability to retrieve relevant context, handle topic switching, perform arithmetic reasoning, and correctly identify unanswerable queries. Use when the user wants to benchmark on Doc2Dial, QuAC, QReCC, TopiOCQA, INSCIT, CoQA, DoQA, ConvFinQA, SQA, HybridDial, or asks about evaluating this task. Reports Average F1/EM across 10 datasets.

researchpythongo
0
3
Chatterbox Mrg EvalA

Evaluates a model's ability to perform multi-round multimodal referring and grounding, requiring logical consistency across dialogue turns while generating accurate text responses and bounding box coordinates for visual instances. Use when the user wants to benchmark on CB-LC, RefCOCOg, COCO 2017, or asks about evaluating this task. Reports BERT(·).

researchpythongo
0
3
Chattime EvalA

Evaluates a multimodal time series foundation model on zero-shot forecasting, context-guided forecasting, and time series question answering. It probes the model's ability to handle discretized numerical time series alongside textual prompts for prediction and feature recognition without task-specific fine-tuning. Use when the user wants to benchmark on Electric, Exchange, Traffic, Weather, and ETT datasets, Context-guided multimodal datasets, Synthesized TSQA dataset, or asks about evaluatin...

researchpythongo
0
3
Chc Comp21 Lia Lin EvalA

Evaluates the effectiveness of bottom-up versus top-down transformations for solving linear constrained Horn clauses (CHCs) using software verification workflows. It measures how many verification tasks a solver can successfully resolve within a strict time limit. Use when the user wants to benchmark on CHC-COMP21 LIA-Lin track, or asks about evaluating this task. Reports solved_tasks.

researchpython
0
3
Chd CmmsA

Evaluates the perceptual quality and human preference alignment of generative image models by comparing their discrete visual token distributions and sequence statistics against human ratings. Use when the user has predictions and gold and needs to compute CHD, CMMS.

researchpythongo
0
3
Chebi 20 Mm EvalA

Evaluates large language models' ability to process, generate, and retrieve molecular information across multiple modalities (SMILES, InChI, SELFIES, graphs, captions, IUPAC names, images). It probes cross-modal compatibility, chemical knowledge acquisition, and molecular property prediction capabilities. Use when the user wants to benchmark on ChEBI-20-MM, or asks about evaluating this task. Reports ROC_AUC.

researchpythongo
0
3
Chemcotb EvalA

Evaluates large language models' ability to perform stepwise chemical reasoning and molecular structure manipulation. It probes capabilities in molecular understanding, functional group/ring recognition, scaffold extraction, and SMILES-based molecule editing and reaction prediction. Use when the user wants to benchmark on ChemCoTBench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Chemllmbench EvalA

Evaluates LLM-based chemistry agents on four core computational chemistry tasks: text-to-molecule generation, molecule-to-text captioning, molecular property prediction, and reaction product prediction. It probes the model's ability to translate between chemical language and structures, predict physicochemical properties, and correct tool errors via hierarchical agent stacking. Use when the user wants to benchmark on ChemLLMBench, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Chemo EvalA

Evaluates multimodal large language models' ability to solve Olympiad-level theoretical chemistry problems requiring visual perception, chemical reasoning, and structured problem-solving. It specifically probes the visual perception bottleneck in chemistry tasks and tests the effectiveness of multi-agent orchestration and structured visual enhancement. Use when the user wants to benchmark on ChemO, or asks about evaluating this task. Reports normalized rubric-based score.

researchpythonperformance
0
3
Chess Puzzle EvalA

Evaluates a model's ability to play chess and solve tactical puzzles without explicit search, testing its capacity for long-horizon planning and generalization to novel board states. Use when the user wants to benchmark on Lichess puzzles, or asks about evaluating this task. Reports Lichess Elo.

researchpythongo
0
3
Chest Xray Classification EvalA

Evaluates a CNN's ability to classify chest X-ray images into disease categories (COVID-19, pneumonia, tuberculosis, normal) using various preprocessing techniques. It probes robustness across different dataset sizes and class distributions. Use when the user wants to benchmark on Multiclass Chest X-ray Dataset, Hamad Medical Corporation Tuberculosis Dataset, Pneumonia Dataset, NIH Chest X-ray Dataset, or asks about evaluating this task. Reports AUC.

researchpythongo
0
3
Chest Xray Report Generation EvalA

This evaluation probes a model's ability to generate clinically coherent and accurate natural language reports from chest X-ray images. It measures both surface-level linguistic similarity to ground-truth radiology reports and the clinical correctness of extracted pathological findings. Use when the user wants to benchmark on Indiana U. Chest X-Ray, MIMIC-CXR, or asks about evaluating this task. Reports BLEU-1.

researchpythongo
0
3
Chest Xray14 EvalA

Evaluates a model's ability to perform multi-label classification of 14 thoracic diseases on chest X-ray images and localize pathological regions using attention maps. It probes whether anatomically grounded feature weighting improves disease detection over global feature fusion or saliency-based methods. Use when the user wants to benchmark on Chest X-ray14, or asks about evaluating this task. Reports AUROC.

researchpythongo
0
3
Chestagentbench EvalA

Evaluates an AI agent's ability to perform multi-step medical reasoning and tool orchestration for chest X-ray interpretation. It probes capabilities across seven clinically relevant categories: detection, classification, localization, comparison, relationship, diagnosis, and characterization. Use when the user wants to benchmark on ChestAgentBench, or asks about evaluating this task. Reports Accuracy (%).

researchpythongo
0
3
Chestx Det EvalA

Evaluates instance-level detection and segmentation of thoracic diseases on chest X-rays. It probes a model's ability to localize and classify 13 disease categories while handling domain-specific challenges like ambiguous boundaries and disease co-occurrence. Use when the user wants to benchmark on ChestX-Det, DR-private, or asks about evaluating this task. Reports AP_50^bb.

researchpythongo
0
3
Chestxray Pneumonia EvalA

Evaluates deep learning models for binary classification of chest X-rays into normal versus pneumonia categories, while also assessing the spatial interpretability of model predictions using Grad-CAM heatmaps. Use when the user wants to benchmark on Chest X-Rays dataset, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Chestxray14 Bias EvalA

Evaluates the effectiveness of attribute-neutralization models in removing demographic bias (sex and age) from chest X-ray images while preserving diagnostic utility for 15 disease findings. It probes the trade-off between demographic leakage suppression and clinical performance across varying edit intensities. Use when the user wants to benchmark on ChestX-ray14, or asks about evaluating this task. Reports AI-Judge AUC.

researchpythonperformance
0
3
Chestxray14 Multi Label EvalA

Evaluates deep learning models for multi-label chest X-ray pathology classification. It probes the capability of CNN architectures to detect 14 specific thoracic diseases, testing the impact of transfer learning, network depth, and fusion of non-image patient demographics (age, gender, view position) on diagnostic accuracy. Use when the user wants to benchmark on ChestX-ray14, or asks about evaluating this task. Reports AUC.

researchpythontesting
0
3
Chestxray14 Plco EvalA

Evaluates a model's ability to detect and localize multiple pathologies in high-resolution chest X-ray images. It specifically probes the model's robustness to severe class imbalance and its capacity to leverage explicit spatial location information for pathology classification. Use when the user wants to benchmark on ChestX-Ray14, PLCO, or asks about evaluating this task. Reports AUC.

researchpythonperformance
0
3
Chestxray8 EvalA

Evaluates weakly-supervised multi-label classification and spatial localization of eight common thoracic diseases on chest X-rays. It probes a model's ability to detect disease presence from image-level labels and localize pathological regions using only bounding box annotations during testing. Use when the user wants to benchmark on ChestX-ray8, or asks about evaluating this task. Reports AUC.

researchpythontesting
0
3
Chex EvalA

Evaluates a vision-language model's ability to perform interactive localization, region classification, and text generation on chest X-rays. It probes zero-shot multitask capabilities, including sentence grounding, pathology detection, and customizable report generation. Use when the user wants to benchmark on MS-CXR, VinDrCXR, NIH8, CIG, MIMIC-CXR, or asks about evaluating this task. Reports mAP.

researchpythongo
0
3
Chexbench EvalA

Evaluates vision-language foundation models on chest X-ray interpretation across two axes: image perception (view classification, disease identification/classification, VQA, reasoning) and textual understanding (findings generation, summarization). It probes clinical reasoning, visual grounding, and medical text generation capabilities. Use when the user wants to benchmark on MIMIC-CXR, CheXpert, SIIM, RSNA, OpenI, SLAKE, Rad-Restruct, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Chexficient EvalA

Evaluates a chest X-ray vision-language foundation model across zero-shot classification, cross-modal retrieval, and adapted downstream tasks (classification, segmentation, report generation). It specifically probes data and compute efficiency, as well as the model's ability to represent long-tailed thoracic diseases without aggressive scaling. Use when the user wants to benchmark on SIIM-PTX, Pneumonia2017, TBX11K, CheXpert, MIMIC-CXR, ChestX-ray14, VinDr-CXR, VinDr-PCXR, or asks about evalu...

researchpythongo
0
3
Chexgenbench EvalA

Evaluates text-to-image generative models for synthetic chest radiograph generation across three dimensions: fidelity (image quality and mode coverage), privacy (memorization and re-identification risks), and utility (downstream clinical performance for classification and report generation). It probes whether synthetic medical images can match real data in clinical tasks while preserving patient privacy. Use when the user wants to benchmark on MIMIC-CXR, or asks about evaluating this task. Re...

researchpythongo
0
3
Chexpert Atelectasis EvalA

Evaluates the diagnostic accuracy of predictive algorithms and human radiologists on chest X-ray images for detecting atelectasis. It specifically probes whether human expertise provides actionable, non-redundant information on input subsets where algorithms are algorithmically indistinguishable. Use when the user wants to benchmark on CheXpert, or asks about evaluating this task. Reports Matthews Correlation Coefficient (MCC).

researchpythongo
0
3
Chexpert Label Extraction EvalA

Evaluates an automated rule-based pipeline's ability to extract clinical observations from free-text radiology reports. It specifically tests the system's capacity to classify mentions as positive, negative, or uncertain, and to aggregate them into structured labels for 14 predefined observations. Use when the user wants to benchmark on CheXpert, or asks about evaluating this task. Reports F1 score.

researchpythongo
0
3
Chexradinet EvalA

Evaluates a multi-task deep learning framework for thorax disease classification and weakly-supervised localization on chest X-rays. It probes the model's ability to detect multiple pathologies and accurately localize abnormal regions without requiring pre-annotated bounding boxes during training. Use when the user wants to benchmark on NIH Chest X-ray, CheXpert, MIMIC-CXR, or asks about evaluating this task. Reports AUC.

researchpythonperformance
0
3
Chexternal EvalA

Evaluates the generalization of chest X-ray deep learning models to two clinically relevant distribution shifts: smartphone photographs of digital X-rays (introducing visual artifacts like blur and glare) and external institutional data. It probes robustness to imaging degradation and cross-institutional heterogeneity in multi-label pathology detection. Use when the user wants to benchmark on CheXpert, NIH, or asks about evaluating this task. Reports MCC.

researchpythongo
0
3
Chextransfer EvalA

This evaluation probes the transferability of ImageNet-pretrained CNN architectures to chest X-ray interpretation, measuring how model size, architecture family, and pretraining affect classification performance and parameter efficiency on the CheXpert dataset. Use when the user wants to benchmark on CheXpert, or asks about evaluating this task. Reports CheXpert AUC.

researchpythongo
0
3
Chickenpox Hungary Forecasting EvalA

Evaluates the ability of recurrent graph neural networks to forecast spatiotemporal epidemiological time series. It probes how well models capture spatial dependencies between adjacent regions and temporal dynamics like seasonality and zero-inflation over multiple forecasting horizons. Use when the user wants to benchmark on Chickenpox Cases in Hungary, or asks about evaluating this task. Reports mean squared error.

researchpythonnode
0
3
Chikha Po EvalA

Evaluates fundamental lexical comprehension and generation capabilities of multilingual LLMs across thousands of languages. It probes word-level translation, context-aware translation, translation-conditioned language modeling, and bag-of-words machine translation tasks to measure basic linguistic competence beyond high-resource languages. Use when the user wants to benchmark on ChiKhaPo, or asks about evaluating this task. Reports language score.

researchpythongo
0
3
Childsafe Safety EvalA

Evaluates LLM safety alignment across four child developmental stages (ages 6–17) using simulated agents grounded in developmental psychology. It probes how models handle sensitive contexts, boundary-testing, and age-specific cognitive limitations in multi-turn interactions. Use when the user wants to benchmark on ChildSafe Dataset, or asks about evaluating this task. Reports semantic_safety_score.

ai-agentspythontesting
0
3
Chime2 Robust Asr EvalA

Evaluates the effectiveness of a time-domain speech enhancement frontend in improving automatic speech recognition performance on noisy and reverberant speech. It probes the model's ability to enhance speech without introducing distortion that degrades downstream ASR accuracy. Use when the user wants to benchmark on CHiME-2, or asks about evaluating this task. Reports WER.

researchpythonfrontend
0
3
Chime3 Se EvalA

This benchmark evaluates speech enhancement models by measuring how well they clean noisy speech features before they are processed by a downstream automatic speech recognition (ASR) system. It probes the model's ability to preserve speech structure and reduce noise in challenging real-world far-field conditions. Use when the user wants to benchmark on CHiME-3, or asks about evaluating this task. Reports WER.

researchpythontesting
0
3
Chime4 Adapter EvalA

Evaluates the robustness of Automatic Speech Recognition (ASR) models under various real-world noise conditions (bus, cafe, pedestrian, street junction) using adapter-based fine-tuning and speech enhancement front-ends. Use when the user wants to benchmark on CHiME-4, or asks about evaluating this task. Reports WER.

researchpythontesting
0
3
Chime4 Aishell4 Asr EvalA

Evaluates multi-channel end-to-end speech recognition systems in noisy, real-world environments using neural beamforming front-ends combined with CTC-CRF acoustic models. It probes the model's ability to leverage single-channel data via pre-training, data scheduling, or simulation to improve robustness and accuracy. Use when the user wants to benchmark on CHiME4, AISHELL-4, or asks about evaluating this task. Reports WER.

researchpythonshell
0
3
Chime4 Ami Asr EvalA

Evaluates automatic speech recognition systems on real-world meeting and close-talking microphone speech. It measures how well acoustic models trained on raw waveforms generalize to challenging multi-microphone environments compared to traditional feature-based baselines. Use when the user wants to benchmark on CHiME4, AMI, or asks about evaluating this task. Reports WER.

researchpythongo
0
3
Chime4 Asr Robustness EvalA

Evaluates the robustness of automatic speech recognition systems in noisy environments by measuring word error rates on enhanced speech from speaker extraction models across matched and mismatched noisy conditions. Use when the user wants to benchmark on CHiME-4, VoiceBank-DEMAND, WHAM!, or asks about evaluating this task. Reports WER.

researchpythongo
0
3
Chime4 EvalA

Evaluates end-to-end speech recognition and speech enhancement performance in noisy, reverberant multi-channel conditions. It probes the model's ability to jointly dereverberate, denoise, and transcribe speech using self-supervised learning representations. Use when the user wants to benchmark on CHiME-4, or asks about evaluating this task. Reports WER.

researchpythongo
0
3
Chimera Extraction EvalA

Evaluates a model's ability to identify scientific idea recombinations in research abstracts, extract the involved scientific concepts (entities), and classify the type of recombination relation (inspiration or blend) between them. Use when the user wants to benchmark on CHIMERA, or asks about evaluating this task. Reports Precision, Recall, F1.

researchpythongo
0
3
Chinesafe EvalA

This benchmark evaluates the safety of large language models in Chinese by testing their ability to correctly classify text as safe or unsafe across multiple sensitive categories. It probes whether models can reliably detect harmful, policy-violating, or sensitive content in a Chinese-language context using both generation-based and perplexity-based strategies. Use when the user wants to benchmark on ChineseSafe, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Chinese Llm Bench EvalA

Evaluates Chinese large language models' world knowledge, academic understanding, and multi-dimensional alignment after pretraining or instruction fine-tuning. It probes the model's ability to follow instructions, reason across domains, and maintain safety and helpfulness standards in Chinese. Use when the user wants to benchmark on C-Eval, CMMLU, Alignbench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Chinese Llm Benchmarks EvalA

Evaluates the knowledge, reasoning, instruction-following, and safety alignment capabilities of Chinese instruction-tuned LLMs across academic, professional, open-ended, and safety-critical domains. Use when the user wants to benchmark on C-Eval, CMMLU, BELLE-EVAL, SafetyBench, or asks about evaluating this task. Reports log-likelihood.

researchpythongo
0
3