All authors
qhjqhj00 avatar

Claude Skills by qhjqhj00

github.com/qhjqhj00
7,636 skillsA× 7,623B× 11C× 1D× 10 installs2,229 views
Med Tiv EvalA

Evaluates a verifier model's ability to distinguish correct from erroneous reasoning traces in medical question-answering tasks. It measures how well tool-integrated reinforcement learning improves factual justification and reduces hallucination compared to static reward models. Use when the user wants to benchmark on MedQA, MedMCQA, MMLU-Med, MedXpertQA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Med Vqa EvalA

Evaluates multimodal models on medical visual question answering across diverse imaging modalities. It probes intrinsic visual reasoning capabilities and extrinsic biomedical knowledge grounding, while measuring the model's ability to minimize clinical hallucinations. Use when the user wants to benchmark on VQA-RAD, SLAKE, ProbMed, or asks about evaluating this task. Reports accuracy/recall (closed/open-ended).

researchpythongo
0
3
Medagents Medical Qa EvalA

Evaluates zero-shot medical reasoning and multiple-choice question answering capabilities of LLMs using a training-free multi-agent collaboration framework. It probes the model's ability to simulate domain expert role-playing and reach consensus without retrieval-augmented generation. Use when the user wants to benchmark on MedQA, MedMCQA, PubMedQA, MMLU Anatomy, MMLU Clinical Knowledge, MMLU College Medicine, MMLU Medical Genetics, MMLU Professional Medicine, MMLU College Biology, or asks ab...

ai-agentspythongo
0
3
Medarabiq EvalA

Evaluates large language models on Arabic medical reasoning and dialogue across multiple-choice, fill-in-the-blank, and open-ended Q&A tasks. It probes factual accuracy, domain-specific knowledge, and robustness to linguistic variations and injected biases in healthcare contexts. Use when the user wants to benchmark on MedArabiQ, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Medbench EvalA

This benchmark evaluates Chinese large language models on clinical knowledge, diagnostic reasoning, and conversational ability. It probes performance across three standardized medical licensing exams and real-world clinical case scenarios, highlighting gaps in multi-hop reasoning, diagnostic precision, and response fluency. Use when the user wants to benchmark on MedBench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Medbert De Med Bench EvalA

Evaluates German medical language models on classification and named entity recognition tasks across radiology reports, clinical discharge notes, surgery reports, and public medical/general benchmarks. Use when the user wants to benchmark on Chest CT, Chest X-Ray, ICD-10 code classification on discharge notes, OPS code classification on discharge notes, OPS code classification on surgery reports, GermEval-18, Wrist NER, GraSCCo, GGPOnc, or asks about evaluating this task. Reports AUROC.

researchpythonperformance
0
3
Medbert Disease Prediction EvalA

Evaluates disease prediction performance using pre-trained contextualized embeddings on structured electronic health records. It probes the model's ability to capture temporal dependencies and long-term patient history from ICD-coded visit sequences, particularly in low-data transfer-learning scenarios. Use when the user wants to benchmark on DHF-Cerner, PaCa-Cerner, PaCa-Truven, or asks about evaluating this task. Reports AUC.

researchpythongo
0
3
Medbrowsecomp EvalA

Evaluates AI agents' ability to perform multi-hop, evidence-grounded medical information retrieval from live, heterogeneous web sources. It probes long-horizon web navigation, tool allocation, source verification, and the capacity to reconcile conflicting or dense biomedical data. Use when the user wants to benchmark on MedBrowseComp-50, MedBrowseComp-605, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Medcalc Bench EvalA

Probes LLMs' ability to perform evidence-based medical calculations by decomposing the task into formula selection, entity extraction, arithmetic computation, and final answer formatting. It evaluates both final numerical accuracy and granular step-wise reasoning to diagnose specific clinical and computational failure modes. Use when the user wants to benchmark on MedCalc-Bench, or asks about evaluating this task. Reports Step-wise LLM Evaluation.

researchpythongo
0
3
Medcalceval EvalA

Evaluates large language models' quantitative reasoning and clinical calculation capabilities across multiple medical specialties. It probes the model's ability to correctly select medical formulas or scoring rules, extract relevant patient attributes from clinical text, and perform accurate multi-step numerical computations. Use when the user wants to benchmark on MedCalc-Eval, MedCalc-Bench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Medcasereasoning EvalA

Evaluates large language models' ability to perform clinical diagnostic reasoning and arrive at correct final diagnoses based on patient case reports. It specifically probes whether models can align their step-by-step reasoning processes with clinician-authored diagnostic traces, rather than just guessing the final answer. Use when the user wants to benchmark on MedCaseReasoning, or asks about evaluating this task. Reports Diagnostic Accuracy.

ai-agentspythongo
0
3
Medclipseg EvalA

Evaluates data-efficient and domain-generalizable medical image segmentation using vision-language adaptation. It probes a model's ability to segment anatomical structures and lesions across diverse imaging modalities with limited supervision, while maintaining robustness to out-of-distribution domain shifts and providing calibrated uncertainty estimates. Use when the user wants to benchmark on BUSI, BTMRI, ISIC, Kvasir-SEG, QaTa-COV19, EUS, BUSUC, BUSBRA, BUID, UDIAT, CVC-ColonDB, CVC-Clinic...

researchpythonperformance
0
3
MedconceptsqaevalA

Evaluates large language models' ability to reason about and identify medical concepts (diagnoses, procedures, drugs) across different semantic hierarchies and difficulty levels. Use when the user wants to benchmark on MedConceptsQA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Meddocbench EvalA

Evaluates a model's ability to parse and reason over real-world medical documents, including laboratory test reports and general medical documents, through tasks like table extraction, simple/complex QA, and free-form scoring. Use when the user wants to benchmark on MedDocBench, or asks about evaluating this task. Reports Field-level micro Precision/Recall/F1 & Macro-Doc F1.

researchpythongo
0
3
Medebench EvalA

Evaluates the reliability and clinical appropriateness of text-guided medical image editing models. It probes anatomical localization precision, preservation of surrounding clinical context, and overall visual realism across diverse medical imaging modalities and anatomical regions. Use when the user wants to benchmark on MedEBench, or asks about evaluating this task. Reports GPT-4o Editing Accuracy, Masked SSIM, GPT-4o Visual Quality.

researchpythonperformance
0
3
Medec EvalA

Evaluates large language models' ability to detect and correct medical errors in clinical text. It probes the model's sensitivity to clinical inaccuracies, its precision in localizing erroneous sentences, and its capacity to generate semantically and lexically accurate corrections using different prompting strategies. Use when the user wants to benchmark on MEDEC, or asks about evaluating this task. Reports AggScore.

ai-agentspythongo
0
3
Medeval EvalA

Evaluates language models on multi-level (sentence/document) and multi-task (NLU/NLG) medical benchmarks across diverse clinical domains. It probes a model's ability to perform clinical text classification, report code prediction, and medical report summarization using both fine-tuned PLMs and prompted LLMs. Use when the user wants to benchmark on MedEval, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Medexqa EvalA

Evaluates medical language models on multiple-choice question answering and the generation of clinically relevant explanations. It probes the model's ability to perform clinical reasoning, avoid hallucinations, and produce coherent, accurate rationales aligned with medical domain knowledge. Use when the user wants to benchmark on MedExQA, or asks about evaluating this task. Reports Classification Accuracy.

researchpythongo
0
3
Medfair EvalA

Evaluates group fairness and bias mitigation in medical imaging models across multiple datasets and modalities. It probes whether models trained with Empirical Risk Minimization (ERM) or explicit bias mitigation algorithms exhibit performance disparities across sensitive subgroups, and how model selection strategies impact worst-case group performance. Use when the user wants to benchmark on MEDFAIR Benchmark Suite, or asks about evaluating this task. Reports worst-case AUC.

researchpythongo
0
3
Medflowseg EvalA

Evaluates medical image segmentation accuracy across diverse imaging modalities (MRI, fundus, histology, ultrasound). Probes the model's ability to delineate anatomical structures and refine boundaries using a deterministic flow-matching framework. Use when the user wants to benchmark on ACDC, BraTS-2021, REFUGE-2, GlaS, CAMUS, or asks about evaluating this task. Reports Dice.

researchpythongo
0
3
Medformer Ur EvalA

Evaluates medical image classification performance and uncertainty calibration across multiple clinical imaging modalities (mammography, ultrasound, histopathology, MRI). Probes the model's ability to distinguish benign from malignant lesions and classify tumor types while providing reliable uncertainty estimates for selective prediction. Use when the user wants to benchmark on CBIS-DDSM, BUSI, Breast Histopathology (IDC), Brain MRI, or asks about evaluating this task. Reports ECE.

researchpythonperformance
0
3
Medframeqa EvalA

This benchmark evaluates multi-image medical visual question answering and clinical reasoning. It probes a model's ability to integrate diagnostic evidence across temporally coherent medical images, detect salient findings, and propagate reasoning chains to answer single-choice questions. Use when the user wants to benchmark on MedFrameQA, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Medfuse EvalA

Evaluates a multi-modal fusion model's ability to predict patient phenotypes and in-hospital mortality using partially paired clinical time-series data and chest X-ray images. It probes robustness to missing modalities and temporal alignment in ICU settings. Use when the user wants to benchmark on MIMIC-IV / MIMIC-CXR, or asks about evaluating this task. Reports AUROC.

researchpythongo
0
3
Medg Krp EvalA

Evaluates large language models' ability to generate causal knowledge graphs from single medical concepts. It probes biomedical reasoning, causal understanding, and factual consistency by comparing generated graphs against human expert judgments and a ground-truth biomedical ontology (BIOS). Use when the user wants to benchmark on MedG-KRP, or asks about evaluating this task. Reports Precision.

ai-agentspythonnode
0
3
Medhal EvalA

Evaluates AI models' ability to detect factual inconsistencies (hallucinations) in medical text and generate grounded explanations for why statements are non-factual. It probes domain-specific factual consistency reasoning and binary classification under clinical constraints. Use when the user wants to benchmark on MedHal, MedNLI, Hegselmann et al. (2024a) Hallucination Dataset, or asks about evaluating this task. Reports F1-score.

researchpythongo
0
3
Medhelm EvalA

Evaluates large language models on a comprehensive taxonomy of real-world clinical workflows, covering tasks like clinical note generation, patient communication, medical research assistance, clinical decision support, and administration/workflow. It probes both closed-ended factual/reasoning tasks and open-ended free-text generation capabilities in medical domains. Use when the user wants to benchmark on MedHELM, or asks about evaluating this task. Reports Macro-average performance.

researchpythongo
0
3
Medi Aug EvalA

Evaluates the impact of six mix-based data augmentation methods on medical image classification performance. It probes how well models generalize to medical domains (brain MRI and eye fundus) when trained with different augmentation strategies and backbones. Use when the user wants to benchmark on Brain Tumor Classification Dataset, Eye Diseases Classification Dataset, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Median Absolute ErrorA

Compute the median_absolute_error metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute median_absolute_error, or asks how to score with median_absolute_error.

documentationpython
0
3
Mediasum EvalA

Evaluates abstractive dialogue summarization models on long-form, multi-speaker media interviews. It measures how well models can condense multi-topic conversations into concise summaries, and assesses transfer learning capabilities to other dialogue domains. Use when the user wants to benchmark on MediaSum, or asks about evaluating this task. Reports ROUGE-1, ROUGE-2, ROUGE-L F1.

researchpython
0
3
Medic EvalA

Evaluates image classification models on disaster-related social media images across four interdependent tasks: disaster type, informativeness, humanitarian relevance, and damage severity. It tests both single-task and multi-task learning capabilities, including multiclass and multilabel classification settings. Use when the user wants to benchmark on MEDIC, or asks about evaluating this task. Reports weighted F1-score.

researchpythongit
0
3
Medical Ai Security EvalA

Probes the vulnerability of medical AI models to jailbreaking and privacy extraction attacks across different clinical specialties. It assesses how models handle synthetic patient data requests for harmful or sensitive information, measuring compliance rates and protected health information (PHI) leakage severity. Use when the user wants to benchmark on Medical AI Security Attack Scenarios, or asks about evaluating this task. Reports Attack Success Rate (ASR).

securitypythonsecurity
0
3
Medical Asr Denoising EvalA

Evaluates the robustness of modern medical ASR models to various noise conditions and assesses whether speech enhancement preprocessing improves or degrades transcription accuracy. Use when the user wants to benchmark on Medical ASR Recordings (unspecified), or asks about evaluating this task. Reports semWER.

researchpythonperformance
0
3
Medical Benchmarks EvalA

Evaluates large language models on medical knowledge, clinical reasoning, and safety alignment using multiple-choice and open-ended healthcare QA tasks. It measures standard accuracy across major medical benchmarks and quantifies unsafe response rates via automated safety classifiers. Use when the user wants to benchmark on MultiMedQA, MedMCQA, MedQA, PubMedQA, MMLU Med., CareQA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Medical Dataset Distillation EvalA

Evaluates the effectiveness of dataset distillation methods for medical imaging by training classification models on synthetic datasets and measuring their accuracy on held-out real test sets. It probes how well distilled images preserve class-discriminative features across diverse medical modalities, resolutions, and class imbalances. Use when the user wants to benchmark on COVID19-CXR, SKIN-HAM, BREAST-ULS, PATHMNIST, OCTMNIST, ORGAN3D, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Medical Diagnosis EvalA

Evaluates an LLM's ability to perform medical diagnosis through interactive symptom inquiry and confidence-based decision making. It probes the model's capacity to ask targeted yes/no questions, quantify diagnostic uncertainty, and accurately predict diseases from a candidate set or open database. Use when the user wants to benchmark on Muzhi, Dxy, DxBench, or asks about evaluating this task. Reports Acc..

researchpythongo
0
3
Medical Dialogue Gen EvalA

This evaluation probes a model's ability to generate clinically accurate, fluent, and context-aware physician responses in multi-turn medical dialogues. It assesses both automatic language quality and semantic relevance, alongside human-rated fluency, knowledge correctness, and overall satisfaction. Use when the user wants to benchmark on KaMed, MedDialog, MedDG, or asks about evaluating this task. Reports BLEU-2.

researchpythongo
0
3
Medical Emr Extraction EvalA

Probes a model's ability to extract structured clinical information from unstructured physician-patient consultation dialogues. It evaluates how accurately the model maps conversational text into predefined medical record fields such as chief complaint, diagnosis, and treatment recommendations. Use when the user wants to benchmark on EMRModel Dataset, or asks about evaluating this task. Reports weighted average F1 score.

researchpythongo
0
3
Medical Entity Linking EvalA

Evaluates a model's ability to predict semantic types for biomedical mentions and to link those mentions to standardized medical concepts. It probes how well type-based candidate filtering improves broad-coverage medical information extraction pipelines. Use when the user wants to benchmark on NCBI Disease Corpus, Bio CDR, ShARE, MedMentions, WIKIMED, PUBMEDDS, or asks about evaluating this task. Reports AUC (Area Under the Precision-Recall curve).

researchpythongo
0
3
Medical Image Classification EvalA

Evaluates the diagnostic accuracy and computational efficiency of CNNs versus multimodal LLMs on medical imaging tasks. It probes whether vision-language models can match traditional convolutional networks in classifying chest X-rays, MRIs, and CT scans, while also measuring prediction calibration and resource consumption. Use when the user wants to benchmark on Chest X-ray, Brain MRI, Chest CT, or asks about evaluating this task. Reports Accuracy.

researchpythonperformance
0
3
Medical Llm Merging EvalA

Evaluates the effectiveness of various model merging techniques for consolidating knowledge in medical large language models. It probes whether merged models can outperform their base and parent models across diverse medical and general reasoning benchmarks. Use when the user wants to benchmark on MedQA, PubMedQA, HellaSwag, MedMCQA, MMLU Professional Medicine, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Medical Mcqa EvalA

Evaluates a model's ability to answer medical multiple-choice questions by selecting the correct option from a set of distractors, including natural differential diagnoses. It tests both internal knowledge retrieval and the impact of synthetic pretraining with cue-masking strategies. Use when the user wants to benchmark on MedQA-USMLE, MedMCQA, DBPedia, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Medical Mllm Safety EvalA

Evaluates the safety robustness and medical capability of multimodal large language models against general and medical-specific threats, including cross-modality jailbreak attacks. It measures the trade-off between restoring safety guardrails and preserving domain-specific accuracy. Use when the user wants to benchmark on HarmBench, CATQA, HEx-PHI, MedSafetyBench, CARES, MedSentry, 3D-Tiny-1K, VQA_RAD, MedQA, PubMedQA, SuperGPQA, CMExam, Medbullets, or asks about evaluating this task. Reports...

researchpythongo
0
3
Medical Multimodal EvalA

Evaluates cross-modal understanding and generation capabilities in medical imaging, specifically image-report retrieval, radiology report generation, and multi-label disease diagnosis from chest X-rays. Use when the user wants to benchmark on MIMIC-CXR, IU-Xray, ChestX-ray 14, or asks about evaluating this task. Reports CIDEr.

researchpythongo
0
3
Medical Object Detection EvalA

Evaluates the ability of object detection models to localize and classify medical structures in 3D imaging data without manual hyperparameter tuning. It probes generalization across diverse anatomical regions and imaging modalities by testing on a held-out pool of datasets. Use when the user wants to benchmark on nnDetection Medical Object Detection Benchmark, or asks about evaluating this task. Reports mAP@0.1.

researchpythontesting
0
3
Medical Ood Detection EvalA

Evaluates the ability of out-of-distribution detection (OoDD) methods to correctly classify medical images as in-distribution or out-of-distribution across three distinct OoD categories: unrelated domains, incorrect image preparation, and unseen medical conditions due to selection bias. Use when the user wants to benchmark on Medical OoD Benchmark ($D_{test}$), or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Medical Qa EvalA

This benchmark evaluates the medical reasoning and question-answering capabilities of language models across multiple-choice and open-ended clinical tasks. It probes the model's ability to retrieve relevant medical knowledge, perform stepwise reasoning, and select or generate correct answers based on clinical guidelines and literature. Use when the user wants to benchmark on MedQA, MedMCQA, MMLU-Med, DDXPlus, AgentClinicNEJM, AgentClinicMedQA, or asks about evaluating this task. Reports accur...

researchpythongo
0
3
Medical Qa Explanation EvalA

Evaluates large language models' ability to answer challenging medical multiple-choice questions and generate step-by-step clinical reasoning explanations. It probes both factual accuracy in clinical decision-making and the quality of model-generated rationales compared to expert-written references. Use when the user wants to benchmark on JAMA Clinical Challenge, Medbullets, or asks about evaluating this task. Reports accuracy.

ai-agentspythongo
0
3
Medical Radiology Similarity EvalA

This evaluation probes the ability of automated metrics to capture deep clinical semantics in radiology reports. It compares LLM-generated similarity scores against traditional lexical overlap metrics, measuring how well each aligns with ground truth annotations derived from clinical NLP tools. Use when the user wants to benchmark on Radiology Report Pairs (CheXpert/NegBio-derived), or asks about evaluating this task. Reports GPT_sim.

researchpythonperformance
0
3
Medical Reasoning Benchmarks EvalA

Evaluates large language models' ability to perform medical reasoning across multiple-choice clinical questions, specialist-level board exams, and general-domain medical subsets. It probes factual knowledge integration, diagnostic accuracy, and reasoning under uncertainty in safety-critical settings. Use when the user wants to benchmark on MedQA (USMLE), MedMCQA (Validation), PubMedQA, GPQA, JMED, ReDis-QA, MedXpertQA, MMLU-Pro, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Medical Reasoning EvalA

Evaluates large language models on complex medical reasoning and knowledge retrieval across multiple-choice and open-ended clinical questions. It probes the model's ability to apply domain-specific knowledge, perform multi-step clinical reasoning, and handle challenging benchmarks that require more than simple fact recall. Use when the user wants to benchmark on MedQA (USMLE), MedMCQA, PubMedQA, MMLU-Pro, GPQA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3