
Claude Skills by qhjqhj00
github.com/qhjqhj00Evaluates the ability of machine learning models to classify brain MRI images into four categories: glioma, meningioma, pituitary tumor, and no tumor. It tests feature extraction and decision fusion capabilities using both deep learning and traditional ML classifiers. Use when the user wants to benchmark on Kaggle Brain Tumor MRI dataset, Figshare Brain Tumor dataset, or asks about evaluating this task. Reports Accuracy.
Evaluates the ability of various CNN architectures (custom, U-Net, Fast R-CNN, and transfer learning models) to accurately classify brain tumors (glioma, meningioma, pituitary) from MRI images. It probes architectural robustness, generalization across data splits, and performance under class imbalance conditions. Use when the user wants to benchmark on Kaggle Brain Tumor Dataset, or asks about evaluating this task. Reports accuracy.
Evaluates a deep learning model's ability to classify brain MRI images into four tumor categories (glioma, meningioma, no tumor, pituitary). It probes multi-class image classification performance, generalization to unseen medical scans, and the model's capacity to balance precision and recall across classes. Use when the user wants to benchmark on Public MRI dataset (unspecified), or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to classify brain MRI scans into four pathological categories (Glioma, Meningioma, Pituitary Tumor, or None) using a hybrid CNN-ViT architecture with adaptive attention gating. Use when the user wants to benchmark on Brain Tumor MRI Dataset, or asks about evaluating this task. Reports accuracy.
Evaluates the accuracy of brain tumor segmentation algorithms on MRI images by comparing predicted tumor masks against radiologist-annotated ground truth. It probes the ability of thresholding and region-growing methods to correctly identify tumor boundaries and distinguish tumor tissue from healthy brain tissue. Use when the user wants to benchmark on Self-made Brain Tumor MRI Dataset, or asks about evaluating this task. Reports F-score.
Evaluates large language models' problem-solving capabilities using narrative-form brainteasers, probing their ability to generate correct final answers and employ creative, insight-based reasoning strategies rather than relying on brute-force or trial-and-error methods. Use when the user wants to benchmark on Braingle Math, Braingle Logic, or asks about evaluating this task. Reports accuracy.
Evaluates the accuracy of a hybrid acoustic simulation pipeline for generating room impulse responses (IRs) against real-world measured data. It specifically probes the model's ability to capture low-frequency diffraction effects and high-frequency energy decay in complex room geometries. Use when the user wants to benchmark on BRAS benchmark, or asks about evaluating this task. Reports frequency response.
Evaluates interactive brain tumor segmentation models by training and testing on a single patient's MRI data to assess within-brain generalization. It measures voxel-wise classification accuracy across different tumor sub-regions using sparse manual labels. Use when the user wants to benchmark on MICCAI-BRATS 2013, or asks about evaluating this task. Reports Dice.
Evaluates 3D brain tumor segmentation accuracy across three sub-regions (whole tumor, core, enhancing) and tests radiomics-based survival prediction performance on multi-modal MRI scans. Use when the user wants to benchmark on BraTS 2017, or asks about evaluating this task. Reports Dice score.
Evaluates 3D medical image segmentation models on intracranial meningioma MRI scans. It probes volumetric accuracy and boundary sharpness across three tumor subregions (enhancing tumor, tumor core, whole tumor) under varying contrast and lesion size conditions. Use when the user wants to benchmark on BraTS 2023 Intracranial Meningioma Challenge, or asks about evaluating this task. Reports DSC.
Evaluates the ability of deep learning models to accurately segment diverse brain tumor sub-regions (enhancing tumor, tumor core, whole tumor) across adult gliomas, pediatric tumors, and sub-Saharan African populations using MRI scans. Use when the user wants to benchmark on BraTS 2023 PED, BraTS 2023 SSA, BraTS-GLI (GLA), or asks about evaluating this task. Reports DSC.
Evaluates 3D medical image segmentation models on post-treatment glioma MRI scans, testing their ability to delineate four clinically relevant tumor sub-regions (enhancing tissue, non-enhancing tumor core, surrounding FLAIR hyperintensity, and resection cavity) under treatment-induced anatomical variability and imaging artifacts. Use when the user wants to benchmark on BraTS 2024 Post-Treatment Glioma, or asks about evaluating this task. Reports Dice Similarity Coefficient (DSC).
Evaluates 3D deep learning models for brain tumor segmentation across three distinct tumor subtypes (pediatric, meningioma, metastasis) using MRI scans. It probes the model's ability to accurately delineate tumor boundaries and generalize across heterogeneous clinical datasets through adaptive post-processing and model ensembling. Use when the user wants to benchmark on BraTS 2024 (PED, MEN-RT, MET), or asks about evaluating this task. Reports lesion-wise Dice score.
Evaluates the accuracy of deep learning models in segmenting brain tumor subregions and boundaries on low-field MRI scans from Sub-Saharan Africa. It probes the model's ability to handle regional imaging protocol limitations and topological deformations in medical image segmentation. Use when the user wants to benchmark on BraTS-Africa, or asks about evaluating this task. Reports Dice Similarity Coefficient (DSC).
Evaluates machine learning models on three clinical neuro-oncology tasks using multi-modal MRI data: multi-compartment brain tumor segmentation, tumor progression assessment, and overall patient survival prediction. Use when the user wants to benchmark on BraTS Challenge, or asks about evaluating this task. Reports Dice score.
Evaluates multi-modal medical reasoning and visual question answering capabilities on brain tumor MRI scans. It probes the model's ability to parse clinical features and answer structured questions across three distinct tumor subtypes: metastases, glioblastoma, and meningioma. Use when the user wants to benchmark on BraTS (MET, GLI, MEN cohorts), or asks about evaluating this task. Reports accuracy.
Evaluates a deep learning model's ability to perform multi-region brain MRI segmentation (tumors and healthy structures) while simultaneously predicting per-voxel uncertainty. It probes the model's segmentation accuracy across anatomical regions and its calibration of confidence estimates against actual voxel-wise errors. Use when the user wants to benchmark on BraTS, OASIS-1, or asks about evaluating this task. Reports DSC.
Volumetric segmentation of pediatric brain gliomas using multi-institutional MRI data. It probes a model's ability to accurately delineate tumor sub-regions (enhancing tumor, peritumoral edema, necrotic/cystic core) in 3D MRI scans. Use when the user wants to benchmark on BraTS-PEDs 2023, or asks about evaluating this task. Reports Dice Score.
Evaluates the robustness and generalization capability of brain tumor segmentation models when faced with distribution shifts, specifically Gaussian noise perturbations in MRI scans. It probes whether high benchmark accuracy translates to reliable performance on clinically realistic, noisy data rather than just overfitting to clean benchmark distributions. Use when the user wants to benchmark on BraTS2018, or asks about evaluating this task. Reports Dice score.
Evaluates the precision of brain tumor segmentation models on multi-modal MRI scans across three clinically relevant regions (Whole Tumor, Tumor Core, Enhancing Tumor). It probes the model's ability to accurately delineate heterogeneous tumor boundaries and correctly identify positive tumor voxels in medical imaging data. Use when the user wants to benchmark on BraTS2019/2020, or asks about evaluating this task. Reports Dice coefficient.
Evaluates a model's ability to perform unsupervised domain adaptation for brain tumor segmentation, specifically transferring segmentation capabilities from T1-weighted MRI scans to T2-weighted MRI scans without target labels. Use when the user wants to benchmark on BraTS'19, or asks about evaluating this task. Reports DSC.
Evaluates the capability of deep learning models to segment brain tumors from multi-modal MRI scans. It probes volumetric overlap accuracy and boundary localization precision across distinct tumor sub-regions (enhancing tumor, tumor core, whole tumor). Use when the user wants to benchmark on BraTS-Glioma (BraTS 2020), TCGA LGG, or asks about evaluating this task. Reports Dice Similarity Coefficient (DSC).
Evaluates the capability of 2D and 3D convolutional neural networks to segment brain tumor sub-regions (enhancing tumor, whole tumor, tumor core) from multi-modal volumetric MRI scans. It specifically probes the effectiveness of ImageNet pretraining and architectural extensions on segmentation accuracy and robustness across benchmark and private clinical data. Use when the user wants to benchmark on BraTS 2020, Syrian-Lebanese Hospital Clinical Dataset, or asks about evaluating this task. Rep...
Evaluates a model's ability to automatically segment brain tumor subtypes (complete, core, enhancing) from multimodal MRI scans. It specifically probes performance on high-grade gliomas (HGG) versus low-grade gliomas (LGG), highlighting challenges with ambiguous boundaries and lack of contrast enhancement. Use when the user wants to benchmark on BRATS 2015, or asks about evaluating this task. Reports Dice Similarity Coefficient (DSC).
Probes 3D brain tumor segmentation capability across multiple MRI sequences, evaluating boundary accuracy and volumetric overlap for hierarchical tumor subregions. Use when the user wants to benchmark on BraTS 2017, or asks about evaluating this task. Reports Dice score.
Evaluates the capability of deep learning models to perform multi-modal brain tumour segmentation from 3D MRI volumes. It specifically probes robustness to architectural choices, loss functions, and intensity normalization preprocessing pipelines by aggregating diverse model configurations. Use when the user wants to benchmark on BRATS 2017, or asks about evaluating this task. Reports Dice score (DSC).
Evaluates the ability of 3D U-Net architectures to segment brain tumors (enhancing tumor, whole tumor, and tumor core) from multimodal MRI scans. It specifically probes how well lesion prior information (VOI maps) can be fused with imaging data to improve volumetric segmentation accuracy and boundary precision. Use when the user wants to benchmark on BraTS 2017, or asks about evaluating this task. Reports DSC (Dice Similarity Coefficient).
Evaluates 3D medical image segmentation models on brain tumor subregions (enhancing tumor, whole tumor, tumor core) using multimodal MRI scans. It probes the model's ability to accurately delineate complex, irregular tumor boundaries and differentiate tumors from surrounding vasculature and edema. Use when the user wants to benchmark on BraTS 2018, or asks about evaluating this task. Reports Dice.
This evaluation probes a model's ability to perform 3D medical image segmentation on brain tumors using MRI scans. It measures how well the architecture delineates tumor boundaries and classifies each voxel as tumor or background across multiple standard segmentation metrics. Use when the user wants to benchmark on BraTS 2020, or asks about evaluating this task. Reports Dice Coefficient.
Evaluates 3D semantic segmentation of brain tumor sub-regions on multi-modal MRI scans. It probes the model's ability to capture long-range spatial dependencies and multi-scale contextual information for precise tumor boundary delineation. Use when the user wants to benchmark on BraTS 2021, or asks about evaluating this task. Reports Dice score.
Evaluates 3D brain tumor segmentation accuracy and robustness under missing MRI modalities using multi-modal MRI scans. It probes a model's ability to delineate tumor sub-regions while maintaining calibration and stability when contrast sequences are corrupted or absent. Use when the user wants to benchmark on BraTS 2021, or asks about evaluating this task. Reports Dice Score.
Evaluates 3D brain tumor segmentation and inpainting on multi-modal MRI scans. It probes the model's ability to accurately delineate tumor sub-regions (enhancing, core, whole) and synthesize realistic healthy tissue to replace tumor areas. Use when the user wants to benchmark on BraTS 2023, or asks about evaluating this task. Reports Lesion-wise DSC.
Evaluates deep learning models for multi-parametric MRI brain tumor segmentation across three clinical tasks: pediatric tumors, meningiomas, and metastases. It probes the model's ability to accurately delineate tumor sub-regions (enhancing tumor, tumor core, whole tumor) using lesion-wise overlap and boundary distance metrics. Use when the user wants to benchmark on BraTS-PED, BraTS-MEN, BraTS-MET, or asks about evaluating this task. Reports Lesion-wise Dice.
Evaluates the zero-shot and fine-tuned performance of promptable and non-promptable 3D medical image segmentation models on brain tumor MRI data. It probes how prompt type (points vs. bounding boxes) and prompt accuracy affect segmentation quality compared to a strong unprompted baseline. Use when the user wants to benchmark on BraTS 2023 Adult Glioma, BraTS 2023 Pediatrics, or asks about evaluating this task. Reports Dice score (DSC).
Evaluates zero-shot medical knowledge and clinical reasoning of LLMs and MLLMs on a Brazilian Portuguese medical residency exam. Probes text-only comprehension versus multimodal image interpretation across five clinical domains. Use when the user wants to benchmark on HCFMUSP Brazilian Portuguese Medical Residency Exam, or asks about evaluating this task. Reports accuracy.
Measures the sensitivity of deep Q-learning performance to various sources of nondeterminism (GPU operations, environment stochasticity, exploration seeds, weight initialization, minibatch sampling) by comparing performance variance across controlled experimental groups. Use when the user wants to benchmark on Atari BREAKOUT, or asks about evaluating this task. Reports mean score.
Evaluates a vision model's ability to classify mammograms as benign or malignant, focusing on fine-grained discrimination of localized malignancies and handling high-resolution medical images. Use when the user wants to benchmark on Public mammogram dataset(s), or asks about evaluating this task. Reports F1 score.
Evaluates deep learning models for breast cancer lesion detection and classification in mammogram images, comparing segmentation and classification performance against conventional and heuristic-optimized baselines. Use when the user wants to benchmark on Dataset 1 & 2, or asks about evaluating this task. Reports accuracy.
Evaluates deep learning models for binary classification of breast cancer (malignant vs. benign/normal) on screening mammograms. It probes the model's ability to generalize across different mammography platforms (film vs. digital) and transfer learned features from patch-level to whole-image classification without requiring costly lesion-level annotations. Use when the user wants to benchmark on CBIS-DDSM, INbreast, or asks about evaluating this task. Reports AUC.
This evaluation probes the classification accuracy and prototype-based interpretability of deep learning models on mammography datasets. It measures how well models predict malignancy while ensuring that their learned prototypes align with domain-specific radiological features (e.g., mass/calcification types and BIRADS descriptors). Use when the user wants to benchmark on CBIS-DDSM, CMMD, VinDr-Mammo, or asks about evaluating this task. Reports F1.
Evaluates multi-view deep learning models for medical image classification and segmentation, specifically testing the ability to model correlations between paired or multi-view medical images (mammograms, chest X-rays, brain MRIs) for tasks like lesion classification and survival prediction. Use when the user wants to benchmark on CBIS-DDSM, INbreast, CheXpert, BraTS19, or asks about evaluating this task. Reports AUC.
Predicts county-level breast cancer screening rates using census-tract-level socioeconomic, demographic, and geospatial features. Evaluates the regression performance of Random Forest, Linear Regression, and Support Vector Machine models to identify which algorithm best captures underlying patterns in screening access disparities. Use when the user wants to benchmark on US Census Tracts Mammography Screening Dataset, or asks about evaluating this task. Reports R^2.
Evaluates a model's ability to classify breast lesions as benign or malignant using paired mammography and ultrasound images. It probes multimodal fusion capabilities by comparing single-modality performance against a simple average of combined modality predictions. Use when the user wants to benchmark on Breast Lesion Dataset (153 pairs), or asks about evaluating this task. Reports AUC.
This evaluation probes a model's ability to predict breast cancer severity (benign vs. malignant) using clinical and radiological features. It measures classification performance across multiple metrics to assess diagnostic reliability and clinical utility. Use when the user wants to benchmark on Mammographic mass dataset, or asks about evaluating this task. Reports Accuracy.
Evaluates the ability of 3D medical image segmentation models to accurately delineate left and right breast tissues in MRI scans. It probes anatomical partitioning robustness and generalization across diverse clinical sources. Use when the user wants to benchmark on Breast MRI Left-Right Segmentation Dataset, or asks about evaluating this task. Reports Dice Similarity Coefficient (DSC).
Evaluates a multimodal framework's ability to fuse patch-level histopathology features with structured EHR data for early breast cancer diagnosis, specifically probing performance on class-imbalanced minority classes like mitosis. Use when the user wants to benchmark on BreCaHAD, MIMIC-IV, or asks about evaluating this task. Reports Macro-average AUC.
Evaluates the phonetic accuracy, audio quality, and speaker similarity of a Taiwanese Mandarin TTS system, with a focus on voice cloning robustness and code-switching scenarios. The benchmark probes the model's ability to handle long-tail speaker variability and context-dependent pronunciation ambiguities in both monolingual and bilingual contexts. Use when the user wants to benchmark on FormosaSpeech (subset), Spontaneous Recordings, Traditional Chinese Monologue Dataset (TCMD), Traditional ...
Compute brian920128/doc_retrieve_metrics via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of brian920128/doc_retrieve_metrics.
Compute BridgeAI-Lab/Sem-nCG via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of BridgeAI-Lab/Sem-nCG.
Compute BridgeAI-Lab/SemF1 via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of BridgeAI-Lab/SemF1.