Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 9,265–9,288 of 21,341 skills
Volumetric segmentation of pediatric brain gliomas using multi-institutional MRI data. It probes a model's ability to accurately delineate tumor sub-regions (enhancing tumor, peritumoral edema, necrotic/cystic core) in 3D MRI scans. Use when the user wants to benchmark on BraTS-PEDs 2023, or asks about evaluating this task. Reports Dice Score.
Evaluates a deep learning model's ability to perform multi-region brain MRI segmentation (tumors and healthy structures) while simultaneously predicting per-voxel uncertainty. It probes the model's segmentation accuracy across anatomical regions and its calibration of confidence estimates against actual voxel-wise errors. Use when the user wants to benchmark on BraTS, OASIS-1, or asks about evaluating this task. Reports DSC.
Evaluates multi-modal medical reasoning and visual question answering capabilities on brain tumor MRI scans. It probes the model's ability to parse clinical features and answer structured questions across three distinct tumor subtypes: metastases, glioblastoma, and meningioma. Use when the user wants to benchmark on BraTS (MET, GLI, MEN cohorts), or asks about evaluating this task. Reports accuracy.
Evaluates machine learning models on three clinical neuro-oncology tasks using multi-modal MRI data: multi-compartment brain tumor segmentation, tumor progression assessment, and overall patient survival prediction. Use when the user wants to benchmark on BraTS Challenge, or asks about evaluating this task. Reports Dice score.
Evaluates the accuracy of deep learning models in segmenting brain tumor subregions and boundaries on low-field MRI scans from Sub-Saharan Africa. It probes the model's ability to handle regional imaging protocol limitations and topological deformations in medical image segmentation. Use when the user wants to benchmark on BraTS-Africa, or asks about evaluating this task. Reports Dice Similarity Coefficient (DSC).
Evaluates 3D deep learning models for brain tumor segmentation across three distinct tumor subtypes (pediatric, meningioma, metastasis) using MRI scans. It probes the model's ability to accurately delineate tumor boundaries and generalize across heterogeneous clinical datasets through adaptive post-processing and model ensembling. Use when the user wants to benchmark on BraTS 2024 (PED, MEN-RT, MET), or asks about evaluating this task. Reports lesion-wise Dice score.
Evaluates 3D medical image segmentation models on post-treatment glioma MRI scans, testing their ability to delineate four clinically relevant tumor sub-regions (enhancing tissue, non-enhancing tumor core, surrounding FLAIR hyperintensity, and resection cavity) under treatment-induced anatomical variability and imaging artifacts. Use when the user wants to benchmark on BraTS 2024 Post-Treatment Glioma, or asks about evaluating this task. Reports Dice Similarity Coefficient (DSC).
Evaluates the ability of deep learning models to accurately segment diverse brain tumor sub-regions (enhancing tumor, tumor core, whole tumor) across adult gliomas, pediatric tumors, and sub-Saharan African populations using MRI scans. Use when the user wants to benchmark on BraTS 2023 PED, BraTS 2023 SSA, BraTS-GLI (GLA), or asks about evaluating this task. Reports DSC.
Evaluates 3D brain tumor segmentation accuracy across three sub-regions (whole tumor, core, enhancing) and tests radiomics-based survival prediction performance on multi-modal MRI scans. Use when the user wants to benchmark on BraTS 2017, or asks about evaluating this task. Reports Dice score.
Evaluates interactive brain tumor segmentation models by training and testing on a single patient's MRI data to assess within-brain generalization. It measures voxel-wise classification accuracy across different tumor sub-regions using sparse manual labels. Use when the user wants to benchmark on MICCAI-BRATS 2013, or asks about evaluating this task. Reports Dice.
Evaluates the accuracy of a hybrid acoustic simulation pipeline for generating room impulse responses (IRs) against real-world measured data. It specifically probes the model's ability to capture low-frequency diffraction effects and high-frequency energy decay in complex room geometries. Use when the user wants to benchmark on BRAS benchmark, or asks about evaluating this task. Reports frequency response.
Evaluates large language models' problem-solving capabilities using narrative-form brainteasers, probing their ability to generate correct final answers and employ creative, insight-based reasoning strategies rather than relying on brute-force or trial-and-error methods. Use when the user wants to benchmark on Braingle Math, Braingle Logic, or asks about evaluating this task. Reports accuracy.
Evaluates the accuracy of brain tumor segmentation algorithms on MRI images by comparing predicted tumor masks against radiologist-annotated ground truth. It probes the ability of thresholding and region-growing methods to correctly identify tumor boundaries and distinguish tumor tissue from healthy brain tissue. Use when the user wants to benchmark on Self-made Brain Tumor MRI Dataset, or asks about evaluating this task. Reports F-score.
Evaluates a model's ability to classify brain MRI scans into four pathological categories (Glioma, Meningioma, Pituitary Tumor, or None) using a hybrid CNN-ViT architecture with adaptive attention gating. Use when the user wants to benchmark on Brain Tumor MRI Dataset, or asks about evaluating this task. Reports accuracy.
Evaluates a deep learning model's ability to classify brain MRI images into four tumor categories (glioma, meningioma, no tumor, pituitary). It probes multi-class image classification performance, generalization to unseen medical scans, and the model's capacity to balance precision and recall across classes. Use when the user wants to benchmark on Public MRI dataset (unspecified), or asks about evaluating this task. Reports accuracy.
Evaluates the ability of various CNN architectures (custom, U-Net, Fast R-CNN, and transfer learning models) to accurately classify brain tumors (glioma, meningioma, pituitary) from MRI images. It probes architectural robustness, generalization across data splits, and performance under class imbalance conditions. Use when the user wants to benchmark on Kaggle Brain Tumor Dataset, or asks about evaluating this task. Reports accuracy.
Evaluates the stability and reliability of global pointwise scores (accuracy, AUC, F1) versus pairwise Bradley-Terry rankings for ordering NLP models across classification and text generation tasks. Use when the user has predictions and gold and needs to compute Bradley-Terry.
Evaluates the ability of audio-language models to align audio with captions and distinguish caption quality across different generation sources (human-human, human-machine, machine-machine). It probes fine-grained semantic and syntactic alignment capabilities under realistic captioning conditions. Use when the user wants to benchmark on BRACE-Main, or asks about evaluating this task. Reports F1-score.
Probes models' robustness in detecting subtle hallucinations in audio captions, specifically those introduced via LLM-driven noun substitution. It measures the ability to identify semantically flawed or factually incorrect descriptions against audio ground truth. Use when the user wants to benchmark on BRACE-Hallucination, or asks about evaluating this task. Reports F1-score.
Evaluates the effectiveness of synthetic 3D MRI tumor ROI generation for data augmentation by measuring downstream binary classification performance on imbalanced brain tumor subtypes. Use when the user wants to benchmark on BraTS 2019, SickKids pLGG, or asks about evaluating this task. Reports AUC.
Evaluates vision-language models' ability to extract structured information (names, types, and connectivity) from Business Process Model and Notation (BPMN) diagrams provided as images. It tests both raw visual understanding and the utility of OCR-enriched inputs for schema-constrained diagram parsing. Use when the user wants to benchmark on BPMN Diagrams (Custom), or asks about evaluating this task. Reports F1 Score.
Evaluates machine translation systems on a contamination-free, multilingual dataset covering diverse domains and registers. It measures translation quality at both sentence and paragraph levels to assess how well models handle linguistic diversity and cultural authenticity across 8 major languages. Use when the user wants to benchmark on BOUQuET, or asks about evaluating this task. Reports CometKiwi.
This evaluation probes a model's ability to forecast and reconstruct the dynamics of partially observed, chaotic geophysical systems. It specifically tests short-term prediction accuracy and long-term topological stability (boundedness) under both in-distribution and out-of-attractor initial conditions. Use when the user wants to benchmark on Lorenz-63, Lorenz-96, or asks about evaluating this task. Reports RMSE.
Evaluates a hierarchical multi-agent reinforcement learning scheduler's ability to optimize task allocation, frequency scaling, and core selection for OpenMP DAG workloads on embedded systems. It probes the trade-off between makespan, energy consumption, and thermal constraints under real-time profiling feedback. Use when the user wants to benchmark on Barcelona OpenMP Tasks Suite (BOTS), or asks about evaluating this task. Reports makespan.