Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 6,217–6,240 of 20,797 skills
Evaluates a robot's ability to navigate to a goal in unknown environments using only local sensor data without a pre-built map. It probes the policy's generalization across varying obstacle densities, passage widths, and real-world conditions, as well as its energy efficiency on neuromorphic hardware. Use when the user wants to benchmark on Gazebo Training Environments, Gazebo Test Environment, Real-world Office Environment, or asks about evaluating this task. Reports success rate.
This evaluation probes the fidelity and predictive accuracy of local explanations generated by MAPLE. It measures how well a local linear model approximates the target model's predictions in the neighborhood of a test point, while also benchmarking overall regression accuracy against standard baselines. Use when the user wants to benchmark on UCI datasets, or asks about evaluating this task. Reports causal metric.
Probes autonomous driving models' ability to extract lane-level traffic regulations from visual inputs and map them to vectorized HD map centerlines. It evaluates both rule extraction from image sequences and bipartite graph construction for rule-lane correspondence reasoning. Use when the user wants to benchmark on MapDR, or asks about evaluating this task. Reports correspondence status.
This evaluation probes the operational utility of data-driven Fire Danger Index (FDI) models for wildfire forecasting. It assesses both point-level classification accuracy and full-map spatial inference performance, explicitly quantifying detection rates and false positive distributions under realistic deployment conditions. Use when the user wants to benchmark on FireCube, or asks about evaluating this task. Reports Map-based Recall Percentiles.
Evaluates the semantic alignment, geometric fidelity, and visual appearance of automatically generated 3D assets against input images or text. It also assesses the accuracy of VLM-generated annotations across five dimensions and the physical validity of simulated grasp poses for robotic manipulation readiness. Use when the user wants to benchmark on ManiTwin-100K, or asks about evaluating this task. Reports CLIP(I-I/T).
Evaluates a robot policy's ability to perform long-horizon manipulation tasks involving deformable soft bodies (e.g., clay, noodles, liquid, plasticine). It probes spatial reasoning, contact dynamics, and precise end-effector control under varying initial conditions. Use when the user wants to benchmark on ManiSkill2 Challenge (Soft-body Track), or asks about evaluating this task. Reports Success Metric.
Evaluates the generalization and robustness of embodied AI manipulation policies across soft-body, rigid-body, and assembly tasks in a simulated environment. Use when the user wants to benchmark on ManiSkill2, or asks about evaluating this task. Reports success rate.
Evaluates low-level robotic manipulation policies for long-horizon home rearrangement tasks. It probes a robot's ability to successfully pick, place, and interact with household objects across cluttered and constrained environments. Use when the user wants to benchmark on ManiSkill-HAB, or asks about evaluating this task. Reports success once rate.
Evaluates real-world robot manipulation capabilities across two complementary tracks: physical skills (sensorimotor execution under contact, clearance, and perceptual constraints) and embodied reasoning (multimodal grounding of natural language and visual instructions into grounded actions). Use when the user wants to benchmark on ManipulationNet Benchmark, or asks about evaluating this task. Reports task success rate.
Evaluates whether learned hierarchical motor skills can transfer across different object geometries, downstream stacking tasks, and observation modalities (state vs. vision). It probes sample efficiency, directed exploration, and performance under varying reward sparsities (dense, staged sparse, fully sparse). Use when the user wants to benchmark on red_on_blue_stacking, all_pairs_stacking, or asks about evaluating this task. Reports reward.
Evaluates physically-grounded image editing by measuring 2D spatial accuracy, depth prediction, 3D geometric consistency, image quality, and VLM-based physical plausibility for object manipulation tasks. Use when the user wants to benchmark on ManipEval, or asks about evaluating this task. Reports Chamfer.
Evaluates the robustness of a dimensionality reduction pipeline (Isomap + Procrustes alignment + TDA clustering) against ambient noise, outliers, and hyperparameter variation. It tests whether the method can consistently recover a low-distortion 2D embedding of a contractible manifold or correctly detect topological failure on non-contractible data. Use when the user wants to benchmark on Swiss roll, Buckyball, or asks about evaluating this task. Reports Persistent homology features ($PH_1$, ...
Evaluates mathematical information retrieval capabilities by testing whether models can match natural language names or LaTeX formulas to their corresponding mathematical identities from a candidate pool. It probes the model's ability to learn structural and notational variations in mathematical expressions through pretraining and fine-tuning. Use when the user wants to benchmark on MAMUT-generated datasets (MF, MT, NMF, MFR), or asks about evaluating this task. Reports nDCG.
Evaluates multimodal reasoning and instruction-following capabilities across single-image, multi-image, and video scenarios. Probes OCR, chart/document understanding, mathematical reasoning, and real-world visual interactions. Use when the user wants to benchmark on AI2D, ChartQA, DocVQA, InfoVQA, MMStar, MMMU, MMMU-Pro, SeedBench, MMBench, MMvet, Mathverse, Mathvista, RealworldQA, WildVision, Llava-Wilder-Small, MuirBench, MEGABench, EgoSchema, PerceptionTest, SeedBench (Video), MLVU, MVBenc...
Evaluates a hybrid CNN-SSM architecture's ability to classify mammography regions of interest (ROIs) as benign or malignant. It probes the model's capacity for local feature extraction, global context modeling, and robust performance under class imbalance in medical imaging. Use when the user wants to benchmark on CBIS-DDSM, or asks about evaluating this task. Reports AUC-ROC.
Evaluates the ability of deep learning models to detect breast masses in digital mammography images across multiple clinical domains with varying scanner manufacturers and imaging protocols. It specifically probes domain generalization capabilities by measuring detection robustness on unseen data distributions. Use when the user wants to benchmark on OPTIMAM Hologic, OPTIMAM Siemens, OPTIMAM GE, OPTIMAM Philips, INbreast, BCDR, or asks about evaluating this task. Reports TPR at 0.75 FPPI.
Evaluates self-supervised learning models for breast cancer detection on screening mammography using a linear evaluation protocol on whole images derived from tiled patches. The protocol extracts fixed encoder features from image patches, pools them using attention-based or average pooling, and trains a linear classifier for final prediction. Use when the user wants to benchmark on Screening mammography dataset, or asks about evaluating this task. Reports linear evaluation.
Evaluates the cross-domain generalisation capability of deep learning models for breast cancer screening using multi-view mammography images. It tests robustness to out-of-distribution data from different vendors, imaging protocols, and centers by training on seen domains and testing on unseen domains. Use when the user wants to benchmark on CBIS, CMMD, INBreast, TOMMY1, TOMMY2, or asks about evaluating this task. Reports AUC.
Evaluates a multi-modal AI system's ability to detect breast cancer and localize malignant lesions using 2D (FFDM, C-View) and 3D (DBT) mammography images. It measures classification performance at the breast and image levels, as well as lesion localization accuracy via bounding boxes across internal and external clinical datasets. Use when the user wants to benchmark on NYU Comprehensive Mammography Dataset (V1), NYU Comprehensive Mammography Dataset (V2), OPTIMAM, CMMD, CSAW-CC, EMBED, CBIS...
Probes a deep learning model's ability to classify mammograms into BI-RADS categories (normal, benign, malignant) and localize suspicious lesions using weakly and semi-supervised learning. It evaluates both image-level diagnostic accuracy and region-level detection performance under clinically relevant operating points. Use when the user wants to benchmark on IMG, INbreast, or asks about evaluating this task. Reports AUROC.
Evaluates deep CNN architectures for classifying mammographic abnormalities (calcifications and masses) and localizing them using class activation maps. It probes the model's ability to learn patch-based features and generalize to full-image localization without explicit spatial supervision. Use when the user wants to benchmark on Mammography dataset (unspecified), or asks about evaluating this task. Reports accuracy.
Evaluates lightweight CNN architectures for pixel-wise lesion segmentation in mammograms. It measures segmentation accuracy and computational efficiency, while also probing cross-dataset generalization under domain shift and the sensitivity of performance metrics to post-processing thresholds. Use when the user wants to benchmark on INbreast, DMID, or asks about evaluating this task. Reports Dice Score.
Evaluates a federated learning framework for quantitative breast density estimation from mammographic images. It probes the model's ability to segment breast and dense tissue, predict percent density, and generalize across different medical institutions while preserving patient privacy. Use when the user wants to benchmark on MC, UPHS, or asks about evaluating this task. Reports PD MAE.
Evaluates a breast-specific foundational model's ability to generalize across in-distribution and out-of-distribution mammographic datasets for zero-shot diagnosis, linear probing, full fine-tuning, and pathology localization. It probes the model's robustness, data efficiency, and representation quality for clinical tasks like cancer detection and risk prediction. Use when the user wants to benchmark on EMBED, VinDr, RSNA, or asks about evaluating this task. Reports AUROC.