Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 7,825–7,848 of 20,853 skills
Evaluates the capability of models to analyze electrocardiogram (ECG) time-series data across four medical tasks: classification, detection, forecasting, and generation. It probes semantic fidelity and diagnostic accuracy in quasi-periodic physiological signals, emphasizing robustness to temporal shifts and class imbalance. Use when the user wants to benchmark on CPSC2018, CPSC2019, CPSC2020, CPSC2021, MITDB, PTBXL, FEPL, DALIA, SST, or asks about evaluating this task. Reports FFD.
Evaluates a CNN's ability to reconstruct missing QRS complexes in ECG signals via self-supervised regression and to classify cardiac arrhythmias. It probes signal reconstruction fidelity and multi-class rhythm recognition under imbalanced conditions. Use when the user wants to benchmark on DS0 dataset (MIT-BIH Arrhythmia), or asks about evaluating this task. Reports NRMSE.
Evaluates a deep learning model's ability to classify cardiac arrhythmias from ECG signals, testing both intra-dataset performance and cross-dataset generalization using demographic attributes. Use when the user wants to benchmark on MITDB, INCARTDB, EDB, or asks about evaluating this task. Reports F1-score.
Evaluates a model's ability to classify cardiac arrhythmias from short ECG signal windows by leveraging transfer learning from pre-trained image CNNs. It probes the effectiveness of converting 1D physiological signals into 2D spectrograms and extracting high-level features for multi-class rhythm discrimination. Use when the user wants to benchmark on Combined MIT-BIH & European ST-T ECG Datasets, or asks about evaluating this task. Reports accuracy.
Machine translation performance on low-resource languages using verse-aligned Bible texts. It probes model robustness across different biblical book genres (Gospels, Epistles, OT books) and the utility of related language data for translation. Use when the user wants to benchmark on eBible Corpus, or asks about evaluating this task. Reports BLEU.
Evaluates the video understanding and reasoning capabilities of multimodal language models after reinforcement learning training. It probes performance across general video comprehension, long-context video understanding, complex reasoning, and STEM knowledge tasks using a standardized greedy decoding protocol. Use when the user wants to benchmark on Video-MME, MVBench, TempCompass, LVBench, LongVideoBench, MLVU, Video-Holmes, MMVU, Video-MMMU, VideoMathQA, or asks about evaluating this task....
Evaluates neural Temporal Point Process models on event sequence prediction tasks, specifically forecasting the timing and categorical type of future events given historical sequences. Use when the user wants to benchmark on Retweet, Taxi, or asks about evaluating this task. Reports TIME RMSE.
Evaluates the adversarial robustness and out-of-distribution (OOD) generalization of vision models on large-scale image classification benchmarks. It measures clean accuracy, robust accuracy against AutoAttack, and corruption error rates across multiple synthetic and real-world distribution shifts. Use when the user wants to benchmark on ImageNet, ImageNet-C, ImageNet-R, ImageNet-A, ImageNet-Sketch, Stylized-ImageNet, ObjectNet, ImageNet-V2, or asks about evaluating this task. Reports Top-1 a...
Evaluates semantic segmentation models on fine-grained face parsing and portrait segmentation. It probes a model's ability to accurately delineate nine distinct facial and occlusion classes in high-resolution indoor portrait images. Use when the user wants to benchmark on EasyPortrait, or asks about evaluating this task. Reports mIoU.
Evaluates real-time, dynamic audio-visual speech enhancement and beamforming systems in noisy, egocentric augmented reality settings. It probes the model's ability to isolate a target speaker's voice from competing talkers and background noise while preserving speech quality and intelligibility across diverse user movements. Use when the user wants to benchmark on EasyCom, or asks about evaluating this task. Reports SNR.
Evaluates high-spatial-resolution remote sensing models on land-cover semantic segmentation and visual question answering. It probes pixel-level object recognition, spatial reasoning, and relational counting capabilities in complex urban scenes. Use when the user wants to benchmark on EarthVLSet, or asks about evaluating this task. Reports mIoU, OA.
This benchmark evaluates the forecasting capability of neural spatio-temporal point processes (NPPs) on earthquake sequences. It probes how well models capture the joint temporal and spatial intensity of seismic events compared to traditional seismological baselines like ETAS. Use when the user wants to benchmark on EarthquakeNPP (ComCat, QTM_SaltonSea, QTM_SanJac, White, SCEDC), or asks about evaluating this task. Reports temporal log-likelihood.
Binary classification of low-magnitude seismic events versus background noise in seismological time-series data. It probes model robustness to varying noise-to-signal ratios and evaluates the trade-off between detection sensitivity and false positive rates in safety-critical monitoring. Use when the user wants to benchmark on Groningen gas field seismic data, or asks about evaluating this task. Reports MCC.
Evaluates speech enhancement models by measuring how well they recover clean speech from noisy mixtures across a wide range of signal-to-noise ratios and speaker demographics. The benchmark covers both controlled training/validation splits and a blind test set with unseen speakers and noise. Use when the user wants to benchmark on EARS-WHAM, or asks about evaluating this task. Reports SI-SDR.
Evaluates dereverberation models by measuring their ability to remove room acoustics effects from speech using real room impulse responses with RT60 up to 2 seconds. The protocol ensures fair comparison by normalizing loudness and removing direct-path delays before convolution. Use when the user wants to benchmark on EARS-Reverb, or asks about evaluating this task. Reports SI-SDR.
Evaluates machine learning models' ability to detect early-stage COVID-19 infection from chest X-ray images, specifically targeting cases with minimal or invisible radiological signs compared to healthy controls. Use when the user wants to benchmark on Early-QaTa-COV19, or asks about evaluating this task. Reports sensitivity.
Evaluates video action recognition models on classifying untrimmed real-world videos of elderly individuals into six daily activity categories. It probes robustness and generalization in wild, uncontrolled settings using a held-out test set. Use when the user wants to benchmark on EAR Challenge Test Set, or asks about evaluating this task. Reports accuracy.
Evaluates vision-language models on document understanding, chart and table reasoning, OCR, diagram comprehension, and general visual question answering. The protocol measures accuracy across a diverse suite of 14 established multimodal benchmarks to assess overall multimodal capability and robustness. Use when the user wants to benchmark on DocVQA, ChartQA, MMMU, MMB1.1, MathVista, or asks about evaluating this task. Reports OpenCompass.
Probes multimodal models' ability to perform complex, long-horizon spatial reasoning and physical consistency checks in dynamic, embodied scenarios. It evaluates object attribute recognition, relational understanding, and robotic manipulation planning across static images and real-world assembly tasks. Use when the user wants to benchmark on eSpatial-Benchmark, or asks about evaluating this task. Reports accuracy.
Probes 5-DoF viewpoint control and active perception in photorealistic 3D scenes. Tests whether vision-language models can navigate, resolve occlusions, and answer questions by strategically selecting viewpoints to gather spatially dependent visual evidence. Use when the user wants to benchmark on E3VS-Bench, or asks about evaluating this task. Reports VLM Judge Score.
E3VQA evaluates a model's ability to perform multi-view visual question answering using synchronized egocentric and exocentric image pairs. It specifically probes whether models can identify relevant regions across views, filter redundant information, and integrate complementary visual cues to answer multiple-choice questions. Use when the user wants to benchmark on E3VQA, or asks about evaluating this task. Reports accuracy.
Evaluates the effectiveness, robustness, and inference efficiency of end-to-end 3D Geometric Foundation Models across sparse-view depth estimation, video depth estimation, and multi-view relative pose estimation. It probes models' ability to generalize across diverse domains including indoor, outdoor, aerial, and highly dynamic scenes under both normalized and metric-scale settings. Use when the user wants to benchmark on DTU, ETH3D, KITTI, Tanks and Temples, ScanNet, Bonn, TUM Dynamics, Sint...
Evaluates a model's ability to perform end-to-end grounded multimodal named entity recognition, jointly identifying entity spans in text, predicting their semantic types, and grounding them to corresponding bounding boxes in an associated image. It probes the model's capacity for multimodal alignment, structured generation, and robustness to annotation noise via chain-of-thought reasoning. Use when the user wants to benchmark on Twitter-GMNER, Twitter-FMNERG, or asks about evaluating this tas...
Evaluates the effectiveness of AI-generated related search queries in an e-commerce setting by measuring their ability to drive user engagement and purchases compared to a production baseline. Use when the user wants to benchmark on eBay user interaction logs, or asks about evaluating this task. Reports click-through rate (CTR).