Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 7,801–7,824 of 20,853 skills
Evaluates the knowledge-based question-answering capabilities of speech-to-speech and speech-to-text models on audio and text inputs. Use when the user wants to benchmark on Llama Questions, Web Questions, TriviaQA, or asks about evaluating this task. Reports accuracy.
Evaluates the quality, comprehensiveness, and evidence support of AI-generated academic peer reviews across multiple dimensions. It also measures the alignment between AI-identified research limitations and human reviewer findings, as well as the impact of citation time spans on review coherence and technical focus. Use when the user wants to benchmark on EchoReview-Bench, or asks about evaluating this task. Reports Overall Quality score.
Probes the robustness of speech deepfake detection models against physical replay attacks and cross-dataset generalization. It evaluates how well anti-spoofing systems distinguish between genuine speech, zero-shot TTS-generated deepfakes, and their physically replayed counterparts under realistic acoustic conditions. Use when the user wants to benchmark on EchoFake, or asks about evaluating this task. Reports EER.
Evaluates how voice assistants handle mid-generation interruptions by testing their ability to revise in-progress responses while maintaining context and switching objectives. It probes state-update reasoning, contextual inertia, interruption amnesia, and objective displacement under full-duplex interaction conditions. Use when the user wants to benchmark on EchoChain, or asks about evaluating this task. Reports pass_fail.
This benchmark probes an image generation model's ability to follow complex, real-world user prompts and produce high-quality outputs that preserve specific attributes like identity and color. It evaluates how well models handle non-standard, community-driven inputs and context-dependent instructions often found in social media discussions. Use when the user wants to benchmark on ECHO, or asks about evaluating this task. Reports quality_label.
Evaluates text-to-image generation models on instruction-following accuracy, surreal/fantasy creativity, and multi-reference composition. It probes the model's ability to align complex textual prompts with visual outputs, handle long-tail attributes, and integrate multiple reference images. Use when the user wants to benchmark on GenEval, DPG-Bench, GenEval++, Imagine-Bench, OmniContext, or asks about evaluating this task. Reports GenEval Overall.
Evaluates multimodal LLMs on interpreting electrocardiogram (ECG) images across classification, clinical report generation, and open-ended QA tasks. Probes robustness to real-world image artifacts, out-of-domain generalization, and clinical reasoning capabilities. Use when the user wants to benchmark on ECGBench, or asks about evaluating this task. Reports Accuracy.
Evaluates the robustness of ECG classification models against six adversarial attack types (FGSM, BIM, PGD, CW, DBB, HSJ) compared to clean data. It measures classification performance and signal generation quality on two public ECG datasets. Use when the user wants to benchmark on PhysioNet MIT-BIH Arrhythmia, PTB Diagnostic ECG Database, or asks about evaluating this task. Reports Accuracy.
Evaluates the ability of generative models to reconstruct standard 12-lead ECG signals from arbitrary single-lead ECG inputs. It probes signal fidelity, physiological feature preservation (heart rate statistics), and downstream diagnostic accuracy for arrhythmia classification. Use when the user wants to benchmark on PTB-XL, CPSC2018, or asks about evaluating this task. Reports MSE, PCC.
Evaluates how ECG signal pre-processing techniques, particularly down-sampling rates, affect the performance of multi-label time-series classification models for diagnosing heart conditions. It probes the trade-off between signal fidelity, computational cost, and diagnostic accuracy across varying sampling frequencies. Use when the user wants to benchmark on Unspecified multi-label ECG datasets, or asks about evaluating this task. Reports MRR.
This benchmark evaluates the robustness of Partial Label Learning (PLL) algorithms for multi-label ECG diagnosis under simulated clinical uncertainty. It probes how well models handle ambiguous candidate label sets generated through random, class-level, and instance-level ambiguity strategies. Use when the user wants to benchmark on PTB-XL, Chapman, or asks about evaluating this task. Reports micro-F1.
Evaluates the ability of foundation models (LLMs, time-series, and ECG-specific) and traditional deep learning models to perform regression and classification tasks on electrocardiogram (ECG) signals across zero-shot, few-shot, and fine-tuned settings. Use when the user wants to benchmark on ECG Multi-task Benchmark, or asks about evaluating this task. Reports MAE, F1 Score, Accuracy (ACC).
This evaluation probes the ability of ECG foundation models to learn robust, generalizable representations from unsupervised pretraining and transfer them to downstream multi-label classification tasks. It specifically tests generalization across different clinical datasets and sampling rates by measuring performance on arrhythmia conditions and rhythm classifications. Use when the user wants to benchmark on PTB-XL, Chapman, or asks about evaluating this task. Reports macro AUC.
Evaluates the quality of self-supervised ECG image representations by measuring classification performance under linear probing and zero-shot settings across multiple clinical ECG datasets. Use when the user wants to benchmark on PTB-XL, CSN, CPSC2018, CODE-test, or asks about evaluating this task. Reports AUC (in %).
Evaluates the ability of encoder-free ECG-language models to process raw ECG signals alongside textual queries for medical question answering and instruction following. Probes whether models genuinely leverage physiological ECG data or rely on language priors and benchmark artifacts. Use when the user wants to benchmark on PTB-XL ECG-QA, PULSE ECG-Bench, ECG-Chat Instruct, or asks about evaluating this task. Reports Accuracy.
Evaluates a deep convolutional neural network's ability to classify ECG heartbeats into arrhythmia categories and detect myocardial infarction using transferable learned representations. The protocol tests both in-domain arrhythmia classification and cross-domain transfer learning for MI detection. Use when the user wants to benchmark on MIT-BIH Arrhythmia Database, PTB Diagnostics, or asks about evaluating this task. Reports accuracy.
Evaluates the ability of deep learning models to classify heart diseases from electrocardiogram (ECG) signals. The benchmark probes multi-level feature extraction by processing ECG data hierarchically (waves, heartbeats, segments) and measures classification performance alongside model complexity and interpretability. Use when the user wants to benchmark on MIT-BIH, PTB-XL, or asks about evaluating this task. Reports accuracy.
Evaluates a multimodal LLM's ability to perform reliable, evidence-based ECG interpretation under full and missing modality conditions. It probes diagnostic accuracy, clinical reasoning fidelity, cross-modal consistency, and real-world clinical utility compared to cardiologist standards. Use when the user wants to benchmark on ECG-Grounding test set, or asks about evaluating this task. Reports Diagnosis Accuracy.
Evaluates the clinical utility and label efficiency of ECG foundation models across diverse tasks including adult/pediatric ECG interpretation, cardiac structure prediction, clinical outcome forecasting, and patient characteristic regression. It probes cross-domain generalization, fine-tuning adaptability, and the quality of frozen/linear representations compared to strong supervised baselines. Use when the user wants to benchmark on PTB-XL, EchoNext, MIMIC-IV (ECG), CPSC2018, PTB, Ningbo, Ge...
Evaluates medical large language models on heart disease diagnosis using expert-validated QA pairs. It probes clinical reasoning, risk-aware decision-making, and patient-centric interaction capabilities across multiple diagnostic sub-tasks. Use when the user wants to benchmark on ECG-Expert-QA, or asks about evaluating this task. Reports BLEU-1.
This benchmark evaluates a model's ability to accurately segment and delineate the onset and offset boundaries of P, QRS, and T waves in electrocardiogram (ECG) signals. It specifically probes robustness across diverse cardiac arrhythmias and tests the effectiveness of classification-guided post-processing in reducing false positive detections during atrial fibrillation and flutter. Use when the user wants to benchmark on Internal dataset, LUDB, QTDB, or asks about evaluating this task. Repor...
Evaluates machine learning models for detecting cardiovascular diseases and arrhythmias from ECG signals, comparing classification performance against computational complexity and energy efficiency. Use when the user wants to benchmark on CinC 2017, CinC 2020, or asks about evaluating this task. Reports F1 score.
Evaluates the efficiency and fidelity of ECG signal compression algorithms, focusing on how well they preserve critical clinical features like R-peaks for heart rate variability analysis. Use when the user wants to benchmark on MIT-BIH arrhythmia database, or asks about evaluating this task. Reports PRD.
Evaluates a model's ability to classify 12-lead electrocardiogram (ECG) signals into multiple clinical diagnoses. It probes the model's robustness to class imbalance and its capacity to learn from long, redundant time-series sequences using self-supervised pre-training and supervised fine-tuning. Use when the user wants to benchmark on Fujiak, PCinC, PTB-XL, or asks about evaluating this task. Reports macro F1 score.