All authors
qhjqhj00 avatar

Claude Skills by qhjqhj00

github.com/qhjqhj00
7,636 skillsA× 7,623B× 11C× 1D× 10 installs2,233 views
E3d Bench EvalA

Evaluates the effectiveness, robustness, and inference efficiency of end-to-end 3D Geometric Foundation Models across sparse-view depth estimation, video depth estimation, and multi-view relative pose estimation. It probes models' ability to generalize across diverse domains including indoor, outdoor, aerial, and highly dynamic scenes under both normalized and metric-scale settings. Use when the user wants to benchmark on DTU, ETH3D, KITTI, Tanks and Temples, ScanNet, Bonn, TUM Dynamics, Sint...

researchpythonperformance
0
3
E3vqa EvalA

E3VQA evaluates a model's ability to perform multi-view visual question answering using synchronized egocentric and exocentric image pairs. It specifically probes whether models can identify relevant regions across views, filter redundant information, and integrate complementary visual cues to answer multiple-choice questions. Use when the user wants to benchmark on E3VQA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
E3vs Bench EvalA

Probes 5-DoF viewpoint control and active perception in photorealistic 3D scenes. Tests whether vision-language models can navigate, resolve occlusions, and answer questions by strategically selecting viewpoints to gather spatially dependent visual evidence. Use when the user wants to benchmark on E3VS-Bench, or asks about evaluating this task. Reports VLM Judge Score.

researchpythongo
0
3
ESpatial EvalA

Probes multimodal models' ability to perform complex, long-horizon spatial reasoning and physical consistency checks in dynamic, embodied scenarios. It evaluates object attribute recognition, relational understanding, and robotic manipulation planning across static images and real-world assembly tasks. Use when the user wants to benchmark on eSpatial-Benchmark, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Eagle2 Vlm EvalA

Evaluates vision-language models on document understanding, chart and table reasoning, OCR, diagram comprehension, and general visual question answering. The protocol measures accuracy across a diverse suite of 14 established multimodal benchmarks to assess overall multimodal capability and robustness. Use when the user wants to benchmark on DocVQA, ChartQA, MMMU, MMB1.1, MathVista, or asks about evaluating this task. Reports OpenCompass.

researchpythongo
0
3
Ear Challenge EvalA

Evaluates video action recognition models on classifying untrimmed real-world videos of elderly individuals into six daily activity categories. It probes robustness and generalization in wild, uncontrolled settings using a held-out test set. Use when the user wants to benchmark on EAR Challenge Test Set, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Early Qata Cov19 EvalA

Evaluates machine learning models' ability to detect early-stage COVID-19 infection from chest X-ray images, specifically targeting cases with minimal or invisible radiological signs compared to healthy controls. Use when the user wants to benchmark on Early-QaTa-COV19, or asks about evaluating this task. Reports sensitivity.

researchpythongit
0
3
Ears Reverb EvalA

Evaluates dereverberation models by measuring their ability to remove room acoustics effects from speech using real room impulse responses with RT60 up to 2 seconds. The protocol ensures fair comparison by normalizing loudness and removing direct-path delays before convolution. Use when the user wants to benchmark on EARS-Reverb, or asks about evaluating this task. Reports SI-SDR.

researchpython
0
3
Ears Wham EvalA

Evaluates speech enhancement models by measuring how well they recover clean speech from noisy mixtures across a wide range of signal-to-noise ratios and speaker demographics. The benchmark covers both controlled training/validation splits and a blind test set with unseen speakers and noise. Use when the user wants to benchmark on EARS-WHAM, or asks about evaluating this task. Reports SI-SDR.

researchpython
0
3
Earthquake Detection EvalA

Binary classification of low-magnitude seismic events versus background noise in seismological time-series data. It probes model robustness to varying noise-to-signal ratios and evaluates the trade-off between detection sensitivity and false positive rates in safety-critical monitoring. Use when the user wants to benchmark on Groningen gas field seismic data, or asks about evaluating this task. Reports MCC.

researchpythongo
0
3
Earthquakenpp EvalA

This benchmark evaluates the forecasting capability of neural spatio-temporal point processes (NPPs) on earthquake sequences. It probes how well models capture the joint temporal and spatial intensity of seismic events compared to traditional seismological baselines like ETAS. Use when the user wants to benchmark on EarthquakeNPP (ComCat, QTM_SaltonSea, QTM_SanJac, White, SCEDC), or asks about evaluating this task. Reports temporal log-likelihood.

researchpythonaws
0
3
Earthvlset EvalA

Evaluates high-spatial-resolution remote sensing models on land-cover semantic segmentation and visual question answering. It probes pixel-level object recognition, spatial reasoning, and relational counting capabilities in complex urban scenes. Use when the user wants to benchmark on EarthVLSet, or asks about evaluating this task. Reports mIoU, OA.

researchpythongo
0
3
Easycom EvalA

Evaluates real-time, dynamic audio-visual speech enhancement and beamforming systems in noisy, egocentric augmented reality settings. It probes the model's ability to isolate a target speaker's voice from competing talkers and background noise while preserving speech quality and intelligibility across diverse user movements. Use when the user wants to benchmark on EasyCom, or asks about evaluating this task. Reports SNR.

researchpythongo
0
3
Easyportrait EvalA

Evaluates semantic segmentation models on fine-grained face parsing and portrait segmentation. It probes a model's ability to accurately delineate nine distinct facial and occlusion classes in high-resolution indoor portrait images. Use when the user wants to benchmark on EasyPortrait, or asks about evaluating this task. Reports mIoU.

researchpythongit
0
3
Easyrobust EvalA

Evaluates the adversarial robustness and out-of-distribution (OOD) generalization of vision models on large-scale image classification benchmarks. It measures clean accuracy, robust accuracy against AutoAttack, and corruption error rates across multiple synthetic and real-world distribution shifts. Use when the user wants to benchmark on ImageNet, ImageNet-C, ImageNet-R, ImageNet-A, ImageNet-Sketch, Stylized-ImageNet, ObjectNet, ImageNet-V2, or asks about evaluating this task. Reports Top-1 a...

researchpythongo
0
3
Easytpp EvalA

Evaluates neural Temporal Point Process models on event sequence prediction tasks, specifically forecasting the timing and categorical type of future events given historical sequences. Use when the user wants to benchmark on Retweet, Taxi, or asks about evaluating this task. Reports TIME RMSE.

researchpythongo
0
3
Easyvideor1 EvalA

Evaluates the video understanding and reasoning capabilities of multimodal language models after reinforcement learning training. It probes performance across general video comprehension, long-context video understanding, complex reasoning, and STEM knowledge tasks using a standardized greedy decoding protocol. Use when the user wants to benchmark on Video-MME, MVBench, TempCompass, LVBench, LongVideoBench, MLVU, Video-Holmes, MMVU, Video-MMMU, VideoMathQA, or asks about evaluating this task....

researchpythongo
0
3
Ebible Benchmarks EvalA

Machine translation performance on low-resource languages using verse-aligned Bible texts. It probes model robustness across different biblical book genres (Gospels, Epistles, OT books) and the utility of related language data for translation. Use when the user wants to benchmark on eBible Corpus, or asks about evaluating this task. Reports BLEU.

researchpythongo
0
3
Ecg Arrhythmia Classification EvalA

Evaluates a model's ability to classify cardiac arrhythmias from short ECG signal windows by leveraging transfer learning from pre-trained image CNNs. It probes the effectiveness of converting 1D physiological signals into 2D spectrograms and extracting high-level features for multi-class rhythm discrimination. Use when the user wants to benchmark on Combined MIT-BIH & European ST-T ECG Datasets, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Ecg Arrhythmia Detection EvalA

Evaluates a deep learning model's ability to classify cardiac arrhythmias from ECG signals, testing both intra-dataset performance and cross-dataset generalization using demographic attributes. Use when the user wants to benchmark on MITDB, INCARTDB, EDB, or asks about evaluating this task. Reports F1-score.

researchpythongo
0
3
Ecg Arrhythmia EvalA

Evaluates a CNN's ability to reconstruct missing QRS complexes in ECG signals via self-supervised regression and to classify cardiac arrhythmias. It probes signal reconstruction fidelity and multi-class rhythm recognition under imbalanced conditions. Use when the user wants to benchmark on DS0 dataset (MIT-BIH Arrhythmia), or asks about evaluating this task. Reports NRMSE.

researchpython
0
3
Ecg Benchmark EvalA

Evaluates the capability of models to analyze electrocardiogram (ECG) time-series data across four medical tasks: classification, detection, forecasting, and generation. It probes semantic fidelity and diagnostic accuracy in quasi-periodic physiological signals, emphasizing robustness to temporal shifts and class imbalance. Use when the user wants to benchmark on CPSC2018, CPSC2019, CPSC2020, CPSC2021, MITDB, PTBXL, FEPL, DALIA, SST, or asks about evaluating this task. Reports FFD.

researchpythongo
0
3
Ecg Classification EvalA

Evaluates the ability of deep learning architectures to accurately classify electrocardiogram (ECG) recordings into predefined physiological or pathological categories. The benchmark probes joint time-frequency feature extraction capabilities by comparing models that embed Fourier analysis directly into convolutional layers against traditional signal processing and baseline CNN approaches. Use when the user wants to benchmark on MIT-BIH, ECG-ID, Apnea-ECG, or asks about evaluating this task. ...

datapythongo
0
3
Ecg Classification Macro F1 EvalA

Evaluates a model's ability to classify 12-lead electrocardiogram (ECG) signals into multiple clinical diagnoses. It probes the model's robustness to class imbalance and its capacity to learn from long, redundant time-series sequences using self-supervised pre-training and supervised fine-tuning. Use when the user wants to benchmark on Fujiak, PCinC, PTB-XL, or asks about evaluating this task. Reports macro F1 score.

researchpythontesting
0
3
Ecg Compression EvalA

Evaluates the efficiency and fidelity of ECG signal compression algorithms, focusing on how well they preserve critical clinical features like R-peaks for heart rate variability analysis. Use when the user wants to benchmark on MIT-BIH arrhythmia database, or asks about evaluating this task. Reports PRD.

researchpythongo
0
3
Ecg Cvd Classification EvalA

Evaluates machine learning models for detecting cardiovascular diseases and arrhythmias from ECG signals, comparing classification performance against computational complexity and energy efficiency. Use when the user wants to benchmark on CinC 2017, CinC 2020, or asks about evaluating this task. Reports F1 score.

researchpythongo
0
3
Ecg Delineation EvalA

This benchmark evaluates a model's ability to accurately segment and delineate the onset and offset boundaries of P, QRS, and T waves in electrocardiogram (ECG) signals. It specifically probes robustness across diverse cardiac arrhythmias and tests the effectiveness of classification-guided post-processing in reducing false positive detections during atrial fibrillation and flutter. Use when the user wants to benchmark on Internal dataset, LUDB, QTDB, or asks about evaluating this task. Repor...

researchpythongo
0
3
Ecg Expert Qa EvalA

Evaluates medical large language models on heart disease diagnosis using expert-validated QA pairs. It probes clinical reasoning, risk-aware decision-making, and patient-centric interaction capabilities across multiple diagnostic sub-tasks. Use when the user wants to benchmark on ECG-Expert-QA, or asks about evaluating this task. Reports BLEU-1.

researchpythongit
0
3
Ecg Fm Benchmark EvalA

Evaluates the clinical utility and label efficiency of ECG foundation models across diverse tasks including adult/pediatric ECG interpretation, cardiac structure prediction, clinical outcome forecasting, and patient characteristic regression. It probes cross-domain generalization, fine-tuning adaptability, and the quality of frozen/linear representations compared to strong supervised baselines. Use when the user wants to benchmark on PTB-XL, EchoNext, MIMIC-IV (ECG), CPSC2018, PTB, Ningbo, Ge...

researchpythonperformance
0
3
Ecg Grounding EvalA

Evaluates a multimodal LLM's ability to perform reliable, evidence-based ECG interpretation under full and missing modality conditions. It probes diagnostic accuracy, clinical reasoning fidelity, cross-modal consistency, and real-world clinical utility compared to cardiologist standards. Use when the user wants to benchmark on ECG-Grounding test set, or asks about evaluating this task. Reports Diagnosis Accuracy.

researchpythongo
0
3
Ecg Heart Disease Classification EvalA

Evaluates the ability of deep learning models to classify heart diseases from electrocardiogram (ECG) signals. The benchmark probes multi-level feature extraction by processing ECG data hierarchically (waves, heartbeats, segments) and measures classification performance alongside model complexity and interpretability. Use when the user wants to benchmark on MIT-BIH, PTB-XL, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Ecg Heartbeat Classification EvalA

Evaluates a deep convolutional neural network's ability to classify ECG heartbeats into arrhythmia categories and detect myocardial infarction using transferable learned representations. The protocol tests both in-domain arrhythmia classification and cross-domain transfer learning for MI detection. Use when the user wants to benchmark on MIT-BIH Arrhythmia Database, PTB Diagnostics, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Ecg Language Models EvalA

Evaluates the ability of encoder-free ECG-language models to process raw ECG signals alongside textual queries for medical question answering and instruction following. Probes whether models genuinely leverage physiological ECG data or rely on language priors and benchmark artifacts. Use when the user wants to benchmark on PTB-XL ECG-QA, PULSE ECG-Bench, ECG-Chat Instruct, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Ecg Linear Zero Shot EvalA

Evaluates the quality of self-supervised ECG image representations by measuring classification performance under linear probing and zero-shot settings across multiple clinical ECG datasets. Use when the user wants to benchmark on PTB-XL, CSN, CPSC2018, CODE-test, or asks about evaluating this task. Reports AUC (in %).

researchpythongo
0
3
Ecg Multi Label EvalA

This evaluation probes the ability of ECG foundation models to learn robust, generalizable representations from unsupervised pretraining and transfer them to downstream multi-label classification tasks. It specifically tests generalization across different clinical datasets and sampling rates by measuring performance on arrhythmia conditions and rhythm classifications. Use when the user wants to benchmark on PTB-XL, Chapman, or asks about evaluating this task. Reports macro AUC.

researchpythonperformance
0
3
Ecg Multitask EvalA

Evaluates the ability of foundation models (LLMs, time-series, and ECG-specific) and traditional deep learning models to perform regression and classification tasks on electrocardiogram (ECG) signals across zero-shot, few-shot, and fine-tuned settings. Use when the user wants to benchmark on ECG Multi-task Benchmark, or asks about evaluating this task. Reports MAE, F1 Score, Accuracy (ACC).

researchpythongo
0
3
Ecg Pll EvalA

This benchmark evaluates the robustness of Partial Label Learning (PLL) algorithms for multi-label ECG diagnosis under simulated clinical uncertainty. It probes how well models handle ambiguous candidate label sets generated through random, class-level, and instance-level ambiguity strategies. Use when the user wants to benchmark on PTB-XL, Chapman, or asks about evaluating this task. Reports micro-F1.

researchpythongo
0
3
Ecg Preprocessing EvalA

Evaluates how ECG signal pre-processing techniques, particularly down-sampling rates, affect the performance of multi-label time-series classification models for diagnosing heart conditions. It probes the trade-off between signal fidelity, computational cost, and diagnostic accuracy across varying sampling frequencies. Use when the user wants to benchmark on Unspecified multi-label ECG datasets, or asks about evaluating this task. Reports MRR.

researchpythongo
0
3
Ecg Reconstruction EvalA

Evaluates the ability of generative models to reconstruct standard 12-lead ECG signals from arbitrary single-lead ECG inputs. It probes signal fidelity, physiological feature preservation (heart rate statistics), and downstream diagnostic accuracy for arrhythmia classification. Use when the user wants to benchmark on PTB-XL, CPSC2018, or asks about evaluating this task. Reports MSE, PCC.

researchpythongo
0
3
Ecg Robustness EvalA

Evaluates the robustness of ECG classification models against six adversarial attack types (FGSM, BIM, PGD, CW, DBB, HSJ) compared to clean data. It measures classification performance and signal generation quality on two public ECG datasets. Use when the user wants to benchmark on PhysioNet MIT-BIH Arrhythmia, PTB Diagnostic ECG Database, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Ecg Ssl EvalA

Evaluates the transferability of self-supervised representations learned from single-lead ECG signals to downstream clinical and activity recognition tasks. It probes the model's ability to extract robust cardiac and physiological features by training linear probes on frozen encoder outputs across classification and regression benchmarks. Use when the user wants to benchmark on PhysioNet 2017, PTB-XL, Human Activity Recognition (HAR), or asks about evaluating this task. Reports Accuracy.

datapythonperformance
0
3
Ecgbench EvalA

Evaluates multimodal LLMs on interpreting electrocardiogram (ECG) images across classification, clinical report generation, and open-ended QA tasks. Probes robustness to real-world image artifacts, out-of-domain generalization, and clinical reasoning capabilities. Use when the user wants to benchmark on ECGBench, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Echo 4o EvalA

Evaluates text-to-image generation models on instruction-following accuracy, surreal/fantasy creativity, and multi-reference composition. It probes the model's ability to align complex textual prompts with visual outputs, handle long-tail attributes, and integrate multiple reference images. Use when the user wants to benchmark on GenEval, DPG-Bench, GenEval++, Imagine-Bench, OmniContext, or asks about evaluating this task. Reports GenEval Overall.

researchpythongo
0
3
Echo EvalA

This benchmark probes an image generation model's ability to follow complex, real-world user prompts and produce high-quality outputs that preserve specific attributes like identity and color. It evaluates how well models handle non-standard, community-driven inputs and context-dependent instructions often found in social media discussions. Use when the user wants to benchmark on ECHO, or asks about evaluating this task. Reports quality_label.

researchpythongo
0
3
Echochain EvalA

Evaluates how voice assistants handle mid-generation interruptions by testing their ability to revise in-progress responses while maintaining context and switching objectives. It probes state-update reasoning, contextual inertia, interruption amnesia, and objective displacement under full-duplex interaction conditions. Use when the user wants to benchmark on EchoChain, or asks about evaluating this task. Reports pass_fail.

researchpythongo
0
3
Echofake EvalA

Probes the robustness of speech deepfake detection models against physical replay attacks and cross-dataset generalization. It evaluates how well anti-spoofing systems distinguish between genuine speech, zero-shot TTS-generated deepfakes, and their physically replayed counterparts under realistic acoustic conditions. Use when the user wants to benchmark on EchoFake, or asks about evaluating this task. Reports EER.

researchpythongit
0
3
Echoreview Bench EvalA

Evaluates the quality, comprehensiveness, and evidence support of AI-generated academic peer reviews across multiple dimensions. It also measures the alignment between AI-identified research limitations and human reviewer findings, as well as the impact of citation time spans on review coherence and technical focus. Use when the user wants to benchmark on EchoReview-Bench, or asks about evaluating this task. Reports Overall Quality score.

researchpythongo
0
3
Echox Speech Qa EvalA

Evaluates the knowledge-based question-answering capabilities of speech-to-speech and speech-to-text models on audio and text inputs. Use when the user wants to benchmark on Llama Questions, Web Questions, TriviaQA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Eci EvalA

This evaluation protocol assesses the capability of NLP models to identify causal relationships between event pairs in text. It probes both sentence-level and document-level reasoning, measuring how well models can distinguish true causal links from mere correlations or coincidental co-occurrences. Use when the user wants to benchmark on CTB, ESL, MAVEN-ERE, MECI, or asks about evaluating this task. Reports F1-score.

researchpythongo
0
3
Ecoli Periodicity Detection EvalA

Evaluates the ability to detect periodic outlier patterns in protein sequence time-series data. It measures the statistical significance and reliability of discovered patterns compared to a baseline algorithm. Use when the user wants to benchmark on E.Coli, or asks about evaluating this task. Reports Surprise score.

researchpythongo
0
3