All authors
qhjqhj00 avatar

Claude Skills by qhjqhj00

github.com/qhjqhj00
7,636 skillsA× 7,623B× 11C× 1D× 10 installs2,217 views
Micro F1A

Evaluates the ability of transformer models to extract Task-Dataset-Metric (TDM) triples from scholarly AI publications. It measures how accurately models can identify leaderboard components and distinguish them from papers that do not report empirical research. Use when the user has predictions and gold and needs to compute micro-F1.

researchpythongo
0
3
Microbiorel Re EvalA

Evaluates generative and discriminative models on document-level relation extraction in the microbiome domain. It probes the ability to correctly classify pairwise relations between biomedical entities (species, diseases, chemicals, etc.) under a low-resource setting. Use when the user wants to benchmark on MicrobioRel, or asks about evaluating this task. Reports Weighted F1-score.

researchpythongo
0
3
Microg 4m EvalA

Evaluates video-based human action recognition, temporal captioning, and visual question answering in microgravity environments. It probes a model's ability to generalize to orientation-invariant motion, floating objects, and lack of ground contact where terrestrial models typically fail. Use when the user wants to benchmark on MicroG-4M, or asks about evaluating this task. Reports mAP@0.5.

researchpythongo
0
3
Microisp EvalA

Evaluates a deep learning-based image signal processing (ISP) model's ability to reconstruct high-resolution RGB images from RAW mobile sensor data. It measures reconstruction fidelity, visual quality, and inference efficiency across various mobile hardware platforms and resolutions. Use when the user wants to benchmark on Fujifilm UltraISP, or asks about evaluating this task. Reports PSNR.

researchpythongo
0
3
Microseismic Fno EvalA

Evaluates a lightweight Fourier Neural Operator model's ability to classify microseismic events versus noise in seismic waveforms. It probes resolution-invariant signal processing, cross-domain generalization, and real-time detection efficiency under varying signal-to-noise ratios. Use when the user wants to benchmark on STEAD, Microseismic dataset, or asks about evaluating this task. Reports F1 score.

researchpythongo
0
3
Microservice Latency Estimation EvalA

This benchmark evaluates the accuracy of machine learning models in estimating microservice latency under diverse workload conditions. It probes the model's ability to capture hierarchical system behaviors and adapt to different operational scenes using non-intrusive service mesh monitoring data. Use when the user wants to benchmark on Online Boutique, Sock Shop, or asks about evaluating this task. Reports MAE.

researchpythonperformance
0
3
Microsoft Malware Classification EvalA

Evaluates the ability of hybrid machine learning models to classify Windows malware into specific family categories by fusing hand-crafted structural features with deep learning-derived representations. It probes robustness against class imbalance and measures both categorical correctness and probabilistic calibration. Use when the user wants to benchmark on Microsoft Malware Classification Challenge Dataset, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Midam Mil Auc EvalA

Evaluates multi-instance learning models for AUC maximization across tabular, histopathological, and medical image datasets. It probes the model's ability to handle large bags of instances via stochastic pooling while optimizing a min-max margin AUC loss. Use when the user wants to benchmark on MUSK1, MUSK2, Elephant, Fox, Tiger, Breast Cancer, Colon Ade., PDGM, OCT, or asks about evaluating this task. Reports testing AUC.

researchpythontesting
0
3
Midog EvalA

Evaluates deep learning models for detecting mitotic figures in histopathology whole-slide images, specifically probing their ability to generalize across different scanner-induced domain shifts such as color distribution, contrast, and depth-of-field variations. Use when the user wants to benchmark on MIDOG, or asks about evaluating this task. Reports F_1 score.

researchpythondocker
0
3
Mieb EvalA

MIEB evaluates the diverse capabilities of image and image-text embedding models across 130 tasks spanning retrieval, document understanding, classification, clustering, compositionality, and visual question answering. It probes zero-shot generalization, multilingual understanding, spatial/depth reasoning, and the model's ability to encode visual representations of text. Use when the user wants to benchmark on MIEB, or asks about evaluating this task. Reports nDCG@10.

ai-agentspythongo
0
3
Mig Bench EvalA

Evaluates a model's ability to perform free-form, multi-image visual grounding by localizing specified objects across multiple input images based on natural language instructions. It probes cross-image reasoning, spatial understanding, and the capacity to follow complex, unstructured queries without relying on chain-of-thought abstractions. Use when the user wants to benchmark on MIG-Bench, or asks about evaluating this task. Reports Acc_0.5.

researchpythonexpress
0
3
Migperf EvalA

Evaluates deep learning training and inference workloads on NVIDIA's Multi-Instance GPU (MIG) technology. It systematically measures performance trade-offs across different partition sizes, batch sizes, and model architectures, while comparing MIG against software-based GPU sharing (MPS) and testing framework compatibility. Use when the user wants to benchmark on Representative DL Models (ResNet18, ResNet50, BERT, ViT, Diffusion), or asks about evaluating this task. Reports tail latency.

researchpythontesting
0
3
Milabench EvalA

Evaluates the real-world computational performance and software stack maturity of AI accelerators across diverse workloads including NLP, computer vision, reinforcement learning, and graph neural networks. Use when the user wants to benchmark on Milabench, or asks about evaluating this task. Reports performance.

researchpythonnode
0
3
Milan Bs Sleep EvalA

Evaluates a deep reinforcement learning framework for dynamic base station sleep control and spatio-temporal traffic forecasting in a real-world cellular network. It probes the model's ability to accurately predict mobile traffic demand across geographical grids and make energy-efficient on/off decisions for base stations while balancing switching costs and quality of service. Use when the user wants to benchmark on Telecom Italia Milan Mobile Traffic Dataset, or asks about evaluating this ta...

researchpythonperformance
0
3
Milebench EvalA

Evaluates Multimodal Large Language Models (MLLMs) on long-context, multi-image comprehension. It probes capabilities like needle-in-a-haystack retrieval, image retrieval, temporal reasoning across multiple images, and semantic understanding in long multimodal contexts. Use when the user wants to benchmark on MileBench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Milpac EvalA

Evaluates the quality of machine translation systems translating English legal text into nine Indian languages. It benchmarks commercial and open-source models against human-translated references to assess domain-specific translation accuracy and metric correlation. Use when the user wants to benchmark on MILPaC, or asks about evaluating this task. Reports BLEU.

researchpythonapi
0
3
Milu EvalA

Evaluates large language models on multi-task understanding across 11 Indic languages, covering 8 domains and 41 subjects. It probes cultural knowledge, region-specific exam data, and multilingual reasoning capabilities. Use when the user wants to benchmark on MILU, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Mimic Cdm EvalA

Evaluates large language models on clinical decision-making tasks by predicting diagnoses from structured patient evidence. It probes whether models possess latent domain-specific reasoning capabilities that are masked by unfamiliarity with benchmark input formats and task definitions. Use when the user wants to benchmark on MIMIC-CDM, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Mimic Cxr Medical EvalA

Evaluates the robustness, consistency, and generalization of a distilled multimodal large language model on medical imaging tasks. It probes the model's ability to perform multi-label chest disease classification and generate diagnostic radiology reports from chest X-ray images, even when trained on limited data from drifting teacher models. Use when the user wants to benchmark on MIMIC-CXR, or asks about evaluating this task. Reports Top-1 accuracy.

researchpythongo
0
3
Mimic Dos EvalA

Evaluates clinical decision-making under conflicting subjective and objective evidence by testing whether agentic reasoning workflows can correctly classify ICU patient states without collapsing to one-sided predictions. Use when the user wants to benchmark on MIMIC-DOS, or asks about evaluating this task. Reports Matthews Correlation Coefficient (MCC).

researchpythongo
0
3
Mimic If Interpretability EvalA

Evaluates the quality of feature importance explanations for deep learning models in healthcare mortality prediction. It measures how well an interpretability method identifies truly predictive features by observing model performance degradation when those features are removed. Use when the user wants to benchmark on MIMIC-IV, or asks about evaluating this task. Reports AUC of performance curve.

researchpythongo
0
3
Mimic Iii Benchmark EvalA

Evaluates clinical risk prediction and intervention forecasting using structured EHR time-series data. Probes a model's ability to handle missingness, temporal gaps, and class imbalance in ICU patient records. Use when the user wants to benchmark on MIMIC-III, or asks about evaluating this task. Reports AUROC.

researchpythonperformance
0
3
Mimic Iii Clinical BenchmarksA

Evaluates clinical time-series models on four critical care prediction tasks: in-hospital mortality, physiologic decompensation, length of stay forecasting, and acute care phenotyping. Tests the model's ability to handle multivariate temporal data, capture long-range dependencies, and perform binary, multi-class, and multi-label classification on real-world ICU records. Use when the user wants to benchmark on MIMIC-III, or asks about evaluating this task. Reports test performance.

researchpythongo
0
3
Mimic Iii Clinical Fact Checking EvalA

Evaluates the factual consistency and logical coherence of LLM-generated clinical discharge summaries against ground-truth Electronic Health Records (EHRs) at a granular propositional level. It measures how accurately a model's extracted propositions align with clinician-validated EHR facts. Use when the user wants to benchmark on MIMIC-III, or asks about evaluating this task. Reports F1-score.

researchpythongo
0
3
Mimic Iii Clinical Fairness EvalA

Probes multimodal clinical NLP models on in-hospital mortality and phenotyping tasks, evaluating both predictive performance (AUC) and group fairness (Equalized Odds) across protected demographic groups. Use when the user wants to benchmark on MIMIC-III, or asks about evaluating this task. Reports AUC ROC.

researchpythongo
0
3
Mimic Iii Clinical Notes EvalA

Evaluates multi-modal deep learning models for predicting patient outcomes (decompensation, in-hospital mortality, phenotyping) using electronic health records (EHR) and clinical notes. It probes the model's ability to fuse tabular/time-series physiological data with unstructured clinical text to improve clinical decision support. Use when the user wants to benchmark on MIMIC-III, or asks about evaluating this task. Reports AUPRC.

researchpythonperformance
0
3
Mimic Iii Healthcare Benchmark EvalA

This benchmark evaluates deep learning and traditional machine learning models on critical care prediction tasks using the MIMIC-III dataset. It probes a model's ability to predict patient mortality across multiple time horizons, classify ICD-9 diagnosis groups, and forecast hospital length of stay from raw clinical time series and tabular data. Use when the user wants to benchmark on MIMIC-III, or asks about evaluating this task. Reports binary classification.

datapythongo
0
3
Mimic Iii Multitask EvalA

Evaluates clinical time series models on four interrelated ICU prediction tasks: in-hospital mortality, physiologic decompensation, length of stay, and phenotype classification. It probes a model's ability to handle heterogeneous multitask learning with varying temporal structures and output types. Use when the user wants to benchmark on MIMIC-III, or asks about evaluating this task. Reports Test score.

researchpythongo
0
3
Mimic Iv Clinical Prediction EvalA

Evaluates clinical prediction models on irregular, sparse time-series data from the ICU module of MIMIC-IV. It probes the ability of models to handle temporal sparsity, missingness, and heterogeneous features for binary classification tasks like in-ICU mortality and length of stay prediction. Use when the user wants to benchmark on MIMIC-IV 2.2, or asks about evaluating this task. Reports AUC-ROC.

researchpythontesting
0
3
Mimic Iv Ecg Classification EvalA

This evaluation probes the ability of transformer-based models to classify cardiac rhythms using time-series features derived from dynamical systems theory (Koopman operator) and signal processing (wavelets). It specifically tests binary and four-class rhythm classification on clinical ECG waveforms, comparing feature-based approaches against raw RNN baselines. Use when the user wants to benchmark on MIMIC-IV-ECG, or asks about evaluating this task. Reports F1 score.

researchpythongo
0
3
Mimic Iv Icd EvalA

Predicts ICD-9 and ICD-10 medical billing codes from clinical discharge notes under extreme multi-label classification settings. It probes a model's ability to handle long-tailed label distributions, high cardinality, and code hierarchy variations in electronic health records. Use when the user wants to benchmark on MIMIC-IV-ICD9, MIMIC-IV-ICD10, MIMIC-IV-ICD9-50, MIMIC-IV-ICD10-50, or asks about evaluating this task. Reports Macro-F1.

researchpythongo
0
3
Mimic Iv Tu Dataset EvalA

Evaluates self-supervised graph representation learning on electronic health records and general graph classification tasks. It probes the model's ability to learn temporal and structural patient representations without task-specific fine-tuning, and its robustness across clinical and non-clinical domains. Use when the user wants to benchmark on MIMIC-IV, TUDataset, or asks about evaluating this task. Reports binary classification accuracy.

testingpythonnode
0
3
Mimic Sepsis EvalA

Evaluates the ability of models to predict clinical outcomes (mortality, length of stay, shock onset) from time-aligned ICU patient trajectories and treatment dynamics. Use when the user wants to benchmark on MIMIC-Sepsis, or asks about evaluating this task. Reports performance.

researchpythongo
0
3
Mimic Vqa EvalA

Evaluates a model's ability to answer clinical questions about chest X-rays, focusing on disease presence, type, location, and severity. It probes multi-modal reasoning by requiring the model to correlate image regions with structured medical knowledge and spatial/semantic relationships. Use when the user wants to benchmark on Mimic-VQA, or asks about evaluating this task. Reports AUC-micro.

researchpythongo
0
3
Mimicking Bench EvalA

Evaluates generalizable humanoid-scene interaction learning by testing motion retargeting, tracking, and imitation learning across six household tasks. It probes the agent's ability to mimic human references to interact with diverse, unseen object geometries while maintaining physical plausibility and energy efficiency. Use when the user wants to benchmark on Mimicking-Bench, or asks about evaluating this task. Reports kinematic success rate.

researchpythongo
0
3
Mimicsql EvalA

Evaluates a model's ability to generate correct SQL queries from natural language clinical questions in the healthcare domain. It specifically probes retrieval-based reasoning, robustness to noisy or ambiguous medical terminology, and generalization under data scarcity conditions. Use when the user wants to benchmark on MIMICSQL, or asks about evaluating this task. Reports Execution Accuracy.

databasespythongo
0
3
Mimii Dg EvalA

Evaluates domain generalization capabilities for anomalous sound detection by measuring how well models trained on a source domain of industrial machine sounds generalize to target domains with shifted operational parameters or background noise. Use when the user wants to benchmark on MIMII DG, or asks about evaluating this task. Reports AUC.

researchpythonperformance
0
3
Mimo Embodied EvalA

Evaluates a cross-embodied foundation model's capabilities in affordance prediction, task planning, spatial understanding, and autonomous driving perception/prediction/planning across diverse visual and language inputs. Use when the user wants to benchmark on RoboRefIt, Where2Place, VABench-Point, Part-Afford, RoboAfford-Eval, EgoPlan2, RoboVQA, Cosmos-Reason1, CV-Bench, ERQA, EmbSpatial, SAT, RoboSpatial, RefSpatial-Bench, CRPE-relation, MetaVQA, VSI-Bench, CODA-LM, DRAMA, MME-RealWorld, IDK...

researchpythongo
0
3
Mimotable EvalA

Evaluates large language models' ability to reason over real-world spreadsheet data, including complex headers, multi-sheet files, and cross-file contexts. It probes capabilities across six meta operations: lookup, edit, calculate, compare, visualize, and reasoning. Use when the user wants to benchmark on MiMoTable, or asks about evaluating this task. Reports accuracy.

datapythongo
0
3
Mimotion EvalA

Evaluates models on predicting future 3D multi-person motion sequences given a short history of interacting subjects. It probes spatial-temporal modeling, interaction awareness, and long-horizon trajectory forecasting under varying scene complexities and prediction horizons. Use when the user wants to benchmark on MI-Motion, or asks about evaluating this task. Reports GJPE, AJPE, RFDE.

researchpythontesting
0
3
Mina Ecg Af EvalA

Evaluates deep learning models for binary classification of atrial fibrillation (AF) versus control using single-lead ECG recordings. It probes the model's ability to capture beat-level morphology, rhythm-level dynamics, and frequency-domain patterns for clinical AF detection. Use when the user wants to benchmark on PhysioNet Challenge 2017, or asks about evaluating this task. Reports PR-AUC.

researchpythonperformance
0
3
Mind Benchmark EvalA

Evaluates an AI co-scientist framework's ability to automatically validate materials science hypotheses using MLIP-based simulations. It measures both binary verification accuracy across energetic, mechanical, and structural categories, and human-rated scientific utility via expert feedback. Use when the user wants to benchmark on MIND MLIP-expert-curated benchmark, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Mind2web EvalA

Evaluates a model's ability to act as a generalist web agent by completing multi-step tasks across diverse, unseen websites and domains. It probes out-of-distribution generalization, web element grounding, and sequential action planning in real-world browser environments. Use when the user wants to benchmark on Mind2Web, or asks about evaluating this task. Reports Step Success Rate.

researchpythongo
0
3
Mind2web Live EvalA

Evaluates a GUI agent's ability to perform web browsing tasks using either HTML tree or image inputs. It measures the agent's capacity to navigate websites and complete user intents across diverse web interfaces. Use when the user wants to benchmark on Mind2Web-Live, or asks about evaluating this task. Reports task success rate.

researchpythongo
0
3
Mindbench EvalA

This benchmark evaluates multimodal large language models on structured document analysis, specifically focusing on mind map parsing and visual question answering. It probes text recognition, spatial awareness, hierarchical relationship discernment, and the ability to reconstruct complex graphical tree structures from high-resolution images. Use when the user wants to benchmark on MindBench, or asks about evaluating this task. Reports TED-based accuracy.

researchpythongo
0
3
Mindgames EvalA

Evaluates large language models' ability to perform higher-order epistemic reasoning and multi-agent belief tracking. It probes whether models can correctly update beliefs based on public announcements and answer True/False questions about agents' knowledge states. Use when the user wants to benchmark on MindGames, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Mine Kg EvalA

Evaluates the ability of LLM-based methods to construct knowledge graphs from raw text that preserve factual information and maintain structural coherence. It measures how well extracted graphs retain ground-truth atomic facts and how densely connected and non-fragmented the resulting graphs are. Use when the user wants to benchmark on MINE, or asks about evaluating this task. Reports Factual Retention Score.

researchpythonnode
0
3
Mined EvalA

Probes large multimodal models' temporal awareness and time-sensitive knowledge across six dimensions: cognition, awareness, trustworthiness, understanding, reasoning, and robustness. It evaluates how well models recall, reason about, and reject outdated or misaligned temporal facts in multimodal queries. Use when the user wants to benchmark on Mined, or asks about evaluating this task. Reports Cover Exact Match (CEM).

researchpythonrust
0
3
Miners EvalA

Evaluates multilingual language models as semantic retrievers across 200+ languages without fine-tuning. It probes bitext mining, retrieval-augmented classification, and in-context learning classification to assess cross-lingual and code-switching capabilities. Use when the user wants to benchmark on MINERS (includes NusaX), or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Minerva Quant Reasoning EvalA

Evaluates large language models on quantitative reasoning and mathematical problem-solving tasks, including arithmetic, multi-step word problems, and standardized test questions, using few-shot prompting and chain-of-thought reasoning. Use when the user wants to benchmark on MATH, MMLU, GSM8k, National Math Exam in Poland, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3