
Claude Skills by qhjqhj00
github.com/qhjqhj00This benchmark evaluates multi-context visual grounding, requiring models to localize target objects across multiple images using open-ended, context-rich text prompts. It probes cross-image reasoning, fine-grained instance localization, and the ability to correctly group and reject irrelevant instances. Use when the user wants to benchmark on MC-Bench, or asks about evaluating this task. Reports AP50.
Evaluates AI-augmented workflow scheduling in mobile edge-cloud environments by comparing execution time, energy consumption, SLA violations, and fairness against state-of-the-art baselines under dynamic workloads and host mobility. Use when the user wants to benchmark on WFCommons (Pegasus workflows: BLAST, Cycles, Montage), or asks about evaluating this task. Reports SLA Violation Rate.
Evaluates the multilingual code generation, explanation, and completion capabilities of LLMs across 40 programming languages. It measures how well models can produce correct code, explain code logic, and complete code snippets in diverse syntaxes. The benchmark highlights performance disparities between closed-source and open-source models, particularly in non-Python languages. Use when the user wants to benchmark on MCEVAL, or asks about evaluating this task. Reports Pass@1 (%).
Evaluates optical flow models on synthetic Particle Image Velocimetry (PIV) datasets with varying particle densities, flow velocities, and turbulence levels. Probes spatio-temporal flow modeling and robustness to sparse particle imagery and high-speed turbulent regimes. Use when the user wants to benchmark on PIV Benchmark (MHD, Isotropic, Mixing, Channel, Boundary Layer), or asks about evaluating this task. Reports NEPE.
Evaluates multimodal models' ability to follow crosslingual instructions on scientific talks, testing speech recognition, translation, question answering, and summarization across short and long contexts in English, German, Italian, and Chinese. Use when the user wants to benchmark on MCIF, or asks about evaluating this task. Reports BERTScore.
Evaluates multimodal continual instruction tuning (MCIT) methods on MLLMs, measuring their ability to learn new multimodal tasks sequentially while mitigating catastrophic forgetting and cross-modal conflict. Use when the user wants to benchmark on MLLM-DCL, UCIT, or asks about evaluating this task. Reports MAA, MFN.
Evaluates an open-vocabulary data selection pipeline for iteratively improving object detection models on rare classes. Specifically, it tests whether adding newly selected and labeled frames of vulnerable road users (pedestrians, cyclists) to a seed dataset improves detection performance on fisheye traffic camera data. Use when the user wants to benchmark on SIP VRU Detection Dataset, or asks about evaluating this task. Reports mAP@0.5.
Evaluates the downstream performance and sample efficiency of the MCL pre-trained language model on general language understanding and reading comprehension benchmarks. It measures how well the model captures multi-perspective semantics and self-corrects during pre-training when fine-tuned on standard NLP tasks. Use when the user wants to benchmark on GLUE benchmark, SQuAD 2.0, or asks about evaluating this task. Reports GLUE Average.
This benchmark evaluates the accuracy and reliability of Markov Chain Monte Carlo (MCMC) light curve fitting routines for detecting and measuring exoplanet secondary eclipses. It probes how well analysis pipelines recover known eclipse depths and phase centers under varying noise conditions, including synthetic white/red noise and real observational systematics. Use when the user wants to benchmark on MCMC Eclipse Benchmark Suite, or asks about evaluating this task. Reports eclipse depth.
Evaluates the end-to-end and component-level latency of retrieving and searching AI/ML model cards across three server architectures (REST, Native MCP, Layered MCP) under local and wide-area network conditions. It measures how protocol overhead, payload size, and network distance impact system responsiveness in edge computing environments. Use when the user wants to benchmark on Patra Model Cards (Pseudo-Synthetic), or asks about evaluating this task. Reports end-to-end latency.
Evaluates the ability of hyperbolic and Euclidean LLMs to answer multiple-choice questions across STEM, general knowledge, and commonsense reasoning domains. It probes how well the models capture semantic hierarchies and perform complex reasoning under few-shot and zero-shot settings. Use when the user wants to benchmark on MMLU, ARC-Challenging, CommonsenseQA, HellaSwag, OpenbookQA, or asks about evaluating this task. Reports accuracy.
Evaluates large language models' ability to perform compositional relation reasoning across multiple languages. It tests whether models can infer a transitive relationship (R∘S) from two given relations (R and S) using multiple-choice questions covering positional, comparative, personal, mathematical, identity, and other logical relations. Use when the user wants to benchmark on Multilingual Compositional Relation (MCR), or asks about evaluating this task. Reports Accuracy (%).
Evaluates a model's ability to reason about temporal commonsense, including event duration, ordering, typical time, frequency, and stationarity. It tests whether systems can correctly classify candidate answers as 'likely' or 'unlikely' given a context sentence and a question. Use when the user wants to benchmark on MCTACO, or asks about evaluating this task. Reports F1.
Evaluates multimodal large language models on fine-grained video captioning by measuring how well generated descriptions capture verified, detailed key points from videos. It probes the model's ability to reason about and describe specific visual details like color, quantity, and position. Use when the user wants to benchmark on MCTS-VCB, or asks about evaluating this task. Reports F1.
Evaluates a model's ability to follow natural language instructions to perform multi-step, spatially grounded tasks in an open-world 3D environment (Minecraft). It specifically probes capabilities in mining, combat, crafting, and smelting under human-like visibility constraints. Use when the user wants to benchmark on MCU Benchmark, or asks about evaluating this task. Reports success rate.
Evaluates a model's ability to generalize compositionally to unseen syntactic structures and cross-lingual settings in semantic parsing. It measures how well models translate natural language questions into correct SPARQL queries across monolingual and zero-shot cross-lingual scenarios. Use when the user wants to benchmark on MCWQ, or asks about evaluating this task. Reports Exact Match (%).
Evaluates large language models on molecular dynamics domain knowledge, LAMMPS scripting syntax comprehension, and automatic generation of executable LAMMPS simulation scripts from natural language instructions. Use when the user wants to benchmark on MD-EvalBench, or asks about evaluating this task. Reports Exec-Success@$k$.
Evaluates the reliability of micro-benchmarks in preserving pairwise model rankings compared to full benchmark performance. It quantifies the minimum accuracy gap required between two models for a small sampled subset to consistently rank them correctly. Use when the user has predictions and gold and needs to compute MDAD.
Evaluates self-supervised monocular depth estimation models by measuring image-based accuracy and 3D pointcloud reconstruction quality. It probes the models' ability to generalize across diverse environments (urban, natural, agricultural, indoor) and highlights the impact of scale ambiguity and oversmoothing on relative object positioning. Use when the user wants to benchmark on SYNS-Patches, or asks about evaluating this task. Reports F-Score (Edges).
Evaluates monocular depth estimation models across diverse real-world environments (natural, agricultural, urban, indoor) using high-quality LiDAR ground truth. Probes zero-shot generalization, boundary interpolation accuracy, and robustness to scene diversity and image artifacts. Use when the user wants to benchmark on SYNS-Patches, or asks about evaluating this task. Reports F-Score.
Evaluates monocular depth estimation models on their ability to predict accurate depth maps from single images. It specifically probes zero-shot generalization across diverse natural and indoor scenes using LiDAR-ground truth. Use when the user wants to benchmark on SYNS-Patches, or asks about evaluating this task. Reports F-Score.
Evaluates models' ability to classify pragmatic gender bias in text across three conversational dimensions: ABOUT (topic-related), AS (speaker role/attitude), and TO (addressee-directed). It probes fine-grained detection of gendered language and contextual bias cues. Use when the user wants to benchmark on MDGENDER, or asks about evaluating this task. Reports percentage accuracy.
This benchmark evaluates a model's ability to generate coherent, contextually appropriate, and lexically diverse dialogue responses across 46 languages. It specifically probes cross-lingual transfer capabilities and measures the performance gap between high-resource and low-resource languages in open-domain conversation. Use when the user wants to benchmark on MDIA, or asks about evaluating this task. Reports sacreBLEU.
Evaluates the inference performance and overhead of six major mobile deep learning libraries across diverse model architectures, tasks, and mobile hardware configurations. It measures how software-level optimizations and library choices impact on-device latency compared to hardware capabilities and algorithmic optimizations like quantization. Use when the user wants to benchmark on MDLBench, or asks about evaluating this task. Reports inference time.
Compute mdocekal/multi_label_precision_recall_accuracy_fscore via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of mdocekal/multi_label_precision_recall_accuracy_fscore.
Compute mdocekal/precision_recall_fscore_accuracy via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of mdocekal/precision_recall_fscore_accuracy.
MDPBench evaluates the capability of document parsing models to accurately extract text, formulas, tables, and layout structures from multilingual document images under real-world conditions. It specifically probes robustness to photographic degradation, non-Latin scripts, right-to-left reading orders, and cross-lingual generalization without prior language or image-type knowledge. Use when the user wants to benchmark on MDPBench, or asks about evaluating this task. Reports accuracy.
Evaluates multimodal deception detection across video, audio, and text modalities, while also probing how individual differences—specifically personality traits and emotional expressivity—influence deceptive behavior and detection accuracy. Use when the user wants to benchmark on MDPE, or asks about evaluating this task. Reports Accuracy.
This benchmark evaluates a multimodal deep learning model's ability to predict 33 distinct ICU clinical outcomes by fusing 10-second 12-lead ECG waveforms with structured tabular clinical data. It probes the model's discriminative capacity and probabilistic calibration across mortality, medication administration, clinical deterioration, and organ dysfunction tasks. Use when the user wants to benchmark on MDS-ICU, or asks about evaluating this task. Reports macro-averaged AUROC.
Compute the mean_absolute_error metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute mean_absolute_error, or asks how to score with mean_absolute_error.
Compute the mean_absolute_percentage_error metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute mean_absolute_percentage_error, or asks how to score with mean_absolute_percentage_error.
Compute the mean_gamma_deviance metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute mean_gamma_deviance, or asks how to score with mean_gamma_deviance.
Compute the mean_pinball_loss metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute mean_pinball_loss, or asks how to score with mean_pinball_loss.
Compute the mean_poisson_deviance metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute mean_poisson_deviance, or asks how to score with mean_poisson_deviance.
Compute the mean_squared_error metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute mean_squared_error, or asks how to score with mean_squared_error.
Compute the mean_squared_log_error metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute mean_squared_log_error, or asks how to score with mean_squared_log_error.
Compute the mean_tweedie_deviance metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute mean_tweedie_deviance, or asks how to score with mean_tweedie_deviance.
Compute the MeanAbsoluteError metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MeanAbsoluteError, or asks how to score with MeanAbsoluteError.
Compute the MeanAbsolutePercentageError metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MeanAbsolutePercentageError, or asks how to score with MeanAbsolutePercentageError.
Compute the MeanMetric metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MeanMetric, or asks how to score with MeanMetric.
Compute the MeanSquaredError metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MeanSquaredError, or asks how to score with MeanSquaredError.
Compute the MeanSquaredLogError metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MeanSquaredLogError, or asks how to score with MeanSquaredLogError.
Evaluates vision-language models on fine-grained visual measurement reading, specifically testing their ability to accurately localize pointers and ticks on instrument scales, map visual cues to numerical values, and recognize measurement units from real-world and synthetic images. Use when the user wants to benchmark on MeasureBench, or asks about evaluating this task. Reports Overall accuracy.
Evaluates vision-language models' ability to ground objects and exhibit mutual exclusivity bias when mapping novel pseudo-labels to unknown items in cluttered scenes. It also measures spatial reasoning capabilities and the model's ability to resolve ambiguity among multiple novel objects. Use when the user wants to benchmark on MEBench, or asks about evaluating this task. Reports ME score.
Evaluates the stability, convergence, and resource efficiency of online computation offloading algorithms in dynamic mobile-edge networks under stochastic task arrivals and time-varying channel conditions. It probes whether an algorithm can maintain queue stability and power constraints while maximizing computation throughput. Use when the user wants to benchmark on Simulated MEC Offloading Environment, or asks about evaluating this task. Reports weighted sum computation rate.
This benchmark evaluates fine-grained audio understanding by testing models on generating detailed, multi-perspective captions and answering probing questions across diverse acoustic domains. It specifically probes a model's ability to distinguish between speech, music, and sound events, reason about acoustic scenes, and assess technical audio quality without relying on generic descriptions. Use when the user wants to benchmark on MECAT, or asks about evaluating this task. Reports DATE.
Evaluates multimodal egocentric video understanding in an industrial-like setting. It probes action recognition, active object detection, human-object interaction, and future action anticipation using synchronized RGB, depth, and gaze signals. Use when the user wants to benchmark on MECCANO, or asks about evaluating this task. Reports Top-1 Accuracy.
This evaluation probes a model's ability to detect pathological anomalies in medical images using a one-class learning setting. It measures how well the model distinguishes between normal and abnormal samples across diverse imaging modalities and anatomical regions without seeing abnormal examples during training. Use when the user wants to benchmark on RSNA Pneumonia, VinDr-CXR, Brain Tumor, LAG, ISIC 2018, Camelyon16, BraTS2021, or asks about evaluating this task. Reports AUC-ROC.
This benchmark evaluates a model's ability to verify factual accuracy in long-form medical texts by recursively decomposing claims into a verification tree. It probes fine-grained fact-checking across six medical domains, requiring the model to distinguish between factual and deliberately falsified claims while accounting for contextual dependencies. Use when the user wants to benchmark on Med-Critics, or asks about evaluating this task. Reports accuracy.
Evaluates medical vision-language models on multi-image reasoning tasks, including temporal understanding, cross-modal comparison, multi-view diagnosis, and co-reference resolution across longitudinal and multi-modality medical imaging data. Use when the user wants to benchmark on Med-MIM Benchmark, or asks about evaluating this task. Reports closed-type accuracy.