All authors
qhjqhj00 avatar

Claude Skills by qhjqhj00

github.com/qhjqhj00
7,636 skillsA× 7,623B× 11C× 1D× 10 installs2,217 views
Ming Moe Medical EvalA

Evaluates large language models on a comprehensive suite of medical natural language processing tasks and medical licensing examinations. It probes the model's ability to process clinical text, perform information extraction, and demonstrate domain-specific knowledge and reasoning for medical exams. Use when the user wants to benchmark on CBLUE (via PromptCBLUE), MedQA, MedMCQA, CMB, CMExam, MMLU (medical subset), C-Eval (medical subset), CMMLU (medical subset), 2023 Chinese National Pharmaci...

researchpythongo
0
3
Mini Behavior EvalA

Probes long-horizon decision-making and multi-state object interaction in a procedurally generated 3D gridworld. Evaluates an agent's ability to plan and execute complex household tasks under sparse and dense reward signals. Use when the user wants to benchmark on Mini-BEHAVIOR, or asks about evaluating this task. Reports success_rate.

researchpythongo
0
3
Mini Crosswords EvalA

Evaluates an LLM's ability to perform algorithmic search and recursive reasoning within a single generation window, without external tree search or iterative prompting. It probes systematic exploration, pruning, and backtracking capabilities in a lexical constraint satisfaction task. Use when the user wants to benchmark on Mini Crosswords, or asks about evaluating this task. Reports Word success rate.

researchpythongo
0
3
Minif2f Pipeline EvalA

Evaluates the end-to-end capability of autoformalizers and theorem provers to translate informal mathematical statements into verified Lean 4 proofs. It probes semantic fidelity during translation and the ability of provers to generate correct, aligned proofs for Olympiad-style problems. Use when the user wants to benchmark on miniF2F, or asks about evaluating this task. Reports effective_accuracy.

researchpythongo
0
3
Minimax Speech EvalA

Evaluates zero-shot and one-shot text-to-speech voice cloning fidelity, multilingual synthesis capability, and cross-lingual generalization. It measures perceptual naturalness and speaker identity preservation through objective transcription and embedding similarity metrics, alongside human preference rankings. Use when the user wants to benchmark on Seed-TTS-eval, Artificial Arena, MiniMax Multilingual Test Set, or asks about evaluating this task. Reports WER, SIM.

researchpython
0
3
Minimwob EvalA

Tests a visual GUI agent's ability to complete web automation tasks by interacting with simplified web environments based on screenshots. It measures task completion success rates across various interactive web widgets. Use when the user wants to benchmark on MiniWob, or asks about evaluating this task. Reports success rate.

researchpythonperformance
0
3
Minisuperb EvalA

Evaluates self-supervised speech models by measuring computational efficiency (forward MACs) and downstream task performance using a lightweight, offline feature extraction protocol. It probes the trade-off between model complexity and representation quality across speech tasks while enabling rapid early-stage model screening. Use when the user wants to benchmark on MiniSUPERB, or asks about evaluating this task. Reports forward MACs.

businesspythongit
0
3
Minivla Nav V1 EvalA

Language-conditioned robot navigation in continuous differential-drive settings. It probes a model's ability to process multi-modal observations (RGB, depth) and language instructions to output continuous control actions to reach a target object within a specified distance. Use when the user wants to benchmark on MiniVLA-Nav v1, or asks about evaluating this task. Reports success.

researchpythongo
0
3
Miniwob Wge EvalA

Evaluates an agent's ability to navigate and interact with semi-structured web interfaces to complete goal-directed tasks. It probes relational reasoning over DOM trees, handling of natural language instructions, and sample efficiency in sparse-reward reinforcement learning settings. Use when the user wants to benchmark on MiniWoB, MiniWoB++, Alaska, or asks about evaluating this task. Reports success rate.

developmentjavascriptpython
0
3
MinkowskidistanceA

Compute the MinkowskiDistance metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MinkowskiDistance, or asks how to score with MinkowskiDistance.

documentationpython
0
3
MinmaxmetricA

Compute the MinMaxMetric metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MinMaxMetric, or asks how to score with MinMaxMetric.

documentationpython
0
3
MinmetricA

Compute the MinMetric metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MinMetric, or asks how to score with MinMetric.

documentationpython
0
3
Mint 1t EvalA

Evaluates the multimodal interleaved reasoning and in-context learning capabilities of large multimodal models (LMMs) across image captioning, visual question answering, and multi-image reasoning tasks. Use when the user wants to benchmark on COCO (Karpathy test), TextCaps, VQAv2, OK-VQA, TextVQA, VizWiz, MMMU, Mantis-Eval, or asks about evaluating this task. Reports scores.

researchpythongo
0
3
Mint EvalA

Evaluates large language models' ability to solve tasks using external tools across multiple interaction turns, and their capacity to leverage natural language feedback to improve performance. It also measures the rate of improvement per turn and identifies failure patterns like formatting issues or training data artifacts. Use when the user wants to benchmark on MINT, or asks about evaluating this task. Reports Success Rate (SR).

researchpythongo
0
3
Mip Solution Prediction EvalA

Evaluates a graph neural network's ability to predict binary variable values in mixed-integer programming (MIP) instances. It also measures how these predictions accelerate primal solution finding and reduce optimality gaps in a Branch-and-Bound solver. Use when the user wants to benchmark on MIP Instances (8 types), or asks about evaluating this task. Reports average precision (AP).

researchpythonnode
0
3
Miqa EvalA

Evaluates how image degradations impact machine vision system (MVS) performance rather than human perception. It measures the correlation between predicted image quality scores and ground-truth machine task metrics (accuracy and consistency) across classification, detection, and segmentation tasks. Use when the user wants to benchmark on MIQD-2.5M, or asks about evaluating this task. Reports SRCC.

researchpythongo
0
3
Mir EvalA

Evaluates multimodal large language models on progressive, interleaved multi-image reasoning tasks. It probes the model's ability to perform structured, step-by-step reasoning across multiple images, including text-to-region alignment, cross-image relationship modeling, and analytical inference. Use when the user wants to benchmark on MIR, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Mir Ref EvalA

This framework evaluates the quality, robustness, and downstream extractability of learned music audio representations. It probes how well representations encode task-relevant information (e.g., instruments, pitch, singer identity) and measures their resilience to real-world audio degradations like noise, gain changes, and compression. Use when the user wants to benchmark on TinySOL, Beatport EDM, VocalSet, or asks about evaluating this task. Reports F1 score.

researchpythongo
0
3
Mirabest Confident EvalA

Evaluates the ability of self-supervised learning models to classify radio galaxy morphologies (e.g., Fanaroff-Riley classes) using standard image classification protocols. It measures how well disentangled generative augmentations and contrastive learning pipelines capture astrophysical structure compared to traditional data augmentations. Use when the user wants to benchmark on MiraBest Confident, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Miracl EvalA

Evaluates multi-lingual ad-hoc retrieval across 18 languages, testing a model's ability to match queries and passages in the same language using dense, sparse, and multi-vector embedding strategies. Use when the user wants to benchmark on MIRACL, or asks about evaluating this task. Reports nDCG@10.

ai-agentspythontesting
0
3
Mirage EvalA

Evaluates multimodal vision-language models on expert-level agricultural reasoning, including grounded entity identification, causal explanation quality, and dialogue management decisions (clarify vs. respond) under partial observability. Use when the user wants to benchmark on MIRAGE-MMST, MIRAGE-MMMT, or asks about evaluating this task. Reports Identification Accuracy.

researchpythongo
0
3
Mirage ScoreA

This metric probes whether multimodal models rely on genuine visual grounding or exploit textual priors and benchmark structures to answer questions without actual image input. It quantifies the mirage effect where models generate confident, visually descriptive answers or achieve high accuracy despite the complete absence of visual data. Use when the user has predictions and gold and needs to compute mirage-score.

researchpythongo
0
3
Miragenews EvalA

Evaluates the ability of models to detect AI-generated news content by analyzing multimodal image-caption pairs. It specifically probes robustness to out-of-distribution generators (e.g., DALL-E 3, SDXL) and publishers (e.g., BBC, CNN) compared to in-domain training data. Use when the user wants to benchmark on MiRAGeNews, or asks about evaluating this task. Reports F-1.

researchpythongo
0
3
Miroeval EvalA

Evaluates multimodal deep research agents on both the quality of their final synthesized reports and the underlying investigative process. It measures adaptive synthesis quality, factual grounding against heterogeneous sources, and process-centric attributes like search breadth, analytical depth, and alignment between intermediate findings and the final report. Use when the user wants to benchmark on MiroEval, or asks about evaluating this task. Reports Adaptive Synthesis Quality (S_quality).

researchpythongit
0
3
Mirror Slate EvalA

Evaluates whether recommender systems manipulate user preferences through slate ranking strategies rather than accurately modeling true preferences. It quantifies the gap between observed click-through rates and clicks on genuinely favored items to detect exploitation of bounded rationality (e.g., decoy effects). Use when the user wants to benchmark on Synthetic Transportation Dataset, TianGong-ST, or asks about evaluating this task. Reports ManiScore.

researchpythongo
0
3
Mirrorbench EvalA

This benchmark evaluates self-centric intelligence and mirror self-recognition in Multimodal Large Language Models (MLLMs) within an embodied simulation. It probes the model's ability to perform self-referential reasoning and navigate tasks under varying cognitive difficulty levels and body configurations (humanoid vs. robotic). Use when the user wants to benchmark on MirrorBench, or asks about evaluating this task. Reports AVG.

researchpythongo
0
3
Misaligned Action Detection EvalA

This evaluation probes a computer-use agent's guardrail capability to detect and correct misaligned actions before execution. It measures how well a system distinguishes between benign, malicious, and task-irrelevant actions using both offline binary classification and online interactive task completion under adversarial and benign conditions. Use when the user wants to benchmark on MisActBench, RedTeamCUA, OSWorld, or asks about evaluating this task. Reports F1, Attack Success Rate (ASR).

researchpythongo
0
3
Misaw Seg EvalA

Evaluates pixel-wise and instance-level segmentation performance for microsurgical instruments, with a specific focus on accurately delineating extremely thin and sparse structures (e.g., wires, needles) under low-contrast, high-magnification conditions. The benchmark probes a model's ability to maintain fine boundaries and avoid background bias when object classes are severely imbalanced. Use when the user wants to benchmark on MISAW-Seg, or asks about evaluating this task. Reports mcIoU.

researchpythongo
0
3
Miss Qa EvalA

Evaluates multimodal foundation models' ability to interpret schematic diagrams in scientific papers and answer information-seeking questions based on visual-textual context. It also probes models' robustness in identifying unanswerable questions when sufficient information is absent. Use when the user wants to benchmark on MISS-QA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Missing Attribute Clustering EvalA

Tests clustering and classification algorithms on tabular datasets containing missing attributes (non-existence attributes). It evaluates how well direct discrepancy-based methods handle missing data under four simulated mechanisms (MCAR, MAR, MNAR-1, MNAR-2) compared to traditional imputation baselines. Use when the user wants to benchmark on Iris, Sonar, Glass, Leaf, Seeds, Libras, Chronic Kidney, Vowel Context, Isolate, Landsat, Breast Tissue, Bank note, or asks about evaluating this task....

researchpythongo
0
3
Missing Indicator EvalA

Evaluates the performance of missing data preprocessing strategies (mean imputation, missForest, Gaussian Copula imputation, with and without Missing Indicator Method) across linear, tree-based, and neural network models. It probes how well these methods handle informative versus uninformative missingness patterns in both low and high-dimensional tabular settings. Use when the user wants to benchmark on Synthetic Low-Dimensional, Synthetic High-Dimensional, OpenML (12 subsets), or asks about ...

researchpythongit
0
3
Mistake Finding EvalA

Evaluates an LLM's ability to detect and locate the first logical error in a multi-step chain-of-thought reasoning trace. It probes whether models can accurately identify specific reasoning steps that contain mistakes or correctly assert that a trace is entirely correct. Use when the user wants to benchmark on BIG-Bench Mistake, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Mit Bih Arrhythmia EvalA

Evaluates a model's ability to classify raw 2-lead ECG signals into five arrhythmia types (normal, supraventricular, ventricular, fusion, unknown) using a multi-class classification setup. It probes temporal feature extraction and attention-based weighting of cardiac cycles without manual preprocessing. Use when the user wants to benchmark on MIT-BIH Arrhythmia Dataset, or asks about evaluating this task. Reports AUC.

researchpythongo
0
3
Mit Bih Ecg Adv Detection EvalA

Evaluates the robustness of ECG arrhythmia classifiers and adversarial detectors on real and synthetically generated adversarial ECG signals. It probes whether models maintain classification accuracy and can distinguish between genuine and adversarial cardiac signals under intra-patient and inter-patient data splits. Use when the user wants to benchmark on PhysioNet MIT-BIH Arrhythmia dataset, or asks about evaluating this task. Reports Accuracy (ACC).

researchpython
0
3
Mit Bih Ecg Class EvalA

Evaluates deep learning architectures for multi-class ECG heartbeat classification, specifically testing their ability to handle severe class imbalance and capture temporal cardiac waveform features. Use when the user wants to benchmark on MIT-BIH Arrhythmia Database, or asks about evaluating this task. Reports F1-score.

researchpythongo
0
3
Mit Bih Ecg Classification EvalA

Evaluates a model's ability to classify individual ECG beats into standard AAMI categories (Normal, Supraventricular Ectopic, Ventricular Ectopic, etc.) on a patient-specific basis. It probes robustness to severe class imbalance and morphological variations in real-time clinical monitoring scenarios. Use when the user wants to benchmark on MIT-BIH arrhythmia database, or asks about evaluating this task. Reports F1-score.

researchpythongo
0
3
Mit Bih Ecg EvalA

Evaluates the classification accuracy and energy efficiency of a hardware-aware spiking neural network (SNN) for real-time ECG beat detection and categorization. Use when the user wants to benchmark on MIT-BIH, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Mit Saliency Benchmark EvalA

Evaluates how well computational saliency models predict human visual attention on natural images. It probes the spatial accuracy and probabilistic alignment of predicted saliency maps against ground-truth eye-tracking fixations. Use when the user wants to benchmark on MIT Saliency Benchmark (MIT300), or asks about evaluating this task. Reports AUC.

researchpythongit
0
3
Mitotic Figure Detection EvalA

Evaluates the ability of object detection models to accurately localize and classify mitotic figures in histopathology images across varying stain domains and scales. It probes model robustness to domain shift and architectural trade-offs between anchor-based and anchor-free detection paradigms. Use when the user wants to benchmark on MIDOG 2025, or asks about evaluating this task. Reports F1 Score.

researchpythongo
0
3
Mixatis Mixsnips EvalA

Evaluates joint intent detection and slot filling on mixed-domain conversational datasets. It probes a model's ability to simultaneously predict multiple intents per utterance and extract corresponding slot entities, measuring both token-level and sentence-level alignment. Use when the user wants to benchmark on MixATIS, MixSNIPS, or asks about evaluating this task. Reports Overall Accuracy.

researchpythongo
0
3
Mixed Bins Bench EvalA

Evaluates a robotic bin-picking framework's ability to estimate 6D object poses, select grasp candidates across parallel jaw and suction grippers, and precisely place objects in cluttered, symmetric, or entangled scenarios. Use when the user wants to benchmark on Mixed bins symmetries, Mixed bins entanglements, or asks about evaluating this task. Reports AP average.

researchpythongo
0
3
Mixeval EvalA

Evaluates LLMs on a dynamically mixed benchmark of real-world web-mined queries and existing datasets to measure alignment with human preferences. It probes a model's general capability, reasoning, and instruction-following across diverse domains, correlating performance with Chatbot Arena Elo scores. Use when the user wants to benchmark on MixEval, MixEval-Hard, or asks about evaluating this task. Reports Spearman's ranking correlation.

ai-agentspythongo
0
3
Mixeval X EvalA

Evaluates multi-modal and agent models across any-to-any generation and action-planning tasks using real-world data mixtures. It probes capabilities in vision-language understanding, audio-language understanding, text-to-media generation, and API-level action planning. Use when the user wants to benchmark on MixEval-X, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Mixqg Qg EvalA

Evaluates a neural question generation model's ability to produce fluent, relevant questions conditioned on a target answer and context. It probes both automatic n-gram/embedding-based similarity metrics and human-rated utility for educational quiz creation. Use when the user wants to benchmark on SQuAD, NQ, Quoref, DROP, or asks about evaluating this task. Reports question approval rate.

researchpythongo
0
3
Mkqa EvalA

Evaluates multilingual open-domain question answering models on their ability to generate or extract short, factoid answers across 26 typologically diverse languages. It specifically probes cross-lingual transfer, handling of unanswerable queries, and robustness to language-specific normalization and threshold tuning for abstention. Use when the user wants to benchmark on MKQA, or asks about evaluating this task. Reports token overlap F1.

researchpythongo
0
3
Ml Accelerator Inference EvalA

Evaluates the inference latency, power consumption, and computational throughput of various machine learning accelerators and CPUs on object detection tasks. It probes the real-world SWaP (Size, Weight, and Power) efficiency and performance discrepancies between advertised and actual hardware capabilities. Use when the user wants to benchmark on Microsoft COCO, or asks about evaluating this task. Reports Avg. Single Image Inference Time (ms).

researchpythongo
0
3
Ml Dev Bench EvalA

This benchmark evaluates AI agents' ability to execute end-to-end machine learning development workflows. It probes capabilities across six categories: dataset handling, model training, debugging, model implementation, API integration, and performance improvement. Success is measured by whether agents can produce fully functional, error-free code and configurations that satisfy the task specifications. Use when the user wants to benchmark on ML-Dev-Bench, or asks about evaluating this task. R...

datapythongo
0
3
Ml Drift Inference EvalA

Measures inference throughput and latency of large generative models across diverse on-device GPU backends. Probes the efficiency of tensor virtualization and runtime shader generation in decoupling logical semantics from physical memory layouts compared to established inference engines. Use when the user wants to benchmark on Stable Diffusion 1.4, Gemma 2B, Gemma2 2B, Llama 3.2 3B, Llama 3.1 8B, or asks about evaluating this task. Reports tokens/s (decode).

researchpythonbackend
0
3
Ml Energy EvalA

Measures and optimizes the inference energy consumption of generative AI models under realistic, production-like serving conditions. It evaluates how different hardware and serving configurations affect the trade-off between latency and energy usage, providing automated recommendations for energy-optimal setups. Use when the user wants to benchmark on ML.ENERGY default request dataset, or asks about evaluating this task. Reports Energy (Joules/request).

researchpythongit
0
3
Ml Superb EvalA

Evaluates multilingual speech processing capabilities across 143 languages on ASR, Language Identification, and joint tasks under normal and few-shot settings. It probes cross-lingual transfer and low-resource adaptation of SSL models. Use when the user wants to benchmark on ML-SUPERB, or asks about evaluating this task. Reports ML-SUPERB score.

researchpythongit
0
3