All authors
qhjqhj00 avatar

Claude Skills by qhjqhj00

github.com/qhjqhj00
7,636 skillsA× 7,623B× 11C× 1D× 10 installs2,230 views
Mammography Cancer Detection EvalA

Evaluates a multi-modal AI system's ability to detect breast cancer and localize malignant lesions using 2D (FFDM, C-View) and 3D (DBT) mammography images. It measures classification performance at the breast and image levels, as well as lesion localization accuracy via bounding boxes across internal and external clinical datasets. Use when the user wants to benchmark on NYU Comprehensive Mammography Dataset (V1), NYU Comprehensive Mammography Dataset (V2), OPTIMAM, CMMD, CSAW-CC, EMBED, CBIS...

researchpythongit
0
3
Mammography Domain Generalisation EvalA

Evaluates the cross-domain generalisation capability of deep learning models for breast cancer screening using multi-view mammography images. It tests robustness to out-of-distribution data from different vendors, imaging protocols, and centers by training on seen domains and testing on unseen domains. Use when the user wants to benchmark on CBIS, CMMD, INBreast, TOMMY1, TOMMY2, or asks about evaluating this task. Reports AUC.

researchpythontesting
0
3
Mammography Linear EvalA

Evaluates self-supervised learning models for breast cancer detection on screening mammography using a linear evaluation protocol on whole images derived from tiled patches. The protocol extracts fixed encoder features from image patches, pools them using attention-based or average pooling, and trains a linear classifier for final prediction. Use when the user wants to benchmark on Screening mammography dataset, or asks about evaluating this task. Reports linear evaluation.

researchpythonperformance
0
3
Mammography Mass Detection EvalA

Evaluates the ability of deep learning models to detect breast masses in digital mammography images across multiple clinical domains with varying scanner manufacturers and imaging protocols. It specifically probes domain generalization capabilities by measuring detection robustness on unseen data distributions. Use when the user wants to benchmark on OPTIMAM Hologic, OPTIMAM Siemens, OPTIMAM GE, OPTIMAM Philips, INbreast, BCDR, or asks about evaluating this task. Reports TPR at 0.75 FPPI.

researchpythongo
0
3
Mammography Report EvalA

Evaluates the ability of local vision-language models to generate clinically styled mammography reports and perform multi-task classification (e.g., BI-RADS, breast density, calcifications) from medical images. It probes the models' robustness under zero-shot, few-shot, Chain-of-Thought prompting, and Retrieval-Augmented Generation (RAG), as well as the impact of parameter-efficient fine-tuning (QLoRA). Use when the user wants to benchmark on VinDr-Mammo, DMID, or asks about evaluating this t...

ai-agentspythongo
0
3
Mammography Roi Classification EvalA

Evaluates a hybrid CNN-SSM architecture's ability to classify mammography regions of interest (ROIs) as benign or malignant. It probes the model's capacity for local feature extraction, global context modeling, and robust performance under class imbalance in medical imaging. Use when the user wants to benchmark on CBIS-DDSM, or asks about evaluating this task. Reports AUC-ROC.

researchpythonperformance
0
3
Mammoth Vl Multimodal EvalA

Evaluates multimodal reasoning and instruction-following capabilities across single-image, multi-image, and video scenarios. Probes OCR, chart/document understanding, mathematical reasoning, and real-world visual interactions. Use when the user wants to benchmark on AI2D, ChartQA, DocVQA, InfoVQA, MMStar, MMMU, MMMU-Pro, SeedBench, MMBench, MMvet, Mathverse, Mathvista, RealworldQA, WildVision, Llava-Wilder-Small, MuirBench, MEGABench, EgoSchema, PerceptionTest, SeedBench (Video), MLVU, MVBenc...

researchpythongo
0
3
Mamut Mir EvalA

Evaluates mathematical information retrieval capabilities by testing whether models can match natural language names or LaTeX formulas to their corresponding mathematical identities from a candidate pool. It probes the model's ability to learn structural and notational variations in mathematical expressions through pretraining and fine-tuning. Use when the user wants to benchmark on MAMUT-generated datasets (MF, MT, NMF, MFR), or asks about evaluating this task. Reports nDCG.

researchpythongo
0
3
Manifold Robustness EvalA

Evaluates the robustness of a dimensionality reduction pipeline (Isomap + Procrustes alignment + TDA clustering) against ambient noise, outliers, and hyperparameter variation. It tests whether the method can consistently recover a low-distortion 2D embedding of a contractible manifold or correctly detect topological failure on non-contractible data. Use when the user wants to benchmark on Swiss roll, Buckyball, or asks about evaluating this task. Reports Persistent homology features ($PH_1$, ...

researchpythonrust
0
3
Manip EvalA

Evaluates physically-grounded image editing by measuring 2D spatial accuracy, depth prediction, 3D geometric consistency, image quality, and VLM-based physical plausibility for object manipulation tasks. Use when the user wants to benchmark on ManipEval, or asks about evaluating this task. Reports Chamfer.

researchpythongo
0
3
Manipulation Transfer EvalA

Evaluates whether learned hierarchical motor skills can transfer across different object geometries, downstream stacking tasks, and observation modalities (state vs. vision). It probes sample efficiency, directed exploration, and performance under varying reward sparsities (dense, staged sparse, fully sparse). Use when the user wants to benchmark on red_on_blue_stacking, all_pairs_stacking, or asks about evaluating this task. Reports reward.

researchpythonperformance
0
3
Manipulationnet EvalA

Evaluates real-world robot manipulation capabilities across two complementary tracks: physical skills (sensorimotor execution under contact, clearance, and perceptual constraints) and embodied reasoning (multimodal grounding of natural language and visual instructions into grounded actions). Use when the user wants to benchmark on ManipulationNet Benchmark, or asks about evaluating this task. Reports task success rate.

researchpythongo
0
3
Maniskill Hab EvalA

Evaluates low-level robotic manipulation policies for long-horizon home rearrangement tasks. It probes a robot's ability to successfully pick, place, and interact with household objects across cluttered and constrained environments. Use when the user wants to benchmark on ManiSkill-HAB, or asks about evaluating this task. Reports success once rate.

researchpythongo
0
3
Maniskill2 EvalA

Evaluates the generalization and robustness of embodied AI manipulation policies across soft-body, rigid-body, and assembly tasks in a simulated environment. Use when the user wants to benchmark on ManiSkill2, or asks about evaluating this task. Reports success rate.

researchpythonperformance
0
3
Maniskill2 Softbody EvalA

Evaluates a robot policy's ability to perform long-horizon manipulation tasks involving deformable soft bodies (e.g., clay, noodles, liquid, plasticine). It probes spatial reasoning, contact dynamics, and precise end-effector control under varying initial conditions. Use when the user wants to benchmark on ManiSkill2 Challenge (Soft-body Track), or asks about evaluating this task. Reports Success Metric.

researchpython
0
3
Manitwin Asset Quality EvalA

Evaluates the semantic alignment, geometric fidelity, and visual appearance of automatically generated 3D assets against input images or text. It also assesses the accuracy of VLM-generated annotations across five dimensions and the physical validity of simulated grasp poses for robotic manipulation readiness. Use when the user wants to benchmark on ManiTwin-100K, or asks about evaluating this task. Reports CLIP(I-I/T).

researchpythongo
0
3
MannwhitneyuA

Compute the mannwhitneyu metric — provided by scipy.stats. Use when the user has predictions and ground-truth and needs to compute mannwhitneyu, or asks how to score with mannwhitneyu.

documentationpythonexpress
0
3
Manueldeprada BeerA

Compute manueldeprada/beer via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of manueldeprada/beer.

developmentpython
0
3
Map Based Fdi EvalA

This evaluation probes the operational utility of data-driven Fire Danger Index (FDI) models for wildfire forecasting. It assesses both point-level classification accuracy and full-map spatial inference performance, explicitly quantifying detection rates and false positive distributions under realistic deployment conditions. Use when the user wants to benchmark on FireCube, or asks about evaluating this task. Reports Map-based Recall Percentiles.

researchpythongo
0
3
Mapdr EvalA

Probes autonomous driving models' ability to extract lane-level traffic regulations from visual inputs and map them to vectorized HD map centerlines. It evaluates both rule extraction from image sequences and bipartite graph construction for rule-lane correspondence reasoning. Use when the user wants to benchmark on MapDR, or asks about evaluating this task. Reports correspondence status.

researchpythongo
0
3
Mapeval EvalA

This benchmark evaluates foundation models' geospatial reasoning capabilities across textual, visual, and API-based interaction modalities. It probes abilities such as place information retrieval, nearby point-of-interest identification, route planning, multi-step trip scheduling, and recognizing unanswerable queries. Use when the user wants to benchmark on MapEval, or asks about evaluating this task. Reports accuracy.

ai-agentspythongo
0
3
Maple EvalA

This evaluation probes the fidelity and predictive accuracy of local explanations generated by MAPLE. It measures how well a local linear model approximates the target model's predictions in the neighborhood of a test point, while also benchmarking overall regression accuracy against standard baselines. Use when the user wants to benchmark on UCI datasets, or asks about evaluating this task. Reports causal metric.

researchpythontesting
0
3
Mapless Navigation EvalA

Evaluates a robot's ability to navigate to a goal in unknown environments using only local sensor data without a pre-built map. It probes the policy's generalization across varying obstacle densities, passage widths, and real-world conditions, as well as its energy efficiency on neuromorphic hardware. Use when the user wants to benchmark on Gazebo Training Environments, Gazebo Test Environment, Real-world Office Environment, or asks about evaluating this task. Reports success rate.

researchpythongo
0
3
Maps Multilingual Agent EvalA

Evaluates the performance and security robustness of agentic AI systems when operating in multilingual settings. It measures how task completion accuracy and vulnerability to adversarial prompts degrade or shift when instructions are translated from English into 11 typologically diverse languages. Use when the user wants to benchmark on GAIA, SWE-bench, MATH, ASB, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Maps X EvalA

Evaluates the performance of explainable multi-robot motion planning algorithms (MAPS-X and Lazy MAPS-X) combined with sampling-based planners like RRT* across custom environments. It measures the trade-off between planning efficiency (runtime, success rate) and explainability (number of trajectory segments), highlighting how segmentation constraints impact computational cost and plan optimality. Use when the user wants to benchmark on MAPS-X custom environments, or asks about evaluating this...

researchpythongo
0
3
Maptrace EvalA

Evaluates fine-grained spatial reasoning and pixel-accurate route tracing on commercial map images. Models must generate precise path coordinates or masks corresponding to text-based navigation queries. Use when the user wants to benchmark on MapTrace, Map-Bench, or asks about evaluating this task. Reports NDTW.

researchpythongo
0
3
Maqiuping59 Table MarkdownA

Compute maqiuping59/table_markdown via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of maqiuping59/table_markdown.

documentationpython
0
3
Marble EvalA

Evaluates pre-trained music audio representation models across a unified taxonomy of 18 downstream tasks spanning acoustic, performance, score, and high-level description levels. It assesses model generalization and representation quality under constrained training settings, including sequence labeling tasks like beat tracking and source separation. Use when the user wants to benchmark on MelodyDB, Muljam, Jamendo, GuitarSet, MUSDB18, NSynth, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Marca EvalA

Evaluates LLMs' ability to perform multilingual web search and extract multiple entities from search results. It probes task decomposition, cross-lingual retrieval, and evidence aggregation under different agentic interaction frameworks. Use when the user wants to benchmark on MARCA, or asks about evaluating this task. Reports Checklist Accuracy.

ai-agentspythongo
0
3
Marcel EvalA

Evaluates molecular property prediction using explicit conformer ensembles versus single-conformer or 1D/2D baselines, probing how 3D structural flexibility and ensemble encoding strategies impact regression accuracy. Use when the user wants to benchmark on MARCEL, or asks about evaluating this task. Reports Mean Absolute Error (MAE).

researchpythongo
0
3
Marco Voice EvalA

Evaluates a unified neural TTS framework's ability to disentangle speaker identity and emotional style, measuring speaker fidelity, emotional expressiveness, and overall speech quality in both English and Mandarin. Use when the user wants to benchmark on LibriTTS, AISHELL-3, CSEMOTIONS, or asks about evaluating this task. Reports Emotional expressiveness.

researchpythonshell
0
3
Marine Hallucination EvalA

Evaluates the ability of Large Vision-Language Models (LVLMs) to mitigate object hallucinations during text generation. It probes visual-text alignment by measuring hallucination rates, recall of existing objects, and accuracy on binary probing questions, alongside GPT-4V-aided assessments of response accuracy and detailness. Use when the user wants to benchmark on MSCOCO val2014, or asks about evaluating this task. Reports CHAIRS.

researchpythongo
0
3
Marineeval EvalA

MarineEval probes the marine domain expertise and visual understanding capabilities of vision-language models. It evaluates tasks including species identification, spatial reasoning, ecological knowledge integration, and precise object localization under real-world marine conditions. Use when the user wants to benchmark on MarineEval, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Marioqa EvalA

Evaluates a model's ability to perform video question answering with varying levels of temporal reasoning complexity. It probes whether models can correctly link visual events in gameplay videos to answer questions that require single-frame, event-level, or multi-step causal/temporal understanding. Use when the user wants to benchmark on MarioQA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Marketgen EvalA

Evaluates embodied agents and multimodal LLMs on long-horizon manipulation tasks in procedurally generated supermarket environments. Specifically, it probes spatial reasoning, occlusion handling, and collision avoidance during checkout unloading and in-aisle item collection. Use when the user wants to benchmark on MarketGen Benchmark, or asks about evaluating this task. Reports Success Rate (SR).

researchpythongo
0
3
Mart Safety EvalA

Evaluates an LLM's ability to refuse harmful or unsafe requests while maintaining helpfulness on benign prompts. It probes safety alignment through automatic reward-model scoring and human flagging of violations across in-distribution and out-of-domain benchmarks. Use when the user wants to benchmark on SafeEval, HelpEval, AlpacaEval, Anthropic Harmless, or asks about evaluating this task. Reports violation_rate.

researchpythongo
0
3
Marvel EvalA

Evaluates multimodal large language models on multidimensional abstract visual reasoning and perceptual grounding. It probes the model's ability to recognize complex geometric and abstract patterns, track temporal/spatial changes, and perform multi-step visual reasoning across diverse puzzle configurations. Use when the user wants to benchmark on MARVEL, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Maryxm Code EvalA

Compute maryxm/code_eval via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of maryxm/code_eval.

developmentpython
0
3
Mas Bench EvalA

Evaluates the ability of mobile GUI agents to complete complex, real-world automation tasks across single-app and cross-app scenarios. It specifically probes how well agents can integrate predefined or self-generated shortcuts (APIs, deep links, RPA scripts) with standard GUI interactions to improve task success, execution efficiency, and cost-effectiveness. Use when the user wants to benchmark on MAS-Bench, or asks about evaluating this task. Reports SR.

toolspythongo
0
3
Masakhaner EvalA

Evaluates named entity recognition (NER) capabilities across ten African languages, probing models' ability to identify PER, ORG, and LOC entities in low-resource, morphologically complex, and culturally specific news text. Use when the user wants to benchmark on MasakhaNER, or asks about evaluating this task. Reports F1 score.

researchpythongo
0
3
Masakhaner20 EvalA

Evaluates named entity recognition (NER) capabilities across 20 typologically and geographically diverse African languages. It probes zero-shot cross-lingual transfer performance and measures how well models generalize to unseen entities and languages when fine-tuned on limited African language data. Use when the user wants to benchmark on MasakhaNER 2.0, or asks about evaluating this task. Reports F1.

researchpythongo
0
3
Masakhanews EvalA

This benchmark evaluates the ability of language models and classical ML algorithms to classify news articles into predefined topics across 16 typologically diverse African languages. It probes multilingual representation quality, script handling, and few-shot/fine-tuning performance in low-resource settings. Use when the user wants to benchmark on MasakhaNEWS, or asks about evaluating this task. Reports weighted F1-score.

businesspythongo
0
3
Mascqa EvalA

Evaluates large language models' domain-specific reasoning and numerical problem-solving capabilities in materials science and metallurgical engineering. It probes their ability to accurately answer multiple-choice, matching, and numerical questions, highlighting gaps in scientific reasoning and computational precision. Use when the user wants to benchmark on MaScQA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
MaskevalA

Evaluates a reference-less, masked language model-based metric's ability to predict human judgments on text summarization and simplification quality. It probes the model's capacity to capture multiple quality dimensions such as fluency, consistency, coherence, relevance, simplicity, and meaning preservation without relying on reference texts. Use when the user has predictions and gold and needs to compute pearson_correlation.

researchpythongo
0
3
Masksql EvalA

Evaluates the privacy-preserving text-to-SQL generation capability of LLMs. It measures execution accuracy against ground-truth SQL while quantifying privacy protection through token abstraction recall and adversarial re-identification resistance. Use when the user wants to benchmark on BIRD, or asks about evaluating this task. Reports Execution Accuracy.

ai-agentspythongo
0
3
Massive EvalA

Evaluates multilingual natural language understanding capabilities, specifically intent classification and slot filling, across 51 typologically diverse languages. It measures model robustness to different scripts, spacing conventions, and zero-shot cross-lingual transfer scenarios. Use when the user wants to benchmark on MASSIVE, or asks about evaluating this task. Reports exact match accuracy.

researchpythongo
0
3
Massive Mimo Scheduler EvalA

Evaluates the ability of a deep reinforcement learning scheduler to allocate wireless resources (users to resource blocks) in massive MIMO networks. It probes the model's capacity to maximize spectral efficiency and user fairness under varying channel conditions (static vs. mobile) and network scales. Use when the user wants to benchmark on QuaDRiGa 3GPP_3D_UMi_LOS, or asks about evaluating this task. Reports normalized spectral efficiency.

researchpythontesting
0
3
Master Set EvalA

Evaluates a model's ability to recommend functionally indispensable (must-cite) papers for a given query paper based only on its title and abstract. It probes scientific retrieval capability by measuring how well systems rank baseline, core-relevant, or frequently mentioned papers from a large candidate pool. Use when the user wants to benchmark on MasterSet-CoreML-v1, or asks about evaluating this task. Reports Recall@K.

researchpythongo
0
3
Matbench EvalA

Evaluates machine learning models on predicting materials properties (e.g., elastic moduli, band gaps, formation energies) from crystal structures or compositions. It probes generalization across diverse data sizes, input types, and property domains using a standardized, pre-cleaned suite of 13 supervised tasks. Use when the user wants to benchmark on Matbench test suite v0.1, or asks about evaluating this task. Reports error estimation (RMSE/Accuracy).

researchpythongo
0
3
Match Compiler EvalA

Evaluates a model-aware compiler framework for deploying deep neural networks on heterogeneous edge microcontrollers. It measures execution latency, hardware utilization efficiency (MACs/cycle), and scheduling robustness under memory constraints across multiple standard DNN architectures. Use when the user wants to benchmark on MLPerf Tiny Benchmark Suite, or asks about evaluating this task. Reports Latency (ms).

developmentpythonbackend
0
3