
Claude Skills by qhjqhj00
github.com/qhjqhj00Evaluates the synthesis speed, intelligibility, and naturalness of a non-autoregressive TTS model trained with conditional flow matching on English speech. It measures how efficiently the model converts text to audio and how closely the output matches human perception of naturalness. Use when the user wants to benchmark on LJ Speech, or asks about evaluating this task. Reports MOS.
Evaluates the matching ability, correspondence sufficiency, and computational efficiency of local feature matchers across short- and wide-baseline image pairs. It measures how well matchers recover camera pose and how many correct correspondences they produce, enabling fair comparison for real-time applications like SLAM. Use when the user wants to benchmark on SfM/SLAM datasets (sequences 01-08), or asks about evaluating this task. Reports AUC (SP curve).
Compute the MatchErrorRate metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MatchErrorRate, or asks how to score with MatchErrorRate.
Evaluates a model's ability to perform few-shot classification by learning to map a small support set of labeled examples to a classifier for unseen classes without fine-tuning. It probes rapid adaptation and generalization across vision and language modalities using an attention-based non-parametric memory mechanism. Use when the user wants to benchmark on Omniglot, ImageNet, miniImageNet, Penn Treebank, or asks about evaluating this task. Reports accuracy.
Evaluates the ability of LLM agents to generate scientifically grounded, constraint-aligned hypotheses for materials discovery under specific application goals. It probes the model's capacity for iterative refinement, consensus-based validation, and adherence to domain-specific feasibility and novelty criteria. Use when the user wants to benchmark on MatDesign, or asks about evaluating this task. Reports Closeness and Quality.
Evaluates the ability of a diffusion-based pipeline to extract physically-based rendering (PBR) materials, specifically spatially varying BRDFs (SVBRDFs), from single real-world images. It measures decomposition accuracy against ground-truth material maps and assesses the perceptual and semantic coherence of the extracted materials compared to known PBR datasets. Use when the user wants to benchmark on AmbientCG, PolyHaven, CGBookcase, OpenSurfaces, TexSD, or asks about evaluating this task. ...
Evaluates multimodal large language models' ability to solve college-level materials science problems that require accurate visual interpretation of scientific figures, such as phase diagrams and stress-strain curves, alongside domain-specific textual reasoning. Use when the user wants to benchmark on MaterialFigBENCH, or asks about evaluating this task. Reports accuracy.
Evaluates large language models' ability to retrieve materials science knowledge and predict continuous physical properties. It probes the fundamental asymmetry in LLM behavior between symbolic tasks (classification, link prediction) and numerical regression tasks, assessing how fine-tuning affects accuracy and output consistency across modalities. Use when the user wants to benchmark on MatKG, Crystal System Classification, Bandgap Prediction, Dielectric Constant Prediction, or asks about ev...
Evaluates the mathematical reasoning and problem-solving capabilities of language models across a spectrum of difficulties, from grade-school arithmetic to advanced competition-level mathematics. Use when the user wants to benchmark on GSM8K, MATH, AMC 2023, AIME 2024, Omni-MATH, or asks about evaluating this task. Reports accuracy.
Evaluates the reliability of reward models for mathematical reasoning by selecting the best solution from a set of sampled candidates using best-of-N search, comparing outcome versus process supervision. Use when the user wants to benchmark on MATH, or asks about evaluating this task. Reports fraction_correct.
Evaluates large language models on mathematical and general reasoning capabilities using a standardized suite of benchmarks. It measures the model's problem-solving accuracy under self-play training conditions, tracking sustained performance gains across multiple evolution iterations. Use when the user wants to benchmark on AMC, Minerva, MATH, GSM8K, Olympiad, AIME25, AIME24, SuperGPQA, MMLU-Pro, BBEH, or asks about evaluating this task. Reports pass@1 accuracy.
This evaluation protocol probes the mathematical reasoning capabilities and solution diversity of large language models. It measures how well models solve grade-school and college-level math problems, and whether reinforcement learning fine-tuning preserves or degrades the variety of generated solution paths. Use when the user wants to benchmark on GSM8K, MATH500, Olympiad Bench, College Math, or asks about evaluating this task. Reports Pass@1 accuracy.
Evaluates the mathematical reasoning capabilities of language models across multiple challenging benchmarks. It measures whether models can correctly solve math problems and follow structured reasoning processes aligned with a teacher model's trace. Use when the user wants to benchmark on MATH-500, MINERVA, OlympiadBench, LiveMathBench, KSAT2025, AIME 2024, AIME 2025, or asks about evaluating this task. Reports Pass@1.
Evaluates multimodal and language-only models on mathematical reasoning, self-judgment/reward accuracy, and general multimodal capabilities. It measures how well a model can solve complex problems, verify its own answers, and generalize across diverse domains without external reward models. Use when the user wants to benchmark on MathVista, GSM8k, RewardBench2, VL-RewardBench, MMBench, MMStar, or asks about evaluating this task. Reports accuracy.
Evaluates multimodal mathematical reasoning capabilities of LLMs and LMMs on problems with visual contexts. Probes geometric invariance, spatial reasoning, and deep mathematical reasoning across 16 disciplines and 5 difficulty levels. Use when the user wants to benchmark on MATH-V, or asks about evaluating this task. Reports accuracy.
Evaluates multimodal mathematical error detection by identifying incorrect steps in student solutions and categorizing the type of error. It probes the model's ability to align visual problem elements with textual reasoning and solution paths. Use when the user wants to benchmark on MathAgent Evaluation Dataset, or asks about evaluating this task. Reports accuracy.
Evaluates large language models' ability to solve mathematical word problems across varying difficulty levels and subjects, including elementary, high school, and collegiate mathematics. It specifically probes the model's capacity for code-interleaved reasoning and execution-aware autoregression. Use when the user wants to benchmark on GSM8K, MATH, SVAMP, Mathematics, SimulEq, or asks about evaluating this task. Reports accuracy.
Evaluates multimodal mathematical reasoning capabilities across diverse benchmarks, focusing on geometry problem solving, multi-step reasoning, and cross-modal alignment between visual diagrams and mathematical concepts. Use when the user wants to benchmark on MATH-Vision, MathVista, MathVerse, GAOKAO-MM, We-Math, or asks about evaluating this task. Reports accuracy.
Evaluates how retrieval quality impacts downstream mathematical problem solving. It compares zero-shot performance against retrieval-augmented settings using either embedding-retrieved or expert-paired problems with their solutions. Use when the user wants to benchmark on MathNet-RAG, or asks about evaluating this task. Reports Retrieval-Augmented Problem Solving Accuracy.
Probes a model's ability to retrieve mathematically equivalent problems from a large corpus using embeddings. It measures whether retrieval systems can recognize structural and symbolic invariance rather than relying on superficial lexical overlap. Use when the user wants to benchmark on MathNet-Retrieve, or asks about evaluating this task. Reports Recall@k.
Evaluates a model's ability to solve Olympiad-level mathematical problems across multiple domains (algebra, geometry, combinatorics, number theory) and modalities (text and images). It measures whether models can produce consistent, correct reasoning rather than just guessing the final answer. Use when the user wants to benchmark on MathNet-Solve, or asks about evaluating this task. Reports Problem Solving Accuracy.
Evaluates large language models' mathematical reasoning capabilities across Olympiad, high school, and university-level problems. It probes multi-step logic, chain-of-thought reasoning, and advanced derivations in algebra, calculus, and number theory. Use when the user wants to benchmark on MathOdyssey, or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to translate complex mathematical word problems into executable Python code that computes a specific numerical answer. It probes semantic grounding of natural language, arithmetic reasoning, and straight-line code generation. Use when the user wants to benchmark on MathQA-Python, or asks about evaluating this task. Reports accuracy.
Evaluates the ability of multimodal large language models to perform K-12 mathematical reasoning on real-world, mobile-captured images. It probes robustness to visual degradation (blur, rotation, handwritten annotations) and perspective variations, measuring how well models extract text and figures to solve math problems under imperfect conditions. Use when the user wants to benchmark on MathReal, or asks about evaluating this task. Reports Loose Accuracy (Acc).
This benchmark evaluates the visual mathematical reasoning capabilities of multi-modal large language models (MLLMs), specifically probing whether they genuinely interpret geometric diagrams or merely rely on textual redundancy. It measures performance across different problem formulations (varying text/image ratios) and mathematical subjects like plane geometry, solid geometry, and functions. Use when the user wants to benchmark on MATHVERSE, or asks about evaluating this task. Reports accur...
This benchmark evaluates the mathematical reasoning capabilities of foundation models (LLMs and LMMs) when processing visual contexts. It probes abilities such as figure interpretation, algebraic and geometric reasoning, and the integration of multimodal inputs like images, OCR text, and captions into mathematical problem-solving. Use when the user wants to benchmark on MATHVISTA, or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to recognize handwritten mathematical expressions and convert them into normalized LaTeX. It probes both offline (rasterized image) and online (ink coordinate) recognition capabilities, measuring how accurately the model reconstructs complex mathematical structures and symbols. Use when the user wants to benchmark on MathWriting, or asks about evaluating this task. Reports CER.
Evaluates the accuracy of a charge-equilibrated equivariant foundation potential in predicting atomic energies, forces, partial charges, and bulk mechanical/thermal properties across diverse crystalline, molecular, and ionic systems. Use when the user wants to benchmark on MatQ, Custom Model Systems (C10H2/C10H3+, Ag3+/−, Na8/9Cl8+, Au2-MgO(001)), or asks about evaluating this task. Reports RMSE, MAE.
Evaluates scientific language models on seven materials science NLP tasks, including named entity recognition, relation classification, event argument extraction, paragraph classification, synthesis action retrieval, sentence classification, and slot filling. It probes the model's ability to extract structured information and classify text from domain-specific scientific literature. Use when the user wants to benchmark on MatSci-NLP, or asks about evaluating this task. Reports micro-F1.
This benchmark evaluates the reasoning capabilities of large language models in materials science, covering six primary fields and 31 sub-fields. It probes domain knowledge, mathematical/formula reasoning, and multimodal visual comprehension through expert-curated problems with three-tier difficulty classifications. Use when the user wants to benchmark on MatSciBench, or asks about evaluating this task. Reports Accuracy Score (%).
Evaluates graph neural networks and equivariant point cloud networks on solid-state materials modeling tasks, including energy/force prediction, bandgap/fermi level regression, and crystal symmetry classification. Probes single-task, multi-task, and multi-dataset generalization capabilities. Use when the user wants to benchmark on OpenCatalyst (OC-20), Materials Project (MP), LiPS, OQMD, NOMAD, CMD, or asks about evaluating this task. Reports MSE.
This benchmark evaluates fine-grained mistake understanding in egocentric videos by attributing errors to specific semantic roles, temporal points of no return, and spatial locations. It probes a model's ability to align video content with instructional text, localize manipulation events, and classify whether actions deviate from intended goals. Use when the user wants to benchmark on Ego4D-M, EPIC-KITCHENS-M, EgoPER, or asks about evaluating this task. Reports F1@0.5, MAE, mIoU.
Evaluates large language models' ability to reason about physicochemical principles in nanomaterial synthesis. It probes whether models can generate scientifically valid hypotheses and understand conceptual mechanisms from literature abstracts, rather than relying on abstract logic alone. Use when the user wants to benchmark on MatterMech, or asks about evaluating this task. Reports accuracy.
Compute the matthews_corrcoef metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute matthews_corrcoef, or asks how to score with matthews_corrcoef.
Compute the MatthewsCorrCoef metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MatthewsCorrCoef, or asks how to score with MatthewsCorrCoef.
Evaluates multimodal large language models' ability to perform fine-grained visual-scientific reasoning in materials science. It probes structure-property-performance relationships through quantitative, comparative, causal, and hypothetical variation tasks, requiring models to integrate visual data from experimental figures with domain-specific knowledge rather than relying on textual shortcuts. Use when the user wants to benchmark on MatVQA, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates a model's ability to perform legal reading comprehension on merger agreements by answering specialized deal point questions. It probes the model's capacity to interpret complex contractual clauses and handle imbalanced classification tasks across various legal categories. Use when the user wants to benchmark on MAUD, or asks about evaluating this task. Reports AUPR.
Evaluates multimodal large language models on video and image understanding tasks, including general knowledge QA, long-video QA, event understanding, temporal reasoning, and captioning, as well as image QA, cognitive understanding, and captioning. Use when the user wants to benchmark on MMWorld, PerceptionTest, Video-MME, MLVU, MVBench, EventHallusion, TempCompass, VinoGround, DREAM-1K, MMMU, MathVista, AI2D, CapsBench, or asks about evaluating this task. Reports score.
Evaluates deepfake detection models on distinguishing real from fake audio-video content under closed-set and open-set conditions. It specifically probes cross-model and cross-lingual generalization by testing on unseen generation methods and languages. Use when the user wants to benchmark on MAVOS-DD, or asks about evaluating this task. Reports mAP.
Evaluates object detection performance on aerial and ground-view imagery, probing how geographic context and multi-view data fusion affect detection accuracy across different object scales. Use when the user wants to benchmark on MAVREC, or asks about evaluating this task. Reports mAP.
Compute the max_error metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute max_error, or asks how to score with max_error.
Evaluates camera pose estimation accuracy in long-term visual localization by measuring the maximum pixel displacement of projected 3D points between a reference and an estimated pose. This indirect measure avoids the non-trivial task of quantifying 6-DoF pose uncertainties while remaining sensitive to camera-to-scene distance variations. Use when the user has predictions and gold and needs to compute Maximum reprojection difference.
This benchmark evaluates multilingual visual question answering (mVQA) by testing models on images paired with questions in seven languages. It probes a model's ability to perform cross-lingual visual reasoning and generate accurate text answers without relying on costly human annotation. Use when the user wants to benchmark on MaXM, or asks about evaluating this task. Reports Exact Match Accuracy.
Compute the MaxMetric metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MaxMetric, or asks how to score with MaxMetric.
Compute maysonma/lingo_judge_metric via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of maysonma/lingo_judge_metric.
Evaluates a model's ability to generate entity-aware, fine-grained text descriptions of basketball videos, specifically requiring accurate prediction of player names and precise action recognition. Use when the user wants to benchmark on MbgVC, or asks about evaluating this task. Reports GDS.
This benchmark evaluates a model's ability to identify various forms of media bias, including linguistic, cognitive, political, racial, gender, and hate speech bias, across diverse text sources like news articles, tweets, and social media comments. It probes whether models can generalize across different bias types and dataset sizes without being skewed by larger datasets. Use when the user wants to benchmark on MBIB, or asks about evaluating this task. Reports macro F1-score.
This evaluation probes a model's ability to detect various forms of media bias in text, including racial, gender, hate speech, fake news, cognitive, and text-level context bias. It tests whether the model can identify subtle, context-dependent linguistic cues and stereotypes across diverse news and social media samples. Use when the user wants to benchmark on Media Bias Identification Benchmark (MBIB), or asks about evaluating this task. Reports accuracy.
Evaluates large language models' ability to detect political bias in media text using in-context learning and chain-of-thought prompting. It probes the model's capacity to distinguish biased from unbiased content across diverse textual chunks without fine-tuning. Use when the user wants to benchmark on Media Bias Identification Benchmark (MBIB), or asks about evaluating this task. Reports Macro-F1.
Evaluates a model's ability to generate correct, self-contained Python functions from natural language problem descriptions. It probes basic programming logic, standard library usage, and semantic grounding of simple algorithmic tasks. Use when the user wants to benchmark on Mostly Basic Programming Problems (MBPP), or asks about evaluating this task. Reports accuracy.