
Claude Skills by qhjqhj00
github.com/qhjqhj00Evaluates whether integrating external knowledge (textual descriptions or embeddings) improves performance on general natural language understanding tasks compared to baseline pre-trained language models. It probes the model's ability to leverage external semantic information to enhance representation learning and decision-making across classification, regression, and sequence labeling benchmarks. Use when the user wants to benchmark on GLUE, Penn Treebank, CoNLL-2003, or asks about evaluatin...
Evaluates cross-domain anomaly detection capability across semantic, near-distribution, and industrial benchmarks. Probes the model's ability to detect pixel-level defects and semantic novelties using a self-supervised transformer discriminator that attends to distorted features without task-specific tuning. Use when the user wants to benchmark on CIFAR-10, CIFAR-100, Fashion-MNIST, View, Aircraft-FGVC, Stanford Cars, MVTec-AD, MVTec-LOCO, VisA, MPDD, or asks about evaluating this task. Repor...
Evaluates a generative ML model's ability to correct detector effects (unfolding) for highly boosted hadronic top-quark decays. It probes the model's capacity to reconstruct high-dimensional kinematic phase space while mitigating simulation-induced model bias and accurately extracting the top_mass_measurement. Use when the user wants to benchmark on CMS benchmark top-pair simulation, or asks about evaluating this task. Reports top_mass_measurement.
Evaluates text-to-image alignment by testing whether models correctly render specified objects, attributes, and spatial relations across six categories: single object, two objects, counting, colors, position, and attribute binding. Use when the user wants to benchmark on GenEval, or asks about evaluating this task. Reports Overall.
Evaluates text-to-image generation by measuring how accurately models follow complex prompts with multiple objects, attributes, and spatial constraints. Use when the user wants to benchmark on GenEval, or asks about evaluating this task. Reports GenEval Overall.
Evaluates a model's ability to generate images that accurately reflect complex, multidisciplinary textual prompts. It probes semantic correctness, visual plausibility (spelling, logical consistency, readability), and the integration of domain knowledge with reasoning during image generation. Use when the user wants to benchmark on GenExam, or asks about evaluating this task. Reports strict score.
This benchmark evaluates vision-language models' ability to synthesize scientifically faithful, visually coherent 'Figure 1' summaries from academic paper text. It probes deep cross-modal reasoning, concept selection, spatial layout planning, and adherence to scientific content without distortion. Use when the user wants to benchmark on GenFig1, or asks about evaluating this task. Reports VLM-as-a-Judge.
This benchmark probes a model's ability to distinguish real from AI-generated images across multiple diffusion and GAN generators, and to identify specific synthetic flaws in generated images. It evaluates standard detection accuracy, cross-generator generalization, and robustness to common image perturbations like blur, rotation, and brightness shifts. Use when the user wants to benchmark on GenImage, GENHARD, GENEXPLAIN, or asks about evaluating this task. Reports Accuracy.
Evaluates how well automated metrics and VLMs align with human judgments on spatially-aware image generation tasks across nine sub-domains. Use when the user wants to benchmark on GenSpace Human Alignment Test Set, or asks about evaluating this task. Reports agreement.
Evaluates a neural semantic parser's ability to map natural language utterances directly to executable SQL queries. It probes compositional generalization and schema grounding by measuring whether predicted queries return the exact same results as gold queries on a target database. Use when the user wants to benchmark on GEO880, ATIS, or asks about evaluating this task. Reports denotation accuracy.
Evaluates the transferability and semantic grounding of remote sensing foundation models on image-level classification and pixel-level semantic segmentation tasks across multiple geospatial benchmarks. Use when the user wants to benchmark on GeoBench, or asks about evaluating this task. Reports F1.
Evaluates height-aware multimodal reasoning in remote sensing, probing models on pixel-level elevation retrieval, object-level relative height ranking, scene-level terrain relief analysis, and fine-grained height-aware mask generation. It specifically tests the model's ability to integrate vertical spatial priors with optical imagery for accurate numerical estimation and spatial segmentation. Use when the user wants to benchmark on GeoHeight-Bench, GeoHeight-Bench+, or asks about evaluating t...
Evaluates a model's ability to predict the geographic coordinates (latitude and longitude) of a query image from worldwide visual data. It probes fine-grained location-aware visual semantics and robustness to geographical heterogeneity across urban, regional, and continental scales. Use when the user wants to benchmark on IM2GPS3k, YFCC4K, or asks about evaluating this task. Reports threshold metric.
Evaluates a CNN's ability to perform semantic segmentation on airborne magnetic data to identify three major lithological groups (dykes, plutons, greywackes). It tests transfer learning from synthetic geostatistical data to real-world geological contexts. Use when the user wants to benchmark on Malartic geological model & synthetic augmentations, or asks about evaluating this task. Reports IOU.
Evaluates the geometric fidelity and surface reconstruction accuracy of neural 3D scene representations (NeRF and Gaussian Splatting variants) against metric-scale laser scan ground truth. Use when the user wants to benchmark on Robotic Manipulation Scenes, or asks about evaluating this task. Reports CD_{P\rightarrow G}.
Evaluates a model's ability to perform geometric matrix completion on multi-network recommendation datasets. It probes how well graph neural networks and low-rank representations can integrate cross-network and within-network features to predict missing user-item ratings. Use when the user wants to benchmark on Douban, Flixster, YahooMusic, ML-100K, ML-1M, or asks about evaluating this task. Reports RMSE.
This benchmark evaluates the geometric reasoning and problem-solving capabilities of multimodal large language models. It tests whether models can accurately interpret geometric diagrams and accompanying text to produce correct final answers or select the right multiple-choice option. Use when the user wants to benchmark on GeoQA, Geometry3K, PGPS9K, MathVista-mini-GPS, or asks about evaluating this task. Reports Top-1 accuracy.
Evaluates monocular depth estimation models for their ability to produce geometry-preserving depth maps and accurate 3D point clouds without requiring explicit 3D annotations. It probes scale-and-shift recovery, generalization across indoor and outdoor domains, and consistency under differentiable rendering. Use when the user wants to benchmark on NYU V2, ScanNet, KITTI, ETH3D, 2D3D, or asks about evaluating this task. Reports AbsRel.
Evaluates expert-level multimodal intelligence in geoscience and remote sensing by testing domain knowledge, perceptual grounding, and spatiotemporal reasoning across diverse sensors, disciplines, and task complexities. Use when the user wants to benchmark on GeoMMBench, or asks about evaluating this task. Reports Micro-averaged accuracy.
This benchmark evaluates a model's ability to perceive and parse geometric diagrams into a structured formal language. It probes fine-grained visual primitive detection (points, lines, circles, planes) and spatial/semantic relations, testing both syntactic correctness and holistic geometric consistency. Use when the user wants to benchmark on GDP-29K, or asks about evaluating this task. Reports F1-score.
Evaluates a model's ability to perform multimodal numerical reasoning on geometric problems by generating executable symbolic programs from text and diagram inputs. The model must fuse cross-modal information to predict step-by-step reasoning programs. These programs are then executed to select the correct multiple-choice answer from the given options. Use when the user wants to benchmark on GeoQA, or asks about evaluating this task. Reports answer accuracy.
Evaluates the quality of object removal and causal visual artifact removal (shadows, reflections) in images, measuring visual fidelity, structural consistency, and artifact suppression. Use when the user wants to benchmark on RORD-Val, RemovalBench, CausRem, or asks about evaluating this task. Reports FID.
Predicts the geothermal gradient (°C/km) across Colombia using geophysical and geological features. Evaluates model generalization on unseen spatial locations and quantifies prediction error against sparse borehole measurements. Use when the user wants to benchmark on Colombia Geothermal Gradient Dataset, or asks about evaluating this task. Reports R2.
Evaluates the quality and linguistic fidelity of an Italian generative language model (GePpeTto) by measuring its perplexity across in-domain and out-of-domain corpora, and profiling its lexical and syntactic complexity against human-written Italian text. Use when the user wants to benchmark on Wikipedia (Italian), ItWac, EUR-Lex Italian Laws, la Repubblica & Il Giornale, Forum Comments, or asks about evaluating this task. Reports Perplexity.
Evaluates Aspect-Based Sentiment Analysis (ABSA) capabilities on German-language restaurant reviews. It probes models on four subtasks: identifying aspect categories, predicting sentiment polarities for aspects, extracting aspect-sentiment pairs, and end-to-end triplet extraction. Use when the user wants to benchmark on GERestaurant, or asks about evaluating this task. Reports F1 Micro.
Evaluates large language models' ability to answer German legal questions accurately in both open-ended and multiple-choice formats. It probes factual grounding in noisy, real-world legal documents (LegalMC4) versus clean statutory text (BGB), and tests robustness to distractor information typical of retrieval-augmented generation (RAG) pipelines. Use when the user wants to benchmark on LegalMC4 QA, BGB QA, LegalMC4 MCQ, BGB MCQ, ARC (Easy/Challenge), ARC-DE, MMLU, or asks about evaluating th...
Evaluates the quality of text embeddings for German-language documents by measuring how well they cluster into predefined topical categories. It probes a model's ability to capture semantic similarity and domain-specific nuances across different text lengths (titles vs. full texts) and sources. Use when the user wants to benchmark on BlurbsClusteringS2S/P2P, TenKGnadClusteringS2S/P2P, SubredditClusteringS2S/P2P, or asks about evaluating this task. Reports V-measure.
Evaluates extractive question answering and dense passage retrieval capabilities in German. It probes a model's ability to locate precise answer spans within a given context and retrieve relevant passages from a large corpus. Use when the user wants to benchmark on GermanQuAD, or asks about evaluating this task. Reports Exact Match (EM).
Evaluates German NLP models on aspect-based sentiment analysis (ABSA) tasks, including relevance classification, document-level polarity, aspect/sentiment classification, and opinion target extraction. Use when the user wants to benchmark on GermEval17, or asks about evaluating this task. Reports micro F1.
Evaluates instruction-tuned decoder-only LLMs on Portuguese language understanding tasks, including natural language inference, paraphrase detection, commonsense reasoning, and multiple-choice question answering. Use when the user wants to benchmark on MRPC, RTE, COPA, ENEM 2022, BLUEX, STS, or asks about evaluating this task. Reports F1 score.
Evaluates geometric deep learning models' out-of-distribution generalization across scientific domains under conditional, covariate, and concept shifts. It probes how different learning paradigms (ERM, domain adaptation, transfer learning, and OOD generalization) perform when provided with varying amounts of target-domain data. Use when the user wants to benchmark on Track (Particle Tracking Simulation), QMOF (Quantum Metal-organic Frameworks), DrugOOD-3D (3D Conformers of Drug Molecules), or...
Evaluates a model's ability to quickly adapt an object detector to novel classes using only a few labeled examples, while preserving performance on previously learned base classes. Use when the user wants to benchmark on MS-COCO (G-FSD benchmark), or asks about evaluating this task. Reports AP.
Evaluates geometric generative reasoning in unified multimodal models by testing their ability to perform stepwise spatial planning, translate visual constraints into executable code or visual sequences, and verify multi-step geometric construction processes. Use when the user wants to benchmark on GGBench, or asks about evaluating this task. Reports VLM-T.
This benchmark probes the susceptibility of vision-language models to tone-induced hallucination under controlled negative-ground-truth conditions. It isolates linguistic prompt intensity as the sole variable to measure both the frequency and severity of unsupported content generation when visual evidence is deliberately absent or illegible. Use when the user wants to benchmark on Ghost-100, or asks about evaluating this task. Reports H-Rate.
This benchmark evaluates vision-language models' ability to detect and predict hidden interactions in mobile user interfaces. It probes whether models can infer concealed gestures (e.g., long presses, swipes) from before-interaction screenshots and task descriptions, localize the interaction target, and anticipate the resulting UI state changes. Use when the user wants to benchmark on GhostUI, or asks about evaluating this task. Reports Accuracy.
Evaluates the geometric complexity, scene diversity, and sim-to-real transfer capability of the Gibson virtual environment. It benchmarks neural rendering pipelines and embodied agents on tasks like depth estimation, scene classification, and navigation. Use when the user wants to benchmark on Gibson, or asks about evaluating this task. Reports Real-World Transfer Error.
Evaluates zero-shot time series forecasting models across diverse domains, frequencies, and prediction horizons. It probes a model's ability to generalize to unseen multivariate and univariate series, identifying strengths and weaknesses in short-term versus long-term forecasting. Use when the user wants to benchmark on GIFT-Eval, or asks about evaluating this task. Reports MAPE.
Evaluates automatic speech recognition (ASR) systems on a large-scale, multi-domain English corpus containing both read and spontaneous speech. It benchmarks transcription accuracy across tiered training subsets and professionally re-transcribed evaluation sets using word error rate. Use when the user wants to benchmark on GigaSpeech, or asks about evaluating this task. Reports WER.
Evaluates automatic speech recognition (ASR) models on low-resource languages (Thai, Indonesian, Vietnamese) to measure transcription accuracy against reference texts. It probes the model's ability to handle domain-shifted audio and varying linguistic structures using character-level or word-level error metrics. Use when the user wants to benchmark on GigaSpeech 2, Common Voice 17.0, FLEURS, or asks about evaluating this task. Reports CER/WER.
Compute ginic/phone_errors via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of ginic/phone_errors.
Evaluates LLMs' ability to extract and verify user interests from interaction histories, focusing on factual grounding, specificity, and strict instruction following across heterogeneous engagement types. Use when the user wants to benchmark on Unspecified real-world engagement datasets, or asks about evaluating this task. Reports IG.
Evaluates visual-linguistic relation understanding by requiring models to classify image-text matching, ground matched objects with bounding boxes, and identify mismatched relations from candidates. It specifically probes data efficiency and length generalization capabilities in out-of-distribution settings. Use when the user wants to benchmark on GITM-MR, or asks about evaluating this task. Reports Match%.
Evaluates semi-supervised gland segmentation performance on histopathology images under limited labeled data (5% or 10%). It probes the model's ability to disentangle stain color and tissue structure while maintaining boundary precision and shape preservation with minimal annotations. Use when the user wants to benchmark on GlaS, CRAG, or asks about evaluating this task. Reports Dice.
Evaluates a weakly supervised segmentation model's ability to accurately delineate glandular structures in colorectal histopathology images using sparse annotations. It also probes cross-domain generalization across different institutional cohorts with varying staining protocols and scanner characteristics. Use when the user wants to benchmark on GlaS, or asks about evaluating this task. Reports mIoU.
Evaluates machine learning models' ability to predict particle-level dynamic propensity and dynamic heterogeneity from static amorphous structural configurations in glass-forming liquids. Use when the user wants to benchmark on GlassBench, or asks about evaluating this task. Reports Pearson correlation coefficient ($\rho_P$).
Compute Glazkov/mars via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of Glazkov/mars.
Evaluates a model's ability to predict glioblastoma molecular subtypes using paired MRI and histopathology data. It probes the model's capacity to fuse heterogeneous imaging modalities, preserve topological structures, and handle missing data scenarios. Use when the user wants to benchmark on Ivy GAP + Cancer Stem Cells ISH Survey, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates a model's ability to predict glioma IDH mutation status (mutant vs. wild-type) by integrating multi-modal MRI data, including anatomical sequences, tumor geometry, and reconstructed brain networks. It probes the model's capacity for cross-modal feature alignment and patient-level binary classification under data-scarce conditions. Use when the user wants to benchmark on TCIA & In-house Glioma Cohort, or asks about evaluating this task. Reports Accuracy.
This evaluation protocol assesses the zero-shot and few-shot capabilities of large bilingual language models across diverse English and Chinese benchmarks. It probes language modeling, multi-choice question answering, reasoning, commonsense, and cross-lingual transfer abilities. Use when the user wants to benchmark on LAMBADA, Pile, MMLU, BIG-bench-lite, CLUE, FewCLUE, or asks about evaluating this task. Reports accuracy.
Evaluates text-to-speech systems on pronunciation accuracy, speaker similarity, and emotional expressiveness across standard Chinese/English benchmarks and challenging internal datasets. It also assesses vocoder quality using objective and subjective audio metrics to measure overall synthesis fidelity. Use when the user wants to benchmark on Seed-TTS-eval, Libri & Chinese Dialects, or asks about evaluating this task. Reports CER.