
Claude Skills by qhjqhj00
github.com/qhjqhj00Evaluates unsupervised visual anomaly detection and segmentation models on real-world supermarket goods. It probes robustness to object misalignment, intra-class appearance variation, and the ability to detect subtle or small anomalies without labeled anomalous training data. Use when the user wants to benchmark on PKU-GoodsAD, or asks about evaluating this task. Reports AUROC, AUPR.
Evaluates Polish and multilingual text embedding models across 28 tasks spanning classification, clustering, pair classification, retrieval, and semantic textual similarity. It measures how well embeddings capture semantic relationships, support downstream classification, cluster documents, and retrieve relevant documents in Polish. Use when the user wants to benchmark on PL-MTEB, or asks about evaluating this task. Reports nDCG@10.
Evaluates NLP systems and large language models on adapting biomedical abstracts to plain language for lay consumers. It probes capabilities in text simplification, term replacement, factual faithfulness, and conciseness while measuring alignment with human expert judgments. Use when the user wants to benchmark on TREC PLABA, or asks about evaluating this task. Reports SARI.
Evaluates an LLM's ability to translate natural language planning task descriptions into valid, semantically equivalent Planning Domain Definition Language (PDDL) code. It specifically probes the model's capacity to accurately capture initial states, goal states, and object relationships while adhering to formal planning semantics. Use when the user wants to benchmark on Planetarium, or asks about evaluating this task. Reports equivalence.
Evaluates pixel-level segmentation capabilities for identifying and localizing plant diseases in real-world, uncontrolled agricultural imagery across 115 disease classes. The benchmark tests a model's ability to handle fine-grained lesion boundaries, overlapping disease symptoms, and high visual diversity typical of field-captured crops. Use when the user wants to benchmark on PlantSeg, or asks about evaluating this task. Reports mIoU.
This benchmark evaluates vision-language models on plant science tasks, ranging from basic species and health identification to detailed symptom verification and higher-order causal or counterfactual reasoning. It probes a model's ability to ground visual attributes, diagnose diseases, and generate descriptive or diagnostic text based on leaf images. Use when the user wants to benchmark on PlantVillageVQA, or asks about evaluating this task. Reports accuracy.
Evaluates multimodal models' ability to generate and edit images that correctly follow planning-oriented instructions (route planning, workflow diagramming, web/UI displaying). It probes procedural reasoning, spatial consistency, and semantic alignment in visual synthesis. Use when the user wants to benchmark on PlanViz, or asks about evaluating this task. Reports Cor.
Evaluates LLM reliability on curated, low-noise subsets of standard benchmarks (VQA v2.0, SQuAD 2.0, HotPotQA, DROP, BIG-bench) by removing ambiguous examples and re-labeling to minimize ground-truth errors, revealing true model failures on elementary reasoning tasks. Use when the user wants to benchmark on VQA v2.0, SQuAD 2.0, HotPotQA, DROP, BIG-bench, or asks about evaluating this task. Reports accuracy.
Evaluates structural pruning methods for pre-trained language models (BERT-base, RoBERTa-base) across eight text classification tasks. It measures the trade-off between model size (parameter count) and task performance (validation error) to identify Pareto-optimal sub-networks. Use when the user wants to benchmark on eight text classification tasks, or asks about evaluating this task. Reports Hypervolume.
Evaluates sequence labeling models on detecting abbreviations and extracting their corresponding long forms in scientific text. It probes domain-specific NER capabilities under challenges like context dependency, sub-abbreviations, and polysemy. Use when the user wants to benchmark on PLOD, SDU@AAAI-22 Shared Task, or asks about evaluating this task. Reports F.
Evaluates multi-modal large language models' ability to visually interpret scientific plots and generate corresponding executable Python (matplotlib) code. It probes fine-grained visual reasoning, text extraction from dense plots, and precise code generation for data visualization. Use when the user wants to benchmark on Plot2Code, or asks about evaluating this task. Reports Pass Rate.
This benchmark evaluates multimodal LLMs on engineering plot reading and visual quantitative reasoning. It probes the model's ability to interpret complex axes (including log scales), read curve values, and compute derived engineering quantities like cutoff frequencies or settling times from rendered plot images. Use when the user wants to benchmark on PlotChain, or asks about evaluating this task. Reports field-level accuracy.
This benchmark evaluates large language models on five core Greek financial NLP tasks: numeric and textual named entity recognition, multiple-choice question answering, abstractive summarization, and financial topic classification. It probes models' ability to handle low-resource language morphology, domain-specific financial terminology, and reasoning within Greek financial contexts. Use when the user wants to benchmark on GRFinNUM, GRFinNER, GRFinQA, GRFNS-2023, GRMultiFin, or asks about ev...
Evaluates multi-modal large language models on medical reasoning tasks involving compound figures, single images, and text-only prompts. It probes the model's ability to synthesize cross-modal information, perform clinical diagnosis, and generate accurate medical explanations across diverse imaging modalities and specialties. Use when the user wants to benchmark on PMC-MI-Bench, or asks about evaluating this task. Reports BLEU@4, Accuracy.
Binary classification of user-independent emotional states (valence and arousal) from electrodermal activity (EDA) signals. It probes the model's ability to generalize across subjects by using subject-specific thresholds and fusing physiological signals with external music benchmarks. Use when the user wants to benchmark on PMEmo, or asks about evaluating this task. Reports accuracy.
This evaluation probes the quality of automatic machine translation between English and 13 Indian languages using a parallel corpus. It measures how well NMT systems can handle diverse linguistic structures, including abugida scripts and agglutinative morphology, across low-resource language pairs. Use when the user wants to benchmark on PMIndia, or asks about evaluating this task. Reports BLEU.
This benchmark evaluates the discriminative performance of machine learning models on tabular classification tasks under data-scarce conditions (sample sizes ≤500). It specifically probes whether complex AutoML and deep learning approaches can consistently outperform simple baselines like logistic regression when training data is limited. Use when the user wants to benchmark on PMLBmini, or asks about evaluating this task. Reports AUC.
Evaluates multilingual capabilities of LLMs across understanding, reasoning, and generation tasks in 10 languages. It probes prompt sensitivity and cross-lingual performance consistency to reveal benchmark origin bias and language-specific scaling trends. Use when the user wants to benchmark on MMMLU, MLogiQA, MGSM, MHellaSwag, XNLI, Flores-200, or asks about evaluating this task. Reports accuracy.
Evaluates the ability of transformer-based models to generate abstractive summaries in Persian. It measures how well generated summaries match reference summaries in terms of lexical overlap and longest common subsequence at the sentence level. Use when the user wants to benchmark on pn-summary, or asks about evaluating this task. Reports ROUGE-1 F-1.
This evaluation probes a retrieval system's ability to identify relevant tabular datasets from a corpus given natural language questions. It measures retrieval accuracy via hit rate, alongside system efficiency metrics including query throughput, offline preparation time, and storage footprint across diverse real-world and benchmark datasets. Use when the user wants to benchmark on ChEMBL, Adventure Works, Public BI, Chicago Open Data, FeTaQA, BIRD, or asks about evaluating this task. Reports...
This benchmark evaluates the zero-shot diagnostic capability of vision-language models on chest X-ray images for binary pneumonia detection. It probes whether models can accurately classify radiological findings without task-specific fine-tuning, relying instead on prompt engineering and pre-trained visual reasoning. Use when the user wants to benchmark on Chest radiographic Images (Pneumonia), or asks about evaluating this task. Reports accuracy.
Evaluates a topology-based pore network model's ability to predict flow-permeable surface area and hydraulic conductance in granular materials from micro-CT images. Use when the user wants to benchmark on Sphere Packing & High-Explosive Micro-CT Samples, or asks about evaluating this task. Reports conductance_ratio.
Evaluates few-shot representation learning under partial observability. Models must match query image views to their underlying source images using only partial support views (≤50% coverage) and viewpoint coordinates. Use when the user wants to benchmark on PO-Meta-Dataset, or asks about evaluating this task. Reports accuracy.
Evaluates the quality of monaural music and speech source separation in podcast audio. It measures both objective signal fidelity using BSS-eval metrics and subjective perceptual quality using standardized listening tests. Use when the user wants to benchmark on PodcastMix, or asks about evaluating this task. Reports SDR.
Evaluates the ability of a hierarchical transformer (PoET) to post-process and calibrate medium-range ensemble weather forecasts for 2m temperature and precipitation, compared to a baseline method (MBM) and raw ensemble outputs. Use when the user wants to benchmark on ECMWF ensemble forecasts, or asks about evaluating this task. Reports CRPS.
Probes a model's ability to identify privacy-sensitive objects in images by reasoning about scene context rather than relying solely on visual appearance. It evaluates whether the system can distinguish between obvious privacy leaks (e.g., faces) and context-dependent sensitive information (e.g., people in specific roles). Use when the user wants to benchmark on MOSAIC, PRIVACY1000, or asks about evaluating this task. Reports F1 Score.
Evaluates anomaly detection performance by granting full credit for all points in an anomalous segment if at least one point is detected, often inflating scores for algorithms that merely hit a segment once. Use when the user has predictions and gold and needs to compute point-adjust F1.
Evaluates vision-language models' embodied reasoning and visual grounding capabilities across three hierarchical stages: referred-object localization, task-driven pointing, and multi-step visual trace prediction in real-world scenarios. Use when the user wants to benchmark on Point-It-Out (PIO), or asks about evaluating this task. Reports score.
Evaluates a hybrid graph attention and 3D point cloud neural network's ability to predict quantum chemical properties and molecular physicochemical traits. It probes the model's capacity to integrate 2D topological graph features with 3D spatial geometry for accurate regression and classification of molecular energies and properties. Use when the user wants to benchmark on MoleculeNet, C10, or asks about evaluating this task. Reports MAE, R².
Evaluates multimodal large language models on fine-grained image understanding and long-form video comprehension tasks. It measures the trade-off between visual token compression efficiency and task accuracy across diverse benchmarks. Use when the user wants to benchmark on MVBench, Video-MME, MLVU, LongVideoBench, MMBench, MMMU_val, or asks about evaluating this task. Reports accuracy.
Evaluates the ability to detect cyber attack campaigns by aligning threat intelligence query graphs with system provenance graphs derived from kernel audit logs. It probes structural pattern matching, causal dependency reasoning, and robustness against malware mutations and benign system noise. Use when the user wants to benchmark on DARPA TC Dataset, Public Malware Reports, or asks about evaluating this task. Reports alignment score.
Evaluates the ability of Graph Neural Networks to perform node classification while mitigating bias related to a protected attribute (Region). It probes the trade-off between predictive accuracy and group fairness across different GNN architectures. Use when the user wants to benchmark on Pokec-n, or asks about evaluating this task. Reports F1 score.
PokeGym evaluates vision-language models' ability to perform long-horizon planning and spatial reasoning in a complex 3D open-world game using only raw RGB observations. It specifically probes visual grounding, autonomous goal decomposition, and physical deadlock recovery, revealing whether models can navigate cluttered environments, interact with objects, and recover from entrapment without explicit state feedback. Use when the user wants to benchmark on PokeGym, or asks about evaluating thi...
This benchmark evaluates a model's ability to detect online polarization in social media text by classifying statements as polarized or non-polarized. It specifically probes the model's capacity to produce interpretable, structured reasoning alongside binary predictions while handling class imbalance and reducing false negatives. Use when the user wants to benchmark on POLAR @ SemEval-2026, or asks about evaluating this task. Reports macro-F1.
Evaluates the ability of machine learning models to distinguish between reference stars and circumstellar exoplanetary disks in high-contrast polarimetric imaging data. It probes representation learning quality through downstream supervised classification and unsupervised clustering tasks. Use when the user wants to benchmark on POLARIS, or asks about evaluating this task. Reports accuracy.
Evaluates training-free multimodal agents on retrieval-augmented generation, general reasoning, and hallucination robustness by testing a polarized latent graph memory that injects logical constraints at inference time. Use when the user wants to benchmark on MRAMG-Bench, MRAG-Bench, Visual-RAG, MMMU, MMStar, HallusionBench, or asks about evaluating this task. Reports performance.
Evaluates the model's ability to retrieve relevant coordination policies that guide task planning based on a high-level progress summary of the current state. Use when the user wants to benchmark on Policy Selection Test Suite, or asks about evaluating this task. Reports F1 Score.
Evaluates the transcription accuracy of various automatic speech recognition (ASR) models on Polish-language audio, contrasting read-speech benchmarks with spontaneous, noisy medical consultations to probe domain generalization. Use when the user wants to benchmark on Mozilla Common Voice (MCV) Polish, Multilingual LibriSpeech (MLS) Polish, Medical interview corpus, or asks about evaluating this task. Reports Word Error Rate (WER).
Evaluates large language models on Polish medical licensing and specialization exams to assess cross-lingual medical knowledge transfer, domain-specific understanding, and specialty-level accuracy compared to human medical graduates. Use when the user wants to benchmark on Polish Medical Exams (LEK/LDEK/PES), or asks about evaluating this task. Reports score.
Evaluates Polish language understanding, summarization, and question answering capabilities of text-to-text models. It probes how well encoder-decoder and decoder-only architectures generalize from multilingual pre-training to monolingual Polish tasks using exact-match generation and ROUGE-based metrics. Use when the user wants to benchmark on KLEJ benchmark, Allegro Articles, Polish Summaries Corpus, or asks about evaluating this task. Reports exact-match accuracy.
Evaluates the ability of LLMs and API-based classifiers to accurately annotate toxicity and incivility in political protest content against a human gold standard. It probes zero-shot classification performance, threshold sensitivity, and output reproducibility across different model sizes and temperatures. Use when the user wants to benchmark on Political protest content dataset, or asks about evaluating this task. Reports F1-score.
Evaluates the alignment and quality of generated image captions relative to reference captions and source images. It measures how well a learned metric correlates with human judgments, specifically probing hallucination robustness and open-vocabulary caption evaluation. Use when the user has predictions and gold and needs to compute Polos.
Evaluates open-domain question answering in Polish by measuring both passage retrieval accuracy and answer generation quality. It probes a model's ability to retrieve relevant evidence from a large corpus and accurately extract or generate answers from those passages. Use when the user wants to benchmark on PolQA, or asks about evaluating this task. Reports fuzzy_match.
Evaluates large language models' ability to detect hallucinations by verifying factual claims across 11 languages. It probes cross-linguistic consistency, topic-aware fact-checking, and resistance to web-resource bias in multilingual settings. Use when the user wants to benchmark on Poly-FEVER, or asks about evaluating this task. Reports accuracy.
Evaluates multimodal language models' ability to answer questions that require reasoning across multiple charts or images. It probes visual decomposition, sub-chart localization, and handling of complex multi-visual contexts versus single-chart inputs. Use when the user wants to benchmark on PolyChartQA, MultiChartQA-RQ1, or asks about evaluating this task. Reports L-Accuracy.
Evaluates the toxicity of LLM-generated continuations across 17 languages using naturally occurring prompts scraped from the web. It probes how model size, language resource availability, and instruction/preference tuning affect the generation of harmful content. Use when the user wants to benchmark on PolygloToxicityPrompts (PTP), or asks about evaluating this task. Reports AT.
Evaluates multi-modal mathematical and cognitive reasoning capabilities on visual puzzles. It probes spatial interpretation, relational understanding, pattern recognition, and long-horizon logical reasoning using diagram-based multiple-choice questions. Use when the user wants to benchmark on POLYMATH, or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to predict polypharmacy side effects (drug-drug interactions) in a multimodal biomedical graph. It probes the model's capacity to learn continuous latent representations for drugs and proteins and generalize to unseen drug pairs across 964 specific side effect types. Use when the user wants to benchmark on Polypharmacy Side Effects Dataset, or asks about evaluating this task. Reports cross-entropy loss.
Evaluates the accuracy and continuity of polyploid haplotype assembly methods by measuring switch errors and read-haplotype conflicts, while also quantifying phasing uncertainty across varying ploidies, coverages, and genomic structures. Use when the user wants to benchmark on Synthetic Polyploid Genomes (S. tuberosum), Experimental Octoploid Strawberry (F. x ananassa), or asks about evaluating this task. Reports Generalized Vector Error Rate (VER).
Evaluates a unified foundation model's ability to perform joint polyp detection, segmentation, classification, and unsupervised tracking on colonoscopy video frames. It tests generalization to unseen clinical datasets and consistency of object association across frames without task-specific fine-tuning. Use when the user wants to benchmark on Kvasir-SEG, CVC-ClinicDB, CVC-ColonDB, ETIS, CVC-300, KUMC, REAL-Colon, or asks about evaluating this task. Reports Dice.