
Claude Skills by qhjqhj00
github.com/qhjqhj00Evaluates the ease of implementation and qualitative complexity of five big-data systems when executing real-world scientific image analytics workloads in astronomy and neuroscience. It probes how well each system handles multidimensional array operations, Python integration, and native support for scientific image formats. Use when the user wants to benchmark on Neuroscience Use Case, Astronomy Use Case, or asks about evaluating this task. Reports Lines of Code (LoC).
Evaluates the robustness and cross-dataset/domain generalization of relation extraction models on scientific abstracts. It probes how annotation discrepancies and domain shifts affect relation classification performance. Use when the user wants to benchmark on SemEval-2018, SciERC, or asks about evaluating this task. Reports Macro F1-score.
Evaluates the ability of language models to accurately classify scientific abstracts into fine-grained disciplinary or sub-disciplinary categories under few-shot and zero-shot conditions. Use when the user wants to benchmark on SDPRA 2021, arXiv, S2ORC, or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to automatically assess K-12 science instructional materials against pedagogical rubrics. It probes domain-aligned reasoning, long-context evidence grounding, and the capacity to generate rubric-consistent scores and justifications. Use when the user wants to benchmark on SciEval, or asks about evaluating this task. Reports Evidence Match Rate (EMR).
Evaluates large language models' scientific intelligence across seven core dimensions, including multimodal perception, understanding, reasoning, knowledge comprehension, code generation, symbolic reasoning, and hypothesis generation. It covers multiple scientific disciplines using both text-only and multimodal inputs to assess real-world scientific workflow capabilities. Use when the user wants to benchmark on SLAKE, MSEarth, SFE, OmniEarth, OmniMedVQA, PhyX, ChemBench, ChemBench4K, LLM4Chem...
Evaluates large language models on solving university-level scientific exams in computer science. It probes capabilities in open-ended reasoning, mathematical proof writing, long-form explanations, and multimodal (image-text) understanding across English and German languages. Use when the user wants to benchmark on SciEx, or asks about evaluating this task. Reports Normalized score (0-100%).
This benchmark evaluates large multimodal models' ability to interpret scientific figures by testing their capacity to match figures to captions and vice versa. It probes fine-grained visual-textual reasoning, attention to scientific details, and robustness against adversarially selected distractors. Use when the user wants to benchmark on SciFIBench, or asks about evaluating this task. Reports accuracy.
Evaluates the logical correctness, structural fidelity, and information utility of AI-generated scientific images. It probes whether generated visuals accurately encode domain-specific facts and geometric relationships, and whether they are indispensable for solving visually grounded scientific quizzes. Use when the user wants to benchmark on SciGenBench, or asks about evaluating this task. Reports inverse_validation_rate.
Evaluates multi-modal large language models' ability to interpret scientific graphs and generate accurate, context-aware answers in a multi-turn conversational setting. It probes open-vocabulary visual reasoning and the model's capacity to leverage auxiliary paper metadata for grounded responses. Use when the user wants to benchmark on SciGraphQA, or asks about evaluating this task. Reports CIDEr.
Probes large language models' scientific knowledge across five progressive cognitive levels: memory, comprehension, reasoning, ethical discernment, and real-world application. Covers four scientific domains (biology, chemistry, physics, materials science) using diverse question formats including multiple-choice, relation extraction, and open-ended protocol design. Use when the user wants to benchmark on SciKnowEval, or asks about evaluating this task. Reports overall normalized score.
Evaluates a model's ability to perform complex, claim-centric reasoning over full scientific documents containing multimodal elements (charts, tables, figures). It probes the model's capacity to localize evidence and answer questions accurately despite long-context noise and distractors. Use when the user wants to benchmark on SciMDR-Eval, or asks about evaluating this task. Reports accuracy.
Evaluates an autonomous agent's capability to generate executable scientific and data-science code by measuring execution reliability and task-goal satisfaction. It probes the model's ability to navigate constrained search budgets while producing outputs that match predefined success criteria across diverse disciplinary benchmarks. Use when the user wants to benchmark on ScienceAgentBench, DA-Code, or asks about evaluating this task. Reports Success Rate (SR).
Evaluates open-ended, closed-book scientific question answering capabilities. It probes a model's ability to generate comprehensive, accurate, and reasonable answers to research-level science questions without external context or reference papers. Use when the user wants to benchmark on SciQAG-24D, SciQ, or asks about evaluating this task. Reports CAR.
Evaluates an LLM's ability to comprehend algorithmic descriptions from academic papers and translate them into executable code. It probes the model's capacity for algorithmic reasoning, dependency resolution, and practical implementation within a repository context. Use when the user wants to benchmark on SciReplicate-Bench, or asks about evaluating this task. Reports Execution Accuracy.
This benchmark evaluates the ability of models to generate extreme, single-sentence summaries (TLDRs) of scientific papers, capturing key contributions while bypassing background details. It tests both automated overlap metrics and human-judged informativeness and correctness under multi-target and multi-input settings. Use when the user wants to benchmark on SCITLDR, or asks about evaluating this task. Reports Rouge-1.
Evaluates long-context language models' ability to perform numerical aggregation, filtering, sorting, and logical operations across extended contexts (up to 1M tokens) using scientific article metadata and full-text articles. Use when the user wants to benchmark on SciTrek, or asks about evaluating this task. Reports exact match.
Evaluates large language models across four dimensions of trustworthiness in scientific contexts: truthfulness, adversarial robustness, scientific safety, and scientific ethics. It probes models' ability to provide accurate scientific information, resist adversarial perturbations, avoid generating harmful content, and make sound ethical judgments in research scenarios. Use when the user wants to benchmark on SciQ, ARC-C, MMLU, GPQA-Diamond, LogiQA, ReClor, LOGICINFERENCE, WMDP, HarmBench, Sci...
Evaluates agentic systems' ability to perform scientific data analysis and visualization workflows. It probes outcome correctness, process behavior, and computational efficiency across real-world scientific domains and tools. Use when the user wants to benchmark on SciVisAgentBench, or asks about evaluating this task. Reports outcome correctness.
Evaluates multimodal LLMs on closed-ended visual and non-visual question answering over scientific figures. It probes recognition of visual attributes (color, shape, position) and reasoning capabilities across diverse chart types. Use when the user wants to benchmark on SciVQA, or asks about evaluating this task. Reports ROUGE-1 F1.
Evaluates a model's ability to perform hierarchical scientific summarization by generating three distinct granularity levels (Abstract, Key Contributions, TL;DR) from a single full-text input. It probes multi-granularity text compression and the model's capacity to maintain coherence across varying compression ratios within a single inference pass. Use when the user wants to benchmark on SciZoom, or asks about evaluating this task. Reports unspecified summarization metric.
Evaluates the fidelity of a particle filter algorithm for generating counterfactual samples from structural causal models by comparing empirical statistics of the generated samples against known ground-truth distributions and correlations. Use when the user wants to benchmark on Synthetic SCM Simulation, or asks about evaluating this task. Reports proportion of unique observations.
Evaluates computational methods for single-cell multi-omics integration by measuring their ability to preserve biological variation, align different omics layers at the cell and single-cell levels, and enable accurate cell type annotation. Use when the user wants to benchmark on SHARE-seq BMMC, SNARE-seq, 10X Genomics Multiome, Human fetal atlas, CITE-seq BMMC S1, CITE-seq BMMC S4, Human brain multi-omics, Human brain 3k, or asks about evaluating this task. Reports biological variation conser...
This evaluation protocol assesses the ability of deep learning models to predict drug-target interactions (DTI) by learning from molecular graphs and protein sequences. It probes the model's capacity to capture cross-domain interaction patterns between small molecules and proteins, particularly in semi-inductive settings where novel compounds are paired with known protein families. Use when the user wants to benchmark on BindingDB, KIBA, Human, SCOPE, or asks about evaluating this task. Repor...
Evaluates dense optical flow estimation accuracy and occlusion detection on standard video benchmarks. Probes the model's ability to predict pixel-wise motion vectors and identify occluded regions under varying motion magnitudes and scene complexities. Use when the user wants to benchmark on Sintel, KITTI, or asks about evaluating this task. Reports End Point Error (EPE).
This evaluation protocol measures the impact of score ties on document ranking repeatability across diverse information retrieval collections. It quantifies how non-deterministic tie-breaking during multi-threaded indexing causes variability in standard ranking metrics, even when using identical queries and ranking models. Use when the user wants to benchmark on TREC 2004 Robust Track (Disks 4 & 5), TREC 2005 Robust Track (AQUAINT), TREC 2017 Common Core Track (NYT Annotated Corpus), TREC 201...
This framework audits medical LLM benchmarks across five lifecycle phases using 46 medically tailored criteria to assess clinical relevance, data integrity, safety-critical capabilities, validity, and governance. Use when the user has predictions and gold and needs to compute score.
Evaluates the scientific reasoning and problem-solving capabilities of LLMs on graduate-level higher education science problems. It measures how well models can parse and solve complex scientific questions involving formulas and equations. Use when the user wants to benchmark on SCP-116K, or asks about evaluating this task. Reports Accuracy.
Evaluates LLM-based web information extraction by measuring structural validity (JSON parseability, schema compliance), key extraction accuracy (precision, recall, F1), and value extraction quality (type-aware exact match, BLEU) on real-world HTML-to-JSON tasks. Use when the user wants to benchmark on ScrapeGraphAI-100k, or asks about evaluating this task. Reports Key F1.
Evaluates LLMs' ability to understand, diagnose, and repair bugs in multimodal, event-driven block-based programming environments (Scratch). It probes functional correctness, structured bug explanation, trigger/mechanism identification, and patch minimality/semantic preservation. Use when the user wants to benchmark on ScratchEval, or asks about evaluating this task. Reports G-Acc.
Evaluates a model's ability to detect, localize, and semantically label all interactable UI elements on a clean screenshot. It probes fine-grained spatial reasoning, handling of dense layouts, and UI semantics understanding. Use when the user wants to benchmark on GUI-360°-Bench, or asks about evaluating this task. Reports F1.
This evaluation probes a vision-language model's ability to localize specific UI elements within graphical user interfaces based on natural language instructions. It tests precise coordinate prediction and cross-resolution generalization across mobile, desktop, and web platforms. Use when the user wants to benchmark on ScreenSpot, ScreenSpot-v2, ScreenSpot-Pro, or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to locate specific GUI elements from a screenshot given a text instruction. It measures both coarse localization accuracy and fine-grained bounding box overlap across desktop, mobile, and web platforms. Use when the user wants to benchmark on ScreenSpot, or asks about evaluating this task. Reports grounding accuracy.
Evaluates a model's ability to automatically generate concise, coherent language summaries of mobile UI screens by fusing visual, structural, and textual modalities. It probes multimodal representation learning and language generation capabilities in the context of human-computer interaction and UI understanding. Use when the user wants to benchmark on Screen2Words, or asks about evaluating this task. Reports BLEU-4.
Evaluates a model's ability to perform fine-grained text dragging interactions on GUI screenshots. It measures whether the model correctly triggers a drag action, accurately selects the target text span, and aligns its predicted coordinates with ground truth. Use when the user wants to benchmark on SCREENDRAG, or asks about evaluating this task. Reports DTR.
Evaluates a model's ability to read and describe the content and layout of a GUI screenshot at a specific pointed location. It probes layout-aware screen reading, spatial reasoning, and the capacity to generate focused descriptions for mobile agent navigation. Use when the user wants to benchmark on ScreenPR, or asks about evaluating this task. Reports Content Acc.
Evaluates a model's ability to locate specific UI elements on screenshots based on natural language instructions. It measures how accurately the model can predict click coordinates that align with ground-truth bounding boxes across mobile, desktop, and web platforms. Use when the user wants to benchmark on ScreenSpot, or asks about evaluating this task. Reports click accuracy.
This benchmark evaluates a model's ability to perform GUI grounding in professional, high-resolution desktop environments. It probes whether vision-language models can accurately locate specific UI elements (both text and icons) based on natural language instructions, highlighting challenges with small targets and complex interfaces. Use when the user wants to benchmark on ScreenSpot-Pro, or asks about evaluating this task. Reports accuracy (center-point).
Evaluates the accuracy of a script identification tool on multilingual web corpora by checking if the predicted writing system matches the admissible scripts for the corpus's assigned language. It also measures script representation coverage in multilingual LLM tokenizers. Use when the user wants to benchmark on Multilingual C4 (mC4), OSCAR 22.01, or asks about evaluating this task. Reports ACC.
Evaluates long-text understanding capabilities across summarization, question answering, and natural language inference tasks. It probes whether models can effectively process and extract information from documents exceeding standard context windows (up to 16K tokens) using chunked encoding and cross-chunk fusion. Use when the user wants to benchmark on SCROLLS, or asks about evaluating this task. Reports Avg SCROLLS score.
Evaluates machine learning models' ability to predict severe convective storm (SCS) frequency and occurrence in European Russia under climate change scenarios. It probes binary classification of SCS events against non-events, as well as regression accuracy on the normalized annual cycle of storm activity using physics-informed deep learning architectures. Use when the user wants to benchmark on CMIP5 RCP8.5 & Meteorological Observations, or asks about evaluating this task. Reports RMSEAC.
Evaluates how well self-supervised learning models learn single-cell representations for three downstream tasks: batch correction, cell type annotation, and missing modality prediction. It probes the trade-off between preserving biological variance and removing technical batch effects, as well as the ability to generalize across uni- and multi-omics data modalities. Use when the user wants to benchmark on PBMC-M, BMMC, PBMC, Pancreas, Immune Cell Atlas, MCA, Lung, Tabula Sapiens, HIC, or asks...
Evaluates the ability of self-distillation augmented masked autoencoders to learn robust visual representations from histopathological images for downstream tasks like classification, segmentation, and detection, particularly in low-class or cross-domain settings. Use when the user wants to benchmark on PatchCamelyon (PCam), NCT-CRC-HE (NCT), MSIIvsMSS, MoNuSeg, Glas, NuCLS, or asks about evaluating this task. Reports top-1 accuracy.
Evaluates a reinforcement learning policy for same-day delivery routing on synthetic geographic settings, measuring the trade-off between overall service utility and regional fairness (minimum service rate). Use when the user wants to benchmark on SDDFCS Simulation, or asks about evaluating this task. Reports r_total.
Evaluates whether a GAN-based oversampling technique (SDG-GAN) improves binary classification performance on imbalanced tabular data compared to traditional and GAN-based baselines. Use when the user wants to benchmark on Credit Card Fraud Dataset, Pima Diabetes Dataset, Breast Cancer Wisconsin (Diagnostic) Dataset, Gambling Fraud Dataset, or asks about evaluating this task. Reports algorithmic performance.
Evaluates the quality and controllability of a stochastic differential music generation and editing model. It probes the model's ability to generate pop piano music from scratch or conditioned on control signals, and perform fine-grained editing tasks like stroke-based generation, inpainting, and style transfer. Use when the user wants to benchmark on ailabs1k7, or asks about evaluating this task. Reports pitch distribution similarity (PD).
Evaluates a model's ability to retrieve relevant Korean public document pages given a text query, comparing text-only parsing against multimodal visual understanding. It probes cross-modal reasoning, layout awareness, and the capacity to interpret tables, charts, and complex visual structures in administrative documents. Use when the user wants to benchmark on SDS KoPub VDR, or asks about evaluating this task. Reports Recall@k.
Evaluates how speech enhancement (SE) artifacts and noise errors affect automatic speech recognition (ASR) performance. It measures Word Error Rate (WER) on enhanced speech signals derived from simulated and real-world reverberant noisy conditions to isolate the impact of artifact components. Use when the user wants to benchmark on Simulated WSJ0+CHiME-3, CHiME-3 et05_real, or asks about evaluating this task. Reports WER [%].
Evaluates the ability of contemporary toxicity detection models to correctly identify toxic language in software engineering contexts, such as code reviews and developer chat logs. It probes whether general-purpose classifiers can handle domain-specific terminology and contextual nuances without significant performance degradation. Use when the user wants to benchmark on Jigsaw Sample, Code Review, Gitter Ethereum, or asks about evaluating this task. Reports F-Score.
Compute SEA-AI/box-metrics via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of SEA-AI/box-metrics.
Compute SEA-AI/panoptic-quality via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of SEA-AI/panoptic-quality.