
Claude Skills by qhjqhj00
github.com/qhjqhj00This benchmark probes the ability of toxicity detection models to identify nuanced, adversarially crafted harmful content (e.g., gaslighting, manipulation, sarcasm) that mainstream tools often miss due to reliance on normative annotations and profanity cues. Use when the user wants to benchmark on SASS, or asks about evaluating this task. Reports F1-Score.
Evaluates the ability of lightweight deep learning models to predict fine-grained class proportions (e.g., vegetation density, population) from satellite image chips. The protocol measures how well models trained on coarse administrative-level label proportions can recover fine-grained spatial distributions, using both proportion regression and pixel-level segmentation accuracy. Use when the user wants to benchmark on esaworldcover, humanpop, or asks about evaluating this task. Reports MAE.
Evaluates a model's ability to retrieve relevant videos given a natural language query in an audio-visual setting. It specifically probes how well speech-aware representations and early vision-audio alignment improve cross-modal matching accuracy across diverse video-text benchmarks. Use when the user wants to benchmark on MSRVTT-9k, MSRVTT-7k, VATEX, Charades, LSMDC, or asks about evaluating this task. Reports SumR.
Evaluates multimodal foundation models' ability to perform observer-centric spatial reasoning, path integration, and camera geometry inference using egocentric videos recorded from smart glasses. It probes sustained tracking of intermediate movements and the distinction between physical translation and camera rotation. Use when the user wants to benchmark on SAW-BENCH, or asks about evaluating this task. Reports accuracy.
Evaluates a session-based recommendation model's ability to predict the next item in a user session using validated and enriched LLM-generated intents. It probes the model's capacity to leverage semantic intent signals alongside sequential interaction patterns for accurate item ranking. Use when the user wants to benchmark on Beauty (Amazon), Yelp, Books (Amazon), or asks about evaluating this task. Reports Hit Rate@10.
Evaluates session-based recommendation models by predicting the next item in a user's session sequence. It probes the model's ability to capture sequential item transitions and leverage global item-transition patterns across sessions to improve ranking accuracy. Use when the user wants to benchmark on Diginetica, Tmall, Nowplaying, or asks about evaluating this task. Reports P@20.
Evaluates LLMs' ability to solve complex, Olympiad-level mathematics problems across Algebra, Combinatorics, Number Theory, and Geometry. It specifically probes multi-turn code generation and execution feedback for iterative problem decomposition and constraint handling. Use when the user wants to benchmark on AIME, AMC-12, MathOdyssey, OlympiadBench, or asks about evaluating this task. Reports accuracy.
Evaluates AI co-scientist agents' ability to autonomously plan, execute, and interpret single-cell biology workflows to answer open-ended research questions (OEQs) and multiple-choice questions (MCQs). It probes hypothesis generation, code execution, data analysis, and scientific reasoning in a domain-specific setting. Use when the user wants to benchmark on sc-HeurekaBench-Lite, or asks about evaluating this task. Reports Correctness [1-5].
Evaluates long-context academic reasoning by testing whether LLMs can correctly identify masked citations within scientific papers. It probes the model's ability to understand semantic context, attributional claims, and descriptive references across varying context lengths and difficulty levels. Use when the user wants to benchmark on SCALAR, or asks about evaluating this task. Reports accuracy.
Compute the ScaleInvariantSignalDistortionRatio metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute ScaleInvariantSignalDistortionRatio, or asks how to score with ScaleInvariantSignalDistortionRatio.
Compute the ScaleInvariantSignalNoiseRatio metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute ScaleInvariantSignalNoiseRatio, or asks how to score with ScaleInvariantSignalNoiseRatio.
This benchmark evaluates the typographic robustness of vision-language models (VLMs) and large vision-language models (LVLMs) by measuring their susceptibility to adversarial handwritten or synthetic text inserted into images. It probes whether models can correctly identify the primary object in an image despite the presence of misleading attack words, revealing vulnerabilities in multimodal alignment and text-visual reasoning. Use when the user wants to benchmark on SCAM, or asks about evalu...
Evaluates the performance of monolingual and multilingual language models across five Scandinavian languages (Danish, Norwegian, Swedish, Icelandic, Faroese) on question answering, linguistic acceptability, and named entity recognition. It also probes cross-lingual transfer capabilities between these languages by measuring performance variance across language groups. Use when the user wants to benchmark on ScandiQA, ScaLA, MIM-GOLD-NER, WikiANN, or asks about evaluating this task. Reports acc...
Evaluates whether translating low-resource language data into English and applying large-scale English/multilingual models outperforms training native monolingual models for sentiment classification. It probes the efficiency and effectiveness of cross-lingual data reuse versus isolated language-specific pre-training. Use when the user wants to benchmark on Sentiment datasets (Swedish, Danish, Norwegian, Finnish, English), or asks about evaluating this task. Reports binary accuracy.
Evaluates multi-modal named entity recognition (MNER) and visual grounding capabilities, specifically probing the model's ability to generalize to unseen entities by leveraging external knowledge (Wikipedia) and image-based features. Use when the user wants to benchmark on MNER, GMNER, or asks about evaluating this task. Reports F1 score.
Evaluates the rendering quality and novel-view synthesis capability of Neural Radiance Field (NeRF) methods on real-world inward-facing object scans. It probes how well models generalize to unseen camera poses when trained with varying image densities and localized acquisition patterns. Use when the user wants to benchmark on ScanNeRF, or asks about evaluating this task. Reports PSNR.
Evaluates the ability of 3D object detectors to localize and classify indoor objects using variable-frame sparse RGB-D inputs. It probes generalization across different input modalities (reconstructed point clouds, multi-view RGB-D, monocular RGB-D) and varying numbers of input views. Use when the user wants to benchmark on ScanNet, or asks about evaluating this task. Reports mAP@0.25.
Evaluates a model's ability to perform 3D visual grounding by jointly reasoning about implicit human instructions and localizing target objects in 3D scenes. It probes spatial, functional, logical, emotional, and safety-related reasoning capabilities alongside precise 3D bounding box localization. Use when the user wants to benchmark on ScanReason, or asks about evaluating this task. Reports matching score.
Evaluates stereo correspondence and depth estimation methods for endoscopic surgical scenes. It probes how accurately models can reconstruct quasi-dense depth maps from stereo image pairs captured with structured light on biological tissue. Use when the user wants to benchmark on SCARED, or asks about evaluating this task. Reports mean absolute error in mm.
Evaluates object detection and instance segmentation capabilities on real-world images of dog feces. It specifically probes model robustness to camouflage, occlusion, varying lighting conditions, and small object detection in outdoor urban environments. Use when the user wants to benchmark on ScatSpotter, or asks about evaluating this task. Reports mAP.
This protocol evaluates neural machine translation quality between English and Thai. It measures translation accuracy on a newly curated 1M-parallel corpus (SCB_1M) and a filtered OPUS corpus, while also testing cross-domain generalization on the IWSLT 2015 Thai-English benchmark. Use when the user wants to benchmark on SCB_1M, MT_OPUS, IWSLT 2015 Thai-English, or asks about evaluating this task. Reports SacreBLEU.
Evaluates the clustering accuracy, label consistency, and biological interpretability of various single-cell RNA-seq analysis methods. It probes how well traditional, deep learning, graph-based, and foundation model algorithms recover known cell type annotations across diverse tissues and dataset sizes. Use when the user wants to benchmark on scCluBench (36 human & mouse scRNA-seq datasets), or asks about evaluating this task. Reports Normalized Mutual Information (NMI).
Evaluates how scenario-induced contextual factors (personality, region, identity) and multilingual settings alter LLM judgments on financial misinformation claims, quantifying behavioral bias as the performance shift relative to a neutral baseline. Use when the user wants to benchmark on Multilingual Financial Misinformation Dataset, or asks about evaluating this task. Reports Bias_scen.
Evaluates the intrinsic diversity of text-to-image generative models by isolating model-driven variation from prompt-driven variation. It uses CLIP embeddings to construct a joint image-text kernel covariance matrix and applies Schur complement decomposition to remove text influence before computing spectral entropy. Use when the user has predictions and gold and needs to compute Scendi score.
Evaluates the factual consistency and scene graph adherence of text-to-image generation models. It probes whether generated images accurately preserve specified objects and their spatial/relational configurations as defined by input scene graphs, rather than just measuring aesthetic quality or text-image alignment. Use when the user wants to benchmark on Visual Genome (VG) test set, MegaSG, or asks about evaluating this task. Reports SGScore.
Evaluates scene graph generation models on predicting subject-predicate-object triplets while mitigating long-tailed training biases. It probes zero-shot generalization and graph-level semantic coherence through sentence-to-graph retrieval. Use when the user wants to benchmark on Visual Genome (VG), MS-COCO Caption (VG Overlap), or asks about evaluating this task. Reports mR@K.
Evaluates a model's ability to modify a source scene graph into a target scene graph conditioned on a natural language query. It probes incremental structure expansion, joint node-edge prediction, and the preservation of unmodified graph components during editing. Use when the user wants to benchmark on User Generated, MSCOCO, GCC, RSICD, or asks about evaluating this task. Reports Graph-level accuracy.
Evaluates text-to-3D indoor scene generation systems on their ability to produce dense, physically plausible, and prompt-faithful environments. It probes both visual realism and simulation-readiness, measuring collision-free layouts and stable physics properties required for robotics policy testing. Use when the user wants to benchmark on SceneSmith Prompt Corpus, or asks about evaluating this task. Reports Realism Win%.
This benchmark evaluates audio forensics models on their ability to discriminate between genuine recordings and audio manipulated via acoustic scene forgery using speech enhancement technologies. It specifically measures threshold-free equal error rate (EER) to assess how well models generalize to unseen attacks without relying on a fixed decision boundary. Use when the user wants to benchmark on SceneFake, or asks about evaluating this task. Reports EER.
Evaluates vision-language models on autonomous driving tasks, including scene understanding, spatial perception, and motion planning. It probes the models' ability to reason about driving scenarios, predict trajectories, and generalize across different geographic regions and traffic conventions. Use when the user wants to benchmark on ScenePilot-Bench, or asks about evaluating this task. Reports Overall Score.
Evaluates autonomous driving agents on their ability to navigate stochastic traffic scenarios while satisfying a hierarchical set of multi-objective specifications. It probes how well agents balance conflicting goals like collision avoidance, road compliance, passenger comfort, and progress under varying priority constraints. Use when the user wants to benchmark on ScenicRules Benchmark, or asks about evaluating this task. Reports Violation Score (VS).
Evaluates medical vision-language models on disease classification and structured visual reasoning. It probes the model's ability to localize lesions, generate clinically faithful chain-of-thought rationales, and produce accurate diagnostic classifications grounded in visual evidence. Use when the user wants to benchmark on S-Chain, or asks about evaluating this task. Reports Accuracy.
Evaluates zero-shot dialogue state tracking across single and multi-domain conversations. It measures the model's ability to predict intents, extract slot values, and maintain accurate dialogue states over long contexts without prior exposure to unseen service domains. Use when the user wants to benchmark on Schema-Guided Dialogue (DSTC8 Track 4), or asks about evaluating this task. Reports Joint Goal Accuracy.
Evaluates the ability of language models to extract structured information from heterogeneous tables (text, LaTeX, HTML, CSV, XML) using only a human-authored JSON schema as supervision. It probes schema-driven information extraction, testing attribute prediction accuracy across diverse domains and input formats without domain-specific labeled data. Use when the user wants to benchmark on MlTables, ChemTables, DisCoMat, SWDE, or asks about evaluating this task. Reports Table-F1.
Tests natural language interfaces for querying scholarly knowledge graphs (DBLP, ORKG) and hybrid multi-source QA, evaluating question-to-SPARQL translation and answer generation accuracy. Use when the user wants to benchmark on Scholarly QALD, or asks about evaluating this task. Reports Exact Match.
Evaluates LLMs' ability to synthesize scientific literature by answering open-ended, multi-domain questions using retrieved papers. It probes long-form generation, factual correctness, citation accuracy, and content quality/organization across single- and multi-paper retrieval setups. Use when the user wants to benchmark on ScholarQABench, or asks about evaluating this task. Reports Citation F1.
Evaluates an LLM's ability to predict human scholarly writing intentions from a LaTeX draft and to iteratively edit the draft according to those intentions. It measures lexical diversity, topic consistency, and intention coverage across a 100-iteration self-writing process. Use when the user wants to benchmark on SCHOLAWRITE, or asks about evaluating this task. Reports intention coverage.
Evaluates a model's ability to perform cross-disciplinary scientific verification by judging the correctness or equivalence of proposed answers to scientific problems. It probes domain-specific logical reasoning, handling of complex mathematical/scientific transformations, and robustness to prompt variations. Use when the user wants to benchmark on SCI-VerifyBench, or asks about evaluating this task. Reports Accuracy.
Evaluates a multi-agent LLM system's ability to solve high-level, cross-disciplinary scientific problems at Olympiad and frontier-exam levels. It probes formal reasoning, proof generation, symbolic derivation, and chemical modeling through adaptive routing and self-verification. Use when the user wants to benchmark on IMO 2025, IMC 2025, IPhO 2024, IPhO 2025, CPhO 2025, IChO 2025, HLE, or asks about evaluating this task. Reports Olympiad Scoring.
This benchmark evaluates large language models' ability to solve college-level scientific problems across mathematics, chemistry, and physics. It probes multi-step quantitative reasoning, unit conversion, physical derivations, and the efficacy of prompting strategies and external computational tools. Use when the user wants to benchmark on SciBench, or asks about evaluating this task. Reports accuracy.
Evaluates multimodal models' ability to verify scientific claims by classifying them as Supported or Refuted based on cross-modal evidence (tables or figures). It probes visual reasoning, table parsing, and resistance to dataset biases or superficial shortcuts. Use when the user wants to benchmark on SciClaimEval, or asks about evaluating this task. Reports macro-F1.
Evaluates cross-modality scientific information extraction by jointly predicting named entities, result entities, and relations from both full-text paragraphs and scientific tables. It probes a model's ability to handle long documents, align entities across modalities, and generalize across different scientific domains. Use when the user wants to benchmark on ScICM, or asks about evaluating this task. Reports F1.
Probes large language models' ability to perform scientific reasoning, domain-specific knowledge recall, and code synthesis on real-world research problems. It evaluates performance on both decomposed subproblems and full main problems under varying conditions of background knowledge and context carry-over. Use when the user wants to benchmark on SciCode, or asks about evaluating this task. Reports pass@1.
Evaluates the quality of sentence alignment and machine translation performance on a trilingual scientific article corpus. It probes cross-lingual translation accuracy and structural alignment precision in a specialized academic domain. Use when the user wants to benchmark on Scielo Parallel Corpus, or asks about evaluating this task. Reports BLEU.
This benchmark evaluates systems on mention-level keyphrase extraction and semantic relation extraction from scientific publications. It probes the model's ability to identify, classify, and link keyphrases across different scientific domains using exact match criteria. Use when the user wants to benchmark on ScienceIE, or asks about evaluating this task. Reports F1-score.
Evaluates multimodal reasoning and scientific question answering by requiring models to process questions, images, and context to select correct multiple-choice answers. It also probes chain-of-thought reasoning capabilities by measuring the quality of generated explanations and lectures. Use when the user wants to benchmark on ScienceQA, or asks about evaluating this task. Reports accuracy.
Evaluates an agent's ability to perform procedural scientific reasoning and navigation within an interactive text-based environment. It probes whether models can execute multi-step experiments (e.g., building circuits, measuring temperatures) rather than just retrieving static facts. Use when the user wants to benchmark on ScienceWorld, or asks about evaluating this task. Reports average_score.
Evaluates static (Word2Vec, FastText) and transformer-based (SciBERT, RoBERTa) embedding models on scientific text using intrinsic (word/sentence similarity) and extrinsic (NER, document classification) tasks. Probes the impact of domain-specific pretraining and sub-word tokenization on representation quality and downstream performance. Use when the user wants to benchmark on UNMSRS, SemEval, Clinical STS 2018, Clinical STS 2019, Conll 2003, CHEMDNER, SciERC, Reuters 12, BioChem 8, or asks ab...
Evaluates a model's ability to perform high-level visual reasoning and domain-specific knowledge grounding on scientific figures within a multiple-choice question answering setting. It specifically probes whether models can resist choice-induced prior bias where text-only answer options incorrectly steer predictions away from visually supported ground truth. Use when the user wants to benchmark on MAC, SciFIBench, MMSci, or asks about evaluating this task. Reports Accuracy.
Evaluates the ability of LLMs to generate novel, feasible, and effective scientific research ideas given a research question. Probes open-ended scientific reasoning and ideation quality under compute-matched inference budgets. Use when the user wants to benchmark on ICLR 2024 & NeurIPS 2025, or asks about evaluating this task. Reports Absolute Novelty.