
Claude Skills by qhjqhj00
github.com/qhjqhj00Evaluates multimodal LLMs' ability to perform integrated agentic reasoning and critical assessment over scientific papers. It probes capabilities in multimodal grounding, experimental interpretation, cross-source evidence synthesis via tool use, and critical evaluation of research claims. Use when the user wants to benchmark on PaperMind, or asks about evaluating this task. Reports F1 score.
Evaluates the accuracy and interaction efficiency of LLM agents performing multi-turn tool-use for scientific paper question-answering. It probes the model's ability to plan, execute tool calls, and extract answers from complex academic documents without excessive interaction or getting stuck in loops. Use when the user wants to benchmark on AirQA-Real, SciDQA, or asks about evaluating this task. Reports Avg., I-Avg.
Evaluates how well a language model's generated responses align with a specific individual's personality traits across the Big Five and Dark Triad dimensions. It measures the distance between the model's predicted personality profile and the target profile, with lower scores indicating better alignment. Use when the user wants to benchmark on PAPI, or asks about evaluating this task. Reports Aligned Score.
Evaluates a privacy-preserving LLM delegation pipeline that sanitizes user queries before sending them to a remote API model. It measures the trade-off between maintaining response quality and minimizing personally identifiable information (PII) leakage in the sanitized prompts. Use when the user wants to benchmark on PUPA-TNB, or asks about evaluating this task. Reports QUAL.
Evaluates how different deep learning model architectures and hyperparameters affect hardware performance across TPU, GPU, and CPU platforms. It probes the interaction between model attributes (size, type, batch size) and hardware bottlenecks like memory bandwidth, compute utilization, and data infeed overhead. Use when the user wants to benchmark on ParaDnn, or asks about evaluating this task. Reports performance.
Evaluates how well speech-to-speech models adapt to and reflect paralinguistic styles (age, emotion, gender, sarcasm) in spoken responses. It measures both content appropriateness and stylistic alignment against ground-truth or human-annotated references. Use when the user wants to benchmark on ParaS2SBench, IEMOCAP, MELD, or asks about evaluating this task. Reports ParaS2SBench score.
This evaluation probes a model's ability to synthesize interpretable surrogate models (decision diagrams) that explicitly trade off prediction fidelity against structural simplicity. It measures how well a method can navigate the Pareto front between accuracy and explainability without collapsing them into a single weighted objective. Use when the user wants to benchmark on Airplane Perception Module (AP), Bank Loan Predictor (BL), Theorem Prover Solvability Predictor (TP), or asks about eval...
Evaluates multilingual and multi-cultural LLM performance across 10 Indic languages using culturally nuanced prompts. It measures model quality via pairwise comparisons (Elo ratings) and direct assessment scores, while also analyzing human-LLM evaluator agreement and various biases (position, verbosity, self-bias). Use when the user wants to benchmark on PARIKSHA, or asks about evaluating this task. Reports Elo rating, Direct Assessment score.
Evaluates autonomous cleaning robots' ability to navigate public park pathways, perceive and collect diverse litter types, avoid obstacles, and operate within strict physical and safety constraints. Use when the user wants to benchmark on Park Cleaning Benchmark, or asks about evaluating this task. Reports collected_items_weight_or_count.
This benchmark evaluates the ability of deep learning models to perform binary semantic segmentation of parking lots from satellite imagery. It specifically probes the model's capacity to generalize across different geographic locations and to leverage near-infrared (NIR) spectral data for improved contrast against vegetation and built environments. Use when the user wants to benchmark on ParkSeg12k, or asks about evaluating this task. Reports mIoU.
Measures the computational, network, and database performance overhead of running containerized workloads inside an AMD SEV-SNP enclave with Parma's attested execution policies compared to a baseline outside the enclave. Use when the user wants to benchmark on nginx (wrk2), redis (redis-benchmark), SPEC2017 intspeed, NVIDIA Triton Inference Server (perf-analyzer), or asks about evaluating this task. Reports performance_overhead.
This benchmark evaluates an LLM's ability to accurately translate SQL queries across different database systems and dialects. It probes the model's capacity to handle system-specific syntax, semantic equivalence, and edge-case safeguards without relying on superficial string matching. Use when the user wants to benchmark on PARROT, or asks about evaluating this task. Reports Acc_EX.
Evaluates the multilingual visual-language understanding capabilities of multimodal large language models (MLLMs) across six languages (English, Chinese, Portuguese, Arabic, Turkish, Russian). It probes how well models align visual features with non-English textual instructions and handle cross-lingual multimodal tasks without relying on naive translation. Use when the user wants to benchmark on MMMB, MMBench, or asks about evaluating this task. Reports Accuracy.
Evaluates Persian language understanding across six distinct NLU tasks, including reading comprehension, textual entailment, sentiment analysis, and machine translation. It measures how well pre-trained monolingual and multilingual models perform on native-speaker annotated Persian data compared to human baselines. Use when the user wants to benchmark on ParsiNLU, or asks about evaluating this task. Reports F1, Accuracy.
This benchmark probes a robot policy's ability to follow fine-grained, part-level natural language instructions for long-horizon manipulation. It specifically tests zero-shot task decomposition, 3D part grounding, and multi-step planning under varying object, part, and task generalization conditions. Use when the user wants to benchmark on PartInstruct, or asks about evaluating this task. Reports success.
Evaluates the transferability and pretraining quality of vision models trained on synthetic domain-specific datasets compared to manually curated and general-domain datasets. It probes the model's ability to generalize to fine-grained classification and object detection tasks within specific domains like birds and food. Use when the user wants to benchmark on CUB-200-2011, NABirds, iNatbirds, Food-101, FoodX-251, Food-2K, or asks about evaluating this task. Reports Top-1 k-NN accuracy.
Evaluates an object detection model's ability to localize and classify objects within images. It measures how well the system predicts bounding boxes and assigns correct class labels across multiple object categories. Use when the user wants to benchmark on PASCAL VOC 2007, or asks about evaluating this task. Reports mAP.
Tests agricultural land cover mapping and crop-type classification by evaluating models on high-resolution satellite imagery combined with optical and radar time series. Use when the user wants to benchmark on PASTIS-HD, or asks about evaluating this task. Reports macro-averaged F1-score.
Evaluates large language models' ability to answer present-anchored temporal questions that require up-to-date world knowledge and multi-hop reasoning, such as identifying the current holder of a position or the previous president. It specifically probes performance degradation due to knowledge obsolescence and complex temporal relations. Use when the user wants to benchmark on PAT-Questions, or asks about evaluating this task. Reports exact-match accuracy (EM).
Evaluates a model's ability to ignore out-of-context patches (patch selectivity) and maintain classification accuracy under simulated occlusion and spatial permutation attacks. Use when the user wants to benchmark on ImageNet-1K val, SMD, NVD, ROD, or asks about evaluating this task. Reports Top-1 accuracy.
This evaluation probes a model's ability to generate clinically accurate diagnostic captions from histopathological image patches. It specifically tests the model's capacity to capture subtype-specific terminology and overall caption fluency using standard and custom n-gram overlap metrics. Use when the user wants to benchmark on PatchGastricADC22, or asks about evaluating this task. Reports BLEU@4.
This benchmark evaluates the quality of generated patent claims against expert-annotated reference claims across five dimensions: feature completeness, conceptual clarity, terminology consistency, logical linkage, and overall quality. It probes a model's ability to capture patent-specific linguistic precision, legal formality, and structural requirements rather than just surface-level text overlap. Use when the user wants to benchmark on Patent-CE, or asks about evaluating this task. Reports ...
Evaluates patent text embedding models across 15 diverse tasks including symmetric/asymmetric retrieval, classification, paraphrase detection, and clustering. It specifically probes domain-specific challenges like cross-domain retrieval, fragment-to-document matching, and temporal citation dynamics. Use when the user wants to benchmark on PatenTEB, or asks about evaluating this task. Reports NDCG@10, Macro-F1, Pearson r, V-measure.
Evaluates a model's ability to make predictions while removing the influence of a sensitive attribute along specific causal pathways, balancing predictive accuracy with path-specific counterfactual fairness constraints. Use when the user wants to benchmark on Berkeley Admission Dataset, UCI Adult Dataset, UCI German Credit Dataset, or asks about evaluating this task. Reports fair accuracy.
Evaluates a multimodal chatbot's ability to interpret real-world pathology images (H&E and IHC) and integrate clinical context to produce accurate diagnoses, terminology, and multimodal reasoning across four anatomical systems. Use when the user wants to benchmark on Pathology Clinical Q&A Dataset, or asks about evaluating this task. Reports diagnosis accuracy.
Evaluates the ability of machine learning and statistical models to forecast daily patient flows at urgent care clinics, specifically testing robustness to concept drift induced by pandemic disruptions using quasi-real-time proxy variables. Use when the user wants to benchmark on Clinic 1 & 2 patient flow data, or asks about evaluating this task. Reports MAPE.
Evaluates the persona fidelity, factual accuracy, and clinical plausibility of an LLM-based patient simulator in doctor-patient dialogues. It measures how well the model adheres to assigned patient profiles, maintains factual consistency, and handles out-of-profile questions plausibly. Use when the user wants to benchmark on PatientSim Profiles, or asks about evaluating this task. Reports Entail (%).
Evaluates the ability of program-by-example (PBE) systems to synthesize correct SQL queries from example input/output tables. It probes query generation accuracy, synthesis speed, and scalability to larger database schemas. Use when the user wants to benchmark on ase13, so-top, so-dev, so-rec, kaggle, or asks about evaluating this task. Reports solve_rate.
This benchmark evaluates a model's ability to identify paraphrases in sentence pairs that share high lexical overlap but differ in meaning due to word order and syntactic structure. It specifically probes sensitivity to non-local contextual information and adversarial word scrambling, revealing whether models rely on superficial word matching rather than true semantic understanding. Use when the user wants to benchmark on PAWS_QQP, PAWS_Wiki, or asks about evaluating this task. Reports classi...
PAWS-X evaluates a model's ability to identify paraphrases across multiple languages, specifically probing sensitivity to word order and syntactic structure under conditions of high lexical overlap. It measures how well models generalize cross-lingually when trained on machine-translated data versus zero-shot settings. Use when the user wants to benchmark on PAWS-X, or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to perform referring expression segmentation across five hierarchical levels of semantic complexity, from basic object recognition to fine-grained attribute binding, OCR-based disambiguation, spatial layout understanding, and relational interactions. It also stress-tests long-context generation and instance stability in crowded scenes with high object counts. Use when the user wants to benchmark on PBench, or asks about evaluating this task. Reports per-level perfo...
Evaluates embodied decision-making capabilities across three dimensions: perception (visual understanding), cognition (reasoning and task breakdown), and action (executing correct steps or decisions). It tests models in autonomous driving, domestic assistance, and game-playing environments. Use when the user wants to benchmark on PCA-EVAL, or asks about evaluating this task. Reports Action Score.
Evaluates a model's ability to classify six specific types of PCB manufacturing defects from cropped defect images. It probes the model's feature extraction and categorization capabilities on a specialized industrial computer vision dataset. Use when the user wants to benchmark on PCB Defect Dataset, or asks about evaluating this task. Reports average_precision_rate.
This benchmark evaluates the ability of docking tools, deep learning models, and meta-modeling ensembles to predict ligand-protein binding affinities. It probes how well different feature representations (physical scores, sequence-based DL outputs, physicochemical properties) generalize to unseen protein-ligand complexes. Use when the user wants to benchmark on PDBbind, or asks about evaluating this task. Reports Pearson correlation coefficient.
Predicts the binding affinity between a protein pocket and a ligand from 3D structural data. It probes the model's ability to quantify molecular interaction strength and generalize across protein sequence identities. Use when the user wants to benchmark on PDBBind, or asks about evaluating this task. Reports RMSE.
Evaluates how well 3D binding affinity models generalise to unseen proteins and novel ligands in low-data regimes. It uses a strict low-Tanimoto-similarity split of the PDBBind dataset to prevent data leakage and benchmark generalisation capabilities. Use when the user wants to benchmark on PDBBind, or asks about evaluating this task. Reports performance.
Evaluates a CNN's ability to classify chest X-ray images into three diagnostic categories: COVID-19, Normal, and Viral Pneumonia. It probes multi-scale feature extraction and robustness to class imbalance in medical imaging. Use when the user wants to benchmark on Custom benchmark dataset, or asks about evaluating this task. Reports Accuracy.
Evaluates open-source PDF information extraction tools across multiple content elements (metadata, references, tables, paragraphs, sections, etc.) on academic documents. It probes how well different tools handle layout-based segmentation, text extraction, and structural recognition in real-world academic PDFs. Use when the user wants to benchmark on DocBank, or asks about evaluating this task. Reports F1 score.
Evaluates the robustness of embedded feature selection methods (LASSO, Ridge, Elastic Net) against training data poisoning attacks in a PDF malware detection setting. It measures how injected malicious samples manipulate feature selection stability and degrade classification performance. Use when the user wants to benchmark on Contagio + Web Benign PDFs, or asks about evaluating this task. Reports classification error.
This evaluation probes the retrieval accuracy of RAG pipelines when processing financial PDFs, specifically testing how different PDF parsers, chunking strategies, and overlap percentages affect the retrieval of relevant pages for both narrative text and structured table queries. Use when the user wants to benchmark on FinanceBench, TableQuest, or asks about evaluating this task. Reports MRR.
Evaluates end-to-end question answering over PDF documents, probing parsing, retrieval, and reasoning capabilities across diverse document types, modalities, and complexity dimensions. Use when the user wants to benchmark on pdfQA, or asks about evaluating this task. Reports G-Eval correctness.
Evaluates learning-based models for PE malware family classification across image, binary, and disassembly input formats. It measures classification accuracy and probes model robustness under concept drift, alongside computational resource overhead. Use when the user wants to benchmark on BIG-15, Malimg, MalwareBazaar, MalwareDrift, or asks about evaluating this task. Reports Macro F1-score ($F1_{macro}$).
Evaluates a separation-of-powers AI agent architecture (PEA) on its ability to prevent unauthorized actions, detect goal drift, and identify implicit coercion in adversarial inputs. Use when the user wants to benchmark on Attack Corpus, Drift Dataset, Coercion Dataset, or asks about evaluating this task. Reports Bypass Rate, Attack Success Rate (ASR).
Compute the PeakSignalNoiseRatio metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute PeakSignalNoiseRatio, or asks how to score with PeakSignalNoiseRatio.
Compute the PeakSignalNoiseRatioWithBlockedEffect metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute PeakSignalNoiseRatioWithBlockedEffect, or asks how to score with PeakSignalNoiseRatioWithBlockedEffect.
Evaluates medium-term weather forecasting capability on a spherical grid by predicting atmospheric variables up to 10 days ahead. It probes the model's ability to capture spatial and temporal dynamics without grid-induced resolution biases. Use when the user wants to benchmark on ERA5-lite, or asks about evaluating this task. Reports ACC.
Evaluates large vision-language models' ability to understand and generate culturally-aware Arabic content across multiple reasoning-centric question types. It probes hypothesis formation, comparative analysis, chronological reasoning, and explicit cultural grounding in both closed-form and open-ended multimodal tasks. Use when the user wants to benchmark on PeARL, or asks about evaluating this task. Reports relaxed-match accuracy (ACC).
Probes the ability of an automated metric to correlate with human judgments of translation quality. Specifically, it tests how well a model's predicted segment-level scores align with Direct Assessment (DA) human scores across multiple language pairs. Use when the user has predictions and gold and needs to compute Pearson correlation coefficient.
Evaluates the validity of automatic machine translation metrics by measuring their linear correlation with human Direct Assessment (DA) scores at both segment and system levels. It emphasizes rigorous validation protocols, including adaptive sample size determination for human judgments and statistical significance testing to compare metric performance. Use when the user has predictions and gold and needs to compute pearson-correlation.
Compute the PearsonCorrCoef metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute PearsonCorrCoef, or asks how to score with PearsonCorrCoef.