
Claude Skills by qhjqhj00
github.com/qhjqhj00Evaluates deep learning architectures for multi-label classification of 12-lead ECG recordings into 23 diagnostic categories. Probes the trade-off between local morphological feature extraction and sequential temporal modeling under severe class imbalance. Use when the user wants to benchmark on PTB-XL, or asks about evaluating this task. Reports Macro AUROC.
Evaluates the robustness of sentence embedding models to token-level variations by measuring performance degradation when test instances are stochastically paraphrased at evaluation time. It probes whether models maintain semantic invariance under LLM-generated paraphrases that preserve meaning but alter surface form. Use when the user wants to benchmark on MTEB (STS & Non-STS tasks), or asks about evaluating this task. Reports Spearman’s rank correlation.
Evaluates multimodal large language models on 3D scene understanding tasks, including visual grounding, dense captioning, and spatial/situated question answering. It specifically probes how different 3D token structures (point-based vs. video-based) and feature fusion strategies impact performance on indoor RGB-D scans. Use when the user wants to benchmark on ScanRefer, Multi3DRefer, Scan2Cap, ScanQA, SQA3D, or asks about evaluating this task. Reports NS.
This benchmark evaluates multimodal large language models' ability to interpret synthetic data visualizations. It probes their capacity to extract quantitative features, identify statistical properties, and answer specific questions about plots like histograms, time series, boxplots, and violin plots without relying on real-world data contamination. Use when the user wants to benchmark on PUB Synthetic Plot Dataset, or asks about evaluating this task. Reports overall score.
This benchmark evaluates automated fact-checking systems on public health claims, measuring both their ability to predict the veracity of a claim and the quality of the generated explanations. It probes domain-specific reasoning, extractive-abstractive explanation generation, and formal coherence properties like consistency and relevance. Use when the user wants to benchmark on PUBHEALTH, or asks about evaluating this task. Reports macroF1.
Evaluates the ability of object detection models to identify and localize document layout elements (text, title, list, table, figure) in scientific PDF pages. It also probes transfer learning capabilities by fine-tuning on out-of-domain documents and table detection tasks. Use when the user wants to benchmark on PubLayNet, or asks about evaluating this task. Reports MAP @ IOU [0.50:0.95].
Evaluates the ability of retrieval and reranking models to surface relevant appellate brief paragraphs for public defender search queries. It probes domain-specific adaptation, query expansion strategies, and the impact of synthetic data generation on legal information retrieval. Use when the user wants to benchmark on PD Dataset, NJ OPD Dataset, BarExam-QA, LePaRD, or asks about evaluating this task. Reports recall@5.
Evaluates a dynamical systems model of scientific publishing by simulating the interplay between AI-accelerated manuscript writing and peer review throughput. It measures how queue pressure drives AI adoption in review, degrades verification quality, and ultimately impacts net scientific knowledge output over a 20-year horizon. Use when the user wants to benchmark on NeurIPS main track submissions, ICLR submissions, arXiv monthly submissions, bioRxiv annual preprints, or asks about evaluating...
This evaluation probes a model's ability to verify scientific claims in a three-way classification setting that includes uncertainty abstention. It measures performance on Supported, Refuted, and NEI (Not Enough Information) labels, testing the model's capacity to avoid overconfident predictions when evidence is insufficient or conflicting. Use when the user wants to benchmark on PubMedFact1k, or asks about evaluating this task. Reports Macro F1.
This benchmark evaluates a model's ability to perform biomedical research question answering by reasoning over structured scientific abstracts. It requires models to infer yes/no/maybe answers to questions derived from paper titles using only the non-conclusion sections of the abstract, without access to the final conclusion. Use when the user wants to benchmark on PubMedQA (PQA-L), or asks about evaluating this task. Reports accuracy.
Evaluates vision-language and specialized models on page-level and document-level table structure recognition, requiring them to extract hierarchical table structures from full pages or multi-page documents. It also probes cross-page table continuation prediction by testing whether models can identify when a table spans two contiguous pages. Use when the user wants to benchmark on PubTables-v2, or asks about evaluating this task. Reports GriTS_Top.
Evaluates models on imputing missing values in pulsative physiological signals (ECG and PPG) under realistic, data-driven missingness patterns. It further assesses clinical utility by measuring downstream performance on heartbeat detection and cardiac classification tasks. Use when the user wants to benchmark on ECG, PPG, or asks about evaluating this task. Reports MSE, F1 Score.
This benchmark evaluates multimodal physiological reasoning by testing whether large language models can accurately answer closed-ended questions conditioned on raw photoplethysmography (PPG) waveforms. It probes the model's ability to align continuous biosignal representations with natural language queries across diverse physiological domains and assesses cross-dataset generalization beyond the training distribution. Use when the user wants to benchmark on PulseLM, or asks about evaluating t...
This benchmark evaluates the ability of Positive Unlabeled (PU) learning algorithms to accurately estimate the true proportion of positive examples ($\alpha$) within an unlabeled dataset, particularly when selection bias violates the SCAR assumption. It also probes the robustness of downstream classification performance and probability calibration under varying degrees of class imbalance and structured selection bias. Use when the user wants to benchmark on Synthetic SCAR, Synthetic SNAR, UCI...
Evaluates pixel-level segmentation capability for distinguishing five histopathological tissue classes (tumour, stroma, necrosis, blood vessels, epidermis) in melanoma H&E images. Use when the user wants to benchmark on PUMA Challenge dataset, or asks about evaluating this task. Reports Dice score.
Evaluates a model's ability to restore punctuation marks (commas, periods, question marks) in English text. It specifically probes robustness on both manually transcribed transcripts and ASR-generated transcripts, ignoring non-punctuation tokens during evaluation. Use when the user wants to benchmark on IWSLT2011, or asks about evaluating this task. Reports F1-score.
Evaluates video-language models on long-form repetition counting and temporal reasoning. It probes whether models can accurately track state changes and count actions across extended video clips, revealing weaknesses in spatio-temporal tracking compared to supervised baselines. Use when the user wants to benchmark on PushupBench, or asks about evaluating this task. Reports Exact Match.
Evaluates medium-range global weather forecasting capability using an autoregressive convolutional network. It measures prediction accuracy over 10-day horizons at 6-hour intervals against reanalysis ground truth. Use when the user wants to benchmark on ERA5, or asks about evaluating this task. Reports MAE.
Evaluates machine learning models for predicting photovoltaic (PV) output power in smart buildings across multiple time frames (30 min, 1 hour, 4 hours) using location-specific meteorological and temporal features. Use when the user wants to benchmark on Tartu, Estonia PV dataset, or asks about evaluating this task. Reports MAPE.
Evaluates a model's ability to generate temporal scene graphs where nodes are grounded with pixel-level panoptic segmentation masks instead of bounding boxes, capturing non-rigid objects, backgrounds, and fine-grained interactions in dynamic videos. Use when the user wants to benchmark on PVSG, or asks about evaluating this task. Reports R/mR@20.
Evaluates the numerical correctness, stability, and parallel performance (CPU/GPU strong scaling and distributed weak scaling) of a matrix-free finite difference geodynamic solver. Use when the user wants to benchmark on Pyroclast Stokes & Advection Benchmarks, or asks about evaluating this task. Reports parallel speedup.
Evaluates machine learning models across clinical trial tasks including patient and trial outcome prediction, trial search, and patient simulation, using standardized tabular and sequential data formats. Use when the user wants to benchmark on Tabular Clinical Trial Patient Datasets, TOP Benchmark, Trial Similarity Dataset, Sequential Trial Patient Data, or asks about evaluating this task. Reports AUROC.
Evaluates open-weight multimodal agentic models on visual search, multimodal mathematical reasoning, multi-turn tool use, and video spatial reasoning. It probes the model's ability to dynamically construct context, invoke tools, and perform long-horizon reasoning with high visual token efficiency. Use when the user wants to benchmark on V*, HRBench-4K, HRBench-8K, MathVerse, MathVision, WeMath, DynaMath, TIR-Bench, VSI-Bench, or asks about evaluating this task. Reports accuracy.
Evaluates ranked retrieval lists using graded relevance assessments, balancing precision and cumulative gain while penalizing lower-ranked relevant documents. Use when the user has predictions and gold and needs to compute Q-measure.
Evaluates the capability of retrieval-augmented generation systems to answer complex, multi-hop, and long-form questions by iteratively retrieving, structuring, and accumulating evidence from documents. Use when the user wants to benchmark on StrategyQA, ASQA, NQ, 2WikiMultiHopQA, HotpotQA, or asks about evaluating this task. Reports EM, F1, ACC.
Evaluates cognition-based hallucination by testing whether LVLMs can leverage world knowledge stored in the LLM to answer entity and relation questions grounded in images. Use when the user wants to benchmark on QA-FB15K, or asks about evaluating this task. Reports Acc.
Evaluates the quality of generated question-answer pairs in an agricultural domain context. It probes a model's ability to produce relevant, accurate, diverse, and fluent Q&A content under varying context conditions. Use when the user wants to benchmark on Agricultural Q&A dataset, or asks about evaluating this task. Reports Relevance.
Evaluates the ability of retrieval models to rank relevant sentences or documents highest for a given question. It probes lexical and semantic matching capabilities in question answering contexts, testing both single-model retrieval and multi-model fusion strategies. Use when the user wants to benchmark on ReQA SQuAD, ReQA NQ, MTEB QA Subset, iapp-wiki-qa-squad, or asks about evaluating this task. Reports MRR.
This benchmark evaluates how well LLM-generated translations preserve the scientific content of original papers. It measures translation fidelity by testing whether a reading model can accurately answer comprehension questions derived from the source text, using only the translated version as context. Use when the user wants to benchmark on Science Across Languages QA Benchmark, or asks about evaluating this task. Reports quiz accuracy.
Evaluates attribute and relation hallucination by asking LVLMs to identify object properties and inter-object relationships in images, probing fine-grained visual understanding beyond basic object detection. Use when the user wants to benchmark on QA-VisualGenome, or asks about evaluating this task. Reports Acc.
Evaluates the ability of a question-answering system to extract missing objects from partially filled relational tuples by searching through a large document corpus. It probes schema-aware information extraction and the model's capacity to leverage relational coherence across multiple questions. Use when the user wants to benchmark on QA-ZRE, or asks about evaluating this task. Reports Exact Match (EM).
Evaluates document-level information extraction by framing it as a question answering task. It probes a model's ability to extract cross-sentence relation triples from large documents using entity-relation queries and knowledge base alignment. Use when the user wants to benchmark on QA4IE, or asks about evaluating this task. Reports Exact Match (EM), F1-score.
Evaluates the factual consistency and information coverage of query-focused summaries for less-resourced languages without reference texts. It probes whether LLMs can preserve source details and align with user intent by comparing answers generated from the source versus answers generated from the candidate summary. Use when the user wants to benchmark on Slovene News Summarization Corpus (MOCHA translation), or asks about evaluating this task. Reports QuestEval F1.
This benchmark evaluates a model's ability to read and reason over full academic research papers to answer information-seeking questions and to identify the specific paragraphs that contain the supporting evidence. It probes document-level comprehension, multi-paragraph reasoning, and handling of diverse answer types including extractive spans, abstractive summaries, yes/no, and unanswerable cases. Use when the user wants to benchmark on QASPER, or asks about evaluating this task. Reports Ans...
Evaluates deep learning models for COVID-19 infected region segmentation and binary detection on chest X-ray images. It probes the model's ability to localize pathological regions at the pixel level and classify whole images as positive or negative for infection. Use when the user wants to benchmark on QaTa-COV19, or asks about evaluating this task. Reports F1-Score.
Evaluates sequential recommendation models on their ability to predict the next item in a user's interaction history using multimodal item features. It probes cross-domain generalization, the effectiveness of discrete semantic tokenization, and robustness in sparse interaction scenarios. Use when the user wants to benchmark on Amazon Product Reviews (Instruments, Arts, Games), or asks about evaluating this task. Reports HR@K, NDCG@K.
Evaluates the impact of ultra-low precision weight quantization (down to 2 bits) on BERTBASE performance across standard NLP tasks. It compares Hessian-guided mixed and group-wise quantization strategies against direct quantization baselines to measure accuracy retention versus model compression. Use when the user wants to benchmark on SST-2, MNLI, CoNLL-03, SQuAD, or asks about evaluating this task. Reports Acc.
Probes vision-language models' ability to interpret quantum calibration plots across diverse visual formats (1D traces, 2D maps, histograms) and perform structured scientific reasoning. It tests capabilities ranging from visual grounding and outcome classification to parameter extraction and operational calibration diagnosis, both in zero-shot and in-context learning settings. Use when the user wants to benchmark on QCalEval, or asks about evaluating this task. Reports accuracy.
Evaluates a hybrid quantum-classical convolutional neural network on MRI-based brain tumor detection. It probes the model's ability to classify medical images into binary (tumor vs. non-tumor) and multiclass (specific tumor types) categories under class imbalance and limited resolution constraints. Use when the user wants to benchmark on Brain MRI Tumor Dataset, or asks about evaluating this task. Reports accuracy.
Evaluates the ability of neural re-ranking models to effectively re-order a candidate set of documents based on complex query semantics and entity relationships. It probes fine-grained semantic matching, entity-aware attention, and late aggregation capabilities in information retrieval tasks across news and complex answer domains. Use when the user wants to benchmark on CODEC, TREC Complex Answer Retrieval (CAR) 2017, TREC Robust 2004, TREC News 2021, TREC Core 2018, or asks about evaluating ...
Probes the practical usability and impact of word-level quality estimation highlights on professional translators' post-editing efficiency, accuracy, and workflow. It measures how different highlight modalities (oracle, supervised, unsupervised, none) affect editing effort, productivity, and final translation quality in real-world domain-specific settings. Use when the user wants to benchmark on QE4PE, or asks about evaluating this task. Reports ESA score.
Evaluates the ability of generative language models to produce paragraph-level questions conditioned on a target answer and a context sentence. It probes domain adaptability and multilingual generalization across diverse extractive QA datasets. Use when the user wants to benchmark on SQuAD v1.1, SQuADShifts, SubjQA, Multilingual QA (JAQuAD, GerQuAD, SberQuAD, KorQuAD, FQuAD, Spanish SQuAD, Italian SQuAD), or asks about evaluating this task. Reports automatic evaluation metrics.
Evaluates the quality of generated questions across seven dimensions: fluency, clarity, conciseness, relevance, consistency, answerability, and answer consistency. It measures how well automatic metrics and LLMs align with human judgments on these dimensions. Use when the user wants to benchmark on QGEval, or asks about evaluating this task. Reports Pearson correlation.
This evaluation probes a unified vision-language model's ability to perform end-to-end document intelligence, including specialized OCR, general text recognition, document understanding, and key information extraction across diverse document types and multilingual scenarios. Use when the user wants to benchmark on Omni-Doc-Bench v1.5, OLMOCRBench, OCRBench, DocVQA, ChartQA, Nanonets KIE, or asks about evaluating this task. Reports normalized accuracy (0-100).
Evaluates multimodal information retrieval systems across search, recommendation, and deep query answering (DQA) tasks using real-world APP-level user sessions. It probes a model's ability to rank heterogeneous content (text, images, videos) and generate accurate answers augmented by retrieved documents. Use when the user wants to benchmark on Qilin, or asks about evaluating this task. Reports MRR@10.
Compute qlemesle/parapluie via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of qlemesle/parapluie.
Evaluates the ability of a tokenization framework to generate valid 3D molecular structures and predict quantum mechanical properties. It probes structural validity, geometric plausibility, conditional controllability, and property prediction accuracy on organic molecules. Use when the user wants to benchmark on QM9, or asks about evaluating this task. Reports Mean Absolute Error (MAE).
Assesses a model's ability to predict scalar quantum chemical properties from molecular structures. It evaluates accuracy on five key electronic and vibrational targets using mean absolute error and average ranking. Use when the user wants to benchmark on QM9, or asks about evaluating this task. Reports MAE.
This benchmark evaluates hybrid quantum-classical neural networks (QMLP and QCNN) on malware classification tasks. It probes the ability of NISQ-era quantum circuits to encode high-dimensional static feature vectors and learn discriminative patterns for binary and multiclass malicious software detection. Use when the user wants to benchmark on API-Graph, EMBER-Domain, AZ-Domain, EMBER-Class, AZ-Class, or asks about evaluating this task. Reports Accuracy.
Evaluates an agent's ability to distill and synthesize key information from high-noise, multi-turn meeting transcripts based on specific user queries. It probes the model's capacity to maintain global context while filtering irrelevant dialogue turns to produce a query-relevant summary. Use when the user wants to benchmark on QMSUM, or asks about evaluating this task. Reports ROUGE-1.