
Claude Skills by qhjqhj00
github.com/qhjqhj00Evaluates the effectiveness of query rewriting models in retrieving relevant documents from a corpus using vector, lexical, and multimodal retrieval systems. It measures how well rewritten queries match the intended source documents across text-only and unstructured visual document benchmarks. Use when the user wants to benchmark on MS MARCO v2.1 testset 1%, MTEB VIDORE V2 benchmark, In-house industrial data, or asks about evaluating this task. Reports NDCG@3.
Evaluates lightly-supervised representation learning and pattern-based extraction for named entity classification. It probes the model's ability to learn custom entity and pattern embeddings via bootstrapping, and to derive an interpretable global decision list for classification without using gold labels during training. Use when the user wants to benchmark on CoNLL-2003, Ontonotes, or asks about evaluating this task. Reports F1-score.
Evaluates large language models' ability to retrieve specific information and perform complex multi-point reasoning within long-context documents. It probes both information-sparse retrieval and information-dense reasoning (Ancestral Trace Challenge) across 32K and 128K token contexts. Use when the user wants to benchmark on NeedleBench, or asks about evaluating this task. Reports Overall.
Probes whether text embedding models suffer from 'negation blindness' by testing their ability to correctly identify paraphrases over negated counterparts, and measures the trade-off with semantic similarity correlation. Use when the user wants to benchmark on STSB, SemAntoNeg, or asks about evaluating this task. Reports Accuracy.
Compute the NegativePredictiveValue metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute NegativePredictiveValue, or asks how to score with NegativePredictiveValue.
Evaluates LLM instruction-following and reasoning capabilities under zero-shot and few-shot settings by appending psychologically grounded negative emotional stimuli to prompts. It probes task accuracy, complex reasoning on beyond-capability tasks, and the truthfulness and informativeness of generated responses. Use when the user wants to benchmark on Instruction Induction, BIG-Bench (curated subset), TruthfulQA, or asks about evaluating this task. Reports accuracy, normalized preferred metric.
Evaluates long-context mathematical reasoning and tool-integrated reasoning capabilities of language models on competition-style and open-domain advanced math problems. It probes symbolic precision, multi-step deduction, and the ability to leverage Python code execution for verification. Use when the user wants to benchmark on Comp-Math-24-25, HLE-Math, or asks about evaluating this task. Reports maj@k.
Evaluates a 12B vision-language model's capabilities across multimodal understanding, long-context reasoning, document/OCR processing, video comprehension, and pure text reasoning. It probes the model's ability to handle diverse visual inputs, follow instructions, and perform complex STEM and code reasoning under varying decoding and reasoning budget constraints. Use when the user wants to benchmark on MMBench V1.1, MMMU, OCRBench, DocVQA, LongVideoBench, MATH-500, GPQA-Diamond, or asks about...
This benchmark evaluates offline reinforcement learning algorithms on seven near real-world environments featuring time delays, external disturbances, safety constraints, and conservative data collection. It probes whether state-of-the-art offline RL methods can improve upon sub-optimal behavior policies without online exploration, highlighting their robustness to realistic dynamics and safety limits. Use when the user wants to benchmark on NeoRL-2, or asks about evaluating this task. Reports...
Evaluates encoder-based language models on core Nepali natural language understanding tasks, including named entity recognition, part-of-speech tagging, text classification, and categorical pair similarity. Use when the user wants to benchmark on Nep-gLUE, or asks about evaluating this task. Reports Nep-gLUE Score.
Evaluates decoder-based language models on abstractive text summarization for Nepali news articles, testing generation quality and context handling. Use when the user wants to benchmark on Nepali Summarization Dataset (Bhandari 2024), or asks about evaluating this task. Reports ROUGE-L.
Evaluates named entity recognition (NER) models under data-scarce conditions, specifically low-resource settings with no human-annotated training labels and few-shot settings with minimal labeled examples. It probes the model's ability to identify and classify entity types in text using automatically generated pseudo-dictionaries and weak supervision. Use when the user wants to benchmark on CoNLL-2003, Wikigold, WNUT-16, NCBI-disease, BC5CDR, CHEMDNER, or asks about evaluating this task. Repo...
This evaluation protocol benchmarks Named Entity Recognition (NER) systems across diverse domains and entity type distributions. It measures how well different architectures (transformers, CRFs, LLMs) identify and classify named entity spans under exact-match conditions. Use when the user wants to benchmark on CoNLL-2003, OntoNotes, WNUT2017, FIN, BioNLP2004, NCBI Disease, BC5CDR, MITRestaurant, Few-NERD, MultiCoNER, or asks about evaluating this task. Reports Macro-averaged F1-score.
This benchmark probes gender bias and temporal drift in Named Entity Recognition (NER) systems. It measures whether models disproportionately misclassify female names compared to male names, and how this bias shifts across 139 years of U.S. census data. The evaluation specifically tests statistical parity in entity recognition under varying contextual templates. Use when the user wants to benchmark on U.S. Census Names (1880-2018), or asks about evaluating this task. Reports error rate.
Evaluates the data-efficiency and labor-cost effectiveness of a trigger-enhanced Named Entity Recognition model compared to a standard baseline. It probes how well the model generalizes when trained on varying fractions of labeled sentences and trigger-annotated data. Use when the user wants to benchmark on CoNLL2003, BC5CDR, or asks about evaluating this task. Reports F1.
Evaluates models on nested named entity recognition and relation extraction in Russian. It probes the ability to identify overlapping/contained entities and classify semantic relations between them, including cross-sentence and nested relations. Use when the user wants to benchmark on NEREL, or asks about evaluating this task. Reports F1.
This evaluation protocol measures the storage efficiency, rendering quality, and inference speed of compressed neural radiance field models. It probes how well a learned codebook preserves high-frequency visual details and accelerates real-time rendering compared to uncompressed baselines. Use when the user wants to benchmark on Synthetic-NeRF, LLFF, or asks about evaluating this task. Reports PSNR.
Evaluates the ability of a neural radiance field to synthesize photorealistic novel views of 3D scenes from a sparse set of input images. It probes geometric reconstruction fidelity, appearance modeling (including non-Lambertian materials), and multi-view consistency across synthetic and real-world captures. Use when the user wants to benchmark on Diffuse Synthetic 360° (DeepVoxels), Realistic Synthetic 360°, Real Forward-Facing, or asks about evaluating this task. Reports PSNR.
Evaluates the ability to reconstruct 3D geometry and synthesize novel views of transparent and specular objects from monocular RGB images and silhouettes. It probes physically accurate light path simulation, including refraction, reflection, and Fresnel effects, using a differentiable rendering framework. Use when the user wants to benchmark on Blender Synthetic Dataset, or asks about evaluating this task. Reports Chamfer Distance (CD).
Evaluates the inference accuracy, computational cost, memory footprint, and switching overhead of a multi-capacity deep learning architecture compared to independent baseline models across six mobile vision classification tasks. It also benchmarks a resource-aware scheduler's ability to maintain accuracy and frame rate under dynamic runtime memory constraints. Use when the user wants to benchmark on CIFAR-10, ImageNet-50, ImageNet-100, GTSRB, Adience-Gender, Places-32, or asks about evaluatin...
Evaluates a model's ability to identify and classify named entities that can overlap or be contained within other entities (nested NER) across multiple domains. It probes the model's span-level understanding and label assignment capabilities in complex textual contexts. Use when the user wants to benchmark on ACE2004, ACE2005, GENIA, KBP2017, or asks about evaluating this task. Reports F1.
This benchmark evaluates nested named entity recognition models on 19th-century Paris trade directories, testing their ability to extract hierarchical entities (e.g., addresses containing street names and numbers) and their robustness to OCR noise. It specifically probes how different sequence tagging formats (IO vs IOB2) and pre-training strategies affect span detection, hierarchical containment, and flat entity recognition. Use when the user wants to benchmark on Paris Trade Directories NER...
Evaluates zero-shot and few-shot named entity typing (NET) and recognition (NER) capabilities of pre-trained auto-regressive language models without fine-tuning. Probes reliance on memorized lexical patterns versus contextual generalization and tests robustness to noisy text and case variations. Use when the user wants to benchmark on CoNLL-2003, WNUT2017, MIT Movie, DBpedia, or asks about evaluating this task. Reports F1.
Evaluates the generalizability and detection performance of machine learning classifiers for network intrusion detection when using a standardized NetFlow feature set across multiple benchmark datasets. It probes whether a common feature representation improves cross-dataset model accuracy and reduces false alarms compared to proprietary or basic NetFlow features. Use when the user wants to benchmark on NF-UNSW-NB15-v2, NF-BoT-IoT-v2, NF-ToN-IoT-v2, NF-CSE-CIC-IDS2018-v2, NF-UQ-NIDS-v2, or as...
Evaluates a generative active adaptation framework for network intrusion detection under concept drift and class imbalance. It probes the model's ability to select informative samples, generate synthetic minority-class data, and maintain high detection performance across shifting temporal and spatial domains with limited labeling budgets. Use when the user wants to benchmark on CIC-IDS (2017/2018), UGR'16, or asks about evaluating this task. Reports F1-score.
Evaluates pre-trained LLMs' networking operations (NetOps) knowledge and reasoning across five technical sub-domains and two languages. It probes both multiple-choice comprehension and open-ended generation capabilities in a domain-specific context. Use when the user wants to benchmark on NetEval, or asks about evaluating this task. Reports accuracy.
Evaluates the ability of machine learning classifiers to detect network intrusions across different traffic datasets, with a focus on handling severe class imbalance and identifying rare attack types. Use when the user wants to benchmark on KDD-99, NSL-KDD, UNSW-NB15, or asks about evaluating this task. Reports Weighted F1-Score.
Evaluates neural cross-language and multilingual information retrieval systems on news collections in Chinese, Persian, and Russian. It probes a model's ability to rank relevant documents when queries are in English and documents are in different languages, as well as its capacity to unify rankings across multiple languages. Use when the user wants to benchmark on NeuCLIR 2023, or asks about evaluating this task. Reports nDCG@20.
Evaluates neural cross-language and multilingual information retrieval systems on news collections in Chinese, Persian, and Russian. It probes a model's ability to rank documents by relevance when queries are in English and documents are in different languages, or when searching across multiple languages simultaneously. Use when the user wants to benchmark on NeuCLIR 1 News Collection, or asks about evaluating this task. Reports nDCG@20.
Evaluates the ranking effectiveness of retrieval and reranking models across monolingual, cross-language, and multilingual information retrieval tasks. It specifically probes how well systems handle language mismatches and multilingual document collections without relying on simple keyword matching. Use when the user wants to benchmark on NeuCLIRBench, or asks about evaluating this task. Reports nDCG@20.
Evaluates the robustness of a run-time adapted neural beamforming system for joint speech dereverberation and denoising under varying acoustic conditions, including different numbers of speakers, reverberation times, and SNRs. Use when the user wants to benchmark on Simulated Librispeech+DEMAND, or asks about evaluating this task. Reports WER.
Evaluates large language models' ability to perform multi-step reasoning across diverse domains including mathematics, commonsense, and expert knowledge. It specifically probes the model's capacity to generate accurate solutions while minimizing computational cost (token usage) through dynamic reasoning path search. Use when the user wants to benchmark on AMC23, ARC-C, GPQA, GSM8K, or asks about evaluating this task. Reports Efficiency Metric ($\eta$).
Evaluates a model's ability to learn algorithmic rules (arithmetic, sequence transformation) from short training sequences and generalize them to arbitrarily long inputs without explicit programming or architectural changes. Use when the user wants to benchmark on Neural GPU Algorithmic Tasks, or asks about evaluating this task. Reports fully_correct_output_rate.
Neural-MedBench probes the clinical reasoning and multimodal synthesis capabilities of vision-language models in neurology diagnostics. It specifically tests whether models can move beyond superficial classification to perform uncertainty resolution, generate clinically justified rationales, and maintain logical coherence when interpreting patient histories and medical imaging. Use when the user wants to benchmark on Neural-MedBench, or asks about evaluating this task. Reports Diagnostic Accu...
This benchmark evaluates the robustness and generalization of multi-agent reinforcement learning policies in a large-scale, open-ended simulation. It probes a model's ability to cooperate with teammates and compete against unknown opponents or fixed baselines across varying difficulty levels and dynamic environments. Use when the user wants to benchmark on Neural MMO, or asks about evaluating this task. Reports TrueSkill.
Evaluates the ability of classifiers and LLMs to detect machine-generated news headlines across four languages. It probes cross-lingual generalization, robustness to zero-shot vs fine-tuned generators, and the effectiveness of linguistic vs transformer-based features for authenticity verification. Use when the user wants to benchmark on Multilingual Neural News Detection Benchmark, or asks about evaluating this task. Reports F1 score.
Evaluates the accuracy and robustness of neural operator models (DeepONet, FNO, and variants) in learning mappings between function spaces for solving partial differential equations across various physical domains and geometries. Use when the user wants to benchmark on PDE Operator Benchmark Suite (Lu et al. 2021), or asks about evaluating this task. Reports L2 relative error.
Evaluates novel view synthesis quality in neural rendering by measuring how well a model reconstructs unseen viewpoints from a set of training images. It probes the model's ability to capture view-dependent appearance, geometric consistency, and texture fidelity under challenging materials and real-world lighting. Use when the user wants to benchmark on Blender, Shiny Blender, Mip-360, or asks about evaluating this task. Reports PSNR.
Evaluates the ability of spatio-temporal point process models to accurately capture complex, history-dependent spatial and temporal distributions of discrete events. It probes how well models can compute exact likelihoods for sequences of events in continuous space and time across diverse domains like seismology, epidemiology, and neuroscience. Use when the user wants to benchmark on PINWHEEL, EARTHQUAKES, COVID-19 CASES, BOLD5000, or asks about evaluating this task. Reports log-likelihood pe...
Evaluates the model's ability to simulate historical global temperature trends and spatial temperature biases over multi-decadal climate simulations, and tests its generalization to warmer climate scenarios. It probes long-term stability, physical consistency, and response to prescribed sea surface temperature forcing. Use when the user wants to benchmark on ERA5, or asks about evaluating this task. Reports RMSB (850hPa temperature).
Evaluates a heterogeneous graph transformer's ability to learn multi-scale biological relationships and predict missing links in a brain knowledge graph. It probes downstream capabilities including genome-wide screen enrichment, pesticide toxicity ranking, and drug repurposing forecasting across neurological diseases. Use when the user wants to benchmark on NeuroKG, or asks about evaluating this task. Reports AUROC.
Evaluates the classification accuracy and energy efficiency of a neuromorphic convolutional network architecture running on Intel's TrueNorth hardware. It probes the model's ability to perform real-time visual and audio recognition while maintaining low power consumption and high throughput. Use when the user wants to benchmark on CIFAR10, CIFAR100, SVHN, GTSRB, Flickr-Logos32, VAD, TIMIT Class., TIMIT Frame, or asks about evaluating this task. Reports accuracy.
Evaluates closed-loop autonomous driving performance in simulated real-world scenarios, measuring safety and planning efficiency through collision avoidance and impact speed mitigation. Use when the user wants to benchmark on nuScenes, or asks about evaluating this task. Reports NeuroNCAP Score (NNS).
Evaluates the efficiency and rendering quality of a neural sample field for novel view synthesis. It probes how well a model can learn ray sampling distributions to reduce computation cost while maintaining high-fidelity image reconstruction compared to baseline NeRF methods. Use when the user wants to benchmark on Realistic Synthetic 360°, Real Forward-Facing, or asks about evaluating this task. Reports PSNR.
Compute nevikw39/specificity via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of nevikw39/specificity.
Evaluates multimodal representation learning for olfaction by testing cross-modal retrieval and classification tasks using paired image and electronic nose signals. Probes the model's ability to generalize from visual supervision to interpret raw or processed olfactory sensor data for scene, object, material, and fine-grained species recognition. Use when the user wants to benchmark on New York Smells, or asks about evaluating this task. Reports classification accuracy.
Evaluates AI's ability to understand humor through multimodal and text-based tasks, including matching captions to cartoons, ranking caption quality, and generating humorous explanations. It probes indirect allusion, cultural context, and visual-linguistic reasoning. Use when the user wants to benchmark on New Yorker Caption Contest, or asks about evaluating this task. Reports Accuracy.
Evaluates the adversarial robustness of tree ensemble models (RF, XGB, LGBM, EBM) on enterprise network intrusion detection using the corrected NewCICIDS dataset. It measures how well models maintain detection performance on benign and malicious traffic when subjected to constrained adversarial perturbations of time-series traffic features. Use when the user wants to benchmark on NewCICIDS, or asks about evaluating this task. Reports F1S.
Evaluates transformer models' ability to classify news articles as biased or unbiased, comparing standard fine-tuning against domain-adapted training. It further probes model decision-making by analyzing word-level SHAP attribution magnitudes and lexical feature importance across true and false predictions. Use when the user wants to benchmark on BABE, or asks about evaluating this task. Reports Binary F1.
Evaluates a model's ability to verify the veracity of real-world news claims by retrieving relevant evidence documents and predicting a multi-class label. It probes the system's capacity for evidence selection, claim decomposition, and fact-checking under black-box LLM constraints. Use when the user wants to benchmark on RAWFC, LIAR-RAW, or asks about evaluating this task. Reports macro-average F1.