
Claude Skills by qhjqhj00
github.com/qhjqhj00Compute ronaldahmed/nwentfaithfulness via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of ronaldahmed/nwentfaithfulness.
Evaluates the robustness of deep learning segmentation models to out-of-distribution MRI data and synthetic corruptions (noise, contrast, resolution, spatial shifts, motion artifacts) across multiple severity levels. It measures performance degradation on anatomical and lesion segmentation tasks compared to clean data. Use when the user wants to benchmark on ROOD-MRI Benchmark (Hippocampus, Ventricle, WMH), or asks about evaluating this task. Reports DSC.
Compute the root_mean_squared_error metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute root_mean_squared_error, or asks how to score with root_mean_squared_error.
Compute the root_mean_squared_log_error metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute root_mean_squared_log_error, or asks how to score with root_mean_squared_log_error.
Compute the RootMeanSquaredErrorUsingSlidingWindow metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute RootMeanSquaredErrorUsingSlidingWindow, or asks how to score with RootMeanSquaredErrorUsingSlidingWindow.
Evaluates the zero-shot generalization capability of instruction-tuned language models by retrieving and applying task-specific soft prompt embeddings at inference time to adapt to unseen tasks. Use when the user wants to benchmark on BIG-bench, SuperGLUE/HellaSwag/StoryCloze/WiC suite, or asks about evaluating this task. Reports accuracy.
Compute the ROUGEScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute ROUGEScore, or asks how to score with ROUGEScore.
Evaluates surface irregularities and topological consistency between predicted and ground-truth 3D medical segmentation masks. It quantifies local surface roughness, relative roughness differences, and average surface distance to detect spikes, holes, and smoothing artifacts. Use when the user has predictions and gold and needs to compute Roughness Index (RI).
This benchmark evaluates a model's ability to detect traffic anomalies in roadside surveillance videos and generate detailed, reasoning-grounded textual summaries of those anomalies. It probes both binary/fine-grained classification accuracy and semantic alignment of generated descriptions with ground-truth event narratives. Use when the user wants to benchmark on Roundabout-TAU, or asks about evaluating this task. Reports 4-cls AP.
Evaluates LLMs' scientific reasoning and proposal writing capabilities through a multi-task accuracy benchmark and a rubric-based narrative generation task. It also assesses the stability of consensus methods and the consistency of AI graders across structured and open-ended scientific domains. Use when the user wants to benchmark on MultiTask scientific tasks, SingleTask scientific proposal writing, or asks about evaluating this task. Reports accuracy.
Evaluates a foundation model's ability to solve and generalize across 48 distinct Vehicle Routing Problem (VRP) variants. It probes constraint satisfaction, route optimization, and zero-shot adaptation to unseen attribute combinations like multi-depots and mixed backhauls. Use when the user wants to benchmark on VRP variants (unified generation), or asks about evaluating this task. Reports optimality gap.
Evaluates a closed-loop LLM routing system's ability to dynamically select between a four-tier model portfolio based on task difficulty, balancing inference cost, response quality, and latency. The benchmark probes how well a router can escalate queries to more capable models only when necessary, while using distillation and conformal cascading to maintain performance at lower cost tiers. Use when the user wants to benchmark on EDGAR (NER), EDGAR (Summarization), BANKING77* (Intent Classifica...
Evaluates the routing algorithm's efficiency, scalability, and classification accuracy on standard NLP and vision benchmarks when used as a classification head over frozen pretrained Transformers. Use when the user wants to benchmark on IMDB, SST-5, SST-2, ImageNet-1K, CIFAR-100, CIFAR-10, or asks about evaluating this task. Reports Accuracy (%).
Evaluates the accuracy and robustness of visual-inertial SLAM systems across diverse outdoor environments, seasons, and lighting conditions. It probes long-term trajectory consistency, scale estimation, and environmental adaptability under challenging visual degradation. Use when the user wants to benchmark on ROVER, or asks about evaluating this task. Reports mATE.
Evaluates the ability of text-to-image models to accurately render specific objects at requested bounding box locations while maintaining prompt fidelity, attribute correctness, and overall aesthetic quality. Use when the user wants to benchmark on ROVI validation set, or asks about evaluating this task. Reports Gen Inst..
This benchmark probes a model's ability to comprehend hand-drawn, object-centric visual instructions (arrows, circles, colors) and translate them into precise spatiotemporal action plans for robotic manipulation. It evaluates both high-level task reasoning and low-level execution accuracy in cluttered, unseen environments. Use when the user wants to benchmark on RoVI Book dataset, SIMPLER, or asks about evaluating this task. Reports action success rate.
This evaluation probes an LLM's ability to generate factually correct responses and mitigate hallucinations under adaptive retrieval-augmented generation. It measures factual accuracy via LLM-based scoring on open-ended questions and exact-match accuracy on multi-step reasoning yes/no questions. Use when the user wants to benchmark on TruthfulQA, StrategyQA, or asks about evaluating this task. Reports FactScore.
Evaluates a model's ability to discover recurring visual patterns in a single image. It measures detection accuracy at both the individual pattern instance level and the whole pattern level against human annotations. Use when the user wants to benchmark on RP-1K, or asks about evaluating this task. Reports RP Instance Recall.
Evaluates the zero-shot image classification and text-to-image retrieval capabilities of Vision-Language Models (VLMs) fine-tuned on remote sensing data. It probes the model's ability to generalize to unseen RS scenes and text queries without task-specific fine-tuning, while also measuring resistance to catastrophic forgetting on general-domain benchmarks. Use when the user wants to benchmark on AID, EuroSAT, fMoW, Million-AID, PatternNet, RESISC, RSI-CB, ImageNet-1K, UCM Captions, RSICD, RSI...
Evaluates the ability of LLM-based evolutionary algorithms to optimize session-based recommendation prompts across multiple objectives (accuracy, diversity, and fairness) simultaneously. Use when the user wants to benchmark on RSBench, or asks about evaluating this task. Reports HV.
Evaluates vision-language models' ability to generate detailed, semantically accurate captions describing changes between bi-temporal remote sensing image pairs, particularly in disaster scenarios. It probes spatiotemporal reasoning, fine-grained environmental change detection, and long-text generation quality. Use when the user wants to benchmark on RSCC, or asks about evaluating this task. Reports ST5-SCS.
Evaluates the optimization capability of evolutionary algorithms on a highly multimodal, nonseparable benchmark function across low (d=5) and high (d=20) dimensional settings. It measures how effectively the algorithm navigates complex, multi-peaked likelihood surfaces to locate the global optimum within a fixed computational budget. Use when the user wants to benchmark on Rotated Schaffers F7 (RSF7), or asks about evaluating this task. Reports mean_max_function_value.
Evaluates a vision-language model's ability to perform zero-shot classification, cross-modal retrieval, visual question answering, and fine-grained spatial grounding (including region-caption retrieval and geo-localization) on remote sensing imagery. It measures how well instruction-conditioned contrastive pretraining aligns multimodal features with geospatial metadata and textual prompts. Use when the user wants to benchmark on AID, Million-AID, RSI-CB, EuroSAT, UCM, PatternNet, RSITMD, RSIC...
This benchmark evaluates large language models' ability to perform fine-grained, region-specific semantic reasoning on remote sensing image pairs. It probes localized change comprehension by asking models to answer binary, multiple-choice, and open-ended questions about specific changes (e.g., new construction, vegetation loss) within satellite imagery. Use when the user wants to benchmark on RSRCC, or asks about evaluating this task. Reports Accuracy (%).
Evaluates the ability of monocular depth estimation and stereo matching models to reconstruct fine-grained road surface profiles and disparities from high-resolution images. It probes the models' accuracy in capturing micro-level road textures and handling near-to-far distance variations under diverse dynamic conditions. Use when the user wants to benchmark on RSRD-dense, RSRD-sparse, or asks about evaluating this task. Reports Abs Rel.
This evaluation probes a semi-supervised temporal intrusion detection system's ability to classify network traffic flows as benign or malicious under label scarcity, adversarial contamination, and temporal distribution shifts across heterogeneous cloud environments. Use when the user wants to benchmark on CIC-IDS2017, CSE-CIC-IDS2018, UNSW-NB15, or asks about evaluating this task. Reports detection accuracy.
Evaluates a model's ability to predict RST-style discourse tree structure and nuclearity relations between elementary discourse units (EDUs). It probes both intra-domain and inter-domain generalization of discourse parsing across different text genres (news, instructions, reviews). Use when the user wants to benchmark on RST-DT, Instr-DT, MEGA-DT, Yelp13-DT, or asks about evaluating this task. Reports Parseval (Par.) / RST-Parseval (R-Par.).
This benchmark evaluates a model's ability to localize specific objects in remote sensing satellite imagery using natural language queries. It probes the model's robustness to scale variations, cluttered backgrounds, and multi-granularity textual descriptions common in aerial/satellite scenes. Use when the user wants to benchmark on RSVGD, or asks about evaluating this task. Reports Pr@0.5.
Evaluates whether data-driven reasoning rubrics improve LLM-based trace correctness classification and serve as effective reward signals for reinforcement learning compared to standard LLM judges and verifiable rewards. Use when the user wants to benchmark on SWE-Bench, NuminaMath, NaturalReasoning, or asks about evaluating this task. Reports Balanced Accuracy.
Compute Ruchin/jaccard_similarity via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of Ruchin/jaccard_similarity.
Probes a model's continual learning capability in a partially observable, non-stationary synthetic environment based on the Rule 110 cellular automaton. It measures how well capacity-constrained agents adapt to gradual distribution shifts induced by increasing prediction horizons and evolving task parameters. Use when the user wants to benchmark on Rule 110 Prediction Environment, or asks about evaluating this task. Reports online accuracy.
This benchmark evaluates long-context language models' ability to retrieve, trace, aggregate, and answer questions across varying context lengths and task complexities. It probes whether models genuinely attend to injected information or rely on parametric knowledge and context copying as sequence length increases. Use when the user wants to benchmark on RULER, or asks about evaluating this task. Reports exact-match accuracy.
This evaluation probes a language model's ability to perform rule-based logical reasoning on both in-distribution and out-of-distribution tasks. It measures how well the model can apply explicit and implicit logical rules to derive correct answers under strict exact-match conditions. Use when the user wants to benchmark on BigBench Hard (BBH), BigBench Extra Hard (BBEH), ProverQA, or asks about evaluating this task. Reports pass@1 (hard exact match).
Evaluates a model's ability to determine the veracity of social media rumours and classify the discourse stance of replies within a conversation tree. It probes contextual discourse analysis, stance detection, and truthfulness judgment in noisy, interactive text. Use when the user wants to benchmark on RumourEval (SemEval-2017 Task 8), or asks about evaluating this task. Reports classification accuracy, macroaveraged accuracy.
Evaluates Russian text embedding models across semantic similarity, classification, retrieval, and reranking tasks to measure their effectiveness in understanding and retrieving Russian language content. Use when the user wants to benchmark on ruMTEB, or asks about evaluating this task. Reports cosine_spearman.
Compute the RunningMean metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute RunningMean, or asks how to score with RunningMean.
Compute the RunningSum metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute RunningSum, or asks how to score with RunningSum.
Evaluates the computational running time, speedup gains, and parallel efficiency of median, standard deviation, and full source-finding algorithms on simulated radio interferometric images. Use when the user has predictions and gold and needs to compute Running time (seconds).
Evaluates the inference latency and computational runtime of transformer models and individual MLX operations across different hardware platforms (Apple Silicon vs NVIDIA GPU) and input configurations. Use when the user has predictions and gold and needs to compute runtime.
Evaluates large language models' ability to solve complex, logic-heavy Chinese natural language puzzles and perform multi-step reasoning on diverse benchmark tasks. It probes the model's capacity for progressive reasoning, self-verification, and adaptability to structured prompt frameworks without manual tuning. Use when the user wants to benchmark on Ruozhiba, BIG-Bench-Hard, or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to perform Word Sense Induction (WSI) by clustering contextual usages of target words into sense groups without prior sense labels. It probes the system's capability to handle morphological complexity and free word order in Russian across different sense granularities. Use when the user wants to benchmark on wiki-wiki, bts-rnc, active-dict, or asks about evaluating this task. Reports ARI.
Evaluates a model's ability to extend an existing semantic hierarchy (RuWordNet) by predicting hypernym relationships for novel Russian words using only contextual corpus information, without relying on explicit word definitions. It probes contextual lexical grounding and unsupervised taxonomy extension capabilities specifically for Slavic languages. Use when the user wants to benchmark on RUSSE'2020 Taxonomy Enrichment Test Set, or asks about evaluating this task. Reports MAP.
Evaluates the quality, phonetic accuracy, and prosodic fidelity of Russian speech datasets and generative models across synthesis, restoration, and denoising tasks. It probes natural intonation, stress accuracy, and audio clarity using standardized subjective ratings and automatic speech quality metrics. Use when the user wants to benchmark on Balalaika (Proposed), M-AILABS Russian, RUSLAN, Russian LibriSpeech, SOVA RuYoutube, Mozilla Common Voice 21.0, or asks about evaluating this task. Rep...
Evaluates an iterative red-blue adversarial framework for automated AI system hardening. It probes the system's ability to autonomously generate defensive patches for code vulnerabilities and optimize guardrail rules against jailbreak attacks through multi-round adversarial interaction. Use when the user wants to benchmark on Pharmacy Management System v1.0, HarmBench, JailBreakBench, AdvBench, SorryBench, XGuard-Train, or asks about evaluating this task. Reports Defense Success Rate (DSR).
Evaluates the robustness of modern voice cloning models under realistic deployment conditions, including input variations (accents, text shifts, long context), cross-lingual synthesis, post-processing degradation, and adversarial perturbations. It probes the trade-offs between generation quality, content fidelity, speaker similarity, and deepfake detectability across diverse acoustic and linguistic stressors. Use when the user wants to benchmark on LibriTTS, VCTK, LibriSpeech, RVCBench, or as...
Evaluates object detectors' robustness to real-world spatial domain shifts across different climate zones in satellite imagery. It measures how well models generalize from in-domain training data to out-of-distribution target domains without fine-tuning. Use when the user wants to benchmark on RWDS-CZ, or asks about evaluating this task. Reports mAP.
Evaluates object detectors' robustness to real-world spatial domain shifts across flood-affected regions in satellite imagery. It measures how well models generalize from in-domain training data to out-of-distribution target domains without fine-tuning. Use when the user wants to benchmark on RWDS-FR, or asks about evaluating this task. Reports mAP.
Evaluates object detectors' robustness to real-world spatial domain shifts across hurricane-affected regions in satellite imagery. It measures how well models generalize from in-domain training data to out-of-distribution target domains without fine-tuning. Use when the user wants to benchmark on RWDS-HE, or asks about evaluating this task. Reports mAP.
Evaluates a model's ability to maintain context and generate coherent responses in multi-turn dialogue, comparing stateful event-driven architectures against standard decoder-only LLMs. Use when the user wants to benchmark on MRL Curriculum Datasets (derived from TinyStories), or asks about evaluating this task. Reports MRL Reward Score.
Evaluates the ability of embodied agents to follow natural language instructions for navigation in photo-realistic 3D environments. It probes multilingual understanding, spatial reasoning, and dense spatiotemporal grounding by measuring how accurately an agent navigates from a start to a target location. Use when the user wants to benchmark on Room-Across-Room (RxR), or asks about evaluating this task. Reports SR, NDTW.