
Claude Skills by qhjqhj00
github.com/qhjqhj00Assesses the quality of automatically mined parallel sentence pairs by training neural machine translation models and measuring their downstream translation accuracy. Use when the user wants to benchmark on WikiMatrix, or asks about evaluating this task. Reports BLEU.
Evaluates a model's ability to detect malicious Wikipedia editors (vandals) using only benign user data for training. It probes one-class anomaly detection and sequential behavior modeling by measuring how well the system distinguishes benign from malicious users based on edit sequences. Use when the user wants to benchmark on UMDWikipedia, or asks about evaluating this task. Reports F1.
Evaluates a model's ability to translate natural language questions into executable SQL queries against a given database schema. It probes semantic parsing, table understanding, and the generation of syntactically and semantically correct structured queries. Use when the user wants to benchmark on WikiSQL, or asks about evaluating this task. Reports Acc_ex.
Evaluates a model's ability to perform semantic parsing over semi-structured tabular data by generating database queries or answers from natural language questions. It measures how well the model aligns textual utterances with table schemas and content to retrieve correct results. Use when the user wants to benchmark on WIKITABLEQUESTIONS, or asks about evaluating this task. Reports execution accuracy.
Evaluates a model's ability to perform complex tabular reasoning to answer open-ended questions based on a provided table. It probes the model's capacity to extract, aggregate, and filter information from structured data to produce short text span answers. Use when the user wants to benchmark on WikiTQ, or asks about evaluating this task. Reports denotation accuracy.
Compute the wilcoxon metric — provided by scipy.stats. Use when the user has predictions and ground-truth and needs to compute wilcoxon, or asks how to score with wilcoxon.
Evaluates the out-of-distribution (OOD) generalization capability of tabular regression models by measuring performance gaps between in-distribution and out-of-distribution test sets. It probes whether advanced OOD training strategies or complex architectures can reliably outperform simple Empirical Risk Minimization (ERM) on unseen data distributions. Use when the user wants to benchmark on VPower_S, VPower_R, Weather, or asks about evaluating this task. Reports MAE.
This benchmark probes the robustness of automatic speech recognition (ASR) systems under realistic, out-of-distribution conditions. It specifically evaluates performance degradation across environmental noise, demographic shifts (accent, age, child speech), and linguistic diversity (short, incomplete, code-switched utterances), while also measuring semantic hallucination rates beyond standard lexical error metrics. Use when the user wants to benchmark on WildASR, or asks about evaluating this...
Evaluates object detection algorithms on drone-captured images of wild berries in cluttered, dynamic forest environments. It probes localization and classification capabilities under severe lighting variations, occlusion, and cross-domain transfer settings (different areas, cameras, and datasets). Use when the user wants to benchmark on WildBe, or asks about evaluating this task. Reports Average Precision (AP).
Evaluates open-vocabulary monocular 3D object detection across diverse real-world and synthetic scenes. It probes the model's ability to localize and regress 3D bounding boxes using text or geometric prompts, measuring generalization to unseen categories and datasets with and without depth cues. Use when the user wants to benchmark on WildDet3D-Bench, Omni3D, Argoverse 2, ScanNet, Stereo4D, or asks about evaluating this task. Reports AP_3D.
This benchmark evaluates automatic speech recognition (ASR) capabilities on Mandarin speech produced by elderly individuals. It probes a model's robustness to real-world acoustic degradation, articulation variability, tremors, and diverse accent strengths under uncontrolled recording conditions. Use when the user wants to benchmark on WildElder, or asks about evaluating this task. Reports Word Error Rate (WER).
Evaluates machine learning and deep learning models' ability to predict wildfire occurrences in Morocco using integrated spatio-temporal environmental and meteorological features. It specifically tests temporal generalization by training on pre-2022 data and validating on post-2022 data. Use when the user wants to benchmark on Morocco Wildfire Dataset, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates a model's ability to forecast the final spatial extent of a wildfire using multi-day spatio-temporal environmental and dynamic features. It probes the model's capacity to capture complex temporal dependencies and spatial patterns in binary segmentation tasks under significant class imbalance. Use when the user wants to benchmark on Mediterranean Wildfire Dataset (2006-2022), or asks about evaluating this task. Reports Dice Score.
This evaluation protocol assesses the safety moderation capabilities of LLMs and dedicated moderation models. It probes their ability to detect harmful content in user prompts, classify harmful or safe model responses, and identify whether a model appropriately refuses unsafe requests across multiple risk categories. Use when the user wants to benchmark on ToxicChat, OpenAI Mod, AegisSafetyTest, SimpleSafetyTests, Harmbench Prompt, Harmbench Resp, BeaverTails, SafeRLHF, XSTest-Resp, WildGuard...
Evaluates the safety and robustness of language models against adversarial jailbreak attacks. It probes whether models can correctly refuse harmful requests while avoiding over-refusal on benign prompts, specifically under stealthy, adversarially composed prompts. Use when the user wants to benchmark on WILDJAILBREAK, or asks about evaluating this task. Reports Attack success rate (ASR).
Evaluates novel view synthesis and motion mask estimation in dynamic environments where both camera and objects move. It probes a model's ability to remove transient objects, complete occluded backgrounds, and preserve scene geometry from sparse input views without 3D supervision or ground-truth poses. Use when the user wants to benchmark on D-RE10K-Mask, D-RE10K-iPhone, or asks about evaluating this task. Reports PSNR.
Evaluates machine learning models' robustness to real-world distribution shifts, specifically domain generalization and subpopulation shifts. It measures how much model performance degrades when tested on out-of-distribution (OOD) data compared to in-distribution (ID) data, highlighting gaps in generalization for real-world deployment. Use when the user wants to benchmark on WILDS, or asks about evaluating this task. Reports ID and OOD performance.
Probes the capability of models to perform semantic segmentation on unstructured, large-scale natural environments using both 2D images and 3D LiDAR point clouds. It evaluates robustness to semantic ambiguity, clutter, and temporal environmental shifts in outdoor traversals. Use when the user wants to benchmark on WildScenes, or asks about evaluating this task. Reports mIoU.
Evaluates language models' scientific reasoning capabilities by testing their ability to answer domain-specific multiple-choice questions derived from peer-reviewed literature and established scientific benchmarks. Use when the user wants to benchmark on WildSci-Val, GPQA-Aug, SuperGPQA, MMLU-Pro, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates multimodal large language models' ability to perform multi-step, context-sensitive reasoning over symbolic musical notation. It probes capabilities in harmonic analysis, rhythmic interpretation, structural form recognition, and expressive markings through multiple-choice questions derived from real-world compositions and forum queries. Use when the user wants to benchmark on WildScore, or asks about evaluating this task. Reports accuracy.
Evaluates the ability of models to correctly identify the language of monolingual text paragraphs. It probes language identification capabilities across a wide range of languages (235) with balanced representation. Use when the user wants to benchmark on WiLI-2018, or asks about evaluating this task. Reports F1.
Evaluates the ability of a Gaussian linear state-space model to accurately forecast short-term wind speeds in the North-East Atlantic using historical observations. It also assesses the model's capacity to reproduce realistic spatiotemporal wind statistics and compares parameter estimation methods (GMM vs ML). Use when the user wants to benchmark on ERA Interim reanalysis data (North-East Atlantic), or asks about evaluating this task. Reports MSPE.
Evaluates the calibration, sharpness, and accuracy of probabilistic wind power forecasts under different ensemble post-processing strategies (raw, weather-only, power-only, and joint weather-power post-processing). It probes whether correcting biases at the weather stage alone is sufficient, or if direct post-processing of the final power ensemble is required to handle non-linear power curve biases. Use when the user wants to benchmark on Benchmark Data, Swedish Data Set, or asks about evalua...
Quantifies the intrinsic predictability of time series data by measuring window-wise pattern complexity in the frequency domain. It establishes a data-driven performance lower bound for forecasting models and identifies whether standard benchmarks have reached saturation. Use when the user has predictions and gold and needs to compute window-wise complexity.
Evaluates large language models on comprehensive medical reasoning, clinical calculation, and general cognitive capabilities. It probes domain-specific knowledge application, diagnostic reasoning, and complex problem-solving in real-world clinical and academic settings. Use when the user wants to benchmark on MedCalc, MedReMCQ, CMMLU, MATH-500, MedQA-USMLE, MedMCQA, PubMedQA, or asks about evaluating this task. Reports accuracy.
Evaluates gender bias in coreference resolution systems by measuring performance disparity between pro-stereotypical and anti-stereotypical sentences. It probes whether models rely on gender stereotypes when resolving coreferences in challenging, Winograd-style contexts. Use when the user wants to benchmark on WinoBias, or asks about evaluating this task. Reports F1.
This benchmark probes systematic gender bias in coreference resolution systems by measuring how often models resolve gendered pronouns to occupations differently based solely on pronoun gender. It evaluates whether models reinforce real-world occupational gender disparities and how performance degrades on counter-stereotypical ('gotcha') examples. Use when the user wants to benchmark on Winogender schemas, or asks about evaluating this task. Reports bias_score.
Probes vision-language models' ability to understand visio-linguistic compositionality and word order sensitivity. The task requires matching images to captions where identical words are rearranged to change the described scene, testing structural grounding rather than lexical overlap. Use when the user wants to benchmark on Winoground, or asks about evaluating this task. Reports image-caption score.
Evaluates anti-LGBTQ+ bias in language models by measuring their tendency to prefer stereotypical completions over counterfactual ones when prompted with identity-specific contexts. Use when the user wants to benchmark on WinoQueer, or asks about evaluating this task. Reports bias score.
Evaluates whether language models disproportionately associate harmful stereotypes with marginalized groups (Jewish people or LGBTQ+ subgroups) compared to non-target groups. It also assesses the quality and reliability of automated versus human annotation for constructing community-sourced fairness benchmarks. Use when the user wants to benchmark on WinoSemitism, WinoQueer, or asks about evaluating this task. Reports WinoSem. Score.
Evaluates Open Information Extraction systems on their ability to accurately extract relational tuples from text. It probes token-level precision and recall by matching predicted arguments and relations against a fine-grained, manually annotated gold standard. Use when the user wants to benchmark on WiRe57, or asks about evaluating this task. Reports token-weighted F1.
Evaluates a robot's ability to generalize visuomotor planning to unseen visual signals, referring expressions, and spatial relationships during cube-picking and placing tasks. It probes semantic generalization and robustness to distribution shifts in embodied instruction following. Use when the user wants to benchmark on WISER Benchmark, or asks about evaluating this task. Reports Success.
Evaluates human activity recognition (HAR) performance using only wireless Channel State Information (CSI) signals under varying action segmentation windows (1s, 2s, 3s). It probes the robustness of classification models in privacy-preserving environments where visual data is occluded or unavailable during testing. Use when the user wants to benchmark on WiVi, or asks about evaluating this task. Reports OA.
Probes instruction-following capability on complex, real-world prompts across diverse domains like coding, math, reasoning, and formatting. It measures how well models handle demanding, multi-step tasks compared to baselines through blind pairwise human comparison. Use when the user wants to benchmark on WizardEval, or asks about evaluating this task. Reports win_rate.
Evaluates Vision-Language Models on atomic world modeling capabilities across perception (spatial, temporal, motion) and prediction (mechanistic simulation, transitive/compositional inference) tasks. It probes whether VLMs possess internal representations of physical causality, dynamics, and multi-step reasoning comparable to human intuition. Use when the user wants to benchmark on WM-ABench, or asks about evaluating this task. Reports accuracy.
Evaluates large language models' knowledge of hazardous topics in biosecurity, cybersecurity, and chemical security, as well as their general knowledge and fluency. It serves as a proxy for measuring dual-use risk and benchmarking unlearning methods. Use when the user wants to benchmark on WMDP, or asks about evaluating this task. Reports WMDP.
Evaluates a model's ability to perform Automatic Post-Editing (APE) by correcting machine-translated German sentences using the original source text. It probes error detection, grammatical correction, and word-copying capabilities in a multi-source sequence-to-sequence setting. Use when the user wants to benchmark on WMT APE, or asks about evaluating this task. Reports case-sensitive BLEU.
Evaluates the quality of an unsupervised web-mined parallel corpus by training neural machine translation models on it and measuring translation performance on standard WMT test sets. It probes whether mined pseudo-parallel data can effectively substitute for human-labeled data in both supervised and unsupervised MT training pipelines. Use when the user wants to benchmark on WMT2014 test set, WMT2016 test set, or asks about evaluating this task. Reports BELU.
This protocol evaluates the machine translation quality of large language models across multiple language pairs. It measures translation accuracy and fluency by comparing model outputs against gold references and state-of-the-art baselines using neural quality estimation metrics. The benchmark probes the model's ability to generalize across diverse language directions and avoid generating near-perfect but flawed translations. Use when the user wants to benchmark on WMT'21 Test Set, WMT'22 Tes...
Evaluates machine translation metrics by measuring their alignment with human judgments at segment and system levels. It probes whether metrics can correctly rank translations and systems based on quality, and tests the robustness of meta-evaluation statistics like correlation and pairwise ranking accuracy. Use when the user wants to benchmark on WMT Test Sets, or asks about evaluating this task. Reports Pearson correlation.
Evaluates the capability of sign language to text translation systems on low-resource datasets. It measures how accurately a model can convert 3D pose sequences of sign language into corresponding spoken language text. Use when the user wants to benchmark on FocusNews, SRF, or asks about evaluating this task. Reports BLEU.
Evaluates the effectiveness of an RNN Encoder-Decoder architecture for statistical machine translation on English-to-French tasks. It measures how well the model learns phrase representations and improves translation quality over a traditional phrase-based baseline system. Use when the user wants to benchmark on WMT'14 English/French, or asks about evaluating this task. Reports BLEU.
Evaluates neural machine translation systems across multiple language pairs in news and biomedical domains. Probes translation quality, domain adaptation, and system combination techniques like ensembling and reranking on held-out parallel test sets. Use when the user wants to benchmark on WMT17 News Task, HimL Biomedical Task, or asks about evaluating this task. Reports BLEU.
Evaluates the training efficiency, inference speed, and translation quality of neural machine translation systems on standard WMT17 benchmarks. It probes the trade-offs between model architecture, hardware acceleration (FP16/INT8), and batching strategies. Use when the user wants to benchmark on WMT17 English-German, WMT17 Russian-English, or asks about evaluating this task. Reports BLEU.
Evaluates how well a model's predicted quality scores correlate with human judgments on machine translation output. It probes semantic equivalence and paraphrase detection capabilities in the context of MT evaluation. Use when the user wants to benchmark on WMT17, or asks about evaluating this task. Reports Pearson |r|.
Evaluates the effectiveness of automatic web data selection for domain-specific machine translation in the news domain. It probes whether a document-level topic classifier can filter noisy parallel data to improve MT system performance compared to state-of-the-art baselines on the WMT-18 benchmark. Use when the user wants to benchmark on WMT-18 News Shared Task, or asks about evaluating this task. Reports BLEU.
Evaluates an online learning framework's ability to dynamically identify the highest-quality machine translation systems from an ensemble using minimal human feedback, and measures the sample efficiency (number of human assessments needed) to converge to the official top-performing systems. Use when the user wants to benchmark on WMT'19 News Translation, or asks about evaluating this task. Reports convergence_to_top3.
Evaluates machine translation quality between similar languages (Czech to Polish) using a multi-encoder transformer trained on out-of-domain data filtered by cross-entropy differences. Probes the model's ability to adapt to low-resource similar language pairs via domain adaptation and data selection. Use when the user wants to benchmark on WMT19 SLT Shared Task dataset, or asks about evaluating this task. Reports BLEU.
Evaluates machine translation quality by scoring system-generated German translations against human reference sentences and expert MQM scores. It measures how well automatic metrics correlate with human judgments across different translation systems. Use when the user wants to benchmark on WMT20 EN-DE, or asks about evaluating this task. Reports COMET.
Evaluates machine translation quality by scoring system-generated English translations against human reference sentences and expert MQM scores. It measures how well automatic metrics correlate with human judgments across different translation systems. Use when the user wants to benchmark on WMT20 ZH-EN, or asks about evaluating this task. Reports COMET.