
Claude Skills by qhjqhj00
github.com/qhjqhj00Evaluates the effectiveness of tree-based ensemble anomaly detectors in human-in-the-loop active learning settings. It measures how quickly an algorithm can discover anomalies by querying a limited budget of instances, comparing batch and streaming data paradigms. Use when the user wants to benchmark on Abalone, ANN-Thyroid-1v3, Cardiotocography, KDD-Cup-99, Mammography, Shuttle, Yeast, Covtype, Electricity, Weather, or asks about evaluating this task. Reports anomaly_discovery_rate.
Evaluates zero-shot information retrieval capabilities in Hindi across diverse domains and tasks. It probes how well multilingual embedding models and baselines rank relevant documents for Hindi queries without language-specific fine-tuning. Use when the user wants to benchmark on Hindi-BEIR, or asks about evaluating this task. Reports NDCG@10.
Evaluates instruction-following, mathematical reasoning, code/function-calling, and retrieval-augmented generation capabilities of LLMs in Hindi. The benchmark specifically probes the models' ability to handle culturally and linguistically nuanced prompts that go beyond direct English translation. Use when the user wants to benchmark on IFEval-Hi, MT-Bench-Hi, GSM8K-Hi, ChatRAG-Hi, BFCL-Hi, or asks about evaluating this task. Reports score.
This benchmark evaluates a model's ability to perform Named Entity Recognition (NER) on Hindi text. It probes the model's capacity to identify and classify entity spans (e.g., Person, Location, Organization, and others) in a language characterized by free word order, lack of capitalization, and spelling variations. Use when the user wants to benchmark on HiNER, or asks about evaluating this task. Reports F1-Score.
Compute the hinge_loss metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute hinge_loss, or asks how to score with hinge_loss.
Compute the HingeLoss metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute HingeLoss, or asks how to score with HingeLoss.
Evaluates machine translation models on code-mixed and noisy Hindi-English and Bengali-English text, measuring robustness to script variations, romanization, and synthetic noise. The protocol tests both in-domain performance on the HINMIX corpus and out-of-domain generalizability on LinCE, SpokenTutorial, and IITB Hi-En. It also assesses zero-shot transfer to unseen code-mixed Bengali-English translation. Use when the user wants to benchmark on HINMIX, or asks about evaluating this task. Repo...
Evaluates cross-script machine translation quality for low-resource Indian languages (Hindi, Gujarati, Tamil) translating to English. It specifically probes how well models leverage back-translation augmented with quality and transliteration hints to handle noisy data and script conversion challenges. Use when the user wants to benchmark on IIT Bombay en-hi Corpus, WMT-2019 gu-en, TED2020, GNOME & Ubuntu, OPUS, WMT-2020 ta-en, GNOME, OPUS, WMT-2014 hi→en test, WMT-2019 gu→en test, WMT-2020 ta...
Evaluates multimodal large language models on vision-language tasks in Hindi and Telugu, measuring performance regression when transitioning from English to these Indian languages. It probes native-language visual question answering, mathematical reasoning, and multiple-choice comprehension across STEM and cultural domains. Use when the user wants to benchmark on HinTel-AlignBench, or asks about evaluating this task. Reports accuracy.
This benchmark probes a model's ability to detect adverse events (total hip replacement dislocation) from unstructured, free-text clinical narratives. It evaluates whether NLP models can correctly classify medical notes into dislocation status categories, handling complex negation, long-range dependencies, and multi-site anatomical references. Use when the user wants to benchmark on Radiology Notes, Telephone Notes, or asks about evaluating this task. Reports Kappa.
Evaluates multilingual historical text systems on person-place relation extraction, requiring temporal and geographical reasoning to classify relations as 'at' or 'isAt' with nuanced evidence levels (true, probable, false). It probes both extraction accuracy and reasoning quality in noisy, sparse corpora while also measuring computational efficiency. Use when the user wants to benchmark on HIPE-2026, or asks about evaluating this task. Reports macro-averaged Recall.
Evaluates multimodal large language models on authentic high school physics Olympiad problems, probing their ability to perform step-level physical reasoning, interpret complex diagrams and data plots, and solve problems across diverse physics subfields under Olympiad-level difficulty. Use when the user wants to benchmark on HiPhO, or asks about evaluating this task. Reports Mean Normalized Score (MNS).
This benchmark evaluates multimodal physical reasoning and advanced problem-solving capabilities on international and regional physics Olympiad exams. It probes a model's ability to interpret complex diagrams, data, and text, perform multi-step logical derivations, and produce accurate solutions under strict, official scoring rubrics. Use when the user wants to benchmark on HiPhO, or asks about evaluating this task. Reports exam score.
Evaluates multimodal agents' ability to reason over personalized, device-scale file systems. It probes long-horizon cross-file retrieval, multimodal perception, and evidence-grounded factual retention under strict profile-isolation constraints. Use when the user wants to benchmark on HippoCamp, or asks about evaluating this task. Reports accuracy.
Evaluates histopathology image retrieval and classification performance using high-order texture features (Gram barcodes) extracted from CNN layers. It probes the model's ability to capture tissue texture patterns for accurate image matching and class prediction. Use when the user wants to benchmark on KimiaPath24, CRC, EMC, or asks about evaluating this task. Reports η_total, Accuracy.
Evaluates the answer quality of Retrieval-Augmented Generation (RAG) systems across specialized domains. It measures how well generated responses address queries in terms of comprehensiveness, empowerment, diversity, and overall performance using pairwise LLM-as-a-judge comparisons. Use when the user wants to benchmark on UltraDomain, or asks about evaluating this task. Reports win rate.
Evaluates machine learning and deep learning models on predicting clinical outcomes from high-resolution ICU time-series data. It probes capabilities in handling class imbalance, long temporal dependencies, and varying data resolutions across stay-level and online monitoring tasks. Use when the user wants to benchmark on HiRID, or asks about evaluating this task. Reports AUPRC.
Evaluates a model's ability to perform 3D human-in-scene multimodal understanding through open-ended question answering. It probes capabilities in activity recognition, spatial relationship reasoning, and human-object interaction analysis within dynamic 3D environments. Use when the user wants to benchmark on HIS-Bench, or asks about evaluating this task. Reports HIS-Bench score.
Evaluates large language models across five hierarchical cognitive stages of scientific inquiry, ranging from foundational factual recall and literature parsing to advanced synthesis, literature review generation, and data-driven scientific discovery. It probes multimodal comprehension, cross-lingual reasoning, and computational problem-solving across six scientific disciplines. Use when the user wants to benchmark on HiSciBench, or asks about evaluating this task. Reports accuracy.
Evaluates a vision-language model's ability to generate accurate and clinically relevant histopathology reports from whole slide images (WSIs). It probes lexical overlap, semantic coherence, and medical entity coverage in generated text. Use when the user wants to benchmark on HistGen, or asks about evaluating this task. Reports BLEU-4.
Evaluates a model's ability to generate clinical histopathology reports from gigapixel whole slide images (WSIs). It probes cross-modal alignment between dense visual patches and concise textual descriptions using standard natural language generation metrics. Use when the user wants to benchmark on TCGA WSI-Report, or asks about evaluating this task. Reports BLEU-4.
Evaluates the ability of language models to recognize and classify named entities (PERSON, ORGANIZATION, LOCATION, PRODUCT, DATE) in historical Romanian newspaper texts across four distinct geographical regions. Use when the user wants to benchmark on HistNERo, or asks about evaluating this task. Reports strict F1-score.
Evaluates vision-language foundation models on diverse histopathology clinical tasks (detection, subtyping, grading, mutation prediction) to assess their robustness to textual/visual perturbations, magnification changes, stain normalization, and model calibration. Use when the user wants to benchmark on CRC-100K, BreakHist, DataBiox, GasHisSDB, Breast IDC, LC25000-lung, or asks about evaluating this task. Reports balanced accuracy.
Evaluates the prognostic and molecular predictive value of 38 automated histomic features extracted from H&E whole-slide images across 21 solid-tumor cancer types. It probes whether purely morphological patterns can recover canonical biology, predict survival outcomes, and correlate with gene expression, pathway activity, and immune subtypes. Use when the user wants to benchmark on TCGA Pan-Cancer H&E Cohort, or asks about evaluating this task. Reports Cox proportional-hazards model.
Evaluates the linguistic quality and cultural appropriateness of the Histoires Morales dataset. It measures reference-free translation accuracy and assesses whether moral norms and actions align with native French speakers' cultural values. Use when the user wants to benchmark on Histoires Morales, or asks about evaluating this task. Reports CometKiwi22.
Evaluates the robustness of vision-language models (VLMs) and test-time adaptation (TTA) methods when applied to histopathology images under realistic domain shifts. It probes how well models maintain classification accuracy when exposed to synthetic corruptions like staining variations, dust, blurring, and noise that mimic real-world clinical imaging artifacts. Use when the user wants to benchmark on NCT-7K, NCT-100K, LC25000, SkinCancer, RenalCell, MHIST, or asks about evaluating this task....
Evaluates a model's ability to generalize to out-of-distribution domains (different hospitals or staining protocols) in histopathology image classification. It measures classification accuracy on held-out OOD validation and test splits, alongside the reconstruction quality of self-supervised generative augmentation. Use when the user wants to benchmark on CAMELYON17-WILDS, Epithelium-Stroma, or asks about evaluating this task. Reports Accuracy (%).
Evaluates graph neural network explainers on histopathology images by measuring how well they identify critical tumor nuclei and preserve model fidelity. It probes the explainer's ability to extract global, class-specific patterns and produce accurate instance-level importance maps for downstream nuclei classification. Use when the user wants to benchmark on BRACS, BACH, BreCaHAD, CRC, or asks about evaluating this task. Reports macro-averaged F1 score.
Evaluates deep learning models on binary classification of histopathology tissue tiles as malignant or benign. It probes the capacity of multi-stream architectures to capture diverse morphological and textural features for medical image grading. Use when the user wants to benchmark on CAMELYON16, Invasive Ductal Carcinoma (IDC), or asks about evaluating this task. Reports accuracy.
Evaluates LLMs' ability to accurately transcribe historical 18th-century Russian documents while preserving period-specific orthography and avoiding anachronistic character insertions. It probes both standard OCR accuracy and historical fidelity under varying input contexts and prompt strategies. Use when the user wants to benchmark on 18th-century Russian Civil Font Texts, or asks about evaluating this task. Reports CER.
Evaluates a unified GAN framework's ability to perform stain-invariant segmentation of glomeruli in renal histopathology. It tests generalization across multiple known staining modalities and unseen stainings, measuring how well the model maintains segmentation accuracy despite domain shifts in histological appearance. Use when the user wants to benchmark on AIDPATH & Custom PAS dataset, or asks about evaluating this task. Reports F1.
Multi-class histopathological image classification for cancer diagnosis across four tissue types (breast, prostate, bone, cervical). It probes the model's ability to extract robust morphological features from stained whole-slide image tiles without data augmentation. Use when the user wants to benchmark on ICIAR2018, SIPAkMeD, SICAPv2, UT-Osteosarcoma, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates graph-based generative models for their ability to produce biologically relevant, drug-like hit molecules rather than just chemically valid structures. It probes the models' capacity to satisfy strict medicinal chemistry constraints, maintain distributional similarity to known bioactive compounds, and achieve strong predicted binding affinity to specific protein targets. Use when the user wants to benchmark on REINVENT Dataset, Hit-like Dataset, Target-Specific Ligand...
Evaluates the ability of LLM-based sequential recommendation models to predict the next item in a user's interaction history. It specifically probes how well models capture temporal dynamics by incorporating irregular time intervals between interactions, and assesses performance under warm and cold-start conditions. Use when the user wants to benchmark on Amazon Reviews (Video Games, CDs and Vinyl, Books), or asks about evaluating this task. Reports Hit Ratio@1.
Evaluates large language models' multilingual comprehension of Hong Kong-specific knowledge, Cantonese linguistic capabilities, and reasoning across STEM, social sciences, and humanities in both Traditional and Simplified Chinese. Use when the user wants to benchmark on HKMMLU, or asks about evaluating this task. Reports accuracy.
Evaluates the quality of self-supervised visual representations learned from histopathology images by measuring downstream classification performance at patch, slide, and patient levels. It probes the model's ability to capture hierarchical pathological structures and align them with clinical diagnostic categories. Use when the user wants to benchmark on OpenSRH, TCGA, or asks about evaluating this task. Reports kNN classification accuracy (ACC).
This benchmark evaluates a robot's ability to autonomously explore an unseen indoor environment and answer multiple-choice questions requiring object identification, counting, spatial reasoning, and multi-goal navigation. It probes the agent's multimodal perception, iterative reasoning, and navigation efficiency in a dynamic, tool-invoking workflow. Use when the user wants to benchmark on HM-EQA, or asks about evaluating this task. Reports Accuracy.
Evaluates an embodied AI agent's ability to navigate indoor 3D environments to find specific object categories using RGB-D observations. It measures both navigation quality (success and path efficiency) and computational efficiency (latency, memory, and skip ratio) on a large-scale dataset. Use when the user wants to benchmark on HabitatMatterport3D (HM3D), or asks about evaluating this task. Reports SPL.
Evaluates the accuracy of optical flow estimation models on synthetic and real-world video sequences. It specifically probes the model's ability to capture fine object contours, handle small or fast-moving targets, and maintain robustness under downscaling and occlusion. Use when the user wants to benchmark on Sintel, KITTI-2015, or asks about evaluating this task. Reports EPE.
Evaluates the trade-off between predictive performance and group fairness when applying causal pre-processing to approximate an unbiased data distribution. It probes whether debiasing techniques can simultaneously satisfy multiple fairness constraints without degrading model accuracy. Use when the user wants to benchmark on HMDA (Wisconsin, 2022), or asks about evaluating this task. Reports AUC.
This benchmark evaluates the quality and robustness of heterogeneous network embedding (HNE) algorithms across diverse real-world graphs. It probes how well learned representations preserve multi-type structural and attribute information, measured via downstream node classification and link prediction tasks. Use when the user wants to benchmark on DBLP, Yelp, Freebase, PubMed, or asks about evaluating this task. Reports macro-F1.
This evaluation probes a model's capability to detect network intrusions by classifying traffic flows as benign or malicious. It assesses the system's ability to learn complex temporal and multi-scale features from network flow data to distinguish between normal activities and various attack types. Use when the user wants to benchmark on Hogzilla Dataset, or asks about evaluating this task. Reports Accuracy.
Evaluates a model's ability to detect Human-Object Interactions (HOIs) by predicting triplets of person, verb, and object along with their bounding boxes. It specifically probes the model's robustness to object bias by measuring performance on rare versus frequent interactions under both standard and object-conditional evaluation protocols. Use when the user wants to benchmark on HICO-DET, HOI-COCO, or asks about evaluating this task. Reports mAP.
Evaluates a video generation model's ability to synthesize high-fidelity videos conditioned on multimodal inputs (reference images, audio, pose, and text) while maintaining reference consistency, audio-visual synchronization, and temporal coherence. Use when the user wants to benchmark on HOIVG-Bench, EMTD, or asks about evaluating this task. Reports NexusScore.
Evaluates the quality, text-motion alignment, and diversity of generated 2D whole-body human motion sequences conditioned on text prompts. It probes the model's ability to capture fine-grained spatial-temporal dynamics and handle occlusions or noisy 2D pose data. Use when the user wants to benchmark on Holistic-Motion2D, or asks about evaluating this task. Reports FID.
Evaluates demographic and intersectional biases in language models by measuring disparities in token likelihoods, generation styles, and offensiveness across a curated set of demographic descriptor terms embedded in sentence templates. Use when the user wants to benchmark on HOLISTICBIAS, or asks about evaluating this task. Reports Full Gen Bias.
Evaluates the security resilience and hardware overhead of Higher-Order Logic Locking (HOLL) against a counterexample-guided inductive synthesis (CEGIS) attack on combinational circuits. It measures how long an attacker takes to recover the secret key relation and the area penalty incurred by the locking mechanism. Use when the user wants to benchmark on ISCAS'85 and MCNC benchmarks, or asks about evaluating this task. Reports attack_time.
Evaluates a deep learning model's ability to reconstruct complex object fields from in-line digital holograms, specifically testing its capacity to suppress twin-image artifacts and maintain reconstruction fidelity under various noise conditions. Use when the user wants to benchmark on Synthetic Inline Holography Dataset, or asks about evaluating this task. Reports reconstruction fidelity.
Compute the homogeneity_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute homogeneity_score, or asks how to score with homogeneity_score.
Compute the HomogeneityScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute HomogeneityScore, or asks how to score with HomogeneityScore.