All authors
qhjqhj00 avatar

Claude Skills by qhjqhj00

github.com/qhjqhj00
7,636 skillsA× 7,623B× 11C× 1D× 10 installs2,406 views
Imagenet Zoom Classification EvalA

Evaluates image classification models' accuracy on standard and out-of-distribution datasets. It specifically probes the impact of spatial zooming and foreground/background signal separation on model performance, revealing how much background cues contribute to classification accuracy. Use when the user wants to benchmark on ImageNet, ImageNet-A, ObjectNet, or asks about evaluating this task. Reports top-1 accuracy.

researchpythongo
0
3
Imagenet32 EvalA

Evaluates image classification performance on downsampled variants of ImageNet to test whether lower-resolution datasets can serve as reliable proxies for full-resolution ImageNet in hyperparameter tuning and architecture search. It probes the stability of optimal hyperparameters and model performance across different spatial resolutions while maintaining the original dataset's class structure and image count. Use when the user wants to benchmark on ImageNet32x32, ImageNet64x64, ImageNet16x16...

researchpythongo
0
3
Imagenethink250k EvalA

Evaluates vision-language models' ability to generate structured, step-by-step reasoning (thinking tokens) and final answers for multimodal inputs. It probes reasoning coherence, logical progression, and alignment with reference synthetic reasoning traces. Use when the user wants to benchmark on ImageNet-Think-250K, or asks about evaluating this task. Reports BERTScore.

researchpythongo
0
3
Imagenetvc EvalA

Evaluates zero- and few-shot visual commonsense reasoning capabilities of language models and visually-augmented language models across 1,000 ImageNet categories using human-annotated QA pairs. Use when the user wants to benchmark on ImageNetVC, or asks about evaluating this task. Reports Top-1 accuracy.

researchpythongo
0
3
Imasc EvalA

Evaluates the perceptual quality and naturalness of synthesized Malayalam speech generated by a multi-speaker TTS model trained on the IMaSC corpus. It probes the model's ability to capture agglutinative morphology, phonemic orthography, and diverse prosodic styles through subjective human listening tests. Use when the user wants to benchmark on IMaSC, or asks about evaluating this task. Reports Mean Opinion Score (MOS).

researchpythongo
0
3
Imdb Sentiment EvalA

Tests the model's capability to generate fixed-length representations for variable-length documents containing multiple sentences. It probes whether the method can scale to longer texts and outperform traditional bag-of-words baselines on a large-scale sentiment classification benchmark. Use when the user wants to benchmark on IMDB dataset, or asks about evaluating this task. Reports error rate.

researchpythongo
0
3
Imdrug EvalA

Evaluates deep learning models for imbalanced and long-tailed classification and regression in AI-aided drug discovery. It probes model robustness to severe class imbalance, open long-tailed distributions, and out-of-distribution chemical splits across graph, sequence, and fingerprint molecular representations. Use when the user wants to benchmark on HIV, SBAP, USPTO-50K, DrugBank, or asks about evaluating this task. Reports Balanced-Acc.

researchpythongit
0
3
Imgedit Bench EvalA

Evaluates text-and-image-to-image editing capabilities, including addition, removal, replacement, motion change, style transfer, background change, object extraction, and hybrid edits. It tests the model's capacity to modify existing images according to natural language instructions while preserving unedited regions. Use when the user wants to benchmark on ImgEdit-Bench, or asks about evaluating this task. Reports ImgEdit-Bench.

researchpythongo
0
3
Imigue Speech EvalA

Evaluates models on recognizing spontaneous emotional states from unscripted speech and text. It probes acoustic prosody through dimensional regression and categorical classification, as well as linguistic sentiment polarity in real-world sports interview contexts. Use when the user wants to benchmark on iMiGUE-Speech, or asks about evaluating this task. Reports Categorical Emotion Classification.

researchpythongo
0
3
Imis EvalA

Evaluates the ability of vision models to perform interactive medical image segmentation using user prompts like clicks, bounding boxes, or text. It probes how well models generalize across different imaging modalities, anatomical structures, and interaction strategies (single vs. multi-turn). Use when the user wants to benchmark on IMed-361M, TotalSegmentator MRI dataset, ISLES dataset, or asks about evaluating this task. Reports Dice score.

researchpythonperformance
0
3
Imo Shortlist EvalA

This benchmark probes an LLM's ability to produce logically sound, step-by-step mathematical reasoning for Olympiad-level problems. It specifically measures the gap between achieving the correct final answer and maintaining rigorous, fallacy-free solution processes. Use when the user wants to benchmark on IMO shortlist problems (2009-2023), or asks about evaluating this task. Reports Final Answer Accuracy (%), Correct|Correct Final Answer (%).

researchpythongo
0
3
Imo2025 EvalA

Evaluates a model's ability to solve Olympiad-level mathematical proof problems by generating rigorous solutions and iteratively refining them through self-verification. Use when the user wants to benchmark on IMO 2025, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Implicit Cot EvalA

Evaluates a model's ability to perform arithmetic and grade-school math reasoning without generating explicit intermediate chain-of-thought steps. It measures both the exact-match accuracy of the final answer and the inference speed relative to a no-CoT baseline. Use when the user wants to benchmark on Multi-digit multiplication, GSM8K, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Impromptu Vla Diagnostic EvalA

Diagnoses VLM capabilities in autonomous driving by evaluating perception, prediction, meta-planning via Q&A accuracy, and planning via trajectory prediction L2 error. Use when the user wants to benchmark on Impromptu VLA, or asks about evaluating this task. Reports Q&A Accuracy.

researchpythongo
0
3
Inatag EvalA

Evaluates multi-class image classification capabilities for agricultural species, genus, family, and crop/weed distinction. Probes fine-grained visual recognition and taxonomic hierarchy understanding in plant identification. Use when the user wants to benchmark on iNatAg, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Inbreast EvalA

Evaluates a deep learning model's ability to detect and classify malignant lesions in mammograms. It measures classification accuracy at the breast level and detection/localization sensitivity against false positive rates. Use when the user wants to benchmark on INbreast, or asks about evaluating this task. Reports AUC.

researchpythonperformance
0
3
Inbreast Mammogram Classification EvalA

Evaluates the ability of deep multi-instance learning models to classify whole mammograms as benign or malignant without requiring region-of-interest (ROI) annotations. It probes patch-level malignancy prediction and whole-image classification robustness under sparse label conditions. Use when the user wants to benchmark on INbreast, or asks about evaluating this task. Reports Accuracy.

researchpythontesting
0
3
Inbreast Mammography EvalA

Probes the ability of deep learning models to detect malignancy in mammograms and generalize across different imaging scanners and patient populations. It specifically tests whether injecting stable, multi-scale topological features improves robustness to domain shifts compared to standard grayscale inputs. Use when the user wants to benchmark on INbreast, or asks about evaluating this task. Reports patient-level AUC.

researchpythongo
0
3
Include EvalA

Evaluates multilingual language understanding and regional/cultural knowledge across 44 languages using native-language exam questions. Probes models' ability to handle region-specific contexts without English bias or translation artifacts, and measures performance variance across languages and prompting strategies. Use when the user wants to benchmark on INCLUDE-base, INCLUDE-lite, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Incompebench EvalA

This benchmark evaluates fine-grained music information retrieval by measuring how well models match audio tracks to diverse text queries. It probes the ability to capture nuanced musical attributes, handle negations, and rank candidates based on graded relevance rather than binary matches. Use when the user wants to benchmark on IncompeBench, or asks about evaluating this task. Reports graded relevance (0-3).

researchpythongo
0
3
India Weather Bench EvalA

Evaluates data-driven regional weather forecasting models over India under varying boundary conditioning strategies. It probes the ability of architectures to accurately predict multi-variable meteorological fields at high resolution and assesses their robustness during extreme weather events like heatwaves. Use when the user wants to benchmark on IndiaWeatherBench, or asks about evaluating this task. Reports RMSE.

researchpythongo
0
3
Indian Air Quality Forecasting EvalA

Evaluates the ability of hybrid time-series models to forecast urban air quality indices and specific pollutant concentrations (PM2.5, O3, CO, NOx) using historical environmental data and temporal features. Use when the user wants to benchmark on CPCB Indian Air Quality Dataset, or asks about evaluating this task. Reports RMSE.

datapythonperformance
0
3
Indic Instruct EvalA

This evaluation probes the multilingual instruction-following, natural language understanding, and generation capabilities of LLMs fine-tuned on 13 Indic languages. It measures performance on standardized academic benchmarks across NLU and NLG tasks, as well as real-world cultural relevance and helpfulness through pairwise LLM-as-a-judge comparisons. Use when the user wants to benchmark on MMLU Indic (MMLU-I), ARC Indic (ARC-I), BoolQ Indic (BoolQ-I), TriviaQA Indic (TVQA-I), BeleBele (Bele),...

researchpythongo
0
3
Indic Nmt EvalA

Evaluates neural machine translation models for Indic languages by measuring translation quality against reference texts across multiple standard benchmarks. It specifically tests the effectiveness of training on the Samanantar parallel corpus compared to commercial systems and existing open-source baselines. Use when the user wants to benchmark on WAT2020 Indic task, WAT2021 Multi-IndicMT task, WMT test sets (2014, 2019, 2020), UFAL Entam, FLORES test set, PMIndia en-as testset, or asks abou...

researchpythongo
0
3
Indic Oov EvalA

This evaluation probes the out-of-vocabulary (OOV) intelligibility and perceptual quality of Indian Text-to-Speech systems. It measures how well models synthesize rare or unseen words while preserving speaker similarity and overall audio fidelity compared to ground-truth recordings. Use when the user wants to benchmark on IndicTTS, or asks about evaluating this task. Reports Intelligibility Error Rate (%).

researchpythongo
0
3
Indicaleval EvalA

Evaluates large language models' reasoning capabilities on authentic Indian high-stakes examination questions across STEM and humanities domains. It specifically probes bilingual reasoning, cross-lingual performance differentials, and the impact of prompting strategies (Zero-Shot, Few-Shot, Chain-of-Thought) on model accuracy. Use when the user wants to benchmark on IndicEval, or asks about evaluating this task. Reports exact-match accuracy.

researchpythongo
0
3
Indicdb EvalA

Evaluates cross-lingual semantic parsing and Text-to-SQL capabilities across Indian languages, probing a model's ability to link natural language queries to complex, high-join-depth relational schemas and generate syntactically and semantically correct SQL. Use when the user wants to benchmark on IndicDB, or asks about evaluating this task. Reports Execution Accuracy (EX).

databasespythongo
0
3
Indicgenbench EvalA

Evaluates the multilingual and cross-lingual generation capabilities of LLMs across 29 Indic languages, covering summarization, machine translation, and question answering. It probes how model performance scales with language resourcedness, in-context learning, and fine-tuning. Use when the user wants to benchmark on CrossSum-In, Flores-In, XQuAD-In, XorQA-In, or asks about evaluating this task. Reports Character-F1 (ChrF), SQuAD-style Token-F1.

researchpythongo
0
3
Indicmmlu Pro EvalA

Evaluates large language models on multi-task language understanding across nine major Indic languages. It probes capabilities in reading comprehension, reasoning, and knowledge retention by adapting the English MMLU-Pro benchmark through machine translation and rigorous quality assurance. Use when the user wants to benchmark on IndicMMLU-Pro, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Indictrans2 EvalA

This benchmark evaluates multilingual machine translation quality across 22 scheduled Indian languages and English. It probes a model's ability to handle diverse domains (news, web, conversation, legal, etc.) and both Indic-to-English and English-to-Indic translation directions in an n-way parallel setting. Use when the user wants to benchmark on IN22, FLORES-200, NTREX, WMT (2014, 2019, 2020), WAT (2020, 2021), UFAL, or asks about evaluating this task. Reports chrF++.

researchpythongo
0
3
Indicxnli EvalA

Evaluates multilingual natural language inference capabilities across 11 Indic languages, probing both intra-lingual reasoning and cross-lingual transfer performance of pre-trained language models. Use when the user wants to benchmark on IndicXNLI, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Indicxtreme EvalA

Evaluates zero-shot cross-lingual transfer capabilities of language models on low-resource Indic languages by testing performance on nine natural language understanding tasks after training exclusively on English data. Use when the user wants to benchmark on IndicXTREME, or asks about evaluating this task. Reports task-specific accuracy/F1.

researchpythongo
0
3
Indimathbench EvalA

This benchmark evaluates the ability of LLMs to autoformalize natural language mathematical problems into correct Lean 4 theorems and subsequently prove them. It probes semantic equivalence, syntactic structural similarity, and automated theorem proving success rates on Olympiad-level geometry and algebra problems. Use when the user wants to benchmark on IndiMathBench, or asks about evaluating this task. Reports BEq.

researchpythongo
0
3
Indolem EvalA

Evaluates Indonesian NLP capabilities across morpho-syntax, semantics, and discourse. It probes token-level labeling (POS, NER), syntactic structure (dependency parsing), text classification (sentiment), generation (summarization), and discourse coherence (next tweet prediction, tweet ordering). Use when the user wants to benchmark on INDOLEM, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Indonesian Pos Tagging EvalA

Evaluates sequence labeling performance on Indonesian text by assigning part-of-speech tags to tokens. It probes morphological feature extraction, contextual understanding, and robustness to annotation inconsistencies and rare lexical categories. Use when the user wants to benchmark on IDN Tagged Corpus, or asks about evaluating this task. Reports F1.

researchpythongo
0
3
Indonlu EvalA

This benchmark evaluates Indonesian natural language understanding across 12 diverse tasks, including single-sentence classification, sentence-pair classification, and sequence labeling/tagging. It probes a model's ability to handle sentiment analysis, aspect-based sentiment, textual entailment, part-of-speech tagging, named entity recognition, keyphrase extraction, and question answering in Indonesian. Use when the user wants to benchmark on IndoNLU, or asks about evaluating this task. Repor...

researchpythongo
0
3
Indoor Lidar EvalA

Evaluates 3D object detection and BEV perception capabilities on indoor robotic platforms using LiDAR point clouds. It probes a model's ability to classify indoor objects and localize them with 3D bounding boxes, specifically highlighting the sim-to-real transfer gap in controlled indoor environments. Use when the user wants to benchmark on INDOOR-LiDAR, or asks about evaluating this task. Reports Mean IoU.

researchpythongo
0
3
Indotabvqa EvalA

Probes cross-lingual visual question answering on document images containing tables. It tests a model's ability to perform factual lookup, numerical comparison, aggregation, and structural reasoning across Bahasa Indonesia, English, Hindi, and Arabic. Use when the user wants to benchmark on IndoTabVQA, or asks about evaluating this task. Reports In-Match Accuracy.

researchpythongo
0
3
Inducer Tuning EvalA

Evaluates parameter-efficient fine-tuning methods on natural language understanding and generation tasks, measuring how well they approximate full fine-tuning performance while using significantly fewer trainable parameters. Use when the user wants to benchmark on MNLI, SST2, WebNLG-challenge, CoQA, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Inductive Link Prediction EvalA

Evaluates a model's ability to predict missing links in knowledge graphs using only topological path information, without relying on entity embeddings. It tests inductive generalization by training on one graph and testing on a disjoint graph with unseen entities. Use when the user wants to benchmark on WN18RR, FB15K-237, NELL-995 (inductive versions v1-v4), or asks about evaluating this task. Reports Hits@1.

researchpythontesting
0
3
Industrial Spill Detection EvalA

This benchmark evaluates an AI system's ability to detect and localize industrial safety hazards (e.g., oil spills, chemical stains) in complex factory environments. It probes the model's spatial grounding precision and its capacity to generalize from synthetic or limited real-world data to unseen operational sites. Use when the user wants to benchmark on Public Spill Data, Proprietary Factory Data, Synthetic Spill (SynSpill) Dataset, or asks about evaluating this task. Reports mean hit rate@...

researchpython
0
3
Industryshapes EvalA

This benchmark evaluates 6D object pose estimation, detection, and segmentation capabilities in realistic industrial environments. It specifically probes a model's ability to handle challenging conditions such as heavy occlusion, background clutter, reflective surfaces, textureless materials, and object symmetry. Use when the user wants to benchmark on IndustryShapes Classic, IndustryShapes Extended, or asks about evaluating this task. Reports Average Recall (AR).

researchpythonperformance
0
3
Ineqmath EvalA

This benchmark evaluates large language models' ability to perform informal mathematical reasoning on Olympiad-level inequality problems. It probes step-wise deductive chain integrity by decomposing proofs into bound estimation and relation prediction subtasks, requiring models to generate logically sound derivations rather than just final answers. Use when the user wants to benchmark on IneqMath, or asks about evaluating this task. Reports LLM-as-judge accuracy.

researchpythongo
0
3
Inference Framework Benchmark EvalA

Evaluates the inference performance of four deep learning frameworks (TensorRT, ONNX Runtime, OpenVINO, TensorFlow XLA) across four CNN architectures on GPU hardware. It probes how configuration settings, graph optimizations, and batch sizes impact inference speed and resource utilization, including co-localized model ensembles. Use when the user wants to benchmark on ImageNet, or asks about evaluating this task. Reports speed.

researchpythongit
0
3
Inference LatencyA

Probes how different CPU microarchitectures (Haswell, Broadwell, Skylake) and cache hierarchies affect the inference latency and throughput of production-scale DNN recommendation models under varying batch sizes and co-location scenarios. Use when the user has predictions and gold and needs to compute inference latency.

researchpythongo
0
3
Infinite Dsprites EvalA

Evaluates continual learning methods on a procedurally generated benchmark of 500 shape classification tasks. It probes a model's ability to learn incrementally over a long horizon without catastrophic forgetting, while maintaining open-set recognition and one-shot generalization capabilities on unseen shapes. Use when the user wants to benchmark on Infinite dSprites (idSprites), or asks about evaluating this task. Reports average test accuracy.

researchpythongo
0
3
Infinity Instruct EvalA

Evaluates the conversational and foundational capabilities of LLMs fine-tuned on instruction datasets, comparing them against proprietary and open-source baselines across multiple standard benchmarks. Use when the user wants to benchmark on AlpacaEval 2.0, Arena-Hard, MT-Bench, MATH, GSM-8K, HumanEval, MBPP, MMLU, LUC-EVAL, or asks about evaluating this task. Reports Overall*.

researchpythongo
0
3
Influence Estimation EvalA

This protocol evaluates the accuracy and efficiency of influence function approximation methods (DataInf, LiSSA, Hessian-free) in matching exact influence values, detecting mislabeled training data, and identifying training points that most impact a test instance's loss across text and image generation tasks. Use when the user wants to benchmark on GLUE (binary classification subsets), Custom Text Generation Datasets, Custom Image Generation Datasets, or asks about evaluating this task. Repor...

researchpythonperformance
0
3
Infobench EvalA

Evaluates large language models' ability to follow complex, multi-constraint instructions by decomposing them into granular criteria (Content, Linguistic, Style, Format, Number) and measuring adherence. It probes fine-grained instruction following rather than holistic response quality. Use when the user wants to benchmark on InFoBench, or asks about evaluating this task. Reports DRFR.

researchpythongo
0
3
InfolmA

Compute the InfoLM metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute InfoLM, or asks how to score with InfoLM.

documentationpythonaws
0
3