All authors
qhjqhj00 avatar

Claude Skills by qhjqhj00

github.com/qhjqhj00
7,636 skillsA× 7,623B× 11C× 1D× 10 installs2,402 views
Iot Nids Poisoning EvalA

This evaluation probes the robustness of supervised machine learning models for IoT intrusion detection when their training data is corrupted by adversarial poisoning attacks. It measures how different model architectures degrade in detection capability under label manipulation, outlier injection, and feature impersonation. Use when the user wants to benchmark on CICIoT2023, Edge-IIoTset, N-BaIoT, or asks about evaluating this task. Reports Accuracy.

datapythongo
0
3
Ipds EvalA

Evaluates large language models' ability to support inpatient clinical decision-making by classifying patient cases into appropriate triage, diagnosis, and treatment pathways. It probes the models' clinical reasoning, diagnostic accuracy, and alignment with real-world physician judgments. Use when the user wants to benchmark on IPDS, or asks about evaluating this task. Reports accuracy.

ai-agentspythongo
0
3
Ipho 2025 Theory EvalA

Evaluates an AI agent's ability to solve complex, multi-part physics theory problems that require integrating visual data extraction, causal reasoning, and self-correction via tool use. The benchmark probes whether agentic architectures with domain-specific tools can match elite human performance on standardized physics competitions. Use when the user wants to benchmark on IPhO 2025 Theory Problems, or asks about evaluating this task. Reports score.

researchpythongo
0
3
Iplotbench EvalA

Evaluates a visualization agent's ability to reconstruct interactive charts from static images and answer binary questions about them. It probes the agent's capacity for spec-grounded introspection and view-grounded interaction to resolve visual ambiguities like overlapping geometries. Use when the user wants to benchmark on iPlotBench, or asks about evaluating this task. Reports Question-level accuracy.

researchpythongo
0
3
Ipqa EvalA

This benchmark evaluates a model's ability to identify core user intents in personalized question answering. It probes whether systems can infer prioritized motivations from a user's historical Q&A interactions and a target question's narrative, rather than relying on explicit user statements. Use when the user wants to benchmark on IPQA, or asks about evaluating this task. Reports IPQA-Eval F1.

researchpythonperformance
0
3
Iquad V1 EvalA

Evaluates an agent's ability to navigate, perceive, and interact with dynamic 3D environments to answer visual questions. It probes spatial reasoning, long-term memory of object locations, and the ability to plan exploration based on question semantics. Use when the user wants to benchmark on iquad v1, or asks about evaluating this task. Reports Top-1 question answering accuracy.

researchpythongo
0
3
Ir Metric Correlation EvalA

Evaluates the correlation and predictability of standard information retrieval metrics across multiple TREC test collections. It probes how well low-cost metrics can predict high-cost ones and how metric values vary between topic-wise and system-wise aggregations. Use when the user wants to benchmark on TREC Web & Robust Tracks (2000-2014), or asks about evaluating this task. Reports MAP.

researchpythongo
0
3
Ir Triplet EvalA

Evaluates the model's ability to capture semantic similarity between paragraphs for information retrieval. It tests whether fixed-length vector representations can effectively distinguish query-related documents from irrelevant ones using a triplet ranking protocol. Use when the user wants to benchmark on Information Retrieval (Paragraph Vectors), or asks about evaluating this task. Reports error rate.

researchpython
0
3
Ir3d Bench EvalA

Evaluates vision-language models' ability to understand 3D scenes by generating executable scene descriptions from a single 2D image. It shifts evaluation from passive captioning to active reconstruction, probing geometric layout, spatial reasoning, object appearance, and semantic attributes. Use when the user wants to benchmark on IR3D-Bench, or asks about evaluating this task. Reports Pixel Distance.

ai-agentspython
0
3
Ircad Liver EvalA

Evaluates medical image segmentation models on liver CT volumes, testing their ability to accurately delineate organ boundaries using interactive or automatic refinement techniques. The protocol measures how well models handle low-contrast boundaries and varying slice geometries in clinical imaging. Use when the user wants to benchmark on IRCAD, or asks about evaluating this task. Reports Dice coefficient.

researchpythongo
0
3
Irefvla EvalA

Evaluates a model's ability to ground referential language in 3D scenes when references are imperfect or ambiguous. It probes whether the model can correctly identify existing objects, detect non-existent references, and generate plausible alternative objects based on spatial and semantic reasoning. Use when the user wants to benchmark on IRef-VLA, or asks about evaluating this task. Reports score_sim.

researchpythongo
0
3
Iris Benchmark EvalA

Probes fairness across understanding and generation tasks in Unified Multimodal Large Language Models (UMLLMs) by measuring Ideal Fairness, Real-world Fidelity, and Bias Inertia & Steerability across demographic attributes. It reveals systemic trade-offs, generation gaps, and personality splits that single-task or single-metric evaluations miss. Use when the user wants to benchmark on IRIS Benchmark, or asks about evaluating this task. Reports IRIS-Score.

researchpythongo
0
3
Irish English St EvalA

Evaluates end-to-end speech translation from Irish to English, specifically probing how synthetic audio data and augmentation techniques (noise, VAD) impact model performance in low-resource settings. Use when the user wants to benchmark on IWSLT-2023, FLEURS, Bitesize, SpokenWords, or asks about evaluating this task. Reports chrF++.

researchpythongit
0
3
Irpapers EvalA

Evaluates the ability of multimodal and text-only models to retrieve relevant scientific paper pages and answer questions based on those pages. It probes retrieval depth, modality complementarity, and the impact of context quantity on RAG performance. Use when the user wants to benchmark on IRPAPERS, or asks about evaluating this task. Reports Recall@1.

researchpythongo
0
3
Irsc EvalA

Evaluates embedding models on multilingual information retrieval tasks across five query types (query, title, part-of-paragraph, keyword, summary). It probes semantic comprehension and cross-lingual retrieval alignment in Retrieval-Augmented Generation (RAG) scenarios. Use when the user wants to benchmark on IRSC Benchmark, or asks about evaluating this task. Reports r@10.

researchpythongo
0
3
Irt2 EvalA

Evaluates neural and baseline models on inductive link prediction and ranking tasks across knowledge graphs of varying scales. It probes the models' ability to map textual entity mentions to graph vertices and rank candidate entities based on combined textual and structural signals, particularly under data scarcity conditions. Use when the user wants to benchmark on IRT2, or asks about evaluating this task. Reports MRR.

researchpythongo
0
3
Isaacsim Kitchen EvalA

Evaluates a robot's ability to decompose high-level language instructions into executable task plans and execute them in a simulated kitchen environment. It jointly measures planning accuracy and low-level control success under strict time and spatial constraints. Use when the user wants to benchmark on IsaacSim Kitchen Benchmark, or asks about evaluating this task. Reports EM.

researchpython
0
3
Isac Lawn Multimodal EvalA

Probes the capability of multimodal fusion and adaptive expert routing for integrated sensing and communication tasks in low-altitude wireless networks. Specifically, it evaluates how well models leverage synchronized visual, lidar, radar, GPS, and RF channel data to predict beam indices, estimate path loss, and track UAV trajectories under dynamic environmental conditions. Use when the user wants to benchmark on Public Multimodal ISAC Dataset for Low-Altitude Scenarios, or asks about evaluat...

researchpythongo
0
3
Isafetybench EvalA

Probes vision-language models' ability to recognize routine and hazardous industrial actions in real-world videos under zero-shot conditions. It tests both single-label precision and multi-label recall in safety-critical contexts, evaluating how well models discriminate between semantically similar distractors and identify multiple concurrent actions. Use when the user wants to benchmark on iSafetyBench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Isdrama EvalA

This evaluation protocol assesses the capability of multimodal speech synthesis models to generate high-fidelity, spatially accurate binaural audio from scripts, poses, and prompts. It probes content accuracy, speaker similarity, prosodic expressiveness, and precise spatial localization (interaural phase/level differences and angle/distance consistency). Use when the user wants to benchmark on MRSDrama, or asks about evaluating this task. Reports IPD MAE.

researchpythonexpress
0
3
Isic Ham Segmentation EvalA

Evaluates dermatologic image segmentation models by measuring how training on real versus synthetic data affects performance on held-out real test sets, and how model accuracy correlates with controllable synthetic image parameters like skin tone and lesion shape. Use when the user wants to benchmark on ISIC, HAM, or asks about evaluating this task. Reports Dice score.

researchpythonperformance
0
3
Isign EvalA

Evaluates the accuracy of English text generation from Indian Sign Language (ISL) videos and pose sequences. It probes multimodal translation capabilities, specifically how well models align visual sign language signals with corresponding natural language references. Use when the user wants to benchmark on iSign, or asks about evaluating this task. Reports BLEU-4.

researchpythonexpress
0
3
Isles24 Segmentation EvalA

Evaluates the ability of models to perform 3D medical image segmentation for stroke lesion (infarct) and vessel occlusion detection using longitudinal multimodal CT and MRI scans. Use when the user wants to benchmark on ISLES'24, or asks about evaluating this task. Reports Dice Similarity Coefficient (DSC).

researchpythongit
0
3
Ist Unbabel 2022 Qe EvalA

Evaluates machine translation quality estimation (QE) by predicting human quality scores at the sentence level and identifying error locations at the word level. It also assesses the model's ability to generate faithful explanations for predicted errors. Use when the user wants to benchmark on IST-Unbabel 2022 QE Shared Task, or asks about evaluating this task. Reports Spearman's rank correlation, Matthew's correlation coefficient (MCC), Recall@K (R@K).

researchpythongo
0
3
Iterations Per MinuteA

Evaluates the computational throughput and hardware scalability of deep learning image generation models by measuring how many training iterations can be completed per minute on CPU versus GPU across varying image resolutions. Use when the user has predictions and gold and needs to compute Iterations per minute.

code-qualitypythongo
0
3
Iteris Merging EvalA

Evaluates the effectiveness of iterative LoRA merging (IterIS) across text-to-image diffusion, vision-language, and large language models. It probes the model's ability to preserve multiple concepts or styles without mutual interference while maintaining generation quality and task-specific performance metrics. Use when the user wants to benchmark on CustomConcept101, DreamBooth, SentiCap, Emotion datasets (Emoint, EC, TEC, ISEAR, SUM), GLUE benchmark, or asks about evaluating this task. Repo...

researchpythongo
0
3
Iu Rr Radiology Report EvalA

Evaluates a model's ability to generate clinically accurate and structurally coherent radiology reports from multi-view chest X-ray images. It probes cross-modal alignment, medical terminology recall, and the model's capacity to synthesize findings and impressions from visual evidence. Use when the user wants to benchmark on IU-RR, or asks about evaluating this task. Reports BLEU-4.

researchpythontesting
0
3
Iu Xray Report Gen EvalA

Evaluates a vision-language model's ability to generate clinically accurate and semantically coherent radiology reports from chest X-ray images. It probes the model's capacity for medical terminology usage, anatomical consistency, and structured clinical text generation. Use when the user wants to benchmark on IU X-ray, or asks about evaluating this task. Reports ROUGE-L.

researchpythonperformance
0
3
Ivy Fake EvalA

This benchmark evaluates multimodal AI-generated content (AIGC) detection and explainable reasoning capabilities. It probes a model's ability to classify images and videos as real or fake, and to generate natural-language explanations that localize and justify synthetic artifacts. Use when the user wants to benchmark on Ivy-Fake, GenImage, Chameleon, GenVideo, or asks about evaluating this task. Reports Accuracy (Acc).

researchpythongo
0
3
Iwslt2017 Nmt EvalA

Evaluates neural machine translation quality of character-level versus subword models across multiple language pairs. It probes morphological generalization, noise robustness, and the impact of sequence length expansion on training and inference efficiency. Use when the user wants to benchmark on IWSLT 2017, or asks about evaluating this task. Reports BLEU.

researchpythonapi
0
3
Iwslt2023 St EvalA

Evaluates automatic speech translation systems on long-form audio across offline, multilingual, and simultaneous conditions. Probes the model's ability to handle segmentation, resegmentation, and translation quality under varying acoustic and linguistic challenges. Use when the user wants to benchmark on IWSLT2023 TED Test Set, IWSLT2023 ACL Test Set, or asks about evaluating this task. Reports COMET.

researchpython
0
3
Jaccard ScoreA

Compute the jaccard_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute jaccard_score, or asks how to score with jaccard_score.

documentationpython
0
3
JaccardindexA

Compute the JaccardIndex metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute JaccardIndex, or asks how to score with JaccardIndex.

documentationpythondocumentation
0
3
Jailbreak Attack EvalA

This protocol evaluates the robustness of large language models against automated jailbreak attacks. It measures how effectively generated or human-crafted prompts can bypass safety filters to elicit prohibited or harmful responses. Use when the user wants to benchmark on 100 questions from two open datasets [6,37], or asks about evaluating this task. Reports Attack Success Rate (ASR).

researchpython
0
3
Jailbreak Audio Bench EvalA

This benchmark probes the safety alignment and jailbreak resilience of Large Audio-Language Models (LALMs). It specifically tests whether manipulating audio-specific hidden semantics—such as tone, intonation, emotion, and background noise—can bypass safety guardrails and elicit harmful responses more effectively than text-only prompts. Use when the user wants to benchmark on Jailbreak-AudioBench, or asks about evaluating this task. Reports Attack Success Rate (ASR).

researchpythonrails
0
3
Jailbreak EvalA

Evaluates the robustness of LLM safety training against various jailbreak attacks by measuring the rate at which models produce harmful (BAD BOT), helpful (GOOD BOT), or ambiguous (UNCLEAR) responses to curated harmful prompts. Use when the user wants to benchmark on curated dataset, or asks about evaluating this task. Reports BAD BOT.

researchpythongo
0
3
Jailbreakbench EvalA

Evaluates the robustness of large language models against adversarial jailbreaking attacks and defenses. It measures how effectively various attack methods can bypass safety filters (attack success rate) and how well defenses mitigate these attacks while maintaining normal functionality on benign prompts. Use when the user wants to benchmark on JBB-Behaviors, or asks about evaluating this task. Reports attack success rate (ASR).

researchpythongit
0
3
Jam Alt EvalA

Evaluates automatic lyrics transcription (ALT) systems on readability-aware formatting, including punctuation, capitalization, line breaks, and background vocal annotations. It distinguishes errors by token type to measure how well models adhere to industry-standard musical semantics and prosodic structure. Use when the user wants to benchmark on Jam-ALT, Schubert Winterreise Dataset (SWD), or asks about evaluating this task. Reports WER.

researchpythonapi
0
3
Jama Clinical Challenge EvalA

Evaluates medical multimodal models on real-world diagnostic reasoning using clinical case images and questions. It probes both factual accuracy in close-ended QA and the model's ability to generate clinically sound reasoning across key points, inference steps, and evidence citation. Use when the user wants to benchmark on JAMA Clinical Challenge, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Jamendo Mt Qa EvalA

Evaluates audio-language models on multi-track comparative reasoning by asking them to compare two music tracks and answer questions. It probes the model's ability to perform grounded, sentence-level comparative explanations versus simple binary or short-answer discrimination. Use when the user wants to benchmark on Jamendo-MT-QA, or asks about evaluating this task. Reports accuracy, LLM-as-a-Judge score.

researchpythongo
0
3
Japanese Bar Exam Legal Reasoning EvalA

Evaluates open-ended legal reasoning capabilities of LLMs in the Japanese legal domain. It assesses their ability to generate structured, legally accurate arguments based on bar exam writing tasks. Use when the user wants to benchmark on Japanese Bar Exam Writing Task, or asks about evaluating this task. Reports expert_score.

researchpythongo
0
3
Japanese Financial Bench EvalA

Evaluates large language models on Japanese financial domain knowledge across five distinct tasks: sentiment analysis, fundamental financial knowledge, CPA auditing, and two levels of financial planner exam questions. It probes the models' ability to understand and reason over domain-specific multiple-choice questions in Japanese. Use when the user wants to benchmark on chabsa, cma Basics, cpa Audit, fp2, security_sales_1, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Japanese Sts Ir EvalA

Evaluates Japanese sentence embeddings on domain-specific semantic textual similarity (STS) and information retrieval (IR) tasks. It probes the model's ability to capture fine-grained semantic similarity in clinical text and retrieve relevant question-answer pairs in an educational domain. Use when the user wants to benchmark on JACSTS, QABot, or asks about evaluating this task. Reports Spearman's rank correlation.

researchpythongo
0
3
Jaquad EvalA

Extractive machine reading comprehension in Japanese. It probes a model's ability to locate exact answer spans in Japanese Wikipedia text given a question, evaluating performance across different answer types, question reasoning types, and answer lengths. Use when the user wants to benchmark on JaQuAD, or asks about evaluating this task. Reports F1 score.

researchpythongo
0
3
Jarod0411 AucprA

Compute jarod0411/aucpr via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of jarod0411/aucpr.

developmentpython
0
3
Jat Rl EvalA

Evaluates a multi-modal transformer agent's ability to perform sequential decision-making across diverse reinforcement learning domains, including Atari games, grid-world navigation, and continuous control tasks, without task-specific fine-tuning. Use when the user wants to benchmark on Atari 57, BabyAI, MuJoCo, Meta-World, or asks about evaluating this task. Reports expert normalized score.

ai-agentspythonexpress
0
3
Jax Mpm Geophysical Benchmarks EvalA

Tests the accuracy and computational efficiency of a differentiable Material Point Method (MPM) simulator on geophysical flow benchmarks. It probes the framework's ability to reproduce free-surface dynamics, granular collapse rheology, and rigid-body contact against analytical or experimental ground truth, while measuring GPU acceleration speedups. Use when the user wants to benchmark on JAX-MPM Geophysical Benchmarks, or asks about evaluating this task. Reports normalized_runout.

researchpythonperformance
0
3
Jcola EvalA

This benchmark evaluates the ability of neural language models to judge the grammatical acceptability of Japanese sentences. It probes deep syntactic knowledge, particularly long-distance dependencies and linguistic phenomena, by measuring performance on both in-domain and out-of-domain acceptability judgments. Use when the user wants to benchmark on JCoLA, or asks about evaluating this task. Reports Matthews Correlation Coefficient (MCC).

researchpythongit
0
3
JensenshannondivergenceA

Compute the JensenShannonDivergence metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute JensenShannonDivergence, or asks how to score with JensenShannonDivergence.

documentationpython
0
3
Jet Classification EvalA

Evaluates ultra-low-latency supervised classification of particle physics jet signatures on edge hardware. It probes the ability to distinguish rare boson/top-quark jets from common quark/gluon jets under strict microsecond latency and pipeline interval constraints. Use when the user wants to benchmark on LHC Jet Classification Dataset, or asks about evaluating this task. Reports classification accuracy.

researchpythongo
0
3