All authors
qhjqhj00 avatar

Claude Skills by qhjqhj00

github.com/qhjqhj00
7,636 skillsA× 7,623B× 11C× 1D× 10 installs2,201 views
Realistic Ood Detection EvalA

Evaluates the robustness of Out-of-Distribution (OOD) detection models under realistic distribution shifts caused by semantic-preserving transformations. It measures how well detectors distinguish between true out-of-distribution samples and inlier samples that have undergone common corruptions or augmentations, revealing performance gaps that standard benchmarks miss. Use when the user wants to benchmark on CIFAR-10-R, CIFAR-100-R, ImageNet-30-R, or asks about evaluating this task. Reports A...

researchpythongo
0
3
Realpdebench EvalA

Evaluates scientific machine learning models on sim-to-real transfer for complex physical systems. It probes a model's ability to predict spatiotemporal dynamics from real-world measurements, leveraging simulated pretraining, and assesses long-term prediction stability under autoregressive rollout. Use when the user wants to benchmark on Cylinder, ControlledCylinder, FSI, Foil, Combustion, or asks about evaluating this task. Reports RMSE.

datapythongo
0
3
Realsi Sst EvalA

Evaluates end-to-end simultaneous speech-to-speech translation quality and latency on long-form, multi-domain continuous speech. Probes the model's ability to maintain semantic accuracy, speaker voice characteristics, and low delay in real-time multilingual dialogue. Use when the user wants to benchmark on RealSI, or asks about evaluating this task. Reports VIP.

researchpythonperformance
0
3
Realtime Simulation Checking EvalA

Evaluates the scalability and performance of a simulation-checking algorithm for timed automata under fairness assumptions. It measures how efficiently the algorithm verifies liveness properties and handles state-space explosion across parameterized real-time system benchmarks. Use when the user wants to benchmark on Fischer's timed mutual exclusion algorithm, CSMA/CD, Timed consumer/producer, Network of TAs, or asks about evaluating this task. Reports CPU time.

researchpythongo
0
3
Realx3d Smoke EvalA

Evaluates the ability of 3D reconstruction models to synthesize novel views of smoke-degraded scenes with high photometric fidelity and structural preservation. It probes view-dependent medium modeling and multi-view consistency under severe scattering conditions. Use when the user wants to benchmark on RealX3D (NTIRE 2026 Track 2 Smoke Subset), or asks about evaluating this task. Reports PSNR.

researchpython
0
3
Reasoning Accuracy EvalA

Evaluates large language models' logical reasoning and problem-solving capabilities across mathematical, algorithmic, and creative tasks. It measures both the correctness of final answers and the computational efficiency of the reasoning process. Use when the user wants to benchmark on Game of 24, BIG-Bench (subset), Python Puzzles, MGSM, Shakespearean Sonnet Writing, or asks about evaluating this task. Reports Acc_logic.

ai-agentspythongo
0
3
Reasoning Benchmarks EvalA

Evaluates language model reasoning and general knowledge across multiple standard benchmarks. It measures the impact of data mixture optimization on model performance using accuracy scores on commonsense, scientific, and factual QA tasks. Use when the user wants to benchmark on PIQA, ARC_C, ARC_E, HellaSwag, WinoGrande, SIQA, MMLU, or asks about evaluating this task. Reports test accuracy.

researchpythongo
0
3
Reasoning Collab Memory EvalA

Evaluates LLM logical and spatial reasoning capabilities under multi-agent collaboration and memory-augmented prompting. It probes how different reasoning styles, exemplar retrieval methods, and answer aggregation strategies impact accuracy on formal logic and object-tracking tasks. Use when the user wants to benchmark on FOLIO, RACO, TSO, or asks about evaluating this task. Reports accuracy.

ai-agentspythongo
0
3
Reasoning Sft EvalA

This evaluation protocol assesses the cross-domain generalization, safety, and instruction-following capabilities of models after reasoning-focused supervised fine-tuning (SFT). It measures in-domain math performance, out-of-domain reasoning in coding and science, general instruction following, and resistance to harmful queries. Use when the user wants to benchmark on MATH500, AIME24, LiveCodeBench v2, GPQA-Diamond, MMLU-Pro, IFEval, AlpacaEval 2.0, HaluEval, TruthfulQA, HEx-PHI, or asks abou...

researchpythongo
0
3
Reasonplan3d EvalA

Probes a model's ability to infer implicit human intentions from natural language and generate multi-step, route-aware activity plans grounded in 3D scene segmentation. It evaluates both textual planning coherence and spatial reasoning over 3D environments. Use when the user wants to benchmark on ReasonPlan3D, or asks about evaluating this task. Reports BLEU-4.

researchpython
0
3
Reasonseg EvalA

Evaluates a model's ability to generate precise segmentation masks from implicit, complex text queries that require reasoning and world knowledge. It specifically probes whether the model can move beyond simple explicit referring expressions to handle multi-step logical deductions and visual grounding simultaneously. Use when the user wants to benchmark on ReasonSeg, refCOCO, refCOCO+, refCOCOg, or asks about evaluating this task. Reports gIoU.

researchpythonexpress
0
3
Rec Auc EvalA

Evaluates the predictive performance of recommendation models on large-scale click-through rate datasets. It specifically probes how model scalability and embedding size affect ranking quality, revealing the phenomenon of embedding collapse when scaling up feature interactions. Use when the user wants to benchmark on Criteo, Avazu, or asks about evaluating this task. Reports AUC.

researchpythongo
0
3
Rec Splitting EvalA

Evaluates how different data splitting strategies (leave-one-last-item, leave-one-last-basket, temporal global) impact the performance ranking of recommendation models on e-commerce datasets. It probes whether evaluation protocols introduce temporal leakage or distribution shifts that confound model comparisons and invalidate cross-paper rankings. Use when the user wants to benchmark on Tafeng Dataset, Dunnhumby Dataset, or asks about evaluating this task. Reports NDCG@10.

researchpythonperformance
0
3
Recall ScoreA

Compute the recall_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute recall_score, or asks how to score with recall_score.

documentationpython
0
3
Recall Throughput EvalA

Evaluates language models on associative recall, information extraction, and question answering from long contexts, while measuring generation throughput and language modeling perplexity. It probes the tradeoff between memory efficiency, recall accuracy, and inference speed across synthetic and real-world benchmarks. Use when the user wants to benchmark on Pile, SWDE, FDA, SQUAD, LM Eval Harness (SuperGLUE, ARC, PIQA, WinoGrande, HellaSwag, LAMBADA), or asks about evaluating this task. Report...

researchpythongo
0
3
RecallA

Compute the Recall metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute Recall, or asks how to score with Recall.

documentationpythondocumentation
0
3
RecallatfixedprecisionA

Compute the RecallAtFixedPrecision metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute RecallAtFixedPrecision, or asks how to score with RecallAtFixedPrecision.

documentationpythondocumentation
0
3
Recaptioning Image Gen EvalA

Evaluates the semantic fidelity, object accuracy, and prompt adherence of text-to-image generation models by comparing automated metrics and human ratings on standard benchmarks. Use when the user wants to benchmark on MS-COCO validation set, DrawBench, or asks about evaluating this task. Reports FID.

researchpythonperformance
0
3
Receiptsense EvalA

Evaluates multilingual receipt understanding across four tasks: question answering, object detection/classification, OCR, and information extraction. Probes model capabilities in handling real-world noise, mixed Arabic-English layouts, and complex formatting. Use when the user wants to benchmark on ReceiptSense, or asks about evaluating this task. Reports exact match.

researchpythongo
0
3
Receiver Placement EvalA

Tests an algorithm's ability to optimally place a receiver in 3D indoor environments to maximize speech intelligibility, measured by the Speech Transmission Index (STI). It evaluates how well the optimization handles complex acoustic properties like reverberation and noise across different scene geometries. Use when the user wants to benchmark on Office, Berlin, Suburban 3D scenes, or asks about evaluating this task. Reports STI.

researchpythongo
0
3
Recipe1mplus Retrieval EvalA

Evaluates cross-modal retrieval between food images and cooking recipes. It probes a model's ability to align visual and textual representations in a shared embedding space to rank relevant recipes given an image, and vice versa. Use when the user wants to benchmark on Recipe1M+, or asks about evaluating this task. Reports medR.

researchpythongo
0
3
Recogdrive EvalA

Evaluates an end-to-end autonomous driving agent's ability to generate safe, comfortable, and efficient driving trajectories using only camera inputs. It probes the model's closed-loop planning capabilities, safety-critical scenario handling, and visual reasoning in complex urban environments. Use when the user wants to benchmark on NAVSIM, Bench2Drive, or asks about evaluating this task. Reports PDMS.

researchpythongo
0
3
Reconstruction Pose EvalA

Evaluates how well different 3D reconstruction methods perform in a downstream object pose estimation task, rather than measuring standalone geometric reconstruction accuracy. It compares pose estimation results using reconstructed 3D models against those using ground-truth CAD models. Use when the user wants to benchmark on YCB-V, or asks about evaluating this task. Reports accuracy of the estimated poses.

researchpythongit
0
3
Recruitview EvalA

This benchmark evaluates multimodal models on predicting continuous personality traits and interview performance scores from video, audio, and text inputs. It probes the model's ability to perform fine-grained behavioral analysis and regression across psychometric targets. Use when the user wants to benchmark on RecruitView, or asks about evaluating this task. Reports Spearman's ρ.

researchpythongo
0
3
Recsys Challenge 2015 EvalA

Evaluates session-based recommendation models by predicting the next item in a user's clickstream sequence. It probes the model's ability to capture temporal dynamics and handle data sparsity in e-commerce sessions. Use when the user wants to benchmark on RecSys Challenge 2015, or asks about evaluating this task. Reports Recall@20.

researchpythonexpress
0
3
Recsys2015 Session Recomm EvalA

Evaluates session-based recommendation models by predicting the next item in a user's browsing sequence. It measures ranking quality and prediction efficiency to assess accuracy and deployability in real-time recommender systems. Use when the user wants to benchmark on RecSys Challenge 2015 dataset, or asks about evaluating this task. Reports Recall@20.

researchpythongo
0
3
Rectom EvalA

Evaluates machine theory of mind in LLM-based conversational recommender systems by testing cognitive inference (fine/coarse intention, belief) and behavioral prediction (prediction, judgement) for both recommender and seeker roles in dialogue scenarios. Use when the user wants to benchmark on RECTOM, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Red Diffeq Fwi EvalA

Evaluates the ability of a diffusion-based regularization framework to reconstruct high-resolution subsurface velocity models from seismic data. It probes robustness under varying data conditions, including clean recordings, Gaussian noise contamination, and missing traces. The benchmark also tests out-of-distribution generalization on complex geological structures. Use when the user wants to benchmark on OpenFWI, Marmousi, or asks about evaluating this task. Reports RMSE.

researchpython
0
3
Red Teaming EvalA

Evaluates the quality of AI-generated adversarial responses across three dimensions: adherence to malicious objectives (toxicity), logical/semantic consistency (coherence), and textual variation (diversity). It assesses whether a model can produce high-quality, diverse, and coherent toxic content for red-teaming without suffering from reward hacking or semantic drift. Use when the user wants to benchmark on Curated Red-Teaming Dataset, or asks about evaluating this task. Reports Toxicity-Util...

researchpythongo
0
3
Red1bluelost Evaluate Genericify CppA

Compute red1bluelost/evaluate_genericify_cpp via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of red1bluelost/evaluate_genericify_cpp.

developmentpython
0
3
Redbench EvalA

Evaluates LLM robustness against adversarial prompts (Attack Success Rate) and their tendency to over-defend on benign prompts (Rejection Rate). It probes safety alignment, refusal behavior, and cross-domain vulnerability across 22 risk categories and 19 domains. Use when the user wants to benchmark on RedBench, or asks about evaluating this task. Reports Rejection Rate (RR), Attack Success Rate (ASR).

researchpythongo
0
3
Redcodeagent EvalA

This evaluation probes the ability of automated red-teaming agents to successfully jailbreak diverse code-generating AI assistants. It measures how effectively an attacker can craft and optimize malicious prompts to bypass safety guardrails and force the execution of harmful code across multiple programming languages and agent architectures. Use when the user wants to benchmark on RedCode-Exec, RedCode-Gen, RMCbench, or asks about evaluating this task. Reports attack success rate (ASR).

ai-agentspythongo
0
3
Reddit Cssrs Screening EvalA

This benchmark evaluates zero-shot large language models on their ability to classify suicide risk severity from Reddit posts using the clinically validated Columbia-Suicide Severity Rating Scale (C-SSRS). It probes the models' ordinal classification capabilities, intent detection, and alignment with human clinical annotations across seven severity levels. Use when the user wants to benchmark on Reddit r/SuicideWatch posts (C-SSRS labeled), or asks about evaluating this task. Reports F1-Score.

researchpythongo
0
3
Reddit Tifu Summarization EvalA

This evaluation probes a model's ability to generate abstractive summaries from informal, user-generated text and formal documents. It measures how well the model captures long-range dependencies and abstracts key information without relying on extractive heuristics. Use when the user wants to benchmark on Reddit TIFU, Newsroom-Abs, XSum, or asks about evaluating this task. Reports ROUGE-1.

researchpythongo
0
3
Redial Movie Recommendation EvalA

Evaluates modular components of a conversational recommendation system, specifically cold-start movie rating prediction and movie opinion sentiment analysis (seen/liked status) from dialogue text. Use when the user wants to benchmark on REDIAL, MovieLens, or asks about evaluating this task. Reports RMSE.

researchpython
0
3
Redis Qa EvalA

Evaluates large language models' ability to answer questions about rare diseases, including diagnosis, symptoms, causes, and related properties. It probes the models' medical knowledge retrieval and reasoning capabilities in a specialized, low-resource domain. Use when the user wants to benchmark on ReDis-QA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Redrft EvalA

Evaluates reinforcement fine-tuning methods for red-teaming LLMs by measuring the toxicity and diversity of generated adversarial prompts across toxic continuation and instruction-following tasks. Use when the user wants to benchmark on toxic continuation, instruction following, or asks about evaluating this task. Reports cumulative toxicity-diversity score.

researchpythonperformance
0
3
Redstar Reasoning EvalA

Evaluates the model's ability to solve complex mathematical, coding, and general reasoning problems. It probes multi-step reasoning, domain-specific knowledge integration, and long-context handling across diverse benchmarks. Use when the user wants to benchmark on Math & Reasoning Benchmarks, Hellobench, SedarEval, Chinese Graduate Entrance Mathematics Test, or asks about evaluating this task. Reports AVG.

researchpythongo
0
3
Redundancy Detection EvalA

Evaluates a model's ability to determine whether pairs of P, I, or O spans in an abstract refer to the same underlying information. Use when the user wants to benchmark on EBM-NLP, or asks about evaluating this task. Reports F-1.

businesspythongo
0
3
Reef Substrate Classification EvalA

Evaluates an AI model's ability to classify underwater substrates for autonomous coral reseeding deployment. It probes both fine-grained patch-level semantic segmentation (distinguishing coral, deploy, and no-deploy zones) and coarse-grained image-level decision making for real-time marine robotics. Use when the user wants to benchmark on Great Barrier Reef ReefScan Dataset, or asks about evaluating this task. Reports Macro F1.

devopspythongo
0
3
Ref Adv EvalA

Probes multimodal large language models' ability to perform visual grounding and complex textual reasoning under challenging conditions. It specifically tests whether models rely on shortcut cues or genuinely comprehend referring expressions when faced with linguistically nontrivial descriptions and hard distractors. Use when the user wants to benchmark on Ref-Adv, or asks about evaluating this task. Reports Acc0.5.

ai-agentspythonexpress
0
3
Ref Avs EvalA

Evaluates a model's ability to segment objects in audio-visual videos based on natural language referring expressions. It tests both seen categories and generalization to unseen categories, as well as handling null references where no object exists. Use when the user wants to benchmark on Ref-AVS Dataset, or asks about evaluating this task. Reports Jaccard Index ($\mathcal{J}$).

researchpythongo
0
3
Refact EvalA

This benchmark evaluates large language models' ability to detect, localize, and correct scientific confabulations in generated answers. It probes fine-grained factuality awareness, span-level error identification, and factual restoration capabilities under domain-specific scrutiny. Use when the user wants to benchmark on ReFACT, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Refaerial EvalA

Evaluates a model's ability to localize a target object in an aerial image based on a fine-grained natural language description. It probes cross-modal alignment, scale-invariant object detection, and handling of complex backgrounds with numerous distractors. Use when the user wants to benchmark on RefAerial, RefCOCO, RefCOCO+, RefCOCOg, or asks about evaluating this task. Reports mP (average P@0.5/0.6/0.7/0.8).

researchpythonexpress
0
3
Refchartqa EvalA

This benchmark evaluates a model's ability to answer questions about chart images while simultaneously localizing the visual evidence (via bounding boxes) that supports the answer. It probes spatial-text alignment, arithmetic and logical reasoning over charts, and hallucination reduction through explicit grounding. Use when the user wants to benchmark on RefChartQA, or asks about evaluating this task. Reports answer accuracy.

researchpythongo
0
3
Refcoco Grounding EvalA

Evaluates fine-grained visual grounding and spatial reasoning by measuring how accurately a multimodal model can localize regions in images corresponding to given referring expressions under varying linguistic and spatial conditions. Use when the user wants to benchmark on RefCOCO, RefCOCO+, RefCOCOg, or asks about evaluating this task. Reports IoU@50 accuracy.

researchpythongo
0
3
Refcocom EvalA

Evaluates a model's ability to perform referring expression segmentation at both object and part levels. It probes fine-grained cross-modal alignment and pixel-level semantic understanding by requiring precise mask prediction for diverse textual references. Use when the user wants to benchmark on RefCOCOm, or asks about evaluating this task. Reports mIoU.

researchpythongo
0
3
Refedit EvalA

Evaluates instruction-based image editing models on referring expressions, measuring how well they align edits with text instructions while preserving background and maintaining perceptual quality. It probes multi-object scene editing, background preservation, and the ability to handle complex spatial grounding without relying on CLIP. Use when the user wants to benchmark on RefEdit-Bench, PIE-Bench, or asks about evaluating this task. Reports VIEScore.

ai-agentspythongo
0
3
Refereebench EvalA

Evaluates Multimodal Large Language Models (MLLMs) on automatic sports refereeing tasks, probing their ability to detect incidents, classify fouls, apply sport-specific rules, and ground decisions temporally across 11 different sports. Use when the user wants to benchmark on RefereeBench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Referit Segmentation EvalA

This benchmark evaluates a model's ability to perform pixel-level image segmentation conditioned on natural language expressions. It probes spatial reasoning, attribute grounding, and fine-grained visual-linguistic alignment by requiring the model to segment specific objects or amorphous regions described in text. Use when the user wants to benchmark on ReferIt, or asks about evaluating this task. Reports prec@0.5.

researchpythongo
0
3