All authors
qhjqhj00 avatar

Claude Skills by qhjqhj00

github.com/qhjqhj00
7,636 skillsA× 7,623B× 11C× 1D× 10 installs2,193 views
Session Rec Rnn EvalA

Evaluates a model's ability to predict the next item in a user session based on sequential click or watch history. It probes the capability to capture short-term sequential dependencies and maintain context within session-based recommendation scenarios. Use when the user wants to benchmark on RSC15, VIDEO, or asks about evaluating this task. Reports recall@20.

researchpythongo
0
3
Severe++ EvalA

Evaluates the generalization and sensitivity of video self-supervised learning models across four factors: domain shift, sample efficiency, action granularity, and task diversity. It probes how well pre-trained representations transfer to diverse downstream datasets, varying finetuning sample sizes, fine-grained actions, and tasks beyond standard action recognition. Use when the user wants to benchmark on UCF-101, NTU-60, FineGym (Gym-99), Something-Something-v2, EPIC-Kitchens-100, Charades, ...

researchpythongo
0
3
Severe Benchmark EvalA

Evaluates the generalization and sensitivity of video self-supervised learning models to domain shifts, downstream sample sizes, action similarity, and task shifts beyond action recognition. Use when the user wants to benchmark on UCF-101, NTU-60, FineGym (Gym-99), Something-Something-v2, EPIC-Kitchens-100, Charades, AVA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Sfiog EvalA

Evaluates large language models' ability to generate coherent, logically structured financial investment opinions based on company context and questions. It probes reasoning over mere knowledge retrieval by testing performance across varying degrees of familiarity and novelty in companies and questions. Use when the user wants to benchmark on sFIOG, or asks about evaluating this task. Reports ROUGE-L.

researchpythontesting
0
3
Sfld Ai Gen Image Detection EvalA

Evaluates AI-generated image detectors on their ability to generalize across diverse generative models (GANs, diffusion) and resist content bias. It probes robustness using conventional benchmarks, a new content-preserving benchmark (TwinSynths), and low-level vision/perceptual benchmarks to measure how well models rely on texture vs. semantic artifacts. Use when the user wants to benchmark on Conventional benchmark, TwinSynths, Low-level vision and perceptual benchmarks, or asks about evalua...

researchpythonperformance
0
3
Sft Generalization EvalA

Probes whether language models trained via supervised fine-tuning (SFT) truly learn reasoning and planning capabilities or merely memorize instruction templates. It tests generalization across unseen action mappings (instruction variations) and increased grid/card complexity (difficulty variations). Use when the user wants to benchmark on Sokoban, General Points, or asks about evaluating this task. Reports exact-match accuracy.

researchpythongo
0
3
Sga Interact EvalA

Evaluates models on group activity recognition (GAR) and temporal group activity localization (TGAL) using 3D skeleton sequences from basketball games. It probes spatio-temporal interaction modeling, long-term dependency handling, and the ability to leverage multi-view motion capture data for complex team tactics. Use when the user wants to benchmark on SGA-INTERACT, or asks about evaluating this task. Reports accuracy (mAcc./Top3-mAcc./oAcc.).

researchpythongo
0
3
Sgaligner EvalA

Evaluates the ability to align 3D scene graphs by matching semantic entities across scenes with varying spatial overlap and environmental changes. It further tests downstream 3D point cloud registration and mosaicking capabilities using the predicted node alignments to initialize geometric correspondence extraction. Use when the user wants to benchmark on 3RScan (generated sub-scene pairs), or asks about evaluating this task. Reports MRR.

researchpythongo
0
3
Sged EvalA

Evaluates a model's ability to recognize human emotions (neutral, negative, positive) from dynamic gesture videos. It specifically probes robustness to class imbalance and performance under low-light/high-motion conditions using multimodal inputs. Use when the user wants to benchmark on SGED, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Sgg Vg150 EvalA

Evaluates a model's ability to generate structured scene graphs from images without predefined object boxes. It probes visual relationship reasoning, object detection accuracy, and the model's capacity to produce structurally valid outputs under strict spatial and categorical matching criteria. Use when the user wants to benchmark on VG150, PSG, or asks about evaluating this task. Reports Recall.

researchpythongo
0
3
Sgmri Vqa EvalA

Evaluates vision-language models on multi-frame spatial reasoning and grounding in volumetric MRI scans. It probes the model's ability to answer clinical questions, generate free-text reasoning, and accurately localize anatomical findings across single slices or full 3D volumes. Use when the user wants to benchmark on SGMRI-VQA, or asks about evaluating this task. Reports A-Score.

researchpythongo
0
3
Sgplan EvalA

Evaluates classical and neural-symbolic planners on 3D scene graph environments by measuring their ability to generate valid action sequences for task-driven goals within a strict time limit. Use when the user wants to benchmark on SGPlan, or asks about evaluating this task. Reports task completion.

researchpythongo
0
3
Sh Bench EvalA

Evaluates audio LLMs' ability to comprehend multi-speaker conversations while selectively focusing on a target speaker and ignoring bystanders for privacy. It measures both general audio understanding and selective hearing capability under different instruction modes. Use when the user wants to benchmark on SH-Bench, or asks about evaluating this task. Reports Selective Efficacy (SE).

researchpythongo
0
3
ShaPO Safety EvalA

Evaluates the robustness of LLM safety alignment under in-distribution, cross-domain, and noisy supervision settings. It probes whether geometry-aware optimization preserves safety performance while resisting distribution shift and corrupted preference labels. Use when the user wants to benchmark on PKU-SafeRLHF-30K, HH-RLHF-Safety, Do-Not-Answer, HarmBench, SaladBench, or asks about evaluating this task. Reports Win Rate (WR).

researchpythongo
0
3
Shake Gnn EvalA

Evaluates the predictive accuracy and training efficiency of a hierarchical graph neural network that uses Kirchhoff Forest-based stochastic coarsening for graph classification. The benchmark probes whether multi-resolution graph decomposition can maintain competitive performance while significantly reducing computational costs across molecular and social network domains. Use when the user wants to benchmark on MolHIV, MolPPA, COLLAB, DD, REDDIT-MULTI-12K, or asks about evaluating this task. ...

researchpythonnode
0
3
Shalakasatheesh Squad V2A

Compute shalakasatheesh/squad_v2 via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of shalakasatheesh/squad_v2.

developmentpython
0
3
Shale EvalA

This benchmark evaluates fine-grained hallucination in Large Vision-Language Models (LVLMs) by testing their faithfulness to visual inputs and factuality against external knowledge. It measures model performance under clean conditions and across hierarchical input perturbations (image, instruction, and combination levels) to assess hallucination resistance. Use when the user wants to benchmark on SHALE, or asks about evaluating this task. Reports accuracy, non-hallucination rate.

researchpythongo
0
3
Shap Diversity Ensemble EvalA

Probes how well SHAP-based feature attribution divergence correlates with prediction/output diversity across unsupervised anomaly detection algorithms, and quantifies how this explanation-driven diversity impacts the robustness and accuracy of model ensembles. Use when the user wants to benchmark on ADBench subset (16 datasets: anthyroid, breastw, glass, Hepatitis, Lymphography, mammography, PageBlocks, Pima, Stamps, thyroid, vertebral, vowels, WBC, Wilt, wine, yeast), or asks about evaluatin...

researchpythongo
0
3
Shapley Explanation LatencyA

Evaluates the computational efficiency (latency and memory) and explanation quality of a Shapley value-based neural network explainer framework against baseline implementations across standard vision models. Use when the user has predictions and gold and needs to compute latency.

researchpythongo
0
3
Share EvalA

Evaluates a model's ability to predict the next item in an anonymous user session based on sequential click history. It probes the model's capacity to capture short-term user intent and higher-order item correlations within dynamic session contexts. Use when the user wants to benchmark on YooChoose, Diginetica, or asks about evaluating this task. Reports Hit@20.

researchpythontesting
0
3
Sharegpt4video EvalA

Evaluates the temporal understanding and video-language alignment capabilities of Large Video-Language Models (LVLMs) across three multi-modal video benchmarks. It probes the model's ability to answer questions about video content, track temporal changes, and comprehend complex video sequences without relying on single-frame cues. Use when the user wants to benchmark on VideoBench, MVBench, TempCompass, or asks about evaluating this task. Reports benchmark accuracy (VideoBench, MVBench, TempC...

researchpythongo
0
3
Shawshank Bench EvalA

Evaluates the vulnerability of embodied AI agents (Vision-Language Models) to indirect environmental jailbreaks, where malicious instructions are physically embedded in the environment (e.g., on walls or tables) rather than provided as direct text prompts. It measures both the agent's susceptibility to harmful behavior and the collateral impact on benign task execution. Use when the user wants to benchmark on Shawshank-Bench, or asks about evaluating this task. Reports ASR.

ai-agentspythongo
0
3
Sheet Music Benchmark EvalA

Evaluates end-to-end optical music recognition (OMR) systems on their ability to transcribe scanned sheet music images into standardized **kern musical notation. It probes layout analysis, staff/page-level transcription accuracy, and robustness across diverse musical textures such as monophony, pianoform, and quartets. Use when the user wants to benchmark on Sheet Music Benchmark (SMB), or asks about evaluating this task. Reports OMR-NED.

researchpythongo
0
3
Sheriff EvalA

Evaluates EFCE solvers on a parametric sequential bargaining game modeling smuggling and inspection. It probes the solver's ability to handle multi-round negotiations, bribery, and deterrence to maximize social welfare. Use when the user wants to benchmark on Sheriff, or asks about evaluating this task. Reports Social Welfare (SW).

researchpythongo
0
3
Shirt Pose Tracking EvalA

Evaluates the accuracy and robustness of a neural network-integrated Unscented Kalman Filter for monocular pose tracking of tumbling noncooperative spacecraft. It probes the system's ability to maintain steady-state position and orientation accuracy under domain gaps between synthetic training data and real hardware-in-the-loop test images. Use when the user wants to benchmark on SHIRT, or asks about evaluating this task. Reports e_pose.

researchpythonangular
0
3
Shopping Queries EvalA

Evaluates e-commerce search models on query-product semantic matching, ranking, and multiclass relevance classification. It probes the system's ability to distinguish between Exact matches, Substitutes, Complements, and Irrelevant products across English, Spanish, and Japanese. Use when the user wants to benchmark on Shopping Queries Dataset, or asks about evaluating this task. Reports nDCG.

researchpythongo
0
3
Shorter Splatting EvalA

Evaluates the training efficiency and reconstruction fidelity of a 3D Gaussian Splatting method that uses scale reset and entropy-constrained alpha blending to reduce Gaussian list lengths. Use when the user wants to benchmark on Mip-NeRF 360, Deep Blending, Tanks and Temples, or asks about evaluating this task. Reports PSNR.

researchpythonperformance
0
3
Shot2story EvalA

Evaluates multi-modal video understanding across three tasks: single-shot captioning, multi-shot summarization, and question answering. It probes a model's ability to process visual frames, optional ASR text, and shot-level structure to generate coherent descriptions or answer temporal, holistic, and audio-related questions. Use when the user wants to benchmark on Shot2Story, or asks about evaluating this task. Reports CIDEr.

researchpythongo
0
3
Shotbench EvalA

Evaluates vision-language models on expert-level cinematic understanding by testing their ability to reason about shot-level visual properties. It probes fine-grained visual reasoning and spatial cognition across eight cinematography dimensions such as shot size, framing, camera angle, lens, lighting, composition, and movement. Use when the user wants to benchmark on ShotBench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Showdown Clicks EvalA

Evaluates a vision-language model's ability to accurately predict the correct UI element to click based on a screenshot and a task instruction. It isolates low-level visual grounding and interaction skills in ambiguous or icon-heavy interfaces. Use when the user wants to benchmark on Showdown-Clicks, or asks about evaluating this task. Reports Top-1 Accuracy.

researchpythongo
0
3
Showui Gui EvalA

Evaluates a vision-language-action model's ability to perform GUI visual grounding and task-oriented navigation across web, mobile, and online environments. It probes the model's capacity to interpret screenshots, locate interactive elements, and execute correct action sequences to complete user instructions. Use when the user wants to benchmark on Screenspot, Mind2Web, AITW, MiniWob, or asks about evaluating this task. Reports Zero-shot grounding accuracy.

researchpythonperformance
0
3
Shpi Recommendation EvalA

Evaluates offline reinforcement learning methods for session-based recommendation systems in optimizing long-term user retention versus short-term clicks. It tests the ability of algorithms to learn from fixed logging policies and generalize to online rollouts across synthetic, simulated, and real-world recommendation environments. Use when the user wants to benchmark on Synthetic recommendation problem, RecoGym, HIV treatment simulator, Private dataset X, or asks about evaluating this task. ...

researchpythongo
0
3
Shredbench EvalA

Evaluates multimodal LLMs' ability to reconstruct shredded documents from fragmented visual inputs. It probes cross-modal semantic reasoning, visual discontinuity alignment, and fine-grained positional continuity awareness across natural language, source code, and tabular data. Use when the user wants to benchmark on ShredBench, or asks about evaluating this task. Reports NED.

researchpythonjava
0
3
Shrutilipi Asr EvalA

Evaluates the quality, diversity, and downstream effectiveness of the Shrutilipi audio-text dataset for low-resource Indian language ASR. It measures how adding mined data improves Word Error Rate (WER) on standard and noisy benchmarks compared to existing datasets. Use when the user wants to benchmark on Shrutilipi, MUCS, Kathbath, CommonVoice, or asks about evaluating this task. Reports WER.

researchpythonperformance
0
3
Shunzh Apps MetricA

Compute shunzh/apps_metric via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of shunzh/apps_metric.

developmentpython
0
3
Sib200 Xlt EvalA

This evaluation probes a model's ability to perform zero-shot and fully-supervised cross-lingual text classification across typologically diverse languages. It specifically measures how well parameter-efficient soft prompt tuning methods transfer knowledge from high-resource source languages to low-performing or unseen target languages without language-specific fine-tuning. Use when the user wants to benchmark on SIB-200, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Sid Generalization EvalA

Evaluates a model's ability to generalize synthetic image detection across diverse generative architectures (GANs, Diffusion Models, DiTs) and real-world image sources, focusing on robustness to unseen generators and varying image resolutions. Use when the user wants to benchmark on SID Generalization Benchmark (ForenSynths, Self-Synthesis, Ojha, GenImage, DiTFake), or asks about evaluating this task. Reports ACC.

researchpythontesting
0
3
Sidon Speech Restoration EvalA

Evaluates multilingual speech restoration quality by measuring acoustic fidelity, speaker preservation, and transcription accuracy on noisy speech. It also assesses downstream utility by training TTS models on cleansed data and measuring synthetic speech quality, alongside inference speed benchmarks. Use when the user wants to benchmark on test-clean/test-other subsets (English), Multilingual test set, TED-LIUM Release 3, or asks about evaluating this task. Reports DNSMOS.

researchpythongo
0
3
Sifthinker EvalA

Evaluates a model's spatial reasoning, fine-grained visual perception, and 3D spatial grounding capabilities. It also measures self-correction ability through dynamic bounding box refinement and general visual-language understanding across multiple benchmarks. Use when the user wants to benchmark on SpatialBench, SAT-Static, CV-Bench, VisCoT_s, V*Bench, RefCOCO, RefCOCO+, RefCOCOg, OVDEval, MME, MMBench, SEED-Bench, VQAv2, POPE, or asks about evaluating this task. Reports Top-1 Accuracy@0.5.

researchpythongo
0
3
Sigmacollab EvalA

Probes an AI system's ability to understand and assist in physically situated, goal-directed collaboration tasks using multimodal egocentric sensing. It evaluates real-time scene understanding, interaction modeling, and proactive guidance in fluid, human-AI collaborative scenarios. Use when the user wants to benchmark on SigmaCollab, or asks about evaluating this task. Reports classification.

researchpythongo
0
3
Sign Recommender EvalA

This protocol evaluates a model's ability to detect beneficial feature interactions for recommendation and graph classification tasks. It measures prediction accuracy and ranking performance while assessing how well the model filters out irrelevant feature pairs to improve generalization. Use when the user wants to benchmark on Frappe, MovieLens-tag, Twitter, DBLP, or asks about evaluating this task. Reports accuracy (ACC).

researchpythongo
0
3
Sign Signwriting SimilarityA

Compute sign/signwriting_similarity via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of sign/signwriting_similarity.

developmentpython
0
3
Signal Quality Auditing EvalA

Evaluates the effectiveness of various Signal Quality Indices (SQIs) and machine learning models for classifying noisy vs. clean ECG signals, detecting outliers, and denoising time-series data. Use when the user wants to benchmark on Physionet 2011 ECG Challenge (PICC) Set A, MIT-BIH Arrhythmia & NSTDB, or asks about evaluating this task. Reports AUC, Mean Squared Error (MSE).

researchpythontesting
0
3
SignaldistortionratioA

Compute the SignalDistortionRatio metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute SignalDistortionRatio, or asks how to score with SignalDistortionRatio.

documentationpython
0
3
Signalmc Med EvalA

This benchmark evaluates biosignal foundation models on synchronized, long-duration single-lead ECG and PPG recordings from emergency department visits. It probes the models' ability to extract clinically meaningful representations for tasks such as age and sex prediction, emergency disposition, laboratory value regression, and ICD-10 diagnosis classification. Use when the user wants to benchmark on SignalMC-MED, or asks about evaluating this task. Reports AUROC.

researchpythongo
0
3
SignalnoiseratioA

Compute the SignalNoiseRatio metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute SignalNoiseRatio, or asks how to score with SignalNoiseRatio.

documentationpython
0
3
Signature OverlapA

This metric quantifies the degree of overlap between different LLM benchmarks by comparing their token-level perplexity signatures derived from in-the-wild pretraining corpora, revealing whether performance correlations stem from shared latent capacity familiarity or benchmark-orthogonal factors like question format. Use when the user has predictions and gold and needs to compute signature_overlap.

datapythongo
0
3
Significance EvalA

Evaluates the ability of ML classifiers to distinguish signal from background events in high-energy physics simulations. It measures classification quality via AUC and quantifies discovery potential using a likelihood-ratio-based significance metric optimized over a probability threshold. Use when the user has predictions and gold and needs to compute significance.

researchpythongo
0
3
Silhouette ScoreA

Compute the silhouette_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute silhouette_score, or asks how to score with silhouette_score.

documentationpython
0
3
Silicone EvalA

Evaluates a model's ability to perform sequence labelling on spoken dialogues, specifically predicting dialog acts (DA) and emotion/sentiment (E/S) labels per utterance within multi-utterance conversations. Use when the user wants to benchmark on SILICONE, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3