All authors
qhjqhj00 avatar

Claude Skills by qhjqhj00

github.com/qhjqhj00
7,636 skillsA× 7,623B× 11C× 1D× 10 installs2,118 views
Survey Sum EvalA

Evaluates a modular retrieval-augmented generation pipeline for multi-document scientific literature summarization. It probes the model's ability to dynamically generate queries, retrieve relevant papers, and synthesize citation-aware, coherent summaries from multiple sources. Use when the user wants to benchmark on SurveySum, or asks about evaluating this task. Reports Ref-F1.

researchpythongo
0
3
Sustainableqa EvalA

Evaluates language models' ability to extract precise factual answers and generate semantically accurate responses from complex corporate sustainability and EU Taxonomy reports. It also benchmarks retrieval systems' capacity to locate relevant regulatory and financial passages in domain-specific, long-form documents. Use when the user wants to benchmark on SustainableQA, or asks about evaluating this task. Reports Exact Match (EM).

ai-agentspythongo
0
3
Susuinteracts EvalA

Evaluates the quality of generated 3D human motions for interactive dialogue, measuring semantic alignment with text/audio, motion quality, audio-motion synchronization, and diversity. Use when the user wants to benchmark on SuSuInterActs, or asks about evaluating this task. Reports R@K.

researchpythonexpress
0
3
Suvach Hindi Qa EvalA

Evaluates Hindi extractive question answering capabilities using multiple-choice questions generated from Wikipedia contexts. It probes a model's ability to comprehend Hindi text, locate relevant information, and select the correct answer from four options under varying context availability settings. Use when the user wants to benchmark on Suvach, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Sv Trusteval C EvalA

Evaluates LLMs' structural and semantic reasoning capabilities in C code vulnerability analysis. It probes whether models rely on memorized patterns or genuinely understand code interdependencies by testing consistency across base, data-flow, control-flow, counterfactual, goal-driven, and predictive scenarios. Use when the user wants to benchmark on SV-TrustEval-C, or asks about evaluating this task. Reports Cons_DFL.

researchpythonrust
0
3
Svamp EvalA

Evaluates whether NLP models can genuinely solve simple math word problems through arithmetic reasoning versus relying on shallow heuristics like bag-of-words matching or positional cues. It probes model brittleness by testing performance on standard datasets alongside carefully perturbed variants that remove questions or alter operator types. Use when the user wants to benchmark on MAWPS, ASDiv-A, SVAMP, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Svbench EvalA

Evaluates large vision-language models' ability to perform sustained temporal reasoning and context tracking across long-form streaming videos. It probes multi-turn dialogue continuity, temporal dependency handling, and complex reasoning skills like counterfactual analysis and spatio-temporal speculation. Use when the user wants to benchmark on SVBench, or asks about evaluating this task. Reports Overall Score (OS).

researchpythongo
0
3
Svc Ongoing EvalA

Evaluates on-line signature verification systems across office (stylus), mobile (finger), and hybrid scenarios. It measures robustness against both skilled and random forgeries, testing generalization across different acquisition devices and intra-user variability. Use when the user wants to benchmark on DeepSignDB, SVC2021_EvalDB, or asks about evaluating this task. Reports EER.

developmentpythongo
0
3
Svcc23 EvalA

Evaluates singing voice conversion systems on in-domain (singing-to-singing) and cross-domain (speech-to-singing) speaker conversion. It probes the model's ability to preserve target speaker identity and musical prosody while converting source audio to the target voice. Use when the user wants to benchmark on SVCC 2023, or asks about evaluating this task. Reports perceptual quality (subjective evaluation).

researchpythongo
0
3
Svcd EvalA

This protocol evaluates a singing voice conversion model's ability to transform a source singer's voice into a target singer's timbre while preserving musical naturalness and pitch accuracy. It measures both subjective perceptual quality and objective acoustic fidelity using human ratings and signal processing metrics. Use when the user wants to benchmark on VCTK, NUS-48E, or asks about evaluating this task. Reports MOS (Naturalness), MOS (Similarity).

researchpythongo
0
3
Svd Pc ImportanceA

This protocol evaluates the ability of singular value decomposition (SVD) and gradient boosting regression trees (GBRT) to decompose unresolved planetary light curves into principal components that physically correspond to specific surface and atmospheric features. It quantifies feature attribution through variance explained, model importance scores, and linear correlations. Use when the user has predictions and gold and needs to compute SVD eigenvalue variance ratio.

researchpythongo
0
3
Svdquant EvalA

Evaluates the visual fidelity and text-image alignment of quantized diffusion models by generating images from text prompts and comparing them against reference outputs. It probes whether low-bit quantization preserves distributional similarity, perceptual quality, and human-preferred aesthetics compared to full-precision baselines. Use when the user wants to benchmark on MJHQ-30K, sDCI, or asks about evaluating this task. Reports FID.

researchpythongo
0
3
Svenwey LogmetricA

Compute svenwey/logmetric via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of svenwey/logmetric.

developmentpython
0
3
Svg Sophia Refinement EvalA

Evaluates a model's ability to refine and correct imperfect SVG code, measuring structural accuracy, visual fidelity, and code efficiency. Use when the user wants to benchmark on SVG-Sophia Code Refinement Benchmark, or asks about evaluating this task. Reports SR.

researchpythongo
0
3
Svhn EvalA

Evaluates a model's ability to classify real-world, cropped street-view house numbers into ten digit classes. It probes robustness to natural scene variations such as background clutter, varying colors, orientations, and focus, which are absent in synthetic datasets like MNIST. Use when the user wants to benchmark on Street View House Numbers (SVHN), or asks about evaluating this task. Reports accuracy.

researchpythongit
0
3
Svld Points Ratio EvalA

Probes a model's ability to predict social engagement (upvote ratio) from multimodal inputs (images, videos, and text). It evaluates cross-modal fusion and regression capabilities on socially grounded, context-rich data. Use when the user wants to benchmark on SVLD, or asks about evaluating this task. Reports Mean L1point ratio prediction error.

researchpythongo
0
3
Svo Probes EvalA

Evaluates how well CLIP aligns text and images by measuring its preference for positive over negative image-text pairs, and analyzes how semantic features (part-of-speech, concreteness, length, frequency, ambiguity) influence this alignment. Use when the user wants to benchmark on SVO-Probes, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Svqa Vqa EvalA

Evaluates a multimodal model's ability to answer visual questions when the query is provided as spoken audio rather than text. It probes speech-vision-language alignment, robustness to synthesized speech variations, and the model's capacity to handle modality-specific prompts. Use when the user wants to benchmark on SEED-Bench, MME, DocVQA, MLS, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Svs EvalA

Evaluates the acoustic and perceptual quality of singing voice synthesis (SVS) models trained on large-scale multi-singer corpora. It probes direct synthesis capability, cross-domain transfer learning, and data augmentation via joint training. Use when the user wants to benchmark on ACE-Opencpop, ACE-KiSing, or asks about evaluating this task. Reports MOS.

researchpythongo
0
3
Swapnet System EvalA

Evaluates a block-swapping middleware for DNN inference on memory-constrained edge AI devices. It probes the system's ability to run large models beyond hardware memory limits while measuring peak memory consumption, inference latency, and classification accuracy compared to direct execution, channel division, and model compression baselines across three real-world application scenarios. Use when the user has predictions and gold and needs to compute memory consumption.

researchpythongo
0
3
Swarmbench EvalA

This benchmark evaluates emergent decentralized coordination in LLM-driven multi-agent systems under strict local perception and communication constraints. It simulates five canonical swarm tasks—Pursuit, Synchronization, Foraging, Flocking, and Transport—in a 2D grid environment to probe whether LLMs can form adaptive group strategies and execute robust long-range planning without global information. Use when the user wants to benchmark on SwarmBench, or asks about evaluating this task. Repo...

ai-agentspythongit
0
3
Swe Bench EvalA

Evaluates an agent's ability to automatically resolve software engineering issues by generating and applying code patches to open-source repositories. It measures functional correctness by running the target repository's test suite after the patch is applied. Use when the user wants to benchmark on SWE-bench, HumanEvalFix, or asks about evaluating this task. Reports pass@1.

researchpythonapi
0
3
Swe Bench Java EvalA

This benchmark evaluates an AI agent's ability to autonomously resolve real-world GitHub issues in Java projects. It probes capabilities in code patch generation, repository navigation, test case reasoning, and handling runtime environment dependencies. Use when the user wants to benchmark on SWE-bench-java-verified, or asks about evaluating this task. Reports Resolved Rate (%).

researchpythonjava
0
3
Swe Bench Repair EvalA

This evaluation probes an LLM-based agent's ability to automatically locate faults and generate correct code patches for real-world software issues. It tests both traditional text-only bug fixing and multimodal reasoning where visual UI behavior must be understood alongside code. Use when the user wants to benchmark on SWE-bench Lite, SWE-bench Multimodal, or asks about evaluating this task. Reports %Resolved.

researchpythondocumentation
0
3
Swe Chat EvalA

Evaluates real-world coding agent interactions by measuring how much agent-generated code survives into final commits, alongside efficiency metrics like token usage, cost, and runtime per committed line. Use when the user wants to benchmark on SWE-chat, or asks about evaluating this task. Reports Code survival rate.

researchpythongit
0
3
Swe Rebench V2 EvalA

Evaluates the ability of LLM-based agents to autonomously resolve software engineering issues by modifying code in real-world repositories. It probes environment setup, code generation, and test execution capabilities across multiple programming languages. Use when the user wants to benchmark on SWE-rebench V2, or asks about evaluating this task. Reports pass@1.

developmentpythonrust
0
3
Swebench EvalA

Evaluates language models' ability to resolve real-world software engineering issues by generating code patches. It probes long-context reasoning, cross-file dependency understanding, and execution-based validation within large, complex codebases. Use when the user wants to benchmark on SWE-bench, or asks about evaluating this task. Reports resolve_rate.

researchpythongit
0
3
Swebench Live EvalA

Evaluates the ability of AI coding agents to autonomously resolve real-world software engineering issues by generating and applying patches to GitHub repositories. It probes cross-file reasoning, dependency management, and robustness against contamination from static benchmarks. Use when the user wants to benchmark on SWE-bench-Live, or asks about evaluating this task. Reports Resolved Rate (%).

researchpythongo
0
3
Swiltra Bench EvalA

Evaluates large language models and specialized translation systems on their ability to accurately translate Swiss legal documents (laws, headnotes, press releases) across four national languages and English. It probes domain-specific translation quality, contextual understanding, and zero-shot versus fine-tuned performance in a legal context. Use when the user wants to benchmark on SwiLTra-Bench, or asks about evaluating this task. Reports GEMBA-MQM.

researchpythonaws
0
3
Swim EvalA

Evaluates real-time instance segmentation performance of lightweight models under strict onboard hardware constraints, measuring inference speed, memory usage, and segmentation accuracy for spacecraft boundary localization. Use when the user wants to benchmark on SWiM, or asks about evaluating this task. Reports RAM_footprint.

businesspythongo
0
3
Swim Ir EvalA

Evaluates multilingual dense retrieval models on cross-lingual and monolingual open retrieval tasks. It measures how effectively synthetic LLM-generated training data scales retrieval performance compared to human-labeled baselines across diverse languages and corpus sizes. Use when the user wants to benchmark on XOR-Retrieve, MIRACL, XTREME-UP, or asks about evaluating this task. Reports Recall@mkt.

researchpythongo
0
3
Swimba Standard Bench EvalA

Evaluates language understanding, reasoning, and knowledge recall capabilities of a Switch Mamba model on standard multiple-choice benchmarks. It also measures inference efficiency (throughput, latency, FLOPs) to assess the computational trade-offs of the parameter-space mixture-of-experts design. Use when the user wants to benchmark on BoolQ, OpenBookQA, RTE, MMLU, PIQA, WinoGrande, HellaSwag, ARC-Challenge, ARC-Easy, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Swiss Judgment Prediction EvalA

Evaluates Legal Judgment Prediction (LJP) models across languages (German, French, Italian), regions, and legal domains. It probes cross-lingual, cross-regional, and cross-domain transfer capabilities, as well as the impact of data augmentation and adapter-based fine-tuning on model fairness and accuracy. Use when the user wants to benchmark on SJP (Swiss Judgment Prediction), or asks about evaluating this task. Reports macro-averaged F1 score.

researchpythongo
0
3
Switch Justdance EvalA

Evaluates whole-body motion tracking policies for humanoid robots by measuring how well they synchronize with reference choreography from the commercial game Just Dance. It probes tracking accuracy, stability over long-horizon motions, and movement smoothness compared to human baselines. Use when the user wants to benchmark on Just Dance Routines, or asks about evaluating this task. Reports Just Dance Score (JDS).

researchpythonperformance
0
3
Swivuriso Asr EvalA

Evaluates automatic speech recognition (ASR) capabilities across seven South African languages. It measures how well pre-trained speech models can transcribe spontaneous and scripted audio in low-resource, domain-specific contexts (agriculture, healthcare, general). Use when the user wants to benchmark on Swivuriso, or asks about evaluating this task. Reports WER.

researchpython
0
3
Swsr EvalA

Probes the capability of NLP models to detect online sexism in Chinese microblogging comments. It evaluates performance across three hierarchical classification tasks: binary sexism identification, fine-grained category classification, and target type classification. Use when the user wants to benchmark on SWSR, or asks about evaluating this task. Reports macro F1.

researchpythongo
0
3
Sygu S Comp 15 EvalA

Evaluates the capability of program synthesis solvers to generate correct functions or expressions that satisfy given logical constraints or specifications. It probes how well solvers handle different grammar restrictions, specification completeness, and problem structures like linear arithmetic or invariant generation. Use when the user wants to benchmark on SyGuS-Comp'15, or asks about evaluating this task. Reports number of benchmarks solved.

researchpythongo
0
3
Sygus Comp 2016 EvalA

Evaluates syntax-guided program synthesis solvers on their ability to generate correct programs from logical constraints and grammars. It probes capabilities in conditional linear integer arithmetic, invariant generation, and programming-by-example with bit-vectors and strings. Use when the user wants to benchmark on SyGuS-Comp 2016, or asks about evaluating this task. Reports number_of_benchmarks_solved.

researchpythongo
0
3
Sygus Comp 2017 EvalA

Evaluates the ability of synthesis solvers to generate correct programs or expressions that satisfy given grammatical and semantic constraints across multiple domain-specific tracks. Use when the user wants to benchmark on SyGuS-Comp 2017, or asks about evaluating this task. Reports correctness.

researchpythonexpress
0
3
Sygus Comp 2018 EvalA

Evaluates syntax-guided synthesis solvers on their ability to generate correct programs or specifications across multiple domains, including general synthesis, conditional linear integer arithmetic, invariant generation, and programming by examples. Use when the user wants to benchmark on SyGuS-Comp 2018, or asks about evaluating this task. Reports correctness.

researchpythonexpress
0
3
Symbench EvalA

Probes an LLM's ability to solve symbolic reasoning and planning tasks by dynamically switching between textual reasoning and code generation. It evaluates robustness on both seen and unseen tasks, as well as the model's generalizability across different architectures and complexity levels. Use when the user wants to benchmark on SymBench, or asks about evaluating this task. Reports Average Normalized Score (AveNorm).

researchpythongit
0
3
Symbolic Math EvalA

This benchmark evaluates a model's ability to perform symbolic mathematical computations, specifically indefinite integration and solving ordinary differential equations. It probes the model's capacity to learn complex algebraic patterns and generate syntactically valid, mathematically equivalent expressions from prefix-encoded inputs. Use when the user wants to benchmark on Symbolic Mathematics (FWD/BWD/IBP/ODE), or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Symbolizer EvalA

This evaluation probes a VLM's ability to ground visual and textual observations into structured symbolic states (objects, predicates, goals) and subsequently use those representations for effective task and motion planning. It measures both the accuracy of the symbolic grounding pipeline and the end-to-end success rate of classical planners operating on the generated PDDL problem files. Use when the user wants to benchmark on ProDG, ViPlan, or asks about evaluating this task. Reports F1.

researchpythongo
0
3
Symile M3 EvalA

Evaluates zero-shot cross-modal retrieval capability by requiring a model to jointly leverage audio and text to identify an image, where neither modality alone contains sufficient information. It tests the model's ability to capture joint information across three distinct high-dimensional data types. Use when the user wants to benchmark on Symile-M3, or asks about evaluating this task. Reports mean accuracy.

researchpythongit
0
3
Symile Mimic EvalA

Tests clinical cross-modal prediction by evaluating whether ECG and blood lab measurements can jointly predict a subsequent chest X-ray in a zero-shot retrieval setting. It probes the model's ability to learn from incomplete training data and generalize to full modality combinations. Use when the user wants to benchmark on Symile-MIMIC, or asks about evaluating this task. Reports mean accuracy.

researchpythongit
0
3
Symile Synthetic EvalA

Probes a model's ability to capture higher-order conditional dependencies between modalities by predicting one modality's representation from two others under varying information dynamics. It specifically tests whether a model can leverage joint information when pairwise mutual information is zero. Use when the user wants to benchmark on Synthetic dataset, or asks about evaluating this task. Reports mean accuracy.

researchpythongit
0
3
Symlink EvalA

Probes the ability to extract fine-grained mathematical symbols and their textual descriptions from LaTeX-formatted scientific documents. It evaluates both named entity recognition for identifying symbols and descriptions, and relation extraction for linking them according to specific semantic types. Use when the user wants to benchmark on Symlink, or asks about evaluating this task. Reports F-score.

researchpythongo
0
3
SymmetricmeanabsolutepercentageerrorA

Compute the SymmetricMeanAbsolutePercentageError metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute SymmetricMeanAbsolutePercentageError, or asks how to score with SymmetricMeanAbsolutePercentageError.

documentationpython
0
3
Symsearch Omnigibson EvalA

Evaluates an agent's ability to perform open-vocabulary interactive object search in indoor environments using relational semantic reasoning over 3D scene graphs. It probes exploration efficiency, reasoning accuracy, and computational cost compared to embedding-based and LLM-based planners. Use when the user wants to benchmark on SymSearch, OmniGibson, or asks about evaluating this task. Reports Success Rate (SR), Success weighted by Path Length (SPL).

researchpythonnode
0
3
SynTSBench EvalA

Evaluates the ability of deep learning models to learn and forecast diverse temporal patterns (trends, periodicities, multivariate dependencies) in time series. It also probes model robustness against varying levels of Gaussian and non-Gaussian noise, as well as resilience to point and pulse anomalies. Use when the user wants to benchmark on SynTSBench, or asks about evaluating this task. Reports MSE.

researchpythongo
0
3