All authors
qhjqhj00 avatar

Claude Skills by qhjqhj00

github.com/qhjqhj00
7,636 skillsA× 7,623B× 11C× 1D× 10 installs2,193 views
Silicone Prep Anomaly EvalA

Evaluates multimodal vision-language models' ability to detect context-dependent visual anomalies in robotic scientific laboratory workflows using first-person imagery and stage-specific textual prompts. Use when the user wants to benchmark on Silicone Preparation Workflow, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Sim1 Tshirt Fold EvalA

Evaluates a robot policy's ability to perform structured deformable manipulation (t-shirt folding) in real-world settings after being trained exclusively on simulation data. It probes sim-to-real transfer, out-of-domain robustness to environmental shifts, and data scaling efficiency. Use when the user wants to benchmark on SIM1 T-shirt Folding, or asks about evaluating this task. Reports success.

researchpythongo
0
3
Sim3d EvalA

Evaluates 3D anomaly detection and segmentation capabilities in industrial settings using multiview and multimodal (image + depth) inputs. It probes a model's ability to identify and localize defects across multiple object categories under both in-domain (real-to-real) and out-of-domain (synthetic-to-real) conditions. Use when the user wants to benchmark on SiM3D, or asks about evaluating this task. Reports I-AUROC.

researchpythongo
0
3
Simba Benchmark Analysis EvalA

Evaluates a framework for analyzing language model performance matrices by identifying dataset-model correlations, discovering minimal representative dataset subsets, and predicting held-out model performance while preserving model rankings. Use when the user wants to benchmark on HELM, MMLU, BigBenchLite, or asks about evaluating this task. Reports coverage ($\eta$).

researchpythonperformance
0
3
Simbarca Traffic Forecasting EvalA

Evaluates the ability of deep learning models to forecast urban traffic speeds at both the individual road segment and regional levels. It probes spatio-temporal forecasting capabilities under varying congestion levels, testing how well models integrate multi-source sensor data (drone trajectories and loop detectors) to predict future traffic states. Use when the user wants to benchmark on SimBarca, or asks about evaluating this task. Reports MAE.

researchpythonnode
0
3
Simbav2 Continuous Control EvalA

Evaluates the sample efficiency, optimization stability, and scaling capabilities of deep reinforcement learning algorithms across diverse continuous control environments. Use when the user wants to benchmark on MuJoCo, DMC Suite, MyoSuite, HumanoidBench, or asks about evaluating this task. Reports Normalized Return.

researchpythongo
0
3
Simmc2.0 EvalA

Evaluates multimodal task-oriented dialogue capabilities, specifically focusing on dialogue state tracking, disambiguation, coreference resolution, and response generation using visual scene representations. Use when the user wants to benchmark on SIMMC 2.0, or asks about evaluating this task. Reports Intent-F1.

researchpythongo
0
3
Simmmdg EvalA

Evaluates a model's ability to generalize across unseen domains in multi-modal action recognition. It probes feature disentanglement, cross-modal translation for missing modalities, and robustness to domain shifts in video, audio, and optical flow inputs. Use when the user wants to benchmark on EPIC-Kitchens, HAC, or asks about evaluating this task. Reports Top-1 accuracy.

researchpythongo
0
3
Simmotion Retrieval EvalA

This benchmark evaluates a model's ability to retrieve videos based on semantic motion similarity, disentangling dynamic behavior from static appearance, camera viewpoint, and scene context. It tests robustness to appearance variations in controlled synthetic settings and unsynchronized, in-the-wild video pairs. Use when the user wants to benchmark on SimMotion-Synthetic, SimMotion-Real-1K, Jester, or asks about evaluating this task. Reports Retrieval accuracy.

researchpythongo
0
3
Simpleqa Verified EvalA

This benchmark evaluates an LLM's parametric factuality and internal knowledge recall on short-form questions. It measures whether models can correctly answer factual queries without relying on external search tools or retrieval augmentations. Use when the user wants to benchmark on SimpleQA Verified, or asks about evaluating this task. Reports F1-Score.

ai-agentspythongo
0
3
Simplerenv EvalA

Evaluates a robot's ability to execute manipulation tasks from visual inputs and textual instructions, testing both atomic skill execution and high-level instruction generalization in simulated and real-world settings. Use when the user wants to benchmark on SimplerEnv, SimplerEnv-Instruct, or asks about evaluating this task. Reports visual matching (VM).

researchpythongo
0
3
Simplestories Diversity EvalA

Evaluates the lexical, semantic, and syntactic diversity of a synthetic story dataset compared to a baseline. It measures n-gram distribution, compression ratio, Self-BLEU, n-gram diversity scores, POS template rates, and model-judged semantic variation to assess how well the dataset avoids formulaic phrasing and redundancy. Use when the user wants to benchmark on SimpleStories, or asks about evaluating this task. Reports compression_ratio.

researchpythongit
0
3
Simpletoolhallubench EvalA

Probes an agent's ability to abstain from tool use when no appropriate tools are available, specifically measuring its tendency to hallucinate tool invocations in two controlled scenarios: when no tools are provided and when irrelevant distractor tools are present. Use when the user wants to benchmark on SimpleToolHalluBench, or asks about evaluating this task. Reports R_NTA.

researchpython
0
3
Simulrag Scientific Qa EvalA

Evaluates long-form scientific question answering by measuring how effectively a model generates informative and factual answers using simulator-based retrieval. It also assesses the efficiency and quality of claim-level verification and updating strategies under varying computational budgets. Use when the user wants to benchmark on Climate modeling dataset, Epidemiological modeling dataset, or asks about evaluating this task. Reports factuality.

ai-agentspython
0
3
Simultaneous S2st EvalA

Evaluates simultaneous speech-to-speech translation models on translation accuracy, latency, speaker voice preservation, and audio naturalness across multiple languages. It probes the model's ability to generate high-quality target speech in real-time without relying on word-level aligned training data. Use when the user wants to benchmark on Audio-NTREX-4L, Europarl-ST, or asks about evaluating this task. Reports ASR-BLEU.

researchpython
0
3
Sincnet Speaker Recognition EvalA

Evaluates text-independent speaker identification and verification on raw audio waveforms, testing the model's ability to extract speaker-specific features and generalize across different corpus sizes and utterance lengths. Use when the user wants to benchmark on TIMIT, Librispeech, or asks about evaluating this task. Reports accuracy.

researchpythontesting
0
3
Singapore Meme Offense Detection EvalA

This benchmark evaluates multimodal large language models' ability to detect offensive memes containing social biases within a Singaporean cultural and linguistic context. It tests both standalone VLM reasoning and a multi-step pipeline combining OCR and translation, while comparing different fine-tuning strategies and data compositions. Use when the user wants to benchmark on Singapore Offensive Memes Dataset, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Singer Identification EvalA

Evaluates the ability of audio embedding and classification models to correctly identify the vocalist performing a song track. It specifically probes robustness to synthetic/deepfake voices and generalization across different music datasets and genre contexts. Use when the user wants to benchmark on Train, Validation, Closed, FMA, MTG, Cloned, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Singing Voice Conversion EvalA

Probes a model's ability to convert speech to high-quality singing while preserving the target speaker's timbre, using only normal speech samples. Evaluates both audio naturalness and speaker similarity through subjective listening tests. Use when the user wants to benchmark on Tencent multi-speaker speech corpus (TSP), Tencent singing corpus (TSG), Separate singing corpus (test), or asks about evaluating this task. Reports MOS (Naturalness).

researchpython
0
3
Singing Voice Separation EvalA

This benchmark evaluates a model's ability to extract singing vocals from mixed audio tracks. It probes the effectiveness of adversarial semi-supervised learning in separating sources without relying on perfectly paired mixture-source training data. Use when the user wants to benchmark on DSD100, iKala, MedleyDB, CCMixter, or asks about evaluating this task. Reports SDR.

researchpython
0
3
Single Cell Clustering EvalA

Evaluates clustering algorithms on single-cell RNA sequencing data to assess their ability to separate distinct cell types and capture transitional states. It combines quantitative clustering metrics with visual validation using marker gene expression profiles. Use when the user wants to benchmark on Breast cancer dataset, Embryo neurone dataset, or asks about evaluating this task. Reports misclustering rate.

researchpythongo
0
3
Singverse EvalA

Evaluates singing voice enhancement models on real-world acoustic scenarios, measuring their ability to improve perceptual quality and content intelligibility of degraded singing vocals without degrading speech capabilities. Use when the user wants to benchmark on SingVERSE, or asks about evaluating this task. Reports perceptual quality.

researchpythongo
0
3
Sinhala Script Benchmark EvalA

Evaluates language models' ability to process and generate text in two distinct Sinhala writing systems: standard Unicode and Romanized transliteration. It probes script-specific perplexity, coherence, and grammatical correctness, highlighting morphological understanding and training data biases. Use when the user wants to benchmark on Sinhala Unicode & Romanized, or asks about evaluating this task. Reports perplexity.

researchpythonperformance
0
3
Sintel Kitti Flow EvalA

Evaluates a neural network's ability to interpolate sparse optical flow matches into dense flow maps. It probes the model's capacity to handle missing pixels, occlusions, and motion boundaries while preserving flow accuracy across diverse scenes. Use when the user wants to benchmark on MPI Sintel, KITTI 2012, KITTI 2015, or asks about evaluating this task. Reports EPE.

researchpythongo
0
3
Siqa EvalA

Evaluates multimodal large language models on their ability to understand and assess the quality of scientific images. It probes two distinct capabilities: factual reasoning about scientific content (SIQA-U) and alignment with human expert judgments on perceptual and knowledge dimensions (SIQA-S). Use when the user wants to benchmark on SIQA, or asks about evaluating this task. Reports accuracy, SRCC.

researchpythongo
0
3
Sirius Robot EvalA

Evaluates the policy success rate and human workload reduction of a human-in-the-loop robot learning framework over multiple deployment rounds on contact-rich manipulation tasks in simulation and real-world settings. Use when the user wants to benchmark on Sirius Robot Manipulation Tasks, or asks about evaluating this task. Reports success rate.

researchpythongo
0
3
Siu3r Scanet EvalA

Evaluates simultaneous 3D scene reconstruction and multi-task scene understanding on sparse multi-view images. It probes geometric accuracy, novel view synthesis quality, and cross-view consistent segmentation across semantic, instance, panoptic, and text-referred tasks. Use when the user wants to benchmark on ScanNet, or asks about evaluating this task. Reports mIoU.

researchpythongo
0
3
Siwarex EvalA

Evaluates LLM-based natural language question answering over heterogeneous data sources by testing the model's ability to generate SQL queries that correctly invoke both database tables and external APIs. It probes complex query planning, API sequencing, and routing across mixed data modalities. Use when the user wants to benchmark on Spider (modified with API-replaced tables), or asks about evaluating this task. Reports execution accuracy.

researchpythongo
0
3
Skeleton Anomaly Detection EvalA

Probes a model's ability to detect abnormal events in pedestrian videos by analyzing skeleton sequences. It measures how well the model distinguishes normal from anomalous motion patterns through reconstruction or prediction errors, evaluated via frame-level anomaly scoring. Use when the user wants to benchmark on Avenue, HR-Avenue, HR-STC, UBnormal, HR-UBnormal, or asks about evaluating this task. Reports AUC.

researchpythonperformance
0
3
Sketch Less Retrieval EvalA

Evaluates early retrieval performance using partial or low-quality sketches combined with text, measuring how quickly the correct target image appears in the ranked list as the sketch is drawn. It probes robustness to incomplete visual inputs and multimodal fusion. Use when the user wants to benchmark on FS2K-SDE1, FS2K-SDE2, User-SDE, or asks about evaluating this task. Reports m@A, m@B.

researchpythongo
0
3
Sketch Of Thought EvalA

Evaluates the reasoning efficiency and accuracy of LLMs under cognitive-inspired prompting constraints. It probes the model's ability to produce structured, concise reasoning chains while maintaining correctness across mathematical, commonsense, logical, multi-hop, scientific, medical, multilingual, and multimodal tasks. Use when the user wants to benchmark on GSM8K, SVAMP, AQUA-RAT, DROP, CommonsenseQA, OpenbookQA, StrategyQA, LogiQA, ReClor, HotPotQA, MuSiQue-Ans, QASC, Worldtree, PubMedQA,...

ai-agentspythongo
0
3
Sketch Rnn EvalA

Evaluates a generative model's ability to produce coherent stroke-based vector sketches through reconstruction, latent space interpolation, and completion of incomplete drawings. It probes how well the model captures conceptual features and organizes them in a continuous latent manifold. Use when the user wants to benchmark on QuickDraw, or asks about evaluating this task. Reports LR.

researchpythongo
0
3
Sketchduo EvalA

This evaluation probes a diffusion model's ability to generate pixel-level sketches that align with detailed text prompts while maintaining stylistic abstraction and human-like drawing characteristics. It measures both perceptual image quality and fine-grained text-to-image semantic alignment across multiple complementary metrics. Use when the user wants to benchmark on SketchDUO, or asks about evaluating this task. Reports TIFAScore.

researchpythongo
0
3
Skillflow EvalA

Evaluates autonomous agents' ability to discover, patch, and evolve reusable skills over time in a sequential, lifelong learning setting. It probes whether models can consolidate successful execution traces into a compact, repairable skill library rather than merely accumulating fragmented task-specific traces. Use when the user wants to benchmark on SkillFlow, or asks about evaluating this task. Reports task completion rate (%comp.).

researchpythongo
0
3
Skilllearnbench EvalA

This benchmark evaluates continual learning methods for generating reusable procedural skills in LLM agents. It probes the quality of generated skills, their alignment with execution trajectories, and the ultimate task-solving accuracy and efficiency of a fixed solving agent. Use when the user wants to benchmark on SkillLearnBench, or asks about evaluating this task. Reports Acc..

researchpythongit
0
3
Skillret EvalA

Evaluates the ability of embedding models and rerankers to accurately retrieve relevant software skills from a large, noisy library based on long, scenario-rich user queries. It probes ranking quality, recall, and completeness in a two-stage retrieve-then-rerank pipeline, highlighting the need for domain-specific fine-tuning over general semantic matching. Use when the user wants to benchmark on SkillRet, or asks about evaluating this task. Reports NDCG@k.

ai-agentspythongo
0
3
Skillrouter EvalA

This benchmark evaluates the ability of LLM-based skill routers to accurately select the most relevant tool or function from a large-scale pool of candidate skills based on a user query. It probes retrieval and reranking capabilities under varying difficulty tiers and tests whether models can handle single-skill versus multi-skill routing scenarios. Use when the user wants to benchmark on SkillRouter Benchmark, or asks about evaluating this task. Reports Hit@1.

ai-agentspythongo
0
3
Sku 110k Detection EvalA

Evaluates object detection and counting capabilities in densely packed scenes, specifically testing a model's ability to localize and count tightly overlapping items without false positives from standard non-maximum suppression. Use when the user wants to benchmark on SKU-110K, CARPK, PUCPR+, or asks about evaluating this task. Reports AP.

researchpythongo
0
3
Skyscenes Aerial Seg EvalA

Evaluates semantic segmentation models trained on synthetic aerial imagery for their ability to generalize to real-world UAV datasets and adapt to varying environmental conditions like weather, time of day, and camera viewpoint. Use when the user wants to benchmark on SKYSCENES, UAVid, AEROSCAPES, ICG DRONE, SYNDrone, or asks about evaluating this task. Reports mIoU.

researchpythonperformance
0
3
Skywork Benchmark EvalA

Evaluates bilingual foundation models on general knowledge, Chinese domain-specific reasoning, mathematical problem-solving, and language modeling capabilities using standardized benchmarks and custom held-out text corpora. Use when the user wants to benchmark on MMLU, CEVAL, CMMLU, GSM8K, Custom Chinese LM Testset, or asks about evaluating this task. Reports 5-shot accuracy.

researchpythongo
0
3
Slake Med Vqa EvalA

Evaluates medical visual question answering capabilities by testing a model's ability to reason over radiology images (CT/MRI/X-ray) to answer vision-only and knowledge-based questions in English and Chinese. It probes multimodal fusion, semantic segmentation utilization, and external medical knowledge graph integration for clinical reasoning. Use when the user wants to benchmark on SLAKE, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Slakh2100 Sep EvalA

Evaluates the ability of generative models to separate individual musical instrument stems from a mixed audio track. It probes how well the model captures inter-source dependencies and reconstructs clean waveforms for Bass, Drums, Guitar, and Piano. Use when the user wants to benchmark on Slakh2100, or asks about evaluating this task. Reports SI-SDR_i.

researchpython
0
3
Slam Omni EvalA

This evaluation protocol assesses end-to-end spoken dialogue models across three core capabilities: instruction understanding, logical reasoning, and open-ended oral conversation. It measures both the semantic quality of the generated responses and the acoustic fidelity of the synthesized speech. Use when the user wants to benchmark on Repeat, Summary, StoralEval, TruthfulEval, MLC, AlpacaEval, CommonEval, WildchatEval, or asks about evaluating this task. Reports ChatGPT Score.

researchpythongo
0
3
Slamming Additional EvalA

Evaluates the speech language model's performance across multiple benchmarks including linguistic acceptability, story completion, audio generation quality, and cross-domain text generation perplexity. Use when the user wants to benchmark on sBLIMP, StoryCloze, People Speech, or asks about evaluating this task. Reports MOSnet.

researchpythongo
0
3
Slcp C2st EvalA

Evaluates the accuracy of simulation-based inference engines in recovering the true posterior distribution of model parameters given synthetic observational data. It probes the ability of implicit likelihood methods to handle complex, multimodal posteriors and varying simulation budgets. Use when the user wants to benchmark on SLCP (Simple Likelihood Complex Posterior), or asks about evaluating this task. Reports C2ST.

researchpythonperformance
0
3
Slideagent EvalA

Evaluates multi-page visual document understanding and question answering. It probes a model's ability to retrieve relevant slides, perform spatial and layout reasoning, and accurately extract numeric values or generate lexical matches for open-ended answers. Use when the user wants to benchmark on SlideVQA, TechSlides, FinSlides, or asks about evaluating this task. Reports Num, Overall.

researchpythongo
0
3
Slidequest EvalA

Evaluates an agentic framework's ability to translate natural language scientific queries into executable Python pipelines for automated histopathology analysis on whole-slide images, requiring multi-step computational reasoning rather than simple knowledge recall or diagnosis. Use when the user wants to benchmark on SlideQuest, or asks about evaluating this task. Reports task_success_rate.

researchpythonazure
0
3
Slmrec EvalA

Evaluates the sequential recommendation capability of a distilled small language model against traditional and LLM-based baselines. It measures ranking accuracy on user-item interaction histories and assesses computational efficiency (training/inference time and parameter count). Use when the user wants to benchmark on Amazon18, or asks about evaluating this task. Reports MRR.

researchpythontesting
0
3
Slovene Superglue EvalA

Evaluates monolingual, cross-lingual, and multilingual NLP models on a human- and machine-translated Slovene version of the SuperGLUE benchmark. It probes how well models handle morphological and grammatical challenges in low-resource language processing, and compares translation quality impacts on downstream task performance. Use when the user wants to benchmark on Slovene SuperGLUE, or asks about evaluating this task. Reports Avg.

researchpythongo
0
3
Slt Pose EvalA

This benchmark evaluates how different pose estimation models impact the quality of sign language translation. It probes the robustness of pose estimators to occlusion, temporal instability, and missing hand keypoints, measuring their downstream effect on translation metrics. Use when the user wants to benchmark on RWTH-PHOENIX-Weather 2014, Signsuisse, or asks about evaluating this task. Reports BLEU.

researchpythonapi
0
3