All authors
qhjqhj00 avatar

Claude Skills by qhjqhj00

github.com/qhjqhj00
7,636 skillsA× 7,623B× 11C× 1D× 10 installs2,193 views
Sea Ai User Friendly MetricsA

Compute SEA-AI/user-friendly-metrics via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of SEA-AI/user-friendly-metrics.

developmentpython
0
3
Sea Helm EvalA

Evaluates Thai language models across eight competencies including instruction following, multi-turn dialogue stability, natural language understanding, generation, reasoning, safety, and code-switching resistance. Use when the user wants to benchmark on SEA-HELM, or asks about evaluating this task. Reports SEA-HELM Average Score.

researchpythongo
0
3
Sea Spoof EvalA

This benchmark evaluates audio deepfake detection models on their ability to distinguish real speech from synthetic speech across six South-East Asian languages. It specifically probes cross-lingual generalization and robustness against diverse open-source and commercial text-to-speech and voice conversion systems. Use when the user wants to benchmark on SEA-Spoof, or asks about evaluating this task. Reports EER (%).

researchpythonexpress
0
3
Sea Vision EvalA

Evaluates multimodal language models on document parsing and text-centric visual question answering across 11 Southeast Asian languages. Probes the models' ability to extract structured information from complex documents and answer questions based on visual-textual alignment in low-resource scripts. Use when the user wants to benchmark on SEA-Vision, or asks about evaluating this task. Reports answer accuracy.

researchpythongo
0
3
Seabench EvalA

Evaluates LLMs' ability to handle open-ended, daily interaction scenarios in Southeast Asian languages. It probes contextual adaptation, instruction following, and safety in real-world multilingual usage. Use when the user wants to benchmark on SeaBench, or asks about evaluating this task. Reports LLM-as-a-Judge Score.

researchpythongo
0
3
Seacrowd Benchmark EvalA

Evaluates the zero-shot capability of LLMs, VLMs, and speech models across 13 NLU/NLG tasks, ASR, and image captioning for Southeast Asian languages. It probes multilingual understanding, generation, and cross-modal alignment in low-resource and indigenous language settings. Use when the user wants to benchmark on SEACrowd NLU, SEACrowd NLG, SEACrowd ASR, SEACrowd VL, or asks about evaluating this task. Reports weighted F1 score, WER.

researchpythongit
0
3
Seaexam EvalA

Evaluates LLMs' ability to answer local, culturally grounded multiple-choice questions in Southeast Asian languages (Indonesian, Thai, Vietnamese). It probes regional knowledge, language comprehension, and alignment with actual local usage compared to translated benchmarks. Use when the user wants to benchmark on SeaExam, or asks about evaluating this task. Reports accuracy (%).

researchpythongo
0
3
Sealqa EvalA

Evaluates a model's ability to reason over noisy, conflicting, and ambiguous real-world search results. It probes complex skills like contradiction resolution, temporal tracking, false-premise detection, and multi-document needle-in-a-haystack retrieval. Use when the user wants to benchmark on SealQA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Seam EvalA

Evaluates vision-language models on their ability to reason across semantically equivalent inputs presented in different modalities (vision vs. language) across domain-specific notation systems. It probes cross-modal consistency, modality-agnostic reasoning, and identifies domain-specific perception and tokenization failure modes. Use when the user wants to benchmark on SEAM, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Seamlessm4t Human EvalA

Probes the semantic preservation and audio naturalness of speech-to-text and speech-to-speech translation systems across 24+ languages. Uses human annotators to score translations on a 1-5 scale for meaning similarity (XSTS) and speech quality/naturalness (MOS). Use when the user wants to benchmark on FLEURS test partition, or asks about evaluating this task. Reports XSTS.

researchpython
0
3
Search3d Lerf EvalA

Evaluates a model's ability to perform open-vocabulary segmentation and localization in 3D radiance fields, specifically testing its capacity to understand and process hierarchical queries (e.g., object parts relative to whole objects) versus simple object-level queries. Use when the user wants to benchmark on Search3D (adapted), LERF dataset, or asks about evaluating this task. Reports mIoU.

researchpythontesting
0
3
Seas Safety EvalA

This evaluation probes the safety alignment and refusal capabilities of LLMs when exposed to harmful or adversarial prompts. It measures the frequency of unsafe model outputs to quantify vulnerability, while simultaneously tracking general instruction-following scores to ensure that safety hardening does not degrade overall utility. Use when the user wants to benchmark on SEAS-Test, BeaverTrail, HH-RLHF, XSTest, or asks about evaluating this task. Reports Attack Success Rate (ASR).

researchpythongo
0
3
Sec Gfd EvalA

Evaluates graph neural networks for fraud detection on real-world transaction and review graphs. It specifically probes a model's robustness to severe class imbalance and heterophily, where connected nodes often belong to different classes. Use when the user wants to benchmark on Amazon, YelpChi, T-Finance, T-Social, or asks about evaluating this task. Reports F1-macro, AUC.

researchpythonnode
0
3
Secbench EvalA

Evaluates large language models' cybersecurity knowledge retention and logical reasoning capabilities across multiple subdomains, languages, and difficulty levels using multiple-choice and short-answer questions. Use when the user wants to benchmark on SecBench, or asks about evaluating this task. Reports correctness percentage.

researchpythongo
0
3
Secretbench EvalA

Evaluates the capability of automated secret detection tools to accurately identify hardcoded secrets (e.g., API keys, passwords, private keys) in source code repositories. It probes the tools' ability to balance high recall for true secrets against low false positive rates to mitigate alert fatigue. Use when the user wants to benchmark on SecretBench, or asks about evaluating this task. Reports Precision.

researchpythonexpress
0
3
Secure EvalA

Evaluates large language models' capabilities in cybersecurity, specifically focusing on Industrial Control Systems (ICS). It probes knowledge extraction, vulnerability understanding, out-of-distribution reasoning, and risk evaluation using real-world threat intelligence sources. Use when the user wants to benchmark on MAET, CWET, KCV, VOOD, RERT, CPST, or asks about evaluating this task. Reports accuracy.

securitypythongo
0
3
Secure Inference LatencyA

Measures the online and offline computation latency and communication bandwidth for cryptographic primitives and neural network operations under secure two-party computation. It evaluates how efficiently packed homomorphic encryption and garbled circuits handle matrix-vector products, convolutions, and activation functions without revealing inputs or model parameters. Use when the user has predictions and gold and needs to compute t_online.

researchpythongo
0
3
Securerouter EvalA

This evaluation probes the accuracy and inference efficiency of an encrypted routing framework for secure Transformer inference. It measures how well a cost-aware router dynamically selects smaller MPC-optimized models from a pool to balance privacy-preserving computation costs with task-specific accuracy requirements. Use when the user wants to benchmark on GLUE, or asks about evaluating this task. Reports Inference Speed-up.

researchpythongo
0
3
Sede EvalA

Evaluates text-to-SQL models on naturally occurring, under-specified user queries from Stack Exchange. Probes the model's ability to handle real-world ambiguity, nested subqueries, parameterized queries, and domain-specific schema knowledge without relying on perfectly specified instructions. Use when the user wants to benchmark on SEDE, or asks about evaluating this task. Reports PCM-F1.

researchpythongo
0
3
Seed Emotion Recognition EvalA

Evaluates a model's ability to classify EEG signals into three affective states (negative, neutral, positive) using a semi-supervised learning framework. It tests representation learning and classification performance on high-dimensional, noisy time-series data with limited labeled sessions. Use when the user wants to benchmark on SEED, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Seed Tts EvalA

Evaluates zero-shot voice conversion systems on linguistic preservation, speaker identity retention, and audio naturalness across English, Chinese, and cross-lingual settings. It also measures computational efficiency and latency for both streaming and offline inference modes. Use when the user wants to benchmark on Seed-TTS-Eval, or asks about evaluating this task. Reports WER (%).

researchpythongo
0
3
Seed X EvalA

Evaluates a multimodal model's ability to understand images and text (comprehension) and generate images from text instructions (generation). It probes fine-grained visual perception, reasoning, and compositional image synthesis. Use when the user wants to benchmark on VQAv2, GQA, POPE, MME, SEED, MMB, MM-Vet, MMMU, GenEval, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Seeds Superpixel EvalA

Evaluates the quality of superpixel segmentation algorithms by measuring how well superpixel boundaries align with ground-truth object boundaries and how accurately superpixels can be used as indivisible units for downstream segmentation tasks. Use when the user wants to benchmark on Berkeley Segmentation Dataset (BSD), or asks about evaluating this task. Reports under-segmentation error (UE).

researchpythongo
0
3
Seegull EvalA

Probes a model's propensity to generate or recognize stereotypical associations across diverse global and state-level identity groups. Evaluates the prevalence and cultural specificity of biases in English NLP models, highlighting regional disparities in stereotype content and offensiveness. Use when the user wants to benchmark on SeeGULL, or asks about evaluating this task. Reports stereotype_prevalence.

researchpythongo
0
3
Seephys EvalA

This benchmark evaluates multimodal LLMs' ability to perform physics reasoning using visual diagrams, text, or both. It probes visual interpretation, diagram-to-reasoning mapping, and the model's reliance on textual shortcuts versus actual visual perception across varying knowledge levels and diagram types. Use when the user wants to benchmark on SeePhys, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Seer EvalA

Evaluates LLMs' ability to identify precise textual spans expressing emotion within single-sentence and multi-sentence contexts, distinguishing emotion evidence from other linguistic elements. Use when the user wants to benchmark on SEER, or asks about evaluating this task. Reports F1.

ai-agentspythongo
0
3
Sega Layout EvalA

Evaluates a model's ability to generate content-aware graphic layouts from background images and instructions. It probes spatial reasoning, adherence to design principles (alignment, overlap, occlusion), and aesthetic quality. Use when the user wants to benchmark on PKU, CGL, Crello, or asks about evaluating this task. Reports Ali.

researchpython
0
3
Segale Doc EvalA

Evaluates whether a document-level machine translation evaluation framework can robustly handle translation anomalies (over-translation, under-translation, boundary shifts) and effectively score long-form texts without predefined sentence boundaries. Use when the user wants to benchmark on SEGALE Test Set, or asks about evaluating this task. Reports correlation with human judgments.

researchpythongo
0
3
Segbook EvalA

Evaluates the transfer learning and fine-tuning capabilities of volumetric medical image segmentation models across diverse imaging modalities, anatomical targets, and dataset sizes. It probes how well pre-trained models generalize to downstream segmentation tasks in clinical imaging scenarios and reveals non-linear performance scaling with dataset scale. Use when the user wants to benchmark on SegBook, or asks about evaluating this task. Reports Dice Score (DSC).

researchpythongo
0
3
Segmentmeifyoucan EvalA

This benchmark evaluates a model's ability to segment anomalous or hazardous objects in driving scenes that were not seen during training. It probes out-of-distribution detection and pixel-wise localization of unknown road obstacles, emphasizing safety-critical detection regardless of object class. Use when the user wants to benchmark on RoadAnomaly21, RoadObstacle21, or asks about evaluating this task. Reports AuPRC.

researchpythongo
0
3
Seisclip EvalA

Evaluates a seismology foundation model's ability to classify seismic event types, localize epicenters and depths, and determine focal mechanisms using multi-modal seismic data. It probes cross-dataset generalization and compares fine-tuned, frozen, and scratch-trained variants against spectrum-based baselines. Use when the user wants to benchmark on PNW dataset, SCSN dataset, or asks about evaluating this task. Reports AUC.

datapythonperformance
0
3
Seismic Event Classification EvalA

Evaluates a model's ability to discriminate between three types of seismic events (earthquakes, quarry blasts, and background noise) using waveform and spectral features. It probes the model's capacity to learn physically meaningful seismological signatures like P/S-wave onsets and spectral decay patterns. Use when the user wants to benchmark on Curated Seismic Dataset, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Seismic Inversion EvalA

Evaluates a model's ability to perform semi-supervised seismic impedance inversion using ultra-sparse well-log labels. It probes voxel-level accuracy, patch-level structural similarity, and percentage error on both synthetic and real-world 3D seismic volumes. Use when the user wants to benchmark on SEAM Phase I, Netherlands F3, Delft, or asks about evaluating this task. Reports MAE.

researchpythonexpress
0
3
Seismic Phase Association EvalA

Evaluates the accuracy and computational efficiency of seismic phase associators on synthetic crustal and subduction zone datasets under varying event densities and noise levels. It probes the models' ability to correctly group seismic picks into events and maintain performance under high-stress conditions. Use when the user wants to benchmark on Synthetic Seismic Scenarios (Crustal & Subduction), or asks about evaluating this task. Reports event-level F1 score.

researchpythonrust
0
3
Seismic Picker EvalA

Evaluates the performance of deep learning and classical seismic phase pickers across three tasks: event detection, phase identification, and onset time picking. It probes cross-domain transfer capabilities and robustness to varying signal-to-noise ratios and waveform characteristics. Use when the user wants to benchmark on LenDB, GEOFON, INSTANCE, SCEDC, STEAD, ETHZ, Iquique, NEIC, or asks about evaluating this task. Reports AUC.

researchpythongo
0
3
Seismic Response EvalA

Evaluates a neural operator's ability to map high-frequency seismic wave excitations to building displacement responses across multiple floors. It specifically probes the model's capacity to capture oscillatory function spaces and handle amplitude-frequency disparities between different structural floors. Use when the user wants to benchmark on Custom seismic building response dataset, or asks about evaluating this task. Reports mean relative L2 error.

researchpythontesting
0
3
Seismic Segmentation Al EvalA

This evaluation probes a model's ability to perform semantic segmentation on 3D seismic data under an active learning regime. It measures how effectively a model generalizes to unseen geological volumes when trained on a sequentially selected subset of annotated sections, rather than a fixed passive dataset. Use when the user wants to benchmark on F3 benchmark, Parihaka, or asks about evaluating this task. Reports mIoU.

researchpythongo
0
3
Seismic Wavefield Ctf EvalA

Evaluates machine learning models for seismic wavefield forecasting, reconstruction, and generalization under realistic constraints like noise and limited data. It probes model robustness and dynamic learning by comparing performance across multiple tasks against naive baselines. Use when the user wants to benchmark on global wavefields, DAS, synthetic 3D crustal wavefields, or asks about evaluating this task. Reports multi-metric scoring.

datapythonrust
0
3
Seist Earthquake Monitoring EvalA

Evaluates a deep learning model's capability to perform multiple earthquake monitoring tasks, including seismic phase picking, detection, polarity classification, and magnitude estimation. It specifically probes cross-regional out-of-distribution generalization by training on Chinese seismic network data and testing on geologically distinct Pacific Northwest data. Use when the user wants to benchmark on DiTing, PNW (ComCat event subset), or asks about evaluating this task. Reports F1-Score.

researchpythongo
0
3
Self Adaptive Curriculum Nlu EvalA

Evaluates whether self-adaptive curriculum learning strategies, which use pre-trained model confidence to estimate example difficulty, improve fine-tuning performance over random and length-based sampling baselines across multiple NLU tasks. Use when the user wants to benchmark on SST-2, SST-5, HSOL, XNLI, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Selfcheckgpt EvalA

Evaluates a model's ability to detect hallucinated versus factual content in generated text using zero-resource consistency metrics across stochastic samples. It probes whether factual knowledge yields coherent, consistent outputs while hallucinated content exhibits divergence across multiple generations. Use when the user wants to benchmark on SelfCheckGPT dataset, or asks about evaluating this task. Reports AUC-PR.

researchpythongo
0
3
Seller Outcome Fairness EvalA

Evaluates the trade-off between platform revenue (GMV) and seller-side exposure fairness in online marketplace recommendation systems using simulated online environments trained on historical interaction data. Use when the user wants to benchmark on Proprietary Dataset, Electronics Event History (EVS) Dataset, or asks about evaluating this task. Reports GMV relative change.

researchpythongo
0
3
Selqa EvalA

Evaluates a model's ability to retrieve relevant answer sentences from a document given a question (selection), and to determine whether a document section contains an answer at all (triggering). It probes open-domain QA robustness against paraphrasing, varying question types, and section lengths. Use when the user wants to benchmark on SelQA, or asks about evaluating this task. Reports MAP.

researchpythongo
0
3
Semantic Change Detection EvalA

Evaluates the ability of contextualized language models to detect diachronic semantic change in words across different time periods. It probes whether models can accurately rank words by their degree of meaning shift compared to human-annotated gold standards. Use when the user wants to benchmark on SemEval-2020 Task 1, GEMS, or asks about evaluating this task. Reports Spearman's ρ.

researchpythongo
0
3
Semantic Helm EvalA

Evaluates reinforcement learning agents' ability to learn and utilize memory mechanisms in partially observable environments. It probes sample efficiency, convergence speed, and the capacity to retain and retrieve semantic information across varying levels of visual complexity and task duration. Use when the user wants to benchmark on MiniGrid, MiniWorld, Avalon, Psychlab (CR task), or asks about evaluating this task. Reports IQM.

ai-agentspythonperformance
0
3
Semantic Kg EvalA

Evaluates the ability of semantic similarity methods to correctly classify pairs of natural language statements as semantically similar (label 1) or dissimilar (label 0). It specifically probes how well models handle controlled semantic variations (node and edge perturbations) across general and domain-specific knowledge domains. Use when the user wants to benchmark on Semantic-KG Benchmark, or asks about evaluating this task. Reports F1-score.

researchpythonnode
0
3
Semantic Syntactic Word EvalA

Evaluates whether trained word vector models can capture semantic and syntactic relationships between words through simple algebraic operations in vector space. It probes the model's ability to solve analogy-style questions by measuring how well the vector arithmetic preserves linguistic regularities. Use when the user wants to benchmark on Semantic-Syntactic Word Relationship test set, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Semantic Textual Similarity EvalA

Evaluates a model's ability to quantify the degree of semantic similarity between pairs of sentences, including multilingual and cross-lingual contexts. It probes fine-grained semantic matching and cross-lingual generalization rather than binary paraphrase detection. Use when the user wants to benchmark on SemEval-2017 STS, or asks about evaluating this task. Reports Pearson correlation.

researchpythongo
0
3
Semanticagent EvalA

Evaluates text-to-SQL generation capabilities across cross-domain parsing, knowledge-intensive reasoning, and enterprise-level SQL workflows. It also assesses the semantic validity, execution correctness, and diversity of synthetically generated training data. Use when the user wants to benchmark on Spider, BIRD, Spider2.0, EHRSQL, ScienceBenchmark, Spider-Syn, Spider-Realistic, Spider-DK, or asks about evaluating this task. Reports test-suite accuracy (TS), execution accuracy (EX).

researchpythongo
0
3
SemascoreA

Evaluates automatic speech recognition (ASR) transcription quality by measuring segment-wise semantic similarity and error weighting. It specifically probes robustness on disordered, noisy, and accented speech, testing alignment with human judgments and downstream NLU task metrics. Use when the user has predictions and gold and needs to compute SeMaScore.

researchpythongo
0
3