All authors
qhjqhj00 avatar

Claude Skills by qhjqhj00

github.com/qhjqhj00
7,636 skillsA× 7,623B× 11C× 1D× 10 installs2,089 views
Tox21 Fsl EvalA

Evaluates graph neural networks for few-shot toxic molecule classification. It probes the model's ability to generalize from very limited labeled examples (shots) and adapt to new query sets using meta-learning and graph augmentation techniques. Use when the user wants to benchmark on Tox21, or asks about evaluating this task. Reports ROC-AUC Score.

researchpythonnode
0
3
Toxic Comment Classification EvalA

Evaluates transformer and RNN models on their ability to classify toxic comments while measuring classification accuracy and inference speed. It specifically probes identity-based bias by measuring how well models distinguish between toxic and normal comments across demographic subgroups. Use when the user wants to benchmark on Civil Comments, or asks about evaluating this task. Reports Macro AUROC.

researchpythongo
0
3
Toxic Comment EvalA

This benchmark evaluates the individual and group fairness of toxicity classifiers on online comments. It probes whether model predictions remain stable when sensitive identity tokens are swapped (individual fairness) and whether prediction accuracy is equitable across different demographic groups (group fairness). Use when the user wants to benchmark on Toxic Comment Classification Challenge, or asks about evaluating this task. Reports Balanced Accuracy (BA).

researchpythongit
0
3
Toxic Language Classification EvalA

Evaluates binary toxic language classification under extreme data scarcity and severe class imbalance. It probes how well classifiers can detect the minority 'threat' class when trained on a very small labeled dataset, and measures the effectiveness of various data augmentation techniques in improving recall and macro-F1. Use when the user wants to benchmark on Seed, or asks about evaluating this task. Reports macro-averaged F1-score.

researchpythongo
0
3
Toxicity Analysis EvalA

Measures the toxicity of text generated by language models conditioned on specific prompts, evaluating how alignment techniques like prompting and context distillation affect harmful content generation. Use when the user wants to benchmark on RealToxicityPrompts, or asks about evaluating this task. Reports mean toxicity score.

ai-agentspythonapi
0
3
Toxicity Detection EvalA

Probes the ability of text generation models to produce non-toxic content by measuring average toxicity scores and comparing them via statistical significance testing. It specifically evaluates how accounting for classifier uncertainty affects the reliability of these comparisons. Use when the user wants to benchmark on BOLD, RealToxicityPrompts, or asks about evaluating this task. Reports Confidence Interval.

researchpythongo
0
3
Toxicity EvalA

Evaluates the toxicity of text sequences (prompts and model continuations) by scoring them with a black-box API. It probes how well models generate non-toxic text and how sensitive toxicity metrics are to API updates and score drift over time. Use when the user wants to benchmark on REALTOXICITYPROMPTS, or asks about evaluating this task. Reports Toxic Fraction.

researchpythonapi
0
3
Toxicity Perspectives EvalA

Evaluates how well automated toxicity classifiers align with diverse human perceptions of harmful content, specifically measuring how demographic background and personal harassment experiences influence toxicity judgments. Use when the user wants to benchmark on Toxicity Perspectives Dataset, or asks about evaluating this task. Reports interrater agreement (Cohen's kappa).

researchpythongo
0
3
Toximol EvalA

Evaluates whether Multimodal Large Language Models (MLLMs) can generate structurally valid, low-toxicity alternative molecules from toxic inputs while adhering to drug-likeness, synthetic feasibility, and structural similarity constraints. It probes the model's ability to perform structure-aware molecular editing and cross-modal scientific reasoning. Use when the user wants to benchmark on ToxiMol, or asks about evaluating this task. Reports Toxicity Repair Success Rate.

researchpythongit
0
3
Toyadmos EvalA

Evaluates unsupervised anomalous sound detection systems on miniature machine operating sounds. It probes the ability of models to learn normal acoustic patterns and identify deviations caused by mechanical faults or environmental variations. Use when the user wants to benchmark on ToyADMOS, or asks about evaluating this task. Reports AUC-ROC.

researchpythongit
0
3
Tpc 268 EvalA

Evaluates class-agnostic counting (CAC) capabilities on fine-grained, taxonomically diverse plant species under few-shot exemplar guidance. It probes models' ability to handle dense occlusion, multi-scale observation, and cross-domain generalization from generic objects to biological imagery. Use when the user wants to benchmark on TPC–268, or asks about evaluating this task. Reports MAE.

testingpythontesting
0
3
Tpc H Tpc C Energy EvalA

Evaluates the energy efficiency and performance trade-offs of a distributed database cluster versus a single high-end server under OLAP and OLTP workloads. It probes the system's ability to maintain energy proportionality through dynamic node scaling and measures the overhead incurred during data migration and cluster reconfiguration. Use when the user wants to benchmark on TPC-H, TPC-C, or asks about evaluating this task. Reports energy consumption per query.

businesspythonnode
0
3
Tpcds Structural EvalA

Evaluates an LLM's ability to generate structurally complex SQL queries for real-world decision-making workloads. It probes the model's capacity to handle deep nesting, multiple joins, diverse column references, and complex filtering conditions compared to simpler benchmarks. Use when the user wants to benchmark on TPC-DS, or asks about evaluating this task. Reports structural_similarity.

researchpythongo
0
3
Tpr@FprA

Evaluates the discriminative capability of a speaker verification model by measuring the true positive rate at fixed false positive rate thresholds. It probes how well the model's embedding space separates same-speaker pairs from different-speaker pairs under controlled error constraints. Use when the user has predictions and gold and needs to compute TPR@FPR.

researchpythongo
0
3
Tpscalcbench EvalA

Evaluates large language models' ability to perform analytical calculations in hypersonic thermal protection system engineering using closed-form formulas and thermodynamic relations, without relying on external simulation tools. Use when the user wants to benchmark on TPS-CalcBench, or asks about evaluating this task. Reports relative_error.

researchpythonrust
0
3
Tpu Workload EvalA

Evaluates the performance and energy efficiency of a Tensor Processing Unit (TPU) and alternative hardware designs across six specific neural network workloads. It probes how architectural parameters like memory bandwidth, clock rate, and matrix multiply unit size impact throughput and power consumption. Use when the user wants to benchmark on TPU Benchmark Workloads (MLP0, MLP1, LSTM0, LSTM1, CNN0, CNN1), or asks about evaluating this task. Reports Watt/die.

researchpythonperformance
0
3
Tqabench EvalA

Evaluates large language models' ability to perform multi-table question answering across varying context lengths (8K–64K tokens) and complex reasoning tasks. It probes cross-table inference, symbolic reasoning, and handling of real-world relational data without Wikipedia bias. Use when the user wants to benchmark on TQA-Bench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Trace Classification EvalA

Evaluates a model's ability to classify OpenTelemetry workflow traces as benign, suspicious, or malicious, and assesses its knowledge of cybersecurity frameworks via multiple-choice questions. Use when the user wants to benchmark on OpenTelemetry Workflow Traces, or asks about evaluating this task. Reports Overall Accuracy.

researchpythongo
0
3
Trace Encoding Benchmark EvalA

Evaluates the quality and efficiency of trace encoding methods for process mining event logs. It probes how well encodings preserve trace similarities (expressivity), their computational cost as data scales (scalability), and their suitability for downstream process mining tasks. Use when the user wants to benchmark on Process Mining Event Log Scenarios (1-5), or asks about evaluating this task. Reports T4.

researchpythonexpress
0
3
Trace EvalA

This protocol evaluates training-free partial audio deepfake detection by analyzing the temporal continuity of frozen speech foundation model embeddings. It probes a model's ability to detect splice boundaries and synthetic insertions in speech without requiring labeled training data or architectural modifications. Use when the user wants to benchmark on PartialSpoof, HalfTruth Audio Deepfake (HAD), ADD 2023 Track 2, LlamaPartialSpoof, or asks about evaluating this task. Reports EER.

researchpythonperformance
0
3
Trace Gmv Prediction EvalA

This benchmark evaluates a model's ability to predict post-click Gross Merchandise Volume (GMV) under delayed feedback conditions. It specifically probes how well models adapt to rapidly evolving label distributions through online streaming training and whether they can effectively handle the distinct statistical properties of single-purchase versus repurchase transactions. Use when the user wants to benchmark on TRACE, or asks about evaluating this task. Reports AUC.

researchpythongo
0
3
Trace Reward Hack Detection EvalA

This benchmark evaluates an LLM's ability to detect and classify reward hacking behaviors in multi-turn code generation trajectories. It specifically probes contrastive anomaly detection capabilities by presenting clusters of mixed benign and malicious trajectories, testing whether models can disentangle subtle semantic and syntactic exploit patterns without prior taxonomy exposure. Use when the user wants to benchmark on TRACE, or asks about evaluating this task. Reports Detection Rate.

researchpythongo
0
3
Track Any State EvalA

Evaluates a model's ability to track objects through appearance-changing state transformations and explicitly model those transformations as a state graph. It probes spatiotemporal continuity, zero-shot object recovery, and semantic reasoning about object interactions. Use when the user wants to benchmark on VOST, VSCOS, M3-VOS, DAVIS 2017, VOST-TAS, or asks about evaluating this task. Reports Jaccard (J).

researchpython
0
3
Trackml Edge Probability EvalA

Probes a GNN's capability to perform edge scoring on highly sparse, irregular scientific graphs by predicting the probability that a directional connection between two 3D space-point measurements originates from the same particle. Use when the user wants to benchmark on TrackML, or asks about evaluating this task. Reports edge probability.

researchpythongo
0
3
Tracknet Tracking EvalA

Evaluates the ability of deep learning models to detect and track high-speed, tiny objects (tennis and badminton balls) in broadcast sports videos. It probes robustness to motion blur, occlusion, and domain shifts by comparing single-frame vs. multi-frame tracking and transfer learning across different sports. Use when the user wants to benchmark on Tennis, Badminton, or asks about evaluating this task. Reports F1-measure.

researchpythonperformance
0
3
Trackrad2025 EvalA

Probes the capability of real-time tumor and surrogate localization in MRI-guided radiotherapy using 2D sagittal cine MRI sequences. It evaluates how well algorithms can track anatomical motion across varying frame rates and multi-vendor MRI-linac hardware under clinically relevant conditions. Use when the user wants to benchmark on TrackRAD2025, or asks about evaluating this task. Reports tracking performance.

researchpythongo
0
3
Tracsum EvalA

Evaluates a model's ability to generate aspect-specific summaries from clinical abstracts and accurately cite the supporting source sentences. It probes factual recall, conciseness, and traceability in a medical domain setting. Use when the user wants to benchmark on TracSum, or asks about evaluating this task. Reports Claim Recall.

researchpythongo
0
3
Trade The Event EvalA

This evaluation probes a model's ability to detect objective corporate events in financial news and translate those detections into actionable, timely trading signals. It measures how effectively the detected events predict short-term stock price movements and generate excess returns compared to a market benchmark. Use when the user wants to benchmark on EDT, or asks about evaluating this task. Reports Winning Rate.

researchpythongit
0
3
Traffic Classification EvalA

Evaluates the ability of deep learning models to classify mobile network traffic into specific application categories (e.g., video, music, games) and background services using extracted packet flow features. It specifically probes the effectiveness of semi-supervised VAE-CNN architectures and model pruning techniques for resource-constrained edge devices. Use when the user wants to benchmark on Private Campus Network Dataset, or asks about evaluating this task. Reports Accuracy.

researchpythonrust
0
3
Traffic Destination Prediction EvalA

Evaluates the accuracy of multi-modal trajectory forecasting models in predicting the final destination of traffic agents (pedestrians and vehicles) over a future time horizon. Use when the user wants to benchmark on SDD, InD, Argoverse, or asks about evaluating this task. Reports Minimum final displacement error.

researchpythongo
0
3
Traffic Flow Forecasting EvalA

Evaluates the ability of spatiotemporal GNN models to forecast future traffic flow (speed or occupancy) based on historical sensor data and road network topology. Use when the user wants to benchmark on METR-LA, PEMS-BAY, PeMS04, or asks about evaluating this task. Reports MAE.

researchpythongo
0
3
Traffic Incident Forecasting EvalA

Evaluates spatiotemporal models' ability to localize traffic collision events in time and space, and to forecast network-level congestion and emissions. It probes multi-horizon forecasting accuracy and spatial-temporal coherence under simulated disruption scenarios. Use when the user wants to benchmark on NYC Broadway corridor, or asks about evaluating this task. Reports containment_performance.

researchpythonperformance
0
3
Traffic Speed Forecasting EvalA

This evaluation protocol assesses a model's ability to forecast future traffic speeds on road networks under varying conditions, including the impact of construction workzones. It probes spatio-temporal dependency modeling by measuring prediction accuracy across multiple forecast horizons (15, 30, and 60 minutes) on real-world highway sensor data. Use when the user wants to benchmark on Tyson's Corner, Los-loop, PEMS-BAY, or asks about evaluating this task. Reports RMSE.

researchpythonnode
0
3
Traffic Workzone Forecasting EvalA

Evaluates the ability of spatio-temporal graph neural networks to forecast traffic speed under normal and construction work zone disruption conditions. It probes how well models integrate heterogeneous work zone data to capture nonlinear spatio-temporal dependencies and maintain accuracy during significant traffic flow deviations. Use when the user wants to benchmark on Richmond, Tyson’s, or asks about evaluating this task. Reports MAE.

researchpythontesting
0
3
Tragesql EvalA

This benchmark evaluates a model's ability to classify the intention of a natural language question relative to a database schema. It probes whether the model can distinguish between answerable queries, questions requiring external knowledge, ambiguous queries, grammatically invalid non-SQL questions, and questions unrelated to the schema. Use when the user wants to benchmark on TRIAGESQL, or asks about evaluating this task. Reports Macro F1.

researchpythongo
0
3
Train O Matic Wsd EvalA

Evaluates the quality of automatically generated multilingual word sense disambiguation (WSD) training corpora by training a supervised WSD system (IMS) on them and measuring performance on standard WSD benchmark datasets. It probes whether synthetic sense-annotated data can match or exceed manually annotated corpora, particularly for low-resource languages. Use when the user wants to benchmark on Senseval-2, Senseval-3, SemEval-2007, SemEval-2013, SemEval-2015, or asks about evaluating this ...

researchpythongo
0
3
Training Free Multi Step Audio Sep EvalA

Evaluates a training-free iterative inference method for audio source separation. It probes the model's ability to progressively refine noisy audio mixtures (speech or music) by optimizing blending ratios across multiple inference steps without retraining or architectural changes. Use when the user wants to benchmark on VCTK-DEMAND, DNS Challenge v3, MUSDB18-HQ, or asks about evaluating this task. Reports PESQ, UTMOS, uSDR.

researchpythongo
0
3
Training Speed EvalA

Evaluates the training efficiency and scalability of AlphaFold-like models on GPU clusters. It measures per-step execution time, overall wall-clock training duration, and convergence speed across different hardware configurations and optimization techniques. Use when the user wants to benchmark on OpenFold dataset, or asks about evaluating this task. Reports step time.

researchpythonperformance
0
3
Training ThroughputA

Evaluates the scalability and efficiency of distributed machine learning systems by measuring how many training tokens each system can process per second across different hardware topologies and model parallelism strategies. Use when the user has predictions and gold and needs to compute training throughput (tokens/second).

datapythongo
0
3
Trajectory Generation EvalA

Evaluates the statistical fidelity and practical utility of synthetic human trajectory generation models by measuring how well generated trajectories perform on downstream mobility tasks compared to real trajectories. It probes whether synthetic data can replace real data without performance degradation across recommendation, prediction, labeling, and simulation tasks. Use when the user wants to benchmark on Foursquare Tokyo (TKY), Foursquare Istanbul (IST), Foursquare New York City (NYC), or...

researchpythongo
0
3
Trans Env EvalA

Evaluates the linguistic robustness of LLMs by measuring performance degradation when standard English prompts are transformed into 38 regional dialects and ESL varieties. It probes whether models maintain accuracy and instruction-following capabilities across non-standard linguistic variations. Use when the user wants to benchmark on MMLU, ARC, TruthfulQA, GSM8K, HellaSwag, WinoGrande, IFEval, AlpacaFarm, MT-Bench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Transbench EvalA

TransBench evaluates machine translation models across three industrial capability levels: basic linguistic quality and robustness, domain-specific proficiency (e-commerce/finance), and cultural adaptation (taboo words and honorifics). It probes whether models maintain translation fidelity under input perturbations, adhere to domain terminology, and correctly handle culturally sensitive expressions without omission or over-translation. Use when the user wants to benchmark on TransBench, or as...

researchpythonexpress
0
3
Transboost EvalA

Evaluates a boosting-tree kernel transfer learning algorithm for financial risk prediction and fraud detection under domain distribution shifts and data sparsity. It measures how well the model adapts from a source domain to a target domain with limited labeled samples, while maintaining computational efficiency and interpretability. Use when the user wants to benchmark on Tencent Mobile Payment Dataset, LendingClub Dataset, Wine Quality Dataset, or asks about evaluating this task. Reports AUC.

researchpythongo
0
3
Transevalnia EvalA

Probes a model's ability to rank machine translation candidates by quality and provide fine-grained, dimensionally structured justifications aligned with MQM standards. It also evaluates the model's robustness to candidate ordering (position bias) when performing comparative translation assessment. Use when the user wants to benchmark on WMT-2024 en-es, WMT-2023 en-de, WMT-2023 zh-en, WMT-2022 en-ru, WMT-2021 en-ja, WMT-2021 ja-en, Hard en-ja, Generic, Haiku 100, Haiku Full, or asks about eva...

researchpythongo
0
3
Transfer Fraud Detection EvalA

Evaluates machine learning models for detecting fraudulent bank transfers by optimizing instance-dependent cost-sensitive objectives. It probes the model's ability to minimize financial losses and maximize expected savings under highly imbalanced transaction data. Use when the user wants to benchmark on Credit Card Transaction Data, Bank data set, or asks about evaluating this task. Reports Expected Savings.

datapythongo
0
3
Transformer Interpretability EvalA

Evaluates the faithfulness and class-specificity of Transformer interpretability methods by measuring how well highlighted input features align with model predictions. It probes explanation quality through pixel/token masking, segmentation overlap, and rationale extraction accuracy. Use when the user wants to benchmark on ImageNet Validation (ILSVRC 2012), ImageNet-Segmentation, Movie Reviews, or asks about evaluating this task. Reports AUC (Positive/Negative Perturbation).

researchpythongo
0
3
Transition Based Parsing EvalA

This evaluation probes a parser's ability to construct accurate projective dependency trees for sentences. It measures how well the model identifies correct syntactic heads and their grammatical relations (labels) under both English and Chinese linguistic conditions. Use when the user wants to benchmark on Penn Treebank (PTB) v5, Chinese Treebank (CTB) v5, or asks about evaluating this task. Reports LAS.

researchpythongo
0
3
Translation Proofreading EvalA

Evaluates a hybrid CNN-BERT model's ability to detect and correct errors in English translations. It measures performance across different architectural hyperparameters (kernel size, batch normalization) and linguistic granularity levels. Use when the user wants to benchmark on WMT English-German Parallel Corpus Dataset, Open Subtitles Dataset, or asks about evaluating this task. Reports F1-Score (%).

researchpythongo
0
3
TranslationeditrateA

Compute the TranslationEditRate metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute TranslationEditRate, or asks how to score with TranslationEditRate.

documentationpython
0
3
Transport Forecasting EvalA

Evaluates spatio-temporal forecasting models on traffic speed, volume, and bike flow prediction tasks. It specifically probes whether simple baselines that account for weekly stationarity (historical average plus linear regression on residuals) can match or outperform complex deep learning architectures across diverse transport datasets. Use when the user wants to benchmark on PeMSD7(M), Urban1, NYC Citi Bike, PeMSD4, SZ-taxi, METR-LA, PEMS-BAY, NYC Bike in- and out-flows, Seattle traffic spe...

researchpython
0
3