All authors
qhjqhj00 avatar

Claude Skills by qhjqhj00

github.com/qhjqhj00
7,636 skillsA× 7,623B× 11C× 1D× 10 installs2,120 views
Taxpraben EvalA

Evaluates LLMs on Chinese real-world tax practice tasks spanning classification, generation, structured prediction, and mixed matching. It probes capabilities across Bloom's taxonomy levels, from factual recall and understanding to complex tax strategy planning and risk prevention. Use when the user wants to benchmark on TaxPraBen, or asks about evaluating this task. Reports Overall Average.

researchpythongo
0
3
Tb Bench EvalA

This benchmark evaluates multi-modal large language models' ability to understand spatio-temporal traffic behaviors from ego-centric dashcam images and videos. It probes eight distinct perception tasks, including road detection, object-lane alignment, turning prediction, and ego-trajectory estimation, requiring both spatial reasoning and temporal tracking. Use when the user wants to benchmark on TB-Bench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Tb Classification EvalA

Evaluates deep learning models' ability to classify chest X-rays as tuberculosis or normal, comparing whole-image vs. lung-segmented inputs. Use when the user wants to benchmark on Kaggle CXR images and lung mask dataset, or asks about evaluating this task. Reports accuracy.

researchpythontesting
0
3
Tbar EvalA

Evaluates the effectiveness of template-based automated program repair systems by applying fix patterns to buggy Java programs. It probes the system's ability to localize faults, generate syntactically valid patches, and pass test suites without breaking existing tests. Use when the user wants to benchmark on Defects4J, or asks about evaluating this task. Reports plausible_patch.

researchpythongo
0
3
Tcab EvalA

Evaluates a model's ability to detect whether a given text instance has been adversarially perturbed (attack detection) and to identify the specific attack method used (attack labeling) across multiple text classification domains. Use when the user wants to benchmark on TCAB, or asks about evaluating this task. Reports balanced accuracy.

researchpythongo
0
3
Tcga Histopathology EvalA

Evaluates the representation quality and generalization of self-supervised histopathology models across diverse patch-level diagnostic tasks and weakly supervised slide-level tasks using linear probing and fine-tuning on TCGA whole slide images. Use when the user wants to benchmark on TCGA Histopathology, or asks about evaluating this task. Reports average AUC.

researchpythongit
0
3
Tcm Best4sdt EvalA

This benchmark evaluates large language models' capabilities in Traditional Chinese Medicine (TCM) clinical reasoning, specifically focusing on syndrome differentiation and treatment decision-making. It probes the model's ability to accurately diagnose pathological patterns, formulate appropriate herbal prescriptions, and adhere to medical ethics and safety guidelines across 27 dimensions. Use when the user wants to benchmark on TCM-BEST4SDT, or asks about evaluating this task. Reports select...

researchpythongo
0
3
Tcmsd Sd EvalA

Evaluates a model's ability to perform syndrome differentiation in Traditional Chinese Medicine by classifying clinical records into one of 148 predefined syndromes. It probes the model's capacity to handle domain-specific medical terminology and imbalanced multi-class classification. Use when the user wants to benchmark on TCM-SD, or asks about evaluating this task. Reports Macro-F1.

businesspythongo
0
3
Tdbench EvalA

Evaluates vision-language models on top-down (aerial) image understanding by testing their ability to answer questions about rotated views. It measures rotational consistency to filter out hallucinations and decomposes performance into true knowledge versus lucky guessing via a probabilistic reliability framework. Use when the user wants to benchmark on TDBench, or asks about evaluating this task. Reports RotationalEval (RE).

researchpythongo
0
3
Tdc Adme Pk EvalA

Assesses pharmacological property prediction across ADME, PK, and toxicity tasks, including regression, classification, and correlation-based evaluation. Use when the user wants to benchmark on TDC Benchmark, or asks about evaluating this task. Reports AUROC / AUPRC.

researchpythongo
0
3
Tdc Admet EvalA

Evaluates molecular property prediction across 22 ADMET tasks. It probes the model's ability to generalize across diverse chemical properties using standardized benchmark splits for both regression and classification. Use when the user wants to benchmark on TDC ADMET Group, or asks about evaluating this task. Reports Classification AUROC.

researchpythonperformance
0
3
Tedigan Text To Face EvalA

Evaluates a model's capability to synthesize high-resolution, diverse, and photorealistic face images conditioned on natural language prompts, and to perform text-guided editing of existing faces while preserving identity and irrelevant attributes. Use when the user wants to benchmark on Multi-Modal CelebA-HQ, or asks about evaluating this task. Reports FID.

researchpythongit
0
3
Telel Accents Unplugged EvalA

Compute TelEl/accents_unplugged_eval via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of TelEl/accents_unplugged_eval.

developmentpython
0
3
Telemedicine Feedback EvalA

Predicts whether a patient will give positive feedback (thumbs-up) for a doctor's response in a Romanian telemedicine platform. It probes the model's ability to leverage clinical communication features, patient/doctor history, and metadata to forecast user satisfaction. Use when the user wants to benchmark on Romanian Telemedicine Platform Dataset, or asks about evaluating this task. Reports ROC-AUC.

datapythongo
0
3
Teleoracle EvalA

Evaluates domain-specific question answering and retrieval-augmented generation capabilities in telecommunications. Probes a model's ability to accurately answer multiple-choice questions about 3GPP standards and adhere to retrieved context without relying on generalized prior knowledge. Use when the user wants to benchmark on TeleQnA, or asks about evaluating this task. Reports Accuracy.

ai-agentspythongo
0
3
Teleqna EvalA

Evaluates large language models' domain-specific knowledge in telecommunications, covering general terminology, research concepts, and complex technical standards. It also benchmarks model performance against human telecom professionals under strict no-search conditions. Use when the user wants to benchmark on TeleQnA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Tembed EvalA

Evaluates the quality and efficiency of tabular embedding models across four granularity levels (cell, row, column, table) and six downstream tasks including similarity search, triplet evaluation, prediction, and retrieval. It probes whether a single embedding approach can generalize universally across diverse structured data applications or if performance is highly task- and granularity-dependent. Use when the user wants to benchmark on TEmBed Benchmark Suite, or asks about evaluating this t...

researchpythongit
0
3
Temmed Bench EvalA

Evaluates large vision-language models' ability to perform temporal reasoning on medical images by analyzing condition changes across multiple clinical visits. It probes capabilities in visual question answering, longitudinal clinical report generation, and selecting relevant image pairs based on temporal context. Use when the user wants to benchmark on TemMed-Bench, or asks about evaluating this task. Reports Avg..

researchpythongo
0
3
Tempo Sum EvalA

Evaluates text summarization models' temporal generalization by testing on datasets split by publication date, specifically probing how well models handle knowledge-conflicting future articles versus in-distribution past data. Use when the user wants to benchmark on BBC, CNN, or asks about evaluating this task. Reports FactCC.

researchpythontesting
0
3
Tempobench EvalA

Evaluates large language models' temporal reasoning capabilities by decomposing performance into trace-based (TTE) and causal (TCE) components. It measures how well models handle structured logical specifications with varying complexity, isolating structural factors like horizon depth and information density. Use when the user wants to benchmark on TempoBench, or asks about evaluating this task. Reports exact-match accuracy.

researchpythongo
0
3
Temporal Degradation EvalA

Evaluates how temporal misalignment between pretraining/fine-tuning data and evaluation data impacts model performance across classification and summarization benchmarks. Use when the user wants to benchmark on PubCLS, NewSum, TwiERC, AIC, PoliAff, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Temporal Domain Generalization EvalA

Evaluates a model's ability to generalize to future, unseen temporal domains without full retraining. It measures out-of-distribution accuracy on sequentially arriving target domains after training on historical source domains. Use when the user wants to benchmark on Yearbook, Rotated MNIST (RMNIST), FMoW, Huffpost, Arxiv, CLEAR-10/100, or asks about evaluating this task. Reports OOD_avg accuracy.

researchpythongo
0
3
Temporal Graph Anomaly EvalA

Evaluates the ability of various data-driven models to detect emerging anomalies in temporal graphs derived from social media interactions. It probes how well different architectures generalize across different social platforms and remain robust to parameter variations and temporal/spatial shifts. Use when the user wants to benchmark on Twitter, Facebook, or asks about evaluating this task. Reports weighted F1 score.

researchpythonperformance
0
3
Temporalbench EvalA

TemporalBench probes LLM-based agents' ability to perform contextual and event-informed temporal reasoning across four distinct task families. It disentangles historical pattern interpretation, context-free forecasting, contextual alignment, and event-conditioned adaptation to reveal whether numerical prediction accuracy correlates with qualitative temporal judgment. Use when the user wants to benchmark on FreshRetailNet, PSML, Causal Chambers, MIMIC, or asks about evaluating this task. Repor...

researchpythongo
0
3
Tempqa Wd EvalA

Probes a system's ability to perform temporal question answering over knowledge bases by generating correct answers or SPARQL queries. It specifically evaluates generalization across different knowledge bases (Wikidata vs. Freebase) and interpretability through fine-grained intermediate annotations like entity/relation linking and λ-expressions. Use when the user wants to benchmark on TempQA-WD, or asks about evaluating this task. Reports F1.

researchpythongo
0
3
Tempreason EvalA

Evaluates large language models' ability to perform temporal reasoning across three complexity levels: time-time relations (L1), time-event relations (L2), and event-event relations (L3). It specifically probes models' robustness to historical and futuristic time periods, as well as their capacity for month-level intra-year reasoning. Use when the user wants to benchmark on TEMPREASON, or asks about evaluating this task. Reports EM.

researchpythongo
0
3
Tempusbench Univariate EvalA

Evaluates time-series foundation models, statistical methods, and machine learning algorithms on univariate forecasting tasks. It probes their ability to handle diverse statistical properties like stationarity, seasonality, sparsity, and noise across real-world and synthetic datasets. Use when the user wants to benchmark on TempusBench Univariate Benchmark, or asks about evaluating this task. Reports MASE.

datapythongo
0
3
Tenrec EvalA

Evaluates recommender systems across multiple tasks including click-through rate (CTR) prediction, sequential recommendation, and top-N item ranking. It probes cross-domain generalization, cold-start handling, and the sensitivity of ranking metrics to negative sampling strategies. Use when the user wants to benchmark on Tenrec, or asks about evaluating this task. Reports AUC.

researchpythongit
0
3
Tenspiler EvalA

Evaluates the correctness and performance of a verified lifting-based compiler that transpiles sequential C++/Python code into tensor operations across various DSLs and hardware accelerators. It measures synthesis efficiency, kernel execution speedup, and end-to-end performance including data transfer overhead. Use when the user wants to benchmark on TENSPILER benchmark suite, or asks about evaluating this task. Reports kernel_performance.

researchpythongo
0
3
Terminal Bench 2.0 EvalA

Evaluates an LLM's ability to execute complex, multi-step terminal commands and tasks in a sandboxed environment. It probes capabilities across software engineering, system administration, data processing, security, and debugging. Use when the user wants to benchmark on Terminal-Bench 2.0, or asks about evaluating this task. Reports TB2.0.

researchpythondebugging
0
3
Terminology Aware Translation EvalA

Evaluates a machine translation system's ability to balance overall translation quality with strict adherence to specified terminology constraints across different languages. It measures how well the model enforces lexical rules without degrading fluency or adequacy, particularly in morphologically complex languages. Use when the user wants to benchmark on EN-{DE,ES,RU} translation test sets, or asks about evaluating this task. Reports BLEU.

researchpythonapi
0
3
TerraLingua EvalA

Evaluates the emergence of open-ended dynamics, sustained novelty, and social organization in a persistent multi-agent LLM ecology. It probes how environmental constraints, agent personality, and artifact persistence shape cumulative cultural evolution and cooperative norms. Use when the user wants to benchmark on TerraLingua Simulation Environment, or asks about evaluating this task. Reports artifact novelty score.

ai-agentspythongit
0
3
Tesseract EvalA

Evaluates the robustness of Android malware classifiers against spatio-temporal experimental bias. It measures how model performance degrades when trained on past application data and tested on future data, while accounting for realistic malware-to-goodware class distributions. Use when the user wants to benchmark on Android malware dataset (2014-2016), or asks about evaluating this task. Reports F1-Score.

researchpythongo
0
3
Test AccuracyA

Evaluates the convergence speed and final test performance of distributed synchronous versus asynchronous stochastic gradient descent algorithms. It probes whether backup workers in synchronous training can mitigate stragglers without degrading accuracy due to gradient staleness. Use when the user has predictions and gold and needs to compute test accuracy.

researchpythongo
0
3
Test Time Fairness EvalA

Evaluates whether a zero-shot prompting method (OOC) improves stratified invariance and counterfactual invariance in LLM text classification predictions across real-world and synthetic datasets, while measuring retention of predictive accuracy. Use when the user wants to benchmark on civilcomments (Toxic Comments), Bios (Occupation), Amazon Fashion Reviews, Discrimination (Synthetic), MIMIC-III/SBDH (Clinical), Semantic Leakage Tasks, or asks about evaluating this task. Reports SI-bias.

researchpythongo
0
3
Test Time Scaling Vlm EvalA

Evaluates the impact of test-time scaling (TTS) inference strategies on Vision-Language Models across multimodal reasoning and perception tasks. It measures how techniques like Chain-of-Thought, Best-of-N, Self-Consistency, and Self-Refinement improve or degrade performance on open-source versus closed-source models. Use when the user wants to benchmark on MathVista, MMMU, MMBench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Text Animator EvalA

Evaluates a text-to-video model's ability to accurately render and animate specific text within a video scene while maintaining visual quality and temporal consistency. It probes character-level text fidelity, resistance to text collapse during motion, and overall video generation quality. Use when the user wants to benchmark on LAION subset, or asks about evaluating this task. Reports Sen. Acc.

researchpythongo
0
3
Text Classification Comparison EvalA

Systematic comparison of generative (AR, MLM, Diffusion) and discriminative (encoder) transformer models on text classification tasks, focusing on sample efficiency, robustness to input noise, and output calibration/ordinality. Use when the user wants to benchmark on AG News, Emotion, SST2, SST5, Multiclass Sentiment Analysis, Twitter Financial News Sentiment, IMDb, Hate Speech Offensive, or asks about evaluating this task. Reports weighted-F1 score.

researchpythongo
0
3
Text Classification Energy EvalA

This benchmark evaluates the trade-off between model accuracy, inference energy consumption, and runtime across diverse text classification models and hardware configurations. It probes whether larger or more complex models consistently outperform smaller or traditional ones in accuracy while highlighting the energy costs of different architectures and deployment strategies. Use when the user wants to benchmark on Text classification test set, or asks about evaluating this task. Reports accur...

devopspythongo
0
3
Text Classification EvalA

Evaluates text classification performance across multiple sentiment, subjectivity, question classification, and topic categorization tasks. It probes the model's ability to capture contextual and syntactic features from sequential text using 2D matrix representations and spatial pooling. Use when the user wants to benchmark on MR, SST-1, SST-2, Subj, TREC, 20Newsgroups, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Text Clustering EvalA

Evaluates the ability of centroid-based clustering algorithms to group unlabeled text documents into semantically coherent clusters. It measures clustering accuracy, label alignment with ground truth, and how closely learned centroids match true cluster centers. Use when the user wants to benchmark on Bank77, CLINC, GoEmo, MASSIVE, StackExchange, or asks about evaluating this task. Reports ACC, NMI.

researchpythongo
0
3
Text Image Retrieval EvalA

Evaluates the model's ability to align facial images with their textual descriptions by retrieving the correct image given a text query, and vice versa. It measures how well the model learns cross-modal semantic correspondence for face-centric data. Use when the user wants to benchmark on CelebA-Caption, MM-CelebA, or asks about evaluating this task. Reports R@5, R@10.

researchpythongo
0
3
Text Perturbation Robustness EvalA

Evaluates the robustness of finetuned transformer models (BERT, GPT-2, T5) to various text perturbations (e.g., dropping nouns/verbs, character changes, adding text) across classification and generation tasks. It measures how much model performance degrades when inputs are syntactically or semantically altered. Use when the user wants to benchmark on GLUE, XSum, CommonGen, SQuAD, or asks about evaluating this task. Reports Accuracy, Robustness Score.

researchpythongo
0
3
Text Queried Audio Separation EvalA

This benchmark evaluates a model's ability to perform text-queried audio source separation, specifically its capacity to isolate target sound events from mixed audio based on natural language instructions. It probes both acoustic fidelity (spectral and signal-level accuracy) and semantic alignment (how well the separated audio matches the textual description). Use when the user wants to benchmark on AudioCaps, Clotho v2, FSD50K, 3 Sets, MUSIC, or asks about evaluating this task. Reports LSD.

researchpythongo
0
3
Text Rendering EvalA

Evaluates a model's ability to generate images with accurate, legible, and layout-controlled text based on text prompts or masked regions. It probes text coherence, character-level rendering fidelity, and alignment between generated text and background imagery. Use when the user wants to benchmark on MARIO-10M, DrawBenchText, or asks about evaluating this task. Reports OCR(F-measure).

researchpythongo
0
3
Text Sanitization Reconstruction EvalA

Evaluates the vulnerability of differential privacy-based text sanitization methods by measuring how accurately an attacker can reconstruct original sensitive or personally identifiable information (PII) tokens from their sanitized counterparts. It probes the effectiveness of Bayesian inference-based reconstruction attacks against state-of-the-art sanitization defenses. Use when the user wants to benchmark on SST-2, AGNEWS, QNLI, Yelp, or asks about evaluating this task. Reports ASR.

researchpythongo
0
3
Text Summarization EvalA

Evaluates the abstractive text summarization capability of large language models by measuring how well they condense news articles into coherent, factually faithful, and linguistically natural summaries compared to human-written references. Use when the user wants to benchmark on CNN/Daily Mail 3.0.0, XSum, or asks about evaluating this task. Reports BERT Score.

ai-agentspython
0
3
Text To Code Customization EvalA

Evaluates the ability of customized small language models to generate correct and domain-aligned Python code. It probes functional correctness on general programming tasks versus specialized library APIs (Scikit-learn, OpenCV) under different customization strategies like few-shot prompting, RAG, and LoRA fine-tuning. Use when the user wants to benchmark on HumanEval, BCSk, BCCV, or asks about evaluating this task. Reports Pass@1.

ai-agentspythonapi
0
3
Text To Motion Generation EvalA

Evaluates a unified framework's ability to generate realistic and semantically aligned 3D human motions from text descriptions, recognize actions from skeleton data, and retrieve matching text-motion pairs. It probes the model's semantic fidelity, distributional realism, and cross-modal alignment capabilities. Use when the user wants to benchmark on HumanML3D, KIT, NTU-60, NTU-120, or asks about evaluating this task. Reports R-Precision.

ai-agentspythongo
0
3
Text To Sql Annotation Error EvalA

Evaluates the reliability of text-to-SQL benchmarks by quantifying annotation error rates and measuring how these errors distort agent execution accuracy and leaderboard rankings. Use when the user wants to benchmark on BIRD, Spider 2.0-Snow, or asks about evaluating this task. Reports annotation error rate.

researchpythongo
0
3