All authors
qhjqhj00 avatar

Claude Skills by qhjqhj00

github.com/qhjqhj00
7,636 skillsA× 7,623B× 11C× 1D× 10 installs2,229 views
Cross Lingual Pronoun Prediction EvalA

Evaluates systems' ability to predict target-language pronoun class labels from source-language pronouns using lemmatized, POS-tagged translations and word alignments. It probes cross-lingual anaphora resolution and functional ambiguity handling in machine translation pipelines. Use when the user wants to benchmark on WMT 2016 Cross-lingual Pronoun Prediction Task, or asks about evaluating this task. Reports macro-averaged recall.

researchpythongo
0
3
Cross Linguistic Activation EvalA

Evaluates cross-linguistic disparities in LLMs by measuring activation gaps via Sparse Autoencoders and benchmark performance across high-resource and medium-to-low resource languages. It probes whether surface-level embedding alignment guarantees equitable model behavior and tests if activation-level fine-tuning can close performance gaps without degrading English capabilities. Use when the user wants to benchmark on ARC-Challenge, HellaSwag, MMLU, or asks about evaluating this task. Reports...

ai-agentspythongo
0
3
Crosscheckgpt EvalA

This evaluation probes the ability of multimodal foundation models to generate factual content without hallucination across text, image, and audio-visual modalities. It measures how well reference-free ranking methods correlate with human judgments or gold-standard references to rank model outputs by hallucination severity. Use when the user wants to benchmark on WikiBio, MHaluBench, AVHalluBench, or asks about evaluating this task. Reports System($ ho$).

researchpythongo
0
3
Crossdocked Sbdd EvalA

Evaluates a model's ability to generate novel, drug-like molecules with high binding affinity for unseen protein pockets in structure-based drug design. It probes the trade-offs between binding energy, molecular properties, and synthesis feasibility. Use when the user wants to benchmark on CrossDocked-100k, or asks about evaluating this task. Reports Vina Dock.

researchpythonperformance
0
3
Crossguard Multimodal Safety EvalA

Evaluates the robustness of multimodal LLMs against explicit and implicit jailbreak attacks while measuring their utility on benign queries. It probes whether a defense model can successfully refuse harmful image-text prompts without over-restricting safe inputs. Use when the user wants to benchmark on JailBreakV, VLGuard, FigStep, MM-SafetyBench, SIUO, MMBench, or asks about evaluating this task. Reports Attack Success Rate (ASR).

researchpythonsecurity
0
3
Crosslingual Mtf EvalA

Evaluates zero-shot crosslingual generalization of multilingual LLMs after multitask finetuning. Probes language-agnostic task understanding, robustness to prompt translation, and scaling behavior across NLU, generative, and code tasks. Use when the user wants to benchmark on XNLI, XCOPA, XStoryCloze, XWinograd, HumanEval, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Crosslingual Speech Text Retrieval EvalA

Evaluates cross-lingual speech-to-text retrieval and intent detection capabilities across multiple datasets, testing how well speech queries can retrieve relevant text documents or classify intents without intermediate ASR or translation pipelines. Use when the user wants to benchmark on Kallaama-Retrieval-Eval, Fleurs-Retrieval-Eval, Urban Bus, WolBanking77, or asks about evaluating this task. Reports nDCG@5.

researchpythongo
0
3
Crossmed EvalA

Evaluates compositional generalization in medical vision-language models across a structured Modality–Anatomy–Task (MAT) schema. It probes zero-shot cross-task transfer, generalization to novel MAT combinations, and robustness under low-data regimes using a unified visual question answering interface. Use when the user wants to benchmark on CrossMed, or asks about evaluating this task. Reports top-1 classification accuracy.

researchpythongo
0
3
Crossmoda EvalA

Evaluates unsupervised cross-modality domain adaptation for medical image segmentation (Vestibular Schwannoma and Cochlea) and tumour grading (Koos classification) from ceT1 to T2 MRI. Use when the user wants to benchmark on crossMoDA, or asks about evaluating this task. Reports DSC.

researchpython
0
3
Crossmodal 3600 EvalA

Evaluates multilingual image captioning models across 36 languages, probing their ability to generate stylistically coherent and culturally representative descriptions without relying on direct translation artifacts. It measures how well models generalize to low-resource and geographically diverse languages. Use when the user wants to benchmark on Crossmodal-3600, or asks about evaluating this task. Reports CIDEr.

researchpythonperformance
0
3
Crossner EvalA

Evaluates cross-domain named entity recognition by measuring how well models adapt from a source domain (CoNLL2003) to five specialized target domains. These domains feature unique, domain-specific entity types that test the model's ability to generalize beyond standard categories. Use when the user wants to benchmark on CrossNER, or asks about evaluating this task. Reports F1 score.

researchpythongo
0
3
Crossnews Ua EvalA

Evaluates cross-lingual semantic similarity between news article pairs across four dimensions (Who, What, Where, When) in Ukrainian, Polish, Russian, and English. It probes a model's ability to align event-level information across languages while ignoring publication dates. Use when the user wants to benchmark on CrossNews-UA, or asks about evaluating this task. Reports macro-averaged F1-score.

researchpythongo
0
3
Crosspoint Bench EvalA

Evaluates Vision-Language Models' ability to perform precise point-level geometric correspondence across multiple viewpoints. It probes fine-grained spatial grounding, visibility reasoning, cross-view correspondence judgment, and continuous 2D coordinate pointing. Use when the user wants to benchmark on CrossPoint-Bench, or asks about evaluating this task. Reports average accuracy.

researchpythongo
0
3
Crosssum Alignment EvalA

Evaluates the quality of automatically induced cross-lingual summary alignments in the CrossSum dataset by measuring human agreement on whether two summaries correspond to the same source article. Use when the user wants to benchmark on CrossSum, or asks about evaluating this task. Reports alignment_accuracy.

researchpythongit
0
3
Crossvoice S2st EvalA

Evaluates cross-lingual speech-to-speech translation (S2ST) systems on translation accuracy and prosody preservation. It measures how well a cascade-based S2ST pipeline preserves speaker identity and naturalness while translating speech across different language pairs. Use when the user wants to benchmark on CVSS-T, Indic-TTS, Fisher, MuST-C, VoxPopuli, or asks about evaluating this task. Reports BLEU.

researchpython
0
3
Crosswoz Dst EvalA

Evaluates the ability of generative dialogue state tracking models to accurately predict and maintain the complete set of user intents (domain-slot-value triples) across dialogue turns. It specifically probes cross-lingual and cross-ontology transfer capabilities by measuring how well models trained on one language or ontology generalize to another. Use when the user wants to benchmark on CrossWOZ-en, or asks about evaluating this task. Reports Joint Goal Accuracy.

researchpythongo
0
3
Crowd Pose Estimation EvalA

Evaluates the ability of pose estimation models to accurately predict 2D keypoints for humans and animals in crowded, occluded, and multi-instance scenarios. It probes robustness to detection ambiguity, overlapping instances, and the transferability of conditional pose inputs from bottom-up detectors to top-down refiners. Use when the user wants to benchmark on CrowdPose, OCHuman, COCO, Multi-Animal (SchoolingFish, Marmosets, Tri-Mouse), or asks about evaluating this task. Reports AP.

researchpythonperformance
0
3
Crowdflow EvalA

Evaluates dense optical flow estimation accuracy and long-term temporal consistency in complex crowd surveillance scenarios, specifically testing robustness to non-rigid, self-occluding motion and small object tracking. Use when the user wants to benchmark on CrowdFlow, or asks about evaluating this task. Reports EPE.

researchpythontesting
0
3
Crowdhuman EvalA

Evaluates object detectors' ability to identify humans in highly crowded and heavily occluded scenes. It covers three annotation levels (full body, visible body, head) and assesses cross-dataset generalization for pedestrian and head detection tasks. Use when the user wants to benchmark on CrowdHuman, or asks about evaluating this task. Reports mMR.

researchpythongo
0
3
Crowdsensing Id Dfl EvalA

Evaluates the capability of decentralized federated learning (DFL) models to detect malware and classify benign states in IoT crowdsensing environments. It probes robustness under varying node counts, peer-to-peer network topologies, and data heterogeneity (IID vs. non-IID Dirichlet splits). Use when the user wants to benchmark on Crowdsensing Intrusion Detection Dataset, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Crowdspeech EvalA

Evaluates algorithms for aggregating multiple noisy, crowdsourced transcriptions of the same audio recording into a single high-quality reference. It probes how well methods handle sequential textual noise, estimate worker reliability, and adapt across different audio quality domains. Use when the user wants to benchmark on CROWDSPEECH, VOXDIY, CROWDWSA2019, or asks about evaluating this task. Reports WER.

researchpythongo
0
3
Crown EvalA

Evaluates conversational passage ranking by measuring how effectively a model ranks relevant documents across multi-turn search queries, balancing term similarity with contextual coherence. Use when the user wants to benchmark on TREC CAsT 2019, or asks about evaluating this task. Reports nDCG.

researchpythonnode
0
3
Crows Pairs EvalA

Measures social biases in masked language models by comparing the likelihood assigned to stereotypical versus anti-stereotypical sentence pairs. It quantifies how strongly models favor historically disadvantaged groups' stereotypes across nine demographic categories. Use when the user wants to benchmark on CrowS-Pairs, or asks about evaluating this task. Reports bias metric.

researchpythongo
0
3
Crumb EvalA

Evaluates information retrieval models on complex, multi-aspect, and logically structured queries across eight diverse domains. It probes the model's ability to handle nuanced document alignments, set-based operations, and context-rich instructions beyond simple keyword matching. Use when the user wants to benchmark on CRUMB, or asks about evaluating this task. Reports nDCG@10.

researchpythongit
0
3
Crumqs EvalA

Evaluates RAG systems on synthetic unanswerable and multi-hop queries to measure their refusal behavior, hallucination rates, and susceptibility to reasoning shortcuts when contexts are disjointed or out-of-distribution. Use when the user wants to benchmark on CRUMQs, UAEval4RAG, MultiHop-RAG, or asks about evaluating this task. Reports cheatability ratio.

researchpythongo
0
3
Cruxeval EvalA

Evaluates a model's ability to reason about and execute short Python functions by predicting outputs given inputs (CRUXEval-I) and predicting inputs given outputs (CRUXEval-O). It probes fundamental code execution and understanding capabilities beyond simple code generation. Use when the user wants to benchmark on CRUXEval, or asks about evaluating this task. Reports pass@1.

researchpythongo
0
3
Cruxevalx EvalA

This benchmark evaluates large language models' ability to perform bidirectional code reasoning across 19 programming languages. It probes whether models can predict missing inputs given outputs, and predict missing outputs given inputs, testing their understanding of code semantics and execution flow. Use when the user wants to benchmark on CRUXEval-X, or asks about evaluating this task. Reports Pass@1.

developmentpythontesting
0
3
Crystal Structure Discovery EvalA

Evaluates an LLM's capacity to discover stable crystal structures by iteratively mutating and crossing over parent structures to minimize deformation energy. Use when the user wants to benchmark on MatBenchbandgap, or asks about evaluating this task. Reports deformation energy.

researchpythongo
0
3
Cs 4k EvalA

Evaluates LLMs on end-to-end computer science research workflows by testing their ability to answer scientific questions grounded in academic papers. It probes domain-specific reasoning, factual recall, and methodological understanding across eight research workflow categories. Use when the user wants to benchmark on CS-4k, or asks about evaluating this task. Reports model response score.

researchpythongo
0
3
Cs Dialogue Asr EvalA

This benchmark evaluates automatic speech recognition (ASR) systems on their ability to accurately transcribe spontaneous, full-length dialogues that alternate between Mandarin and English. It probes a model's robustness to language alternation, phonetic mismatches, and contextual dependencies in naturalistic code-switching scenarios. Use when the user wants to benchmark on CS-Dialogue, or asks about evaluating this task. Reports MER.

researchpythonperformance
0
3
Cs Kws EvalA

Evaluates a model's ability to detect and localize multiple spoken keywords within continuous, untrimmed audio streams, distinguishing target keywords from background speech and silence. Use when the user wants to benchmark on LibriTop-20, CMAK-7, or asks about evaluating this task. Reports mAP.

researchpythongo
0
3
Csaw M EvalA

Evaluates models on ordinal classification of mammographic masking potential (levels 1–8) and their clinical utility in predicting interval and large invasive cancers. It probes the model's ability to respect ordinal relationships in breast tissue obscuration and correlate these estimates with cancer outcomes. Use when the user wants to benchmark on CSAW-M, or asks about evaluating this task. Reports average mean absolute error (AMAE).

researchpythongo
0
3
Csegg EvalA

Evaluates continual learning capabilities in scene graph generation by measuring how models retain prior object-relationship knowledge while learning new tasks, handle long-tailed data distributions, and generalize to unseen objects and relationships across incremental learning scenarios. Use when the user wants to benchmark on CSEGG, or asks about evaluating this task. Reports Avg. R@20.

researchpythongo
0
3
Csi Bert2 EvalA

Evaluates a transformer-based framework for Channel State Information (CSI) time-series prediction and wireless sensing classification. It probes the model's ability to recover missing data, predict future CSI sequences, and classify human actions or environmental states from Wi-Fi signals. Use when the user wants to benchmark on WiGesture, WiFall, WiCount, CommPre, or asks about evaluating this task. Reports Accuracy.

researchpython
0
3
Csmbench EvalA

Evaluates large multimodal models' ability to perceive, interpret, and reason about scientific figures across four hierarchical physical scales (atomic, micro, meso, macro) in materials science. It probes both discriminative visual matching and open-ended scientific narrative generation. Use when the user wants to benchmark on CSMBench, or asks about evaluating this task. Reports accuracy.

ai-agentspythongo
0
3
Csr Bench EvalA

Evaluates LLM agents' ability to autonomously deploy computer science research repositories by executing multi-stage workflows including setup, downloading dependencies, training models, running evaluations, and performing inference. It probes instruction comprehension, command generation, and iterative error correction in complex software environments. Use when the user wants to benchmark on CSR-Bench, or asks about evaluating this task. Reports success rate.

ai-agentspythonshell
0
3
Csr L EvalA

Evaluates the robustness of information retrieval models when processing code-switched queries (English mixed with Mandarin Chinese or Japanese). It probes whether multilingual retrievers and rerankers suffer embedding divergence or performance degradation compared to monolingual English queries across argument, code, biomedical, and instruction-following retrieval tasks. Use when the user wants to benchmark on Touché 2020, HumanEval, TRECCOVID, FollowIR, or asks about evaluating this task. R...

researchpythonperformance
0
3
Csrec Sequential Rec EvalA

This evaluation protocol assesses the ranking performance and robustness of sequential recommendation models trained with confident soft labels. It measures whether predicted item sequences align with actual user interactions and verifies if recommendations correspond to genuinely positive user preferences using explicit rating thresholds. Use when the user wants to benchmark on Last.FM, Yelp, Amazon Electronics, Amazon Movies and TV, or asks about evaluating this task. Reports Recall@n, NDCG@n.

researchpythongo
0
3
Css10 Tts EvalA

Evaluates the quality of synthesized speech from single-speaker TTS models trained on the CSS10 datasets across 10 languages. It probes how well models can reproduce natural-sounding audio and accurate pronunciation for held-out test sentences. Use when the user wants to benchmark on CSS10, or asks about evaluating this task. Reports MOS.

researchpythongo
0
3
Csvqa EvalA

This benchmark evaluates the scientific reasoning and domain-grounded visual question answering capabilities of Vision-Language Models (VLMs) in Chinese. It probes the ability to integrate multimodal STEM evidence across physics, chemistry, biology, and mathematics with domain knowledge to solve both multiple-choice and open-ended questions. Use when the user wants to benchmark on CSVQA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Csymr EvalA

Evaluates compositional symbolic music reasoning by requiring models to chain atomic analyses across multiple musical dimensions (e.g., rhythm, harmony, key, structure) to answer multiple-choice questions derived from expert forums and professional exams. Use when the user wants to benchmark on CSyMR-Bench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Ct Brain Segmentation EvalA

Evaluates the ability of segmentation models to accurately delineate brain tissue, cerebrospinal fluid (CSF), and subdural hematomas in post-operative CT scans of hydrocephalic infants. It probes robustness to intensity overlap, anatomical distortion, and limited training data in a real-world clinical setting. Use when the user wants to benchmark on CURE Children's Hospital of Uganda CT Brain Dataset, or asks about evaluating this task. Reports dice-overlap coefficient.

researchpython
0
3
Ct Rate Zero Shot EvalA

This benchmark evaluates the zero-shot multi-abnormality detection capability of a visual-language foundation model on 3D chest CT volumes. It probes the model's ability to generalize to unseen data distributions and classify multiple pathologies simultaneously without task-specific supervised training. Use when the user wants to benchmark on CT-RATE, RAD-ChestCT, or asks about evaluating this task. Reports AUROC.

researchpythongit
0
3
Cti Plausibility EvalA

Evaluates whether neural machine translation models correctly rely on contextual cues when generating target tokens. It compares model-extracted cue-target pairs against human-annotated discourse-level expectations to measure the plausibility of context reliance. Use when the user wants to benchmark on SCAT+, or asks about evaluating this task. Reports Macro F1.

researchpythongo
0
3
Ctibench EvalA

Evaluates large language models on five cyber threat intelligence (CTI) tasks, including knowledge recall, vulnerability mapping, CVSS scoring, attack technique extraction, and threat attribution. It probes factual accuracy, logical reasoning, and contextual understanding within a domain-specific cybersecurity context. Use when the user wants to benchmark on CTIBench, or asks about evaluating this task. Reports accuracy.

securitypythongo
0
3
Ctr Prediction Dhan EvalA

Evaluates a model's ability to predict the next item a user will click or review based on their historical interaction sequence. It probes hierarchical user interest modeling across multiple product dimensions and abstraction levels in recommendation systems. Use when the user wants to benchmark on Amazon Review (Six-Category, Kindle Shop, Electronics), or asks about evaluating this task. Reports AUC.

researchpythongo
0
3
Ctr Prediction EvalA

Evaluates the ability of deep learning models to predict click-through rates (CTR) from sparse, high-dimensional categorical features in advertising and recommendation scenarios. It probes how well models capture multi-scale semantic interactions and handle large-scale, imbalanced binary classification tasks typical of real-world ad systems. Use when the user wants to benchmark on Avazu, MovieLens, Weibo, or asks about evaluating this task. Reports AUC.

researchpythongo
0
3
Ctr Welfare EvalA

Evaluates click-through rate (CTR) prediction models for their ability to maximize economic welfare in simulated and real-world ad auction settings, while also measuring standard classification performance. Use when the user wants to benchmark on Synthetic Dataset, Criteo Display Advertising Challenge, or asks about evaluating this task. Reports test-time welfare.

researchpythonperformance
0
3
Ctta Text Understanding EvalA

Evaluates continual test-time adaptation (CTTA) for text understanding across sequential, unobserved domains. It probes a model's ability to adapt to shifting domains using only unlabeled test data while mitigating error accumulation and maintaining cross-domain generalization. Use when the user wants to benchmark on CTTA-Text-Understanding-Benchmark, or asks about evaluating this task. Reports exact match (EM), F1 score.

researchpythongo
0
3
Cuad EvalA

Evaluates a model's ability to identify and extract relevant text spans from legal contracts corresponding to specific clause categories. It probes domain-specific information extraction and needle-in-a-haystack detection under severe class imbalance. Use when the user wants to benchmark on CUAD, or asks about evaluating this task. Reports Precision@80% Recall.

researchpythongo
0
3