All authors
qhjqhj00 avatar

Claude Skills by qhjqhj00

github.com/qhjqhj00
7,636 skillsA× 7,623B× 11C× 1D× 10 installs2,049 views
Carla Counterfactual EvalA

Evaluates the quality and feasibility of model-agnostic counterfactual explanations generated for tabular data. It probes whether generated counterfactuals successfully flip classifier predictions while maintaining sparsity, proximity, actionability (immutable constraints), and plausibility across multiple binary classification tasks. Use when the user wants to benchmark on adult, COMPAS, Give Me Some Credit, HELOC, Irish, Saheart, Titanic, Wine, or asks about evaluating this task. Reports Su...

researchpythonrust
0
3
Carla Dos EvalA

Evaluates end-to-end autonomous driving agents on route completion, safety/infractions, and occlusion-aware perception in simulated urban environments. Use when the user wants to benchmark on CARLA Town 05 Long, DOS Benchmark, or asks about evaluating this task. Reports Driving Score (DS).

researchpythonperformance
0
3
Carla Leaderboard EvalA

Evaluates autonomous driving agents on their ability to navigate complex urban environments while balancing route completion, safety, and rule compliance. It probes how well models handle dynamic traffic interactions, edge-case scenarios, and physical constraints in a closed-loop simulation. Use when the user wants to benchmark on CARLA Leaderboard 2.0 scenarios, CARLA 42 Routes, Town05, or asks about evaluating this task. Reports Driving Score (DS).

researchpythonperformance
0
3
Carla Longest6 Town05long EvalA

Evaluates an autonomous driving model's ability to navigate complex urban and highway environments under varying weather and lighting conditions, focusing on route completion, safety (infraction avoidance), and overall driving performance. Use when the user wants to benchmark on CARLA Longest6, CARLA Town05 Long, or asks about evaluating this task. Reports Driving Score (DS).

researchpythongo
0
3
Carla No Crash EvalA

Evaluates autonomous driving agents' robustness and generalization in a simulated urban environment, specifically testing performance on familiar and unseen town layouts. Use when the user wants to benchmark on CARLA NoCrash benchmark, or asks about evaluating this task. Reports Driving Score.

ai-agentspythongo
0
3
Carla Real Traffic Scenarios EvalA

Evaluates the ability of reinforcement learning policies to execute tactical driving maneuvers (e.g., lane changes, roundabout navigation) in a simulated environment mapped from real-world traffic data. It probes how observation modalities and reward structures impact policy generalization and success rates across diverse driving scenarios. Use when the user wants to benchmark on NGSIM, openDD, or asks about evaluating this task. Reports success rate.

researchpythongo
0
3
Carla Urban Driving EvalA

Evaluates an autonomous driving agent's ability to navigate urban environments, avoid dynamic and static obstacles, and handle road blockages under varying weather conditions and unseen towns. Use when the user wants to benchmark on CARLA urban driving benchmark, or asks about evaluating this task. Reports Success Rate.

researchpythongo
0
3
Carlaocc EvalA

Evaluates 3D occupancy prediction models on their ability to reconstruct semantic and instance-level voxel grids from camera inputs. It probes geometric completeness, occlusion reasoning, and instance discrimination in complex autonomous driving scenes. Use when the user wants to benchmark on CarlaOcc, or asks about evaluating this task. Reports mIoU.

researchpython
0
3
Carmem EvalA

Evaluates an LLM's ability to extract, maintain, and retrieve long-term user preferences in an in-car voice assistant context using a predefined category-bound schema. It probes structured information extraction, state maintenance via function calling, and semantic retrieval accuracy under privacy-preserving constraints. Use when the user wants to benchmark on CarMem, or asks about evaluating this task. Reports F1-score.

ai-agentspythongo
0
3
Carpatch EvalA

Evaluates the reconstruction quality of Neural Radiance Field (NeRF) models on synthetic vehicle components. It probes both 2D appearance fidelity and 3D geometric accuracy, specifically testing robustness to reflective surfaces, transparent materials, and varying numbers of training viewpoints. Use when the user wants to benchmark on CarPatch, or asks about evaluating this task. Reports PSNR.

researchpythonangular
0
3
Carpe EvalA

Evaluates the visual classification and vision-language understanding capabilities of large vision-language models (LVLMs) under context-aware ensemble prompting. It probes fine-grained visual recognition, scientific question answering, text-rich VQA, hallucination detection, and multimodal reasoning across diverse benchmarks. Use when the user wants to benchmark on ImageNet, Caltech101, Flower102, Food101, ScienceQA (image subset), TextVQA, POPE, MME, MMBench, CV-Bench, MMVP, or asks about e...

researchpythongo
0
3
Carte Tabular EvalA

Evaluates a graph-based neural architecture for tabular learning on single and multiple tables, testing its ability to handle mixed numerical and categorical features without requiring schema or entity matching. Use when the user wants to benchmark on TabLLM datasets, Entity Matching datasets, or asks about evaluating this task. Reports performance.

researchpythongo
0
3
Casefacts EvalA

Evaluates LLMs and retrieval models on verifying colloquial legal claims against U.S. Supreme Court precedents, measuring both verdict prediction accuracy and the quality of retrieved supporting case evidence. Use when the user wants to benchmark on CaseFacts, or asks about evaluating this task. Reports Verdict Score.

researchpythonrust
0
3
Caselawqa EvalA

This benchmark probes a model's ability to perform fine-grained legal text classification and annotation. It tests whether models can accurately extract specific legal features, such as precedent alteration, issue areas, or ideological valence, from lengthy court opinions using multiple-choice prompts. Use when the user wants to benchmark on CaselawQA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Caser Sequential Rec EvalA

Evaluates a model's ability to capture sequential user behavior patterns for personalized top-N item recommendation. It probes the model's capacity to model temporal dependencies, skip behaviors, and union-level sequential patterns from historical interactions to predict future items. Use when the user wants to benchmark on MovieLens, Gowalla, Foursquare, Tmall, or asks about evaluating this task. Reports MAP.

researchpythongo
0
3
Casesumm EvalA

Evaluates the ability of language models to generate accurate, concise, and legally faithful summaries of long U.S. Supreme Court opinions. It probes the alignment between automatic NLP metrics and expert human judgment in a high-stakes, domain-specific summarization task. Use when the user wants to benchmark on CaseSumm, or asks about evaluating this task. Reports ROUGE.

researchpython
0
3
Casif EvalA

This evaluation protocol assesses a model's ability to perform session-based next-item recommendation by predicting the subsequent item a user will click based on their recent interaction history. It probes the model's capacity to capture both short-term sequential dependencies and long-term contextual patterns within a session without relying on explicit user profiles. Use when the user wants to benchmark on Yoochoose1/64, Yoochoose1/4, Diginetica, or asks about evaluating this task. Reports...

researchpythongo
0
3
Casp 2016 Dti EvalA

Evaluates deep learning models for drug-target interaction prediction across four tasks: scoring (affinity correlation), ranking (pose affinity ordering), docking (native pose retrieval from decoys), and screening (true binder identification among random molecules). Probes both regression accuracy and virtual screening generalization. Use when the user wants to benchmark on CASF-2016, CSAR NRC-HiQ, or asks about evaluating this task. Reports Docking power (Top-1 Success Rate), Screening power...

researchpythongit
0
3
Castle EvalA

This benchmark evaluates large language models' ability to detect and mitigate educational safety risks in a student-tailored manner. It probes how well models adapt their responses to individual student profiles across 15 risk domains and 14 psychological and educational attributes. Use when the user wants to benchmark on CASTLE, or asks about evaluating this task. Reports safety score.

researchpython
0
3
Cat Benchmark EvalA

Evaluates whether augmenting code search models with neural machine translation-generated AST representations improves retrieval accuracy over raw code tokens. Probes the capability of NMT to translate natural language queries into compact abstract syntax tree non-terminal sequences and measures the downstream impact on code retrieval performance. Use when the user wants to benchmark on TLC, CSN, Funcom, PCSD, or asks about evaluating this task. Reports MRR.

ai-agentspythonperformance
0
3
Cataract Lmm EvalA

Evaluates a model's ability to recognize and classify temporal surgical phases in cataract surgery videos, specifically testing classification accuracy and robustness to domain shift across different clinical centers. Use when the user wants to benchmark on Cataract-LMM Phase Recognition Subset, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
CatmetricA

Compute the CatMetric metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute CatMetric, or asks how to score with CatMetric.

documentationpython
0
3
Catwalk Default MetricsA

Evaluates language models across multiple dataset categories using standardized metrics attached to datasets rather than model implementations. Probes classification/entailment accuracy, question-answering fidelity, and language modeling fluency/probability calibration. Use when the user has predictions and gold and needs to compute Accuracy & relative improvement over random baseline, SQuAD metric.

researchpythongo
0
3
Causal2needles EvalA

Evaluates Video-Language Models' ability to perform joint retrieval and causal reasoning over two causally separated video clips connected by a 'bridge entity'. It also probes causal world modeling by asking models to identify cause-effect relationships in human behaviors within long videos. Use when the user wants to benchmark on Causal2Needles, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Causal3d EvalA

This benchmark evaluates a model's ability to discover and reason about causal structures from visual and tabular data. It probes whether models can infer correct causal graphs from observational data, learn disentangled representations from images, and perform valid causal interventions with limited visual samples. Use when the user wants to benchmark on Causal3D, or asks about evaluating this task. Reports correctness of inferred causal structures.

researchpythongo
0
3
Causalverse EvalA

Probes the ability of causal representation learning (CRL) models to recover ground-truth latent variables from high-fidelity visual simulations. It evaluates both component-wise and block-wise identifiability under realistic conditions where theoretical assumptions may be violated. Use when the user wants to benchmark on CausalVerse, or asks about evaluating this task. Reports Mean Correlation Coefficient (MCC).

researchpython
0
3
Cavbench EvalA

Evaluates the computational performance and resource efficiency of edge computing platforms for connected and autonomous vehicle workloads. It probes how well hardware handles real-time vision, deep learning, and diagnostic tasks under varying resource constraints. Use when the user wants to benchmark on CAVBench, or asks about evaluating this task. Reports Matching Factor (MF).

researchpythonperformance
0
3
Cbd Guidance Older Adults EvalA

Evaluates whether retrieval-augmented LLMs can generate safe, clinically grounded cannabidiol (CBD) dosage and titration recommendations tailored to older adults with varying cognitive and clinical risk profiles. Use when the user wants to benchmark on Parametric CBD Scenario Set, or asks about evaluating this task. Reports llm_judge_rubric_score.

ai-agentspythonperformance
0
3
Cbm Concept Accuracy EvalA

Evaluates whether Concept Bottleneck Models learn semantically meaningful concept representations from input images under varying annotation granularity and concept correlation structures. Measures how well the model predicts intermediate concepts and downstream tasks compared to standard neural networks. Use when the user wants to benchmark on Playing cards, CheXpert, or asks about evaluating this task. Reports concept accuracy.

researchpythongo
0
3
Cbr Encrypted Traffic EvalA

Evaluates an ANN-based adaptive classifier's ability to classify encrypted network traffic into known categories such as malware families, operating systems, browsers, and applications. It specifically probes the model's capacity to dynamically adapt to new or out-of-distribution classes without retraining, while measuring any performance degradation on existing classes compared to traditional baselines. Use when the user wants to benchmark on BOA, MTA, or asks about evaluating this task. Rep...

researchpythongo
0
3
Cbt Bench EvalA

Evaluates LLMs on cognitive behavioral therapy (CBT) assistance across basic knowledge recall and cognitive model understanding. The benchmark probes multiple-choice knowledge acquisition and multi-label classification of cognitive distortions and core beliefs to measure therapeutic applicability and fine-grained clinical reasoning. Use when the user wants to benchmark on CBT-QA, CBT-CD, CBT-PC, CBT-FC, or asks about evaluating this task. Reports Accuracy, F1.

researchpythongo
0
3
Cc Cliff EvalA

Evaluates whether large language models can effectively learn and utilize spatial coordinate information versus categorical/compositional data for property prediction. It quantifies the systematic performance degradation (the 'Coordinate-Category Cliff') when tasks require geometric reasoning rather than simple type matching, and tests whether scaling model size or dataset volume mitigates this deficit. Use when the user wants to benchmark on Synthetic coordinate-category datasets, Materials ...

researchpythongo
0
3
Ccc EvalA

Evaluates continual semi-supervised learning on crowd counting by measuring how well a model adapts to evolving unlabeled data streams across sequential sessions. Use when the user wants to benchmark on Continual Crowd Counting (CCC), or asks about evaluating this task. Reports Mean Absolute Error (MAE).

researchpythongo
0
3
Ccf Aatc 2025 Speech Restoration EvalA

Evaluates speech restoration models on realistic, multi-stage degradations including acoustic noise/reverberation, codec compression artifacts, and secondary processing artifacts from upstream enhancement models. Probes the trade-off between signal fidelity/intelligibility and perceptual quality while measuring computational efficiency. Use when the user wants to benchmark on CCF AATC 2025 Test Set, or asks about evaluating this task. Reports WAcc.

researchpythongit
0
3
Ccfqa EvalA

This benchmark evaluates the factual accuracy and consistency of multimodal large language models (MLLMs) when answering questions in text or speech modalities across eight languages. It specifically probes cross-lingual transfer capabilities and cross-modal alignment by measuring how well models maintain factual correctness when switching between languages or between text and audio inputs. Use when the user wants to benchmark on CCFQA, or asks about evaluating this task. Reports F1 score.

researchpythongo
0
3
Cci30 Hq EvalA

Evaluates the effectiveness of a high-quality Chinese pre-training corpus by training a 0.5B language model from scratch and measuring its zero-shot generalization across standard English and Chinese knowledge benchmarks. The protocol also includes a comparative analysis of data quality classifiers using macro F1-score on a held-out test set. Use when the user wants to benchmark on Standard Benchmarks (ARC-C, ARC-E, HellaSwag, Winograd, MMLU, OpenbookQA, PIQA, SIQA, CEval, CMMLU), or asks abo...

researchpythongo
0
3
Ccks2017 Cner EvalA

Evaluates a model's ability to identify five types of clinical named entities (diseases, symptoms, exams, treatments, body parts) in Chinese medical texts. It specifically probes character-level sequence labeling performance and the impact of integrating external dictionary features. Use when the user wants to benchmark on CCKS-2017 Task 2, or asks about evaluating this task. Reports F1-score.

researchpythongo
0
3
Ccl Slu EvalA

Evaluates intent classification robustness under noisy ASR conditions by measuring how well a model aligns noisy transcripts with clean references and preserves semantic consistency. Use when the user wants to benchmark on SLURP, Timers, FSC, SNIPS, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Ccmatrix EvalA

Evaluates the quality of mined parallel sentence pairs by training machine translation systems on them and measuring translation performance. It probes the effectiveness of global, margin-based bitext mining in a multilingual embedding space. The benchmark measures how well the mined data generalizes across different language families and scripts. Use when the user wants to benchmark on CCMatrix, or asks about evaluating this task. Reports BLEU.

researchpythongo
0
3
Ccnet Dataset EvalA

Evaluates the quality of a large-scale monolingual web corpus by measuring downstream performance on standard linguistic analogy tasks and a cross-lingual natural language inference benchmark. Use when the user wants to benchmark on CCNet, XNLI, or asks about evaluating this task. Reports XNLI.

researchpythongo
0
3
Cctvbench EvalA

Probes multimodal LLMs' ability to answer binary questions about traffic videos while maintaining logical consistency across counterfactual video-question pairs. It specifically diagnoses failure modes like positive omission, negative hallucination, and mutual-exclusivity violations by enforcing a strict quadruple-level decision rule. Use when the user wants to benchmark on CCTVBench, or asks about evaluating this task. Reports QuadAcc.

researchpythongo
0
3
Cd Fer Benchmark EvalA

Evaluates cross-domain facial expression recognition (CD-FER) models by measuring how well they transfer learned features from a labeled source dataset to an unlabeled target dataset. It probes the model's ability to learn domain-invariant representations and adapt to distribution shifts across different facial expression datasets. Use when the user wants to benchmark on RAF-DB, AFE, CK+, JAFFE, SFEW2.0, FER2013, ExpW, or asks about evaluating this task. Reports accuracy.

researchpythonnode
0
3
Cde Regression EvalA

Evaluates tabular foundation models and traditional baselines on conditional density estimation for regression tasks. It probes density accuracy, probabilistic calibration, prediction sharpness, and computational efficiency across varying training sample sizes and diverse real-world domains. Use when the user wants to benchmark on OpenML & SDSS DR18 regression datasets, or asks about evaluating this task. Reports CDE loss.

testingpythongo
0
3
Cdfsl V EvalA

Cross-domain few-shot video action recognition. It probes a model's ability to adapt to new video domains using a large source dataset and unlabeled target videos, with evaluation restricted to a 5-way 5-shot setting where only five target classes are tested with five labeled support examples each. Use when the user wants to benchmark on Kinetics-100, Kinetics-400, UCF101, HMDB51, Something-SomethingV2, Diving48, RareAct, or asks about evaluating this task. Reports 5-way 5-shot accuracy.

researchpythonreact
0
3
Cdh Bench EvalA

This benchmark probes a vision-language model's ability to maintain visual fidelity when explicit visual evidence conflicts with strong commonsense priors. It specifically measures whether models override entrenched prior-driven expectations with counterfactual image-grounded claims, isolating hallucination from generic perception errors. Use when the user wants to benchmark on CDH-Bench, or asks about evaluating this task. Reports CFAD.

researchpythongo
0
3
Cdi Dti Dti Prediction EvalA

Evaluates a multi-modal deep learning framework's ability to predict drug-target binding interactions across standard, cross-domain, and cold-start scenarios. It probes the model's capacity to integrate textual, structural, and functional biological features for robust binary classification under distribution shifts and unseen entities. Use when the user wants to benchmark on BindingDB, Davis, or asks about evaluating this task. Reports AUROC.

researchpythongo
0
3
Cds Search EvalA

Evaluates clinical information retrieval systems on their ability to rank relevant medical documents for clinical decision support queries. It probes the effectiveness of query and document processing techniques such as negation detection, concept extraction, and pseudorelevance feedback in a standardized biomedical search setting. Use when the user wants to benchmark on TREC CDS'16, or asks about evaluating this task. Reports infNDCG.

researchpythongo
0
3
Ce Recommender Benchmark EvalA

Evaluates the quality, sparsity, and computational efficiency of counterfactual explanation methods for recommender systems. It probes how effectively explanations can alter recommendation rankings, how interpretable the generated explanations are, and the cost of generating them across different input formats and perturbation scopes. Use when the user wants to benchmark on Recommender system interaction datasets, or asks about evaluating this task. Reports POS-P@K.

researchpythongo
0
3
Ceb Fairness EvalA

This benchmark evaluates how large language models exhibit social bias across different tasks, bias types, and social groups. It systematically measures bias through direct classification tasks and indirect text generation tasks, using standardized metrics to enable cross-dataset fairness comparisons. Use when the user wants to benchmark on CEB, or asks about evaluating this task. Reports Micro-F1.

researchpythongo
0
3
Cebench EvalA

Evaluates vision-language-action (VLA) models on cross-embodiment robotic manipulation tasks, including single-arm, bimanual, and mobile manipulation. It probes spatial reasoning, visual generalization under domain randomization, and the ability to unify navigation and manipulation in a single policy. Use when the user wants to benchmark on CEBench, or asks about evaluating this task. Reports success_rate.

researchpythongo
0
3