All authors
qhjqhj00 avatar

Claude Skills by qhjqhj00

github.com/qhjqhj00
7,636 skillsA× 7,623B× 11C× 1D× 10 installs2,206 views
Papermind EvalA

Evaluates multimodal LLMs' ability to perform integrated agentic reasoning and critical assessment over scientific papers. It probes capabilities in multimodal grounding, experimental interpretation, cross-source evidence synthesis via tool use, and critical evaluation of research claims. Use when the user wants to benchmark on PaperMind, or asks about evaluating this task. Reports F1 score.

researchpythongo
0
3
Paperqa EvalA

Evaluates the accuracy and interaction efficiency of LLM agents performing multi-turn tool-use for scientific paper question-answering. It probes the model's ability to plan, execute tool calls, and extract answers from complex academic documents without excessive interaction or getting stuck in loops. Use when the user wants to benchmark on AirQA-Real, SciDQA, or asks about evaluating this task. Reports Avg., I-Avg.

researchpythongo
0
3
Papi Personality Alignment EvalA

Evaluates how well a language model's generated responses align with a specific individual's personality traits across the Big Five and Dark Triad dimensions. It measures the distance between the model's predicted personality profile and the target profile, with lower scores indicating better alignment. Use when the user wants to benchmark on PAPI, or asks about evaluating this task. Reports Aligned Score.

researchpythongit
0
3
Papillon EvalA

Evaluates a privacy-preserving LLM delegation pipeline that sanitizes user queries before sending them to a remote API model. It measures the trade-off between maintaining response quality and minimizing personally identifiable information (PII) leakage in the sanitized prompts. Use when the user wants to benchmark on PUPA-TNB, or asks about evaluating this task. Reports QUAL.

researchpythongo
0
3
Paradnn Hardware BenchA

Evaluates how different deep learning model architectures and hyperparameters affect hardware performance across TPU, GPU, and CPU platforms. It probes the interaction between model attributes (size, type, batch size) and hardware bottlenecks like memory bandwidth, compute utilization, and data infeed overhead. Use when the user wants to benchmark on ParaDnn, or asks about evaluating this task. Reports performance.

researchpythonnode
0
3
Paras2s EvalA

Evaluates how well speech-to-speech models adapt to and reflect paralinguistic styles (age, emotion, gender, sarcasm) in spoken responses. It measures both content appropriateness and stylistic alignment against ground-truth or human-annotated references. Use when the user wants to benchmark on ParaS2SBench, IEMOCAP, MELD, or asks about evaluating this task. Reports ParaS2SBench score.

researchpythongo
0
3
Pareto Interpretation EvalA

This evaluation probes a model's ability to synthesize interpretable surrogate models (decision diagrams) that explicitly trade off prediction fidelity against structural simplicity. It measures how well a method can navigate the Pareto front between accuracy and explainability without collapsing them into a single weighted objective. Use when the user wants to benchmark on Airplane Perception Module (AP), Bank Loan Predictor (BL), Theorem Prover Solvability Predictor (TP), or asks about eval...

researchpythongo
0
3
Pariksha EvalA

Evaluates multilingual and multi-cultural LLM performance across 10 Indic languages using culturally nuanced prompts. It measures model quality via pairwise comparisons (Elo ratings) and direct assessment scores, while also analyzing human-LLM evaluator agreement and various biases (position, verbosity, self-bias). Use when the user wants to benchmark on PARIKSHA, or asks about evaluating this task. Reports Elo rating, Direct Assessment score.

researchpythongo
0
3
Park Cleaning BenchmarkA

Evaluates autonomous cleaning robots' ability to navigate public park pathways, perceive and collect diverse litter types, avoid obstacles, and operate within strict physical and safety constraints. Use when the user wants to benchmark on Park Cleaning Benchmark, or asks about evaluating this task. Reports collected_items_weight_or_count.

researchpython
0
3
Parkseg12k EvalA

This benchmark evaluates the ability of deep learning models to perform binary semantic segmentation of parking lots from satellite imagery. It specifically probes the model's capacity to generalize across different geographic locations and to leverage near-infrared (NIR) spectral data for improved contrast against vegetation and built environments. Use when the user wants to benchmark on ParkSeg12k, or asks about evaluating this task. Reports mIoU.

researchpythongo
0
3
Parma Performance EvalA

Measures the computational, network, and database performance overhead of running containerized workloads inside an AMD SEV-SNP enclave with Parma's attested execution policies compared to a baseline outside the enclave. Use when the user wants to benchmark on nginx (wrk2), redis (redis-benchmark), SPEC2017 intspeed, NVIDIA Triton Inference Server (perf-analyzer), or asks about evaluating this task. Reports performance_overhead.

devopspythondatabase
0
3
Parrot EvalA

This benchmark evaluates an LLM's ability to accurately translate SQL queries across different database systems and dialects. It probes the model's capacity to handle system-specific syntax, semantic equivalence, and edge-case safeguards without relying on superficial string matching. Use when the user wants to benchmark on PARROT, or asks about evaluating this task. Reports Acc_EX.

researchpythongo
0
3
Parrot Multilingual EvalA

Evaluates the multilingual visual-language understanding capabilities of multimodal large language models (MLLMs) across six languages (English, Chinese, Portuguese, Arabic, Turkish, Russian). It probes how well models align visual features with non-English textual instructions and handle cross-lingual multimodal tasks without relying on naive translation. Use when the user wants to benchmark on MMMB, MMBench, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Parsinlu EvalA

Evaluates Persian language understanding across six distinct NLU tasks, including reading comprehension, textual entailment, sentiment analysis, and machine translation. It measures how well pre-trained monolingual and multilingual models perform on native-speaker annotated Persian data compared to human baselines. Use when the user wants to benchmark on ParsiNLU, or asks about evaluating this task. Reports F1, Accuracy.

researchpythongo
0
3
Partinstruct EvalA

This benchmark probes a robot policy's ability to follow fine-grained, part-level natural language instructions for long-horizon manipulation. It specifically tests zero-shot task decomposition, 3D part grounding, and multi-step planning under varying object, part, and task generalization conditions. Use when the user wants to benchmark on PartInstruct, or asks about evaluating this task. Reports success.

researchpythongo
0
3
Pas Dataset EvalA

Evaluates the transferability and pretraining quality of vision models trained on synthetic domain-specific datasets compared to manually curated and general-domain datasets. It probes the model's ability to generalize to fine-grained classification and object detection tasks within specific domains like birds and food. Use when the user wants to benchmark on CUB-200-2011, NABirds, iNatbirds, Food-101, FoodX-251, Food-2K, or asks about evaluating this task. Reports Top-1 k-NN accuracy.

researchpythonperformance
0
3
Pascal Voc Detection EvalA

Evaluates an object detection model's ability to localize and classify objects within images. It measures how well the system predicts bounding boxes and assigns correct class labels across multiple object categories. Use when the user wants to benchmark on PASCAL VOC 2007, or asks about evaluating this task. Reports mAP.

researchpythongo
0
3
Pastis Hd EvalA

Tests agricultural land cover mapping and crop-type classification by evaluating models on high-resolution satellite imagery combined with optical and radar time series. Use when the user wants to benchmark on PASTIS-HD, or asks about evaluating this task. Reports macro-averaged F1-score.

researchpythongo
0
3
Pat Questions EvalA

Evaluates large language models' ability to answer present-anchored temporal questions that require up-to-date world knowledge and multi-hop reasoning, such as identifying the current holder of a position or the previous president. It specifically probes performance degradation due to knowledge obsolescence and complex temporal relations. Use when the user wants to benchmark on PAT-Questions, or asks about evaluating this task. Reports exact-match accuracy (EM).

researchpythongo
0
3
Patch Selectivity EvalA

Evaluates a model's ability to ignore out-of-context patches (patch selectivity) and maintain classification accuracy under simulated occlusion and spatial permutation attacks. Use when the user wants to benchmark on ImageNet-1K val, SMD, NVD, ROD, or asks about evaluating this task. Reports Top-1 accuracy.

researchpythongo
0
3
Patchgastricadc22 EvalA

This evaluation probes a model's ability to generate clinically accurate diagnostic captions from histopathological image patches. It specifically tests the model's capacity to capture subtype-specific terminology and overall caption fluency using standard and custom n-gram overlap metrics. Use when the user wants to benchmark on PatchGastricADC22, or asks about evaluating this task. Reports BLEU@4.

researchpythongit
0
3
Patent Ce EvalA

This benchmark evaluates the quality of generated patent claims against expert-annotated reference claims across five dimensions: feature completeness, conceptual clarity, terminology consistency, logical linkage, and overall quality. It probes a model's ability to capture patent-specific linguistic precision, legal formality, and structural requirements rather than just surface-level text overlap. Use when the user wants to benchmark on Patent-CE, or asks about evaluating this task. Reports ...

researchpythongo
0
3
Patenteb EvalA

Evaluates patent text embedding models across 15 diverse tasks including symmetric/asymmetric retrieval, classification, paraphrase detection, and clustering. It specifically probes domain-specific challenges like cross-domain retrieval, fragment-to-document matching, and temporal citation dynamics. Use when the user wants to benchmark on PatenTEB, or asks about evaluating this task. Reports NDCG@10, Macro-F1, Pearson r, V-measure.

researchpythongo
0
3
Path Specific Fairness EvalA

Evaluates a model's ability to make predictions while removing the influence of a sensitive attribute along specific causal pathways, balancing predictive accuracy with path-specific counterfactual fairness constraints. Use when the user wants to benchmark on Berkeley Admission Dataset, UCI Adult Dataset, UCI German Credit Dataset, or asks about evaluating this task. Reports fair accuracy.

researchpythongo
0
3
Pathology Vqa EvalA

Evaluates a multimodal chatbot's ability to interpret real-world pathology images (H&E and IHC) and integrate clinical context to produce accurate diagnoses, terminology, and multimodal reasoning across four anatomical systems. Use when the user wants to benchmark on Pathology Clinical Q&A Dataset, or asks about evaluating this task. Reports diagnosis accuracy.

researchpython
0
3
Patient Flow Forecasting EvalA

Evaluates the ability of machine learning and statistical models to forecast daily patient flows at urgent care clinics, specifically testing robustness to concept drift induced by pandemic disruptions using quasi-real-time proxy variables. Use when the user wants to benchmark on Clinic 1 & 2 patient flow data, or asks about evaluating this task. Reports MAPE.

datapythongo
0
3
Patientsim EvalA

Evaluates the persona fidelity, factual accuracy, and clinical plausibility of an LLM-based patient simulator in doctor-patient dialogues. It measures how well the model adheres to assigned patient profiles, maintains factual consistency, and handles out-of-profile questions plausibly. Use when the user wants to benchmark on PatientSim Profiles, or asks about evaluating this task. Reports Entail (%).

researchpythongo
0
3
Patsql Sql Synthesis EvalA

Evaluates the ability of program-by-example (PBE) systems to synthesize correct SQL queries from example input/output tables. It probes query generation accuracy, synthesis speed, and scalability to larger database schemas. Use when the user wants to benchmark on ase13, so-top, so-dev, so-rec, kaggle, or asks about evaluating this task. Reports solve_rate.

researchpythongo
0
3
Paws EvalA

This benchmark evaluates a model's ability to identify paraphrases in sentence pairs that share high lexical overlap but differ in meaning due to word order and syntactic structure. It specifically probes sensitivity to non-local contextual information and adversarial word scrambling, revealing whether models rely on superficial word matching rather than true semantic understanding. Use when the user wants to benchmark on PAWS_QQP, PAWS_Wiki, or asks about evaluating this task. Reports classi...

researchpythongo
0
3
Paws X EvalA

PAWS-X evaluates a model's ability to identify paraphrases across multiple languages, specifically probing sensitivity to word order and syntactic structure under conditions of high lexical overlap. It measures how well models generalize cross-lingually when trained on machine-translated data versus zero-shot settings. Use when the user wants to benchmark on PAWS-X, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Pbench EvalA

Evaluates a model's ability to perform referring expression segmentation across five hierarchical levels of semantic complexity, from basic object recognition to fine-grained attribute binding, OCR-based disambiguation, spatial layout understanding, and relational interactions. It also stress-tests long-context generation and instance stability in crowded scenes with high object counts. Use when the user wants to benchmark on PBench, or asks about evaluating this task. Reports per-level perfo...

researchpythongo
0
3
Pca EvalA

Evaluates embodied decision-making capabilities across three dimensions: perception (visual understanding), cognition (reasoning and task breakdown), and action (executing correct steps or decisions). It tests models in autonomous driving, domestic assistance, and game-playing environments. Use when the user wants to benchmark on PCA-EVAL, or asks about evaluating this task. Reports Action Score.

ai-agentspythongit
0
3
Pcb Defect Classification EvalA

Evaluates a model's ability to classify six specific types of PCB manufacturing defects from cropped defect images. It probes the model's feature extraction and categorization capabilities on a specialized industrial computer vision dataset. Use when the user wants to benchmark on PCB Defect Dataset, or asks about evaluating this task. Reports average_precision_rate.

researchpythongo
0
3
Pdbbind Binding Affinity EvalA

This benchmark evaluates the ability of docking tools, deep learning models, and meta-modeling ensembles to predict ligand-protein binding affinities. It probes how well different feature representations (physical scores, sequence-based DL outputs, physicochemical properties) generalize to unseen protein-ligand complexes. Use when the user wants to benchmark on PDBbind, or asks about evaluating this task. Reports Pearson correlation coefficient.

researchpythongo
0
3
Pdbbind Lba EvalA

Predicts the binding affinity between a protein pocket and a ligand from 3D structural data. It probes the model's ability to quantify molecular interaction strength and generalize across protein sequence identities. Use when the user wants to benchmark on PDBBind, or asks about evaluating this task. Reports RMSE.

researchpython
0
3
Pdbbind Low Similarity EvalA

Evaluates how well 3D binding affinity models generalise to unseen proteins and novel ligands in low-data regimes. It uses a strict low-Tanimoto-similarity split of the PDBBind dataset to prevent data leakage and benchmark generalisation capabilities. Use when the user wants to benchmark on PDBBind, or asks about evaluating this task. Reports performance.

researchpythongo
0
3
Pdcovidnet EvalA

Evaluates a CNN's ability to classify chest X-ray images into three diagnostic categories: COVID-19, Normal, and Viral Pneumonia. It probes multi-scale feature extraction and robustness to class imbalance in medical imaging. Use when the user wants to benchmark on Custom benchmark dataset, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Pdf Extraction EvalA

Evaluates open-source PDF information extraction tools across multiple content elements (metadata, references, tables, paragraphs, sections, etc.) on academic documents. It probes how well different tools handle layout-based segmentation, text extraction, and structural recognition in real-world academic PDFs. Use when the user wants to benchmark on DocBank, or asks about evaluating this task. Reports F1 score.

researchpythongo
0
3
Pdf Malware Poisoning EvalA

Evaluates the robustness of embedded feature selection methods (LASSO, Ridge, Elastic Net) against training data poisoning attacks in a PDF malware detection setting. It measures how injected malicious samples manipulate feature selection stability and degrade classification performance. Use when the user wants to benchmark on Contagio + Web Benign PDFs, or asks about evaluating this task. Reports classification error.

researchpythonperformance
0
3
Pdf Parsing Chunking EvalA

This evaluation probes the retrieval accuracy of RAG pipelines when processing financial PDFs, specifically testing how different PDF parsers, chunking strategies, and overlap percentages affect the retrieval of relevant pages for both narrative text and structured table queries. Use when the user wants to benchmark on FinanceBench, TableQuest, or asks about evaluating this task. Reports MRR.

researchpythontesting
0
3
Pdfqa EvalA

Evaluates end-to-end question answering over PDF documents, probing parsing, retrieval, and reasoning capabilities across diverse document types, modalities, and complexity dimensions. Use when the user wants to benchmark on pdfQA, or asks about evaluating this task. Reports G-Eval correctness.

researchpythongit
0
3
Pe Malware Classification EvalA

Evaluates learning-based models for PE malware family classification across image, binary, and disassembly input formats. It measures classification accuracy and probes model robustness under concept drift, alongside computational resource overhead. Use when the user wants to benchmark on BIG-15, Malimg, MalwareBazaar, MalwareDrift, or asks about evaluating this task. Reports Macro F1-score ($F1_{macro}$).

researchpythongit
0
3
Pea Architecture EvalA

Evaluates a separation-of-powers AI agent architecture (PEA) on its ability to prevent unauthorized actions, detect goal drift, and identify implicit coercion in adversarial inputs. Use when the user wants to benchmark on Attack Corpus, Drift Dataset, Coercion Dataset, or asks about evaluating this task. Reports Bypass Rate, Attack Success Rate (ASR).

researchpythongo
0
3
PeaksignalnoiseratioA

Compute the PeakSignalNoiseRatio metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute PeakSignalNoiseRatio, or asks how to score with PeakSignalNoiseRatio.

documentationpython
0
3
PeaksignalnoiseratiowithblockedeffectA

Compute the PeakSignalNoiseRatioWithBlockedEffect metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute PeakSignalNoiseRatioWithBlockedEffect, or asks how to score with PeakSignalNoiseRatioWithBlockedEffect.

documentationpython
0
3
Pear Weather EvalA

Evaluates medium-term weather forecasting capability on a spherical grid by predicting atmospheric variables up to 10 days ahead. It probes the model's ability to capture spatial and temporal dynamics without grid-induced resolution biases. Use when the user wants to benchmark on ERA5-lite, or asks about evaluating this task. Reports ACC.

researchpython
0
3
Pearl EvalA

Evaluates large vision-language models' ability to understand and generate culturally-aware Arabic content across multiple reasoning-centric question types. It probes hypothesis formation, comparative analysis, chronological reasoning, and explicit cultural grounding in both closed-form and open-ended multimodal tasks. Use when the user wants to benchmark on PeARL, or asks about evaluating this task. Reports relaxed-match accuracy (ACC).

researchpythongo
0
3
Pearson Correlation CoefficientA

Probes the ability of an automated metric to correlate with human judgments of translation quality. Specifically, it tests how well a model's predicted segment-level scores align with Direct Assessment (DA) human scores across multiple language pairs. Use when the user has predictions and gold and needs to compute Pearson correlation coefficient.

researchpythongo
0
3
Pearson CorrelationA

Evaluates the validity of automatic machine translation metrics by measuring their linear correlation with human Direct Assessment (DA) scores at both segment and system levels. It emphasizes rigorous validation protocols, including adaptive sample size determination for human judgments and statistical significance testing to compare metric performance. Use when the user has predictions and gold and needs to compute pearson-correlation.

testingpythongo
0
3
PearsoncorrcoefA

Compute the PearsonCorrCoef metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute PearsonCorrCoef, or asks how to score with PearsonCorrCoef.

documentationpython
0
3