All authors
qhjqhj00 avatar

Claude Skills by qhjqhj00

github.com/qhjqhj00
7,636 skillsA× 7,623B× 11C× 1D× 10 installs2,124 views
Emu Edit EvalA

Evaluates image editing capability by measuring instruction following (text similarity) and source image preservation (image similarity). It tests the model's ability to modify images based on textual instructions while maintaining relevant visual elements. Use when the user wants to benchmark on EMU-Edit, or asks about evaluating this task. Reports CLIP-T.

researchpython
0
3
Emu35 T2i X2i EvalA

Evaluates a multimodal model's capability to generate images from text prompts and edit existing images based on natural language instructions. It probes semantic alignment, fine-grained text rendering accuracy, and instruction-following fidelity across diverse visual tasks. Use when the user wants to benchmark on GenEval, DPG-bench, OneIG-Bench, TIIF-Bench mini, LeX-Bench, CVTG-2K, LongText-Bench, ImgEdit, GEdit-Bench, OmniContext, ICE-Bench, or asks about evaluating this task. Reports Word ...

researchpythongo
0
3
En Ca Biomedical Translation EvalA

Evaluates English-to-Catalan machine translation quality in the biomedical domain using a two-stage cascade pivot strategy (English→Spanish→Catalan) versus direct translation, measuring lexical overlap and fluency via BLEU scores on domain-specific test sets. Use when the user wants to benchmark on WMT Biomedical test set, El Periódico test set, or asks about evaluating this task. Reports BLEU.

researchpython
0
3
En Fr Translation BenchA

Evaluates machine translation quality and computational efficiency on a curated set of English-to-French sentences spanning simple, technical, and complex domains. It measures linguistic accuracy alongside inference latency and hardware resource consumption under consumer-grade GPU constraints. Use when the user wants to benchmark on Custom EN-FR Test Set, or asks about evaluating this task. Reports BLEU score.

researchpythongo
0
3
Encoder Adaptation EvalA

This evaluation protocol assesses the capability of decoder-based language models adapted into encoder-only architectures to perform diverse downstream tasks, including text classification, scoring, and information retrieval. It specifically probes how architectural modifications like bidirectional attention, pooling strategies, and dropout affect performance on standard benchmarks. Use when the user wants to benchmark on GLUE, SuperGLUE, MS MARCO, or asks about evaluating this task. Reports ...

researchpythonperformance
0
3
Encqa EvalA

Evaluates vision-language models on their ability to interpret visual encodings (position, length, area, color, shape) and perform chart-specific analytic tasks (e.g., value retrieval, anomaly detection, correlation estimation). It probes fine-grained visual perception, reasoning under different encoding constraints, and whether model capabilities scale with size or prompting strategies. Use when the user wants to benchmark on EncQA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Endo Depth Robustness EvalA

This benchmark evaluates the robustness of monocular depth estimation models when processing endoscopic images degraded by realistic surgical artifacts. It probes how well models maintain depth prediction accuracy and consistency under varying severities of illumination changes, optical blurs, visual obstructions, sensor noise, and compression artifacts. Use when the user wants to benchmark on Endoscopic Depth Estimation Dataset (Synthetically Corrupted), or asks about evaluating this task. R...

researchpythongit
0
3
Enem Vision EvalA

Evaluates multimodal and text-only language models on Brazilian university admission exams (ENEM), specifically probing their ability to comprehend visual information, interpret tables/figures, and perform mathematical reasoning in a multiple-choice format. Use when the user wants to benchmark on ENEM 2022/2023, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Energaizer Gpu Power EvalA

Evaluates the accuracy of a lightweight analytical framework in predicting GPU latency and dynamic power consumption for AI workloads across different hardware architectures, operating frequencies, and algorithm configurations. Use when the user wants to benchmark on EnergAIzer Kernel Database & AI Workloads, or asks about evaluating this task. Reports MAPE.

researchpythongo
0
3
Energy EfficiencyA

Evaluates the trade-off between energy efficiency and physical-layer security in an IRS-assisted MISO network with cooperative jamming. Probes how varying transmit power constraints and secrecy rate thresholds impact system performance compared to baseline beamforming strategies. Use when the user has predictions and gold and needs to compute Energy Efficiency.

researchpythongo
0
3
Energy First Arch EvalA

Evaluates the classification accuracy and training energy efficiency of biologically-inspired and physics-guided neural architectures against conventional baselines across diverse data modalities. It probes whether action-principle regularization yields modality-specific performance gains and reduced internal activation energy without accuracy loss. Use when the user wants to benchmark on Fashion-MNIST, CIFAR-10, DVS Gesture, SHD, SSC, WESAD, DREAMER, SEED-IV, 20newsgroups, or asks about eval...

researchpythongo
0
3
Energy EfficiencyA

Evaluates the energy efficiency of a wireless power transmission system where an energy source learns optimal transmit power levels for energy-harvesting nodes using a stochastic multi-armed bandit algorithm without channel state information. Use when the user has predictions and gold and needs to compute EE.

developmentpythongo
0
3
Enginead EvalA

Probes the ability of one-class anomaly detection algorithms to identify incipient engine faults in real-world, multivariate vehicle sensor telemetry. It specifically evaluates cross-vehicle generalization and robustness to distributional shifts in normal operating conditions across a commercial fleet. Use when the user wants to benchmark on EngineAD, or asks about evaluating this task. Reports F1-score (anomaly class).

researchpythongo
0
3
Ens 10 EvalA

Probes the ability of deep learning and statistical models to correct biases in long-term ensemble weather forecasts. It evaluates how well models can post-process raw ensemble members to produce calibrated predictive distributions for surface and atmospheric variables. Use when the user wants to benchmark on ENS-10, or asks about evaluating this task. Reports CRPS.

datapythonperformance
0
3
Enterprise Benchmarks EvalA

Evaluates LLMs on domain-specific enterprise tasks across finance, legal, climate, and cybersecurity. It probes capabilities like numerical reasoning, named entity recognition, document relevance ranking, and long-document summarization using real-world industry data. Use when the user wants to benchmark on Earnings Call Transcripts, News Headline, Credit Risk Assessment (NER), KPI-Edgar, FiNER-139, Opinion-based QA (FiQA), Sentiment Analysis (FiQA SA), Insurance QA, ConvFinQA, Financial Text...

researchpythongo
0
3
Enterprise Sql Kg Qa EvalA

Evaluates large language models' ability to generate correct SQL or SPARQL queries for natural language questions over an enterprise insurance database. It measures how well zero-shot prompting with raw schema versus knowledge graph augmentation improves factual grounding and query execution accuracy. Use when the user wants to benchmark on Enterprise SQL & KG QA Benchmark, or asks about evaluating this task. Reports execution accuracy.

researchpythonsql
0
3
Entities Of The Union EvalA

Evaluates the ability of models to disambiguate named entity mentions in historical and modern newswire texts, and to cluster coreferent mentions across documents. It specifically probes handling of out-of-knowledgebase individuals common in historical contexts. Use when the user wants to benchmark on Entities of the Union, MSNBC, ACE2004, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Entity Canonicalization EvalA

Evaluates a model's ability to cluster entity mentions (noun phrases) into canonical forms within open knowledge graphs. It probes unsupervised representation learning and clustering capabilities without relying on manually annotated ground truth for training. Use when the user wants to benchmark on Base, Ambiguous, ReVerb45K, CanonicNell, or asks about evaluating this task. Reports Macro F1.

researchpythongo
0
3
Entity Hallucination EvalA

Evaluates large language models' ability to generate factually correct answers to complex and factual questions while mitigating entity-level hallucinations. It also measures the effectiveness of a real-time hallucination detection mechanism in identifying fabricated or low-confidence entities during generation. Use when the user wants to benchmark on WikiBio GPT-3 dataset, 2WikiMultihopQA, StrategyQA, NQ, or asks about evaluating this task. Reports AUC.

researchpythongo
0
3
Entity Linking EvalA

Evaluates end-to-end entity linking systems on their ability to detect entity mentions and correctly disambiguate them to knowledge base entities. It specifically probes for systemic benchmark biases, such as overreliance on named entities, ambiguous disambiguation choices, and underrepresented entity types, by introducing fairer evaluation protocols. Use when the user wants to benchmark on Existing and new EL benchmarks, or asks about evaluating this task. Reports Micro F1.

researchpythongo
0
3
Entity6k EvalA

Evaluates open-domain entity recognition capabilities across four visual grounding and understanding tasks: object detection, zero-shot image classification, image captioning, and dense captioning. It measures how well models can localize, classify, and describe specific real-world entities in images without fine-tuning. Use when the user wants to benchmark on Entity6K, or asks about evaluating this task. Reports AP.

researchpythongo
0
3
Entropy Adaptive Merging EvalA

Evaluates the robustness of an entropy-adaptive model merging framework under heterogeneous domain shifts. It probes whether a merged model can adapt to unseen target domains using only forward passes and entropy-based coefficients, without backpropagation or labeled target data. Use when the user wants to benchmark on MiDog Atypical, Organs, Histopantum, ISIC Skin, Messidor, PACS, VLCS, Office-Home, TerraIncognita, or asks about evaluating this task. Reports accuracy.

researchpythonperformance
0
3
Entropy Minimization Reasoning EvalA

Evaluates the reasoning capabilities of LLMs on mathematical and coding benchmarks by measuring accuracy under various inference-time scaling and unsupervised entropy minimization techniques. It probes whether reducing output uncertainty improves correctness without labeled data or parameter updates. Use when the user wants to benchmark on AMC, AIME, Minerva, LeetCode Live Contest, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Enviroexam EvalA

This benchmark evaluates large language models' domain-specific knowledge in environmental science using multiple-choice questions derived from university curricula. It measures both raw accuracy and performance consistency across different course topics, revealing how well models retain and apply specialized scientific concepts. Use when the user wants to benchmark on EnviroExam, or asks about evaluating this task. Reports composite_index.

researchpythongo
0
3
Eo Bench EvalA

Evaluates a model's ability to reason about embodied interactions, including spatial understanding, physical commonsense, task planning, and state estimation from robot vision and text inputs. Use when the user wants to benchmark on EO-Bench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Epd EvalA

Evaluates how well image quality assessment models correlate with actual robotic task performance under various image distortions. It probes whether traditional human-centric visual quality metrics align with the perception needs of embodied robots performing push and pick tasks. Use when the user wants to benchmark on EPD, or asks about evaluating this task. Reports PLCC.

researchpythongo
0
3
Epic Kitchens 100 Mqa EvalA

Evaluates multi-modal large language models' ability to recognize and distinguish between similar human actions in egocentric videos through multiple-choice question answering. It specifically probes fine-grained action discrimination using hard, semantically and visually similar distractors generated by action recognition models. Use when the user wants to benchmark on EPIC-KITCHENS-100-MQA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Epidemiological Benchmark EvalA

Evaluates spatio-temporal graph models on forecasting, stability, and denoising tasks using synthetic epidemiological data generated from PDEs. Probes the model's ability to predict future states, resist noise/dropout, and recover clean signals from corrupted graph time-series. Use when the user wants to benchmark on Epidemiological Synthetic Dataset, or asks about evaluating this task. Reports RMSE.

researchpythonnode
0
3
Epps Singleton 2sampA

Compute the epps_singleton_2samp metric — provided by scipy.stats. Use when the user has predictions and ground-truth and needs to compute epps_singleton_2samp, or asks how to score with epps_singleton_2samp.

documentationpythongo
0
3
EpsilonA

Evaluates the correlation between a zero-cost NAS metric (epsilon) and actual training accuracy across different neural architecture search spaces, testing the metric's ability to rank architectures without training. It probes whether output dispersion from constant weight initializations can serve as a reliable, hyperparameter-free proxy for architecture selection. Use when the user has predictions and gold and needs to compute Spearman ρ (global).

researchpythongo
0
3
Epsos Nids EvalA

Evaluates the ability of machine learning and deep learning classifiers, particularly Decision Trees optimized with Enhanced Particle Swarm Optimization, to accurately detect and classify multiple types of network intrusions in high-dimensional traffic data. Use when the user wants to benchmark on CSE-CIC-IDS-2018, LITNET-2020, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Ept Benchmark EvalA

Assesses the trustworthiness of large language models within a Persian-Islamic cultural context across six dimensions: Ethics, Fairness, Privacy, Robustness, Safety, and Truthfulness. It measures how well model responses align with culturally specific ethical principles and linguistic nuances. Use when the user wants to benchmark on EPT Benchmark, or asks about evaluating this task. Reports compliance metric.

researchpythonrust
0
3
Ept15 Weather Bench EvalA

Evaluates the accuracy of AI weather forecasting models against established numerical models and ground-truth observations. It probes the model's ability to predict atmospheric variables (e.g., wind speed, solar radiation) at hourly resolution over 20-day lead times. Use when the user wants to benchmark on ERA5, IFS HRES IC, Weather Stations, or asks about evaluating this task. Reports Skill Score (SS).

researchpythonperformance
0
3
Eq Bench EvalA

Evaluates large language models' ability to understand and rate emotional intensity in conflict-driven dialogue scenarios. It probes emotional intelligence through automated scoring of model-generated ratings, avoiding subjective human interpretation. Use when the user wants to benchmark on EQ-Bench, or asks about evaluating this task. Reports EQ-Bench Score.

researchpythongo
0
3
Er Reason EvalA

Evaluates LLMs on longitudinal clinical reasoning across five emergency room workflow stages, including acuity assessment, case summarization, treatment planning, final diagnosis, and patient disposition. It probes the models' ability to integrate sparse clinical notes, perform rule-out differential diagnosis, and align outputs with real-world clinical decision-making and safety constraints. Use when the user wants to benchmark on ER-Reason, or asks about evaluating this task. Reports Accurac...

researchpythongo
0
3
Era5 Weatherbench Forecasting EvalA

Spatio-temporal forecasting of atmospheric temperature using deep learning models. It probes the ability to capture long-range spatial-temporal dependencies and predict future weather states from historical multi-feature sequences. Use when the user wants to benchmark on ERA5 Turkey, WeatherBench, or asks about evaluating this task. Reports RMSE.

researchpythonexpress
0
3
Erase EvalA

Evaluates machine unlearning algorithms in recommender systems across collaborative filtering, session-based, and next-basket recommendation tasks. It measures how well models retain recommendation quality after unlearning sensitive or malicious user interactions, while maintaining computational efficiency and effectiveness compared to full retraining. Use when the user wants to benchmark on ERASE Benchmark (9 datasets), or asks about evaluating this task. Reports utility.

researchpythongo
0
3
Eraser Benchmark EvalA

Evaluates NLP models' ability to generate faithful, task-appropriate rationales for predictions, measuring both alignment with human annotations and causal faithfulness via token perturbation. Use when the user wants to benchmark on Movies, FEVER, CoS-E, eSNLI, or asks about evaluating this task. Reports AUPRC.

researchpythonperformance
0
3
Erc EvalA

Evaluates a model's ability to recognize and classify emotional states from conversational dialogue context. It probes multi-turn emotional understanding, speaker-aware reasoning, and generalization across diverse domain settings and speaker demographics. Use when the user wants to benchmark on IEMOCAP, MELD, EmoryNLP, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Erdes Oculardetachment EvalA

Evaluates the ability of spatiotemporal deep learning models to classify ocular ultrasound videos into binary diagnostic categories: detecting retinal detachment (RD) versus non-RD, and classifying macular status (intact vs. detached). It probes the model's capacity to learn subtle spatiotemporal patterns in medical ultrasound while handling real-world class imbalance. Use when the user wants to benchmark on ERDES, or asks about evaluating this task. Reports F1-Score.

researchpythongo
0
3
Ernie 5.0 EvalA

Evaluates a trillion-parameter multimodal foundation model across text, vision, audio, and generation tasks to measure factual knowledge, reasoning, coding, instruction following, and agent capabilities. Use when the user wants to benchmark on PreciseWikiQA, MMLU-Pro, MATH, LiveCodeBench, MMMU-Pro, MathVista, GenEval, VBench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Erntkn Dice CoefficientA

Compute erntkn/dice_coefficient via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of erntkn/dice_coefficient.

developmentpython
0
3
Error Detection Hmc EvalA

Evaluates a model's ability to detect classification errors and recover hierarchical multi-label constraints without prior knowledge. It probes the system's capacity to generate interpretable logical rules from failure patterns and improve downstream model consistency. Use when the user wants to benchmark on Military Vehicles, ImageNet50, OpenImage36, or asks about evaluating this task. Reports F1-score.

researchpythongo
0
3
ErrorrelativeglobaldimensionlesssynthesisA

Compute the ErrorRelativeGlobalDimensionlessSynthesis metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute ErrorRelativeGlobalDimensionlessSynthesis, or asks how to score with ErrorRelativeGlobalDimensionlessSynthesis.

documentationpython
0
3
Ersb EvalA

This benchmark evaluates the environmental resilience of discrete speech codecs by measuring how reconstruction quality and downstream task performance degrade under varying signal-to-noise ratios, loudness levels, and real-world acoustic conditions. It probes both signal fidelity and semantic/intelligibility consistency after codec compression and subsequent speech enhancement or recognition. Use when the user wants to benchmark on Environment-Resilient Speech Codec Benchmark (ERSB), or asks...

researchpythonperformance
0
3
Ervqa EvalA

This benchmark evaluates the clinical readiness of Large Vision Language Models (LVLMs) for emergency room monitoring tasks. It probes their ability to generate accurate, clinically cautious, and semantically entailed long-form answers from medical images, while identifying specific failure modes like hallucinations and overconfidence. Use when the user wants to benchmark on ERVQA, or asks about evaluating this task. Reports Entailment Score.

researchpythongo
0
3
Es Memeval EvalA

This benchmark evaluates conversational agents' long-term memory capabilities in personalized emotional support contexts. It probes five core competencies—information extraction, temporal reasoning, conflict detection, abstention, and user modeling—across question answering, summarization, and dialogue generation tasks. Use when the user wants to benchmark on ES-MemEval, or asks about evaluating this task. Reports F1-Score, LLM-as-Judge.

ai-agentspythongit
0
3
Esc50 Sep EvalA

Evaluates zero-shot language-queried audio source separation on environmental sound classes. The benchmark tests the model's ability to isolate a target sound from a mixed audio mixture using text labels. Use when the user wants to benchmark on ESC-50, or asks about evaluating this task. Reports SDRi.

researchpythongo
0
3
Esci Similarity And Token Class EvalA

Evaluates e-commerce language understanding through masked token recovery on product texts and graded semantic similarity between search queries and products. Also assesses general natural language understanding capabilities via the GLUE benchmark. Use when the user wants to benchmark on Amazon ESCI, GLUE, or asks about evaluating this task. Reports top-k accuracy, Spearman correlation.

researchpythongo
0
3
Esmm Cvr Ctcvr EvalA

Evaluates a model's ability to estimate post-click conversion rate (CVR) and post-click-and-conversion rate (CTCVR) in recommendation systems. It specifically probes how well the model handles sample selection bias and data sparsity by comparing performance on clicked-only impressions versus the entire impression space. Use when the user wants to benchmark on Public Dataset, or asks about evaluating this task. Reports AUC.

researchpythonperformance
0
3