All authors
qhjqhj00 avatar

Claude Skills by qhjqhj00

github.com/qhjqhj00
7,636 skillsA× 7,623B× 11C× 1D× 10 installs2,121 views
Econlogicqa EvalA

This benchmark evaluates large language models' ability to perform economic sequential reasoning by logically ordering interconnected business and supply chain events. It probes multi-event causality and temporal reasoning beyond simple chronological sorting, requiring models to understand complex economic narratives. Use when the user wants to benchmark on EconLogicQA, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Economic Value Of Water EvalA

Evaluates the economic value of hydropower reservoir operations based on sub-seasonal to seasonal precipitation forecasts under varying energy price differentials. It probes how forecast horizon and reservoir storage capacity interact with market prices to determine operational profitability. Use when the user wants to benchmark on 10-year reservoir timeseries, or asks about evaluating this task. Reports overall value of water.

businesspythonperformance
0
3
Ecosched Hpc Scheduling EvalA

Evaluates an online co-scheduling framework's ability to jointly optimize GPU count selection and job packing to minimize energy consumption while maintaining performance across diverse multi-GPU HPC workloads. It probes the scheduler's capacity to handle non-linear scaling, NUMA-aware placement, and dynamic workload packing without prior knowledge of exact runtimes. Use when the user wants to benchmark on Multi-GPU Benchmark Suite, or asks about evaluating this task. Reports Energy Saving.

researchpythongo
0
3
Ecovnet Covidx EvalA

Evaluates the ability of deep convolutional neural networks to classify chest X-ray images into three categories: COVID-19, normal, and pneumonia. It probes robustness under class imbalance and tests the effectiveness of ensemble learning strategies (hard vs. soft voting) combined with data augmentation. Use when the user wants to benchmark on COVIDx, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Ecvr Prediction EvalA

Evaluates the ability to predict click-through, conversion, and effective conversion rates in a large-scale e-commerce recommender system. It specifically probes how well models handle cascade delayed feedback, sample selection bias, and data sparsity when predicting user purchase and refund behaviors. Use when the user wants to benchmark on Alibaba Production Dataset, or asks about evaluating this task. Reports AUC.

researchpythonperformance
0
3
Edacc EvalA

Evaluates automatic speech recognition (ASR) models on naturalistic, conversational English speech with diverse international accents. It probes the robustness of state-of-the-art ASR systems to real-world speaking conditions and accent variation compared to read-speech benchmarks like LibriSpeech. Use when the user wants to benchmark on EdAcc, or asks about evaluating this task. Reports WER.

researchpython
0
3
Edge Ai Platform Inference EvalA

Evaluates the inference performance of heterogeneous edge AI platforms (CPU, GPU, NPU) across fundamental linear algebra primitives and diverse neural network models. It probes hardware efficiency in compute-bound versus memory-bound workloads, batch processing scalability, and quantization support. Use when the user wants to benchmark on Matrix Multiplication, Matrix-Vector Multiplication, Dot Product, MobileNetV2, LSTM, TinyLlama, or asks about evaluating this task. Reports Latency (ms).

businesspythonperformance
0
3
Edge Case Detection EvalA

Tests the model's ability to classify whether a respondent's message represents an edge case that falls outside the scope of existing coordination policies and requires user escalation. Use when the user wants to benchmark on Edge Case Detection Test Suite, or asks about evaluating this task. Reports Accuracy.

researchpythonrust
0
3
Edge Llm Energy Accuracy EvalA

Evaluates the trade-offs between quantization levels, model architectures, and task characteristics on energy efficiency and output accuracy for LLMs deployed on edge hardware. Use when the user wants to benchmark on bigbenchhard, commonsenseqa, gsm8k, humaneval, truthfulqa, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Edge Llm Inference BenchmarkA

Evaluates the trade-offs between token throughput, latency, energy efficiency, and physical footprint when deploying compact LLMs on various IoT-grade single-board computers with different hardware accelerators (CPU, NPU, GPU). Use when the user has predictions and gold and needs to compute Throughput (tokens/s), Time-to-first-token (TTFT), Energy per million tokens (MJ/Mtok).

researchpythongo
0
3
Edge Lm Inference EvalA

Evaluates the feasibility and performance trade-offs of running small generative language models on edge hardware. It probes memory constraints, inference latency, token throughput, and energy efficiency across different quantization schemes and system configurations. Use when the user wants to benchmark on None (system-level inference benchmark), or asks about evaluating this task. Reports generation_throughput.

researchpythongo
0
3
Edge Llm Inference EvalA

Evaluates the feasibility and performance of deploying various LLMs on CPU-only edge hardware (Raspberry Pi 5 clusters) by measuring inference speed, resource consumption, and reasoning accuracy under constrained conditions. Use when the user wants to benchmark on OpenAssistant/oasst1 (subset), Winogrande, or asks about evaluating this task. Reports accuracy.

researchpythonperformance
0
3
Edgesnn Evaluation ProtocolA

Evaluates the performance, efficiency, and robustness of Spiking Neural Networks (SNNs) deployed on edge hardware or simulated on conventional processors. It probes hardware-independent algorithmic complexity and system-level execution metrics under resource-constrained, latency-sensitive conditions. Use when the user has predictions and gold and needs to compute accuracy / mAP / MSE.

businesspythongo
0
3
Ediref Erc EfrA

Evaluates emotion recognition and emotion-flip reasoning in multi-party conversations, specifically identifying trigger utterances that cause emotional shifts in both code-mixed (Hindi-English) and monolingual English dialogues. Use when the user wants to benchmark on E-MaSaC, MELD-FR, or asks about evaluating this task. Reports weighted F1, F1 score for trigger utterances.

researchpythongit
0
3
EditdistanceA

Compute the EditDistance metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute EditDistance, or asks how to score with EditDistance.

documentationpythongo
0
3
Editeval EvalA

Evaluates instruction-based text editing capabilities across modular tasks such as simplification, fluency, coherence, paraphrasing, neutralization, and information updating. It measures how well models follow prompts to improve or modify text while preserving intended meaning. Use when the user wants to benchmark on EditEval Benchmark, or asks about evaluating this task. Reports SARI.

researchpythongit
0
3
Editreward EvalA

Evaluates the quality and human alignment of instruction-guided image editing models. It measures how well generated images match user instructions and visual realism, as well as how accurately models rank pairs of edited images according to human preferences. Use when the user wants to benchmark on ImagenHub, GenAI-Bench, AURORA-Bench, EditReward-Bench, GEdit-Bench, or asks about evaluating this task. Reports Spearman rank correlation, Pair-wise prediction accuracy.

researchpythongo
0
3
Editverse Bench EvalA

Evaluates instruction-based video editing capabilities, including text alignment, temporal consistency, and editing faithfulness across diverse resolutions and orientations. It probes the model's ability to follow complex editing prompts while preserving unedited regions and maintaining high video quality. Use when the user wants to benchmark on EditVerseBench, or asks about evaluating this task. Reports VLM evaluation (Editing Quality).

researchpythongo
0
3
Ee Power Control EvalA

This evaluation probes the energy efficiency and feasibility of power control algorithms in 5G massive MIMO and relay-assisted interference networks. It measures how well centralized and distributed algorithms maximize Global Energy Efficiency (GEE) while satisfying minimum per-user rate constraints under hardware impairments and Rayleigh fading. Use when the user wants to benchmark on Hardware-Impaired Massive MIMO System, Relay-assisted OFDMA interference network, or asks about evaluating t...

researchpythongo
0
3
Eee Bench EvalA

Probes multimodal reasoning and visual diagram interpretation in electrical and electronics engineering. It tests whether models can integrate complex circuit and system diagrams with textual problem descriptions to apply domain-specific knowledge and perform accurate calculations or logical deductions. Use when the user wants to benchmark on EEE-Bench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Eefsuva EvalA

Evaluates LLMs' ability to solve nonstandard mathematical Olympiad problems from Eastern European and former Soviet Union competitions. It probes genuine mathematical reasoning and adaptability by testing whether models can solve problems from first principles rather than relying on cached solutions or pattern matching from familiar Western benchmarks. Use when the user wants to benchmark on EEFSUVA, or asks about evaluating this task. Reports pass rate.

researchpythongo
0
3
Eeg Asr Noisy Speech EvalA

Evaluates end-to-end continuous speech recognition models using only electroencephalography (EEG) signals, and assesses robustness to background noise by fusing EEG with acoustic features. Use when the user wants to benchmark on Database A, Database B, or asks about evaluating this task. Reports Word Error Rate (WER).

researchpythondatabase
0
3
Eeg Ssl Emotion EvalA

Evaluates semi-supervised EEG-based emotion recognition under extreme label scarcity. It probes the model's ability to leverage unlabeled data via representation alignment while maintaining classification performance across subject-dependent and subject-independent protocols. Use when the user wants to benchmark on SEED, SEED-IV, SEED-V, AMIGOS, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
EerA

Compute the EER metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute EER, or asks how to score with EER.

documentationpythondocumentation
0
3
Ef4inca Nowcast EvalA

Evaluates the capability of spatiotemporal Transformer models to nowcast convective precipitation up to 90 minutes ahead using multi-source meteorological data. It probes the model's ability to fuse satellite infrared, radar, and NWP inputs to accurately predict the initiation, location, and intensity of rapidly evolving convective cells. Use when the user wants to benchmark on Austria convective precipitation dataset, or asks about evaluating this task. Reports Critical Success Index (CSI).

researchpythongit
0
3
Effective DimensionalityA

Effective Dimensionality (ED) quantifies the number of independent signals or latent axes captured by a benchmark, measuring how much redundancy exists across its tasks. It probes whether a benchmark's claimed breadth actually reflects diverse evaluation dimensions or merely correlated task performance. Use when the user has predictions and gold and needs to compute Effective Dimensionality (ED).

researchpythongo
0
3
Efficient Bert EvalA

Evaluates the performance of efficiently trained BERT models (via Mixture-of-Supernets) on downstream natural language understanding tasks. It probes the trade-off between model size, training compute, and accuracy compared to standalone pretraining and other NAS baselines. Use when the user wants to benchmark on GLUE benchmark, or asks about evaluating this task. Reports Avg. GLUE.

researchpythonperformance
0
3
Efok Cqa EvalA

Evaluates knowledge graph complex query answering models on existential first-order (EFO) queries with multiple free variables and complex structures (cycles, multi-hop), testing their ability to handle combinatorially hard queries beyond simple set operations. Use when the user wants to benchmark on EFO_k-CQA, or asks about evaluating this task. Reports MRR.

researchpythongo
0
3
Ege Math Assessment EvalA

This benchmark evaluates vision-language models' ability to assess handwritten mathematical solutions against a standardized educational rubric. It probes the models' capacity for error diagnosis, step-by-step reasoning alignment, and accurate grade assignment under varying levels of contextual guidance. Use when the user wants to benchmark on EGE-Math Solutions Assessment Benchmark, or asks about evaluating this task. Reports final_score.

researchpythongo
0
3
Egida Safety EvalA

Evaluates the robustness of LLMs against jailbreaking attacks after safety alignment. It measures how well models refuse harmful prompts across diverse topics and attack styles, while also tracking unintended side effects like over-refusal and general capability degradation. Use when the user wants to benchmark on Egida, or asks about evaluating this task. Reports ASR.

researchpythonperformance
0
3
Egmm Corpus EvalA

Probes zero-shot visual-language alignment and cultural recognition capabilities of vision-language models on Egyptian cultural concepts. It measures how well models can classify images into specific cultural categories and retrieve matching text descriptions without fine-tuning. Use when the user wants to benchmark on EgMM-Corpus, or asks about evaluating this task. Reports Acc@1.

researchpythongo
0
3
Ego Instructor EvalA

This evaluation protocol assesses a retrieval-augmented egocentric video captioning framework. It probes the model's ability to perform cross-view video-text and video-video retrieval, answer multiple-choice questions based on video-text alignment, and generate accurate egocentric video captions using retrieved exocentric instructional videos as references. Use when the user wants to benchmark on EK100 MIR, EgoMCQ, SummMCQ, YouCook2-Clip, YouCook2-Video, CharadesEgo, EgoLearner-MCQ, Ego4d coo...

researchpythongo
0
3
Ego Walk Nav EvalA

Evaluates the ability of visual navigation models to predict future robot trajectories from egocentric video frames and context history. It probes scale-invariant trajectory prediction and alignment with human navigation behavior under domain shift conditions. Use when the user wants to benchmark on EgoWalk, or asks about evaluating this task. Reports MSE.

researchpythongo
0
3
Ego3d Bench EvalA

Evaluates 3D spatial reasoning and multi-view understanding in Vision-Language Models, specifically testing ego-centric distance estimation, object localization, motion tracking, travel time estimation, and relative location reasoning across multiple camera views. Use when the user wants to benchmark on Ego3D-Bench, or asks about evaluating this task. Reports Accuracy (%), RMSE.

researchpythongo
0
3
Egoavu Bench EvalA

Evaluates multimodal large language models' ability to perform joint audio-visual reasoning on egocentric videos, including action/object/sound recognition, temporal reasoning, hallucination detection, and dense audio-visual narration. It specifically probes whether models can correctly associate environmental sounds with their visual sources and maintain temporal alignment without relying heavily on visual cues. Use when the user wants to benchmark on EgoAVU-Bench, or asks about evaluating t...

researchpythongo
0
3
Egohumans EvalA

Evaluates the ability of models to perform robust multi-human tracking and identity association in unconstrained egocentric 3D environments. It probes how well algorithms handle severe occlusions, dynamic activities, and camera-agnostic spatial reasoning when fusing egocentric and secondary views. Use when the user wants to benchmark on EgoHumans, or asks about evaluating this task. Reports IDF1.

researchpythongo
0
3
Egomem EvalA

Evaluates a lifelong memory agent's ability to perform real-time audiovisual user retrieval, detect dialog session boundaries in continuous streams, and generate personalized, fact-consistent responses in full-duplex omnimodal interactions. Use when the user wants to benchmark on LFW, VoxCeleb, EgoMem Custom Text Retrieval, EgoMem Episodic Trigger, or asks about evaluating this task. Reports pass@5, Fact Score.

researchpythongo
0
3
Egonormia EvalA

Evaluates vision-language models' ability to understand and reason about physical-social norms in egocentric video scenarios. It probes whether models can correctly select normative actions, justify them, and identify plausible alternatives in conflict-prone situations. Use when the user wants to benchmark on EgoNormia, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Egoschema EvalA

Probes long-term visual memory and temporal reasoning in video-language models by requiring them to answer multiple-choice questions about very long-form videos. It measures the model's ability to retain and retrieve information across extended durations without relying on short clip analysis. Use when the user wants to benchmark on EgoSchema, or asks about evaluating this task. Reports QA Accuracy.

researchpythongo
0
3
Egoscreen Emotion EvalA

Evaluates a model's ability to predict human emotional responses to movie scenes from an egocentric, first-person screen-view perspective. It probes multimodal long-context reasoning by combining visual frames, audio cues, and narrative summaries to handle domain shifts from cinematic to realistic viewing conditions. Use when the user wants to benchmark on EgoScreen-Emotion (ESE), or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Egotraj Bench EvalA

Evaluates the robustness of trajectory prediction models when historical observations are corrupted by realistic ego-view perception noise (occlusions, ID switches, ego-motion drift) compared to clean bird's-eye-view ground truth. Use when the user wants to benchmark on EgoTraj-TBD, or asks about evaluating this task. Reports minADE@K, minFDE@K.

researchpythongo
0
3
Egoxtreme EvalA

Evaluates the robustness of 6D object pose estimation models under extreme real-world visual conditions, including severe motion blur, dynamic lighting, and smoke. It also benchmarks temporal tracking strategies in highly dynamic egocentric scenarios to assess motion-aware inference capabilities. Use when the user wants to benchmark on EgoXtreme, or asks about evaluating this task. Reports ADD(-S) recall.

researchpythongo
0
3
Egtr Sgg EvalA

Evaluates a model's ability to detect objects and predict relational triplets (subject-predicate-object) in natural images. It probes both object detection accuracy and scene graph generation quality under graph constraints and standard recall/mAP metrics. Use when the user wants to benchmark on Visual Genome, Open Image V6, or asks about evaluating this task. Reports Recall@k (R@k), micro-R@50.

researchpythongo
0
3
Ehr Clinical Outcome Prediction EvalA

This benchmark evaluates clinical outcome prediction models across three distinct EHR data representations (multivariate time-series, event streams, and textual event streams). It probes how well different architectures handle sparse, irregular longitudinal patient data and varying feature missingness rates in both acute ICU and long-term care settings. Use when the user wants to benchmark on MIMIC-IV, EHRSHOT, or asks about evaluating this task. Reports F1 score, AUROC, AUPRC.

researchpythongit
0
3
Ehrnoteqa EvalA

Evaluates large language models' ability to perform patient-specific clinical reasoning by synthesizing information from multiple electronic health record (EHR) discharge summaries to answer medical questions. It specifically tests multi-document clinical analysis and automated medical model evaluation using structured multi-choice or free-text formats. Use when the user wants to benchmark on EHRNoteQA, or asks about evaluating this task. Reports score.

researchpythongo
0
3
Ehrr1 EvalA

Evaluates a language model's ability to perform clinical decision-making and risk prediction using longitudinal electronic health record (EHR) data. It probes the model's capacity for multi-label entity recommendation, binary outcome forecasting, and generalization across different healthcare systems and diagnostic granularities. Use when the user wants to benchmark on EHR-Bench, MIMIC-IV-CDM, EHRSHOT, or asks about evaluating this task. Reports F1 score, AUROC.

researchpythongo
0
3
Ehrscl 2024 EvalA

Evaluates a system's ability to translate natural language clinical questions into executable SQL and retrieve accurate results from a specialized electronic health record database. It probes complex temporal reasoning, clinical constraint handling, and semantic equivalence in text-to-SQL generation. Use when the user wants to benchmark on EHRSQL 2024, or asks about evaluating this task. Reports execution accuracy.

researchpythonsql
0
3
Eicap Bench EvalA

Evaluates large language models' emotional intelligence (EI) capabilities across a four-layer taxonomy: emotional tracking, cause inference, appraisal, and emotionally appropriate response generation. It probes fine-grained subcategories including cultural sensitivity, valence judgment, and uncertainty calibration using multi-turn conversational contexts. Use when the user wants to benchmark on EICap-Bench, or asks about evaluating this task. Reports macro-average accuracy.

researchpythongo
0
3
Eicu Crd Clinical Bench EvalA

Evaluates machine learning models on four critical care prediction tasks using the multi-centre eICU-CRD dataset: in-hospital mortality, remaining length of stay, patient phenotyping, and physiologic decompensation. It probes the models' ability to handle longitudinal clinical data, compare categorical vs numerical feature representations, and generalize across multi-centre settings. Use when the user wants to benchmark on eICU-CRD, or asks about evaluating this task. Reports AUROC.

researchpythongo
0
3
Elastic Scaling EvalA

Evaluates a deep learning job scheduler's ability to dynamically adjust GPU allocations and batch sizes to maximize cluster throughput and minimize job completion times. It probes how well the system handles compute-bound, communication-bound, and non-elastic workloads under varying job arrival patterns. Use when the user wants to benchmark on CIFAR100, Food101, or asks about evaluating this task. Reports SJS Efficiency.

researchpythongo
0
3