All authors
qhjqhj00 avatar

Claude Skills by qhjqhj00

github.com/qhjqhj00
7,636 skillsA× 7,623B× 11C× 1D× 10 installs2,206 views
Openie6 EvalA

Evaluates Open Information Extraction systems on their ability to accurately identify and extract relational triples from natural language sentences. It measures precision, recall, and F1 using multiple reference-matching protocols, alongside throughput speed and confidence-threshold robustness (AUC). Use when the user wants to benchmark on CaRB, or asks about evaluating this task. Reports F1.

researchpythongo
0
3
Opening EvalA

Evaluates the quality of open-ended interleaved image-text generation by comparing model outputs against human-annotated references. It probes multimodal coherence, visual fidelity, and text-image alignment through pairwise battles and automated judge agreement. Use when the user wants to benchmark on OpenING, or asks about evaluating this task. Reports agreement.

researchpythonperformance
0
3
Openmedia EvalA

This benchmark evaluates the performance and cross-framework compatibility of deep learning algorithms for medical image analysis across classification, segmentation, localization, and detection tasks. It specifically probes how model accuracy and inference efficiency vary when implementations are ported between PyTorch and MindSpore on heterogeneous hardware (NVIDIA GPUs vs. Huawei Ascend NPUs). Use when the user wants to benchmark on OpenMedIA Benchmark Suite, or asks about evaluating this ...

researchpythongo
0
3
Openml Cc18 EvalA

Evaluates machine learning classifiers on a curated collection of standardized classification tasks. It probes the reproducibility and comparability of algorithm performance across diverse datasets under consistent, machine-readable evaluation protocols. Use when the user wants to benchmark on OpenML-CC18, or asks about evaluating this task. Reports accuracy_score.

researchpythongo
0
3
Openmp Energy EvalA

Evaluates the energy efficiency and performance of OpenMP loop transformations (tiling, unrolling) and parallel constructs across different compilers and workloads. Use when the user wants to benchmark on Matrix Multiplication, 2D Stencil, Barcelona OpenMP Task Suite (BOTS), NAS Parallel Benchmarks, PARSEC benchmark, or asks about evaluating this task. Reports Energy (J).

researchpythonc++
0
3
Openner 1.0 EvalA

Evaluates named entity recognition (NER) capabilities across 52 languages and 36 distinct corpora. It probes cross-lingual generalization, robustness to varying entity type ontologies, and the ability of both encoder-based models and LLMs to handle multilingual text with diverse annotation guidelines. Use when the user wants to benchmark on OpenNER 1.0, or asks about evaluating this task. Reports micro-averaged mention-level F1.

researchpythongo
0
3
Openp5 Rec EvalA

This benchmark evaluates the recommendation capability of LLM-based systems on sequential and straightforward recommendation tasks. It probes how well models leverage user interaction histories and different item indexing strategies to predict relevant items across multiple public datasets. Use when the user wants to benchmark on Movielens-1M, Amazon Beauty, LastFM, or asks about evaluating this task. Reports HR@k, NDCG@k.

researchpythongo
0
3
Openpecha BleurtA

Compute openpecha/bleurt via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of openpecha/bleurt.

developmentpython
0
3
Opens2v EvalA

Evaluates Subject-to-Video (S2V) generation models on their ability to maintain subject identity consistency, produce natural temporal dynamics, and align with text prompts. It covers open-domain, human-specific, and single-subject scenarios to expose common failure modes like copy-paste artifacts and fidelity degradation. Use when the user wants to benchmark on OpenS2V-Eval, or asks about evaluating this task. Reports NexusScore, NaturalScore, GmeScore.

researchpythonexpress
0
3
Openscan EvalA

Evaluates 3D vision models' ability to perform open-vocabulary scene understanding by querying for fine-grained object attributes (e.g., material, affordance, synonym) rather than standard object classes. It measures both 3D instance segmentation and 3D semantic segmentation capabilities on attribute-based queries across eight linguistic aspects. Use when the user wants to benchmark on OpenScan, or asks about evaluating this task. Reports AP, mIoU.

researchpythonperformance
0
3
Opensdi EvalA

This benchmark evaluates the ability of vision models to detect and localize diffusion-generated images in an open-world setting. It probes cross-domain generalization across multiple diffusion architectures (SD1.5, SD2.1, SDXL, SD3, Flux.1) and tests robustness against common image degradations like Gaussian blur and JPEG compression. Use when the user wants to benchmark on OpenSDID, or asks about evaluating this task. Reports F1.

researchpythongit
0
3
Opensec EvalA

Probes incident response agent calibration under adversarial prompt injection. It measures how well models distinguish true threats from false positives, resist injection attacks, and execute containment actions without indiscriminately exhausting the available action space. Use when the user wants to benchmark on OpenSec Standard-Tier Episodes, or asks about evaluating this task. Reports Containment rate.

securitypythongo
0
3
Openseeker EvalA

Evaluates the capability of web search agents to perform multi-step navigation, complex deep research planning, and precise information retrieval across English and Chinese web environments. It probes the model's ability to synthesize information from noisy, long-horizon browsing trajectories and extract exact answers or reliable summaries. Use when the user wants to benchmark on BrowseComp, BrowseComp-ZH, xbench-DeepSearch, WideSearch, or asks about evaluating this task. Reports accuracy / F...

researchpythongo
0
3
Opensep EvalA

Evaluates open-world audio source separation by measuring how well a model disentangles multiple audio sources from a mixed input. It probes the model's ability to generalize to seen and unseen audio classes and handle complex natural mixtures without manual intervention. Use when the user wants to benchmark on MUSIC, VGGSound, AudioCaps, or asks about evaluating this task. Reports SDR.

researchpythontesting
0
3
Opensrh Classification EvalA

Evaluates deep learning models on patch-based and patient-level multiclass classification of brain tumor histology images. It probes the ability of CNNs and vision transformers to distinguish between different tumor types and normal tissue using intraoperative stimulated Raman histology data. Use when the user wants to benchmark on OpenSRH, or asks about evaluating this task. Reports top-1 accuracy.

researchpythongit
0
3
Opentom EvalA

This benchmark probes Theory-of-Mind (ToM) reasoning in LLMs by testing their ability to infer psychological mental states (e.g., beliefs, attitudes, intentions) and track physical object locations across naturally generated narratives. It specifically evaluates first- and second-order ToM capabilities under varying narrative lengths and question types. Use when the user wants to benchmark on OpenToM, or asks about evaluating this task. Reports macro-averaged F1 score.

researchpythongo
0
3
Openturingbench EvalA

Evaluates the capability of models to detect machine-generated text and attribute it to specific authors or models across diverse scenarios, including mixed human-machine text, out-of-domain content, and outputs from unseen LLMs. Use when the user wants to benchmark on OpenTuringBench, or asks about evaluating this task. Reports F1-score.

researchpythongo
0
3
Openve Bench EvalA

Evaluates instruction-guided video editing models on spatially-aligned and non-spatially-aligned editing tasks, probing temporal consistency, spatial fidelity, and instruction following. Use when the user wants to benchmark on OpenVE-Bench, or asks about evaluating this task. Reports overall score.

researchpythongo
0
3
Openvid 1m EvalA

Evaluates text-to-video generation models on visual aesthetics, technical quality, text-video alignment, and temporal consistency using a standardized set of 700 prompts. Use when the user wants to benchmark on Liu et al. (2023b) Benchmark, or asks about evaluating this task. Reports VQAA.

researchpythongo
0
3
Openvlthinkerv2 EvalA

Evaluates a multimodal reasoning model's capability across diverse visual tasks, including general and mathematical VQA, document understanding, spatial reasoning, and visual grounding. The protocol tests the model's ability to balance fine-grained perception with multi-step reasoning under a unified RL training framework. Use when the user wants to benchmark on MMMU, MMBench, MMStar, ChartQA, DocVQA, OCRBench, InfoVQA, EmbSpatial, RefSpatial, RoboSpatial, RefCOCO, RefCOCO+, RefCOCOg, or asks...

researchpythongo
0
3
Openxai EvalA

Evaluates the faithfulness, stability, and fairness of post-hoc feature attribution explanation methods (e.g., LIME, SHAP, gradient-based) on tabular datasets to enable reproducible and transparent comparisons. Use when the user wants to benchmark on Popular tabular datasets for XAI and fairness research, or asks about evaluating this task. Reports faithfulness.

researchpythongo
0
3
Ophnet EvalA

Evaluates video understanding models on ophthalmic surgical workflows, including recognizing primary surgery types, classifying hierarchical surgical phases and operations, localizing phase boundaries in time, and anticipating upcoming surgical phases. Use when the user wants to benchmark on OphNet, or asks about evaluating this task. Reports Top-1 Accuracy.

researchpythongo
0
3
Ophthalmic Multimodal EvalA

Evaluates multimodal large language models on clinical ophthalmic image interpretation, specifically diagnosing retinal and macular diseases from fundus photographs and optical coherence tomography (OCT) scans. It probes the models' ability to recognize normal conditions and identify specific pathological states across diverse disease categories. Use when the user wants to benchmark on Ophthalmic Multimodal Benchmark, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Opinion Summarization EvalA

Evaluates a model's ability to synthesize multiple product reviews into a single, coherent opinion summary. It probes the model's capacity to capture diverse aspects, maintain factual accuracy, and adhere to domain-specific stylistic constraints without relying on surface-level lexical overlap. Use when the user wants to benchmark on Amazon, Oposum+, Flipkart, or asks about evaluating this task. Reports human evaluation.

researchpythongo
0
3
Opt Iml Bench EvalA

Evaluates instruction-tuned language models on generalization across 1,991 NLP tasks spanning 100+ categories. It probes zero-shot and few-shot (5-shot) performance on held-out categories, unseen tasks within seen categories, and fully supervised tasks, measuring both generation quality and classification accuracy. Use when the user wants to benchmark on OPT-IML Bench, or asks about evaluating this task. Reports Rouge-L.

researchpythongo
0
3
Optbench EvalA

Evaluates formal theorem proving capabilities specifically within the undergraduate optimization domain. It probes a model's ability to generate syntactically correct and semantically progressive Lean 4 proof steps or full scripts under strict verifier constraints, while measuring robustness against catastrophic forgetting on general math benchmarks. Use when the user wants to benchmark on OptBench, MiniF2F-test, ProofNet-test, or asks about evaluating this task. Reports Pass@32.

researchpythongo
0
3
Optical Flow Epe EvalA

Evaluates the accuracy of predicted optical flow fields against ground truth displacements between consecutive video frames. It probes a model's ability to handle occlusions, non-rigid motion, and large displacements in both synthetic and real-world driving scenarios. Use when the user wants to benchmark on Sintel, KITTI 2012, or asks about evaluating this task. Reports end point error.

researchpython
0
3
Optical Flow Estimation EvalA

Evaluates the accuracy of predicted optical flow fields against ground truth motion vectors between consecutive image frames. It probes a model's ability to estimate dense pixel-wise displacement in both synthetic cinematic scenes and real-world driving environments. Use when the user wants to benchmark on FlyingChairs, Sintel, KITTI12, KITTI15, Middlebury, or asks about evaluating this task. Reports AEE.

researchpython
0
3
Optical Flow EvalA

Evaluates dense optical flow estimation by predicting pixel-wise displacement vectors between consecutive frames. It probes robustness to large motions, occlusions, blur, and atmospheric effects across synthetic and real-world driving scenes. Use when the user wants to benchmark on MPI Sintel, KITTI, Middlebury, or asks about evaluating this task. Reports AEE (Average Endpoint Error).

researchpython
0
3
Optical Flow Kitti Sintel EvalA

This evaluation protocol measures the accuracy of predicted optical flow fields against ground truth displacement vectors across varying motion magnitudes and occlusion conditions. It probes a model's ability to handle both small, fine-grained movements and large, robust displacements using standard video sequence benchmarks. Use when the user wants to benchmark on KITTI2012, KITTI2015, MPI-Sintel, or asks about evaluating this task. Reports Out-Noc, EPE.

researchpythonperformance
0
3
Optical Network Anomaly Detection EvalA

Evaluates the ability of an encoder-decoder LSTM model combined with statistical hypothesis testing to detect unexpected anomalies in optical network quality-of-transmission metrics. It probes whether predicted soft-failure trajectories can distinguish predictable degradation from sudden, anomalous BER deviations in real-time. Use when the user wants to benchmark on Synthetic Optical Network PLM Dataset, or asks about evaluating this task. Reports Accuracy.

datapythongo
0
3
Optiloop 5g Energy EvalA

Evaluates the energy efficiency and operational performance of a 5G network orchestration framework under dynamic traffic conditions. It probes the system's ability to jointly optimize virtual network function (VNF) placement, traffic routing, and network element activation to minimize power consumption while maintaining connectivity and processing capacity. Use when the user wants to benchmark on Real-world mobile operator traffic snapshot, or asks about evaluating this task. Reports energy ...

researchpythonnode
0
3
Optimam Mammography Cad EvalA

Evaluates a deep learning object detection and classification system for identifying malignant lesions in mammographic images. It probes the model's ability to handle high-resolution medical imaging data and diverse lesion morphologies (e.g., microcalcifications, masses) using deformable convolutions. Use when the user wants to benchmark on Optimam, or asks about evaluating this task. Reports AUC.

researchpythongo
0
3
Optimam Mammography EvalA

Binary classification of high-resolution mammography images to detect malignant breast tissue. It probes the model's ability to distinguish between malignant and non-malignant cases using both localized patches and full-resolution inputs. Use when the user wants to benchmark on OPTIMAM, or asks about evaluating this task. Reports AUC.

researchpythonperformance
0
3
Opus 100 Nmt EvalA

Evaluates massively multilingual neural machine translation models on translation quality and language accuracy across 100 languages, including zero-shot translation between unseen language pairs. Use when the user wants to benchmark on OPUS-100, or asks about evaluating this task. Reports BLEU_94.

researchpythongit
0
3
Opus Mt EvalA

Evaluates the translation quality of OPUS-MT models across diverse language pairs using standard automatic metrics and specialized linguistic test suites. It probes general-purpose translation capability, lexical ambiguity disambiguation, and cross-lingual robustness. Use when the user wants to benchmark on Flores, Tatoeba, MuCoW, or asks about evaluating this task. Reports BLEU.

researchpython
0
3
Orad 3d EvalA

Evaluates off-road autonomous driving capabilities across perception, planning, and world modeling. It probes 2D free-space detection, 3D semantic occupancy prediction, GPS-guided trajectory planning, VLM-based scene understanding and path planning, and future video generation in unstructured, variable-terrain environments. Use when the user wants to benchmark on ORAD-3D, or asks about evaluating this task. Reports mIoU.

researchpythongo
0
3
Orak EvalA

Evaluates LLM agents' long-horizon decision-making and gameplay capabilities across 12 diverse video games spanning six genres. It probes the effectiveness of different agentic strategies (zero-shot, reflection, planning, skill-management) and the impact of multimodal inputs (text vs. visual) on action inference. The benchmark also assesses generalization to unseen in-game scenarios, out-of-distribution games, and non-game tasks like math and web navigation. Use when the user wants to benchma...

businesspythonexpress
0
3
Orangesum EvalA

Evaluates abstractive French summarization quality by measuring lexical overlap, semantic similarity, and human judgments of accuracy, informativeness, and fluency. Use when the user wants to benchmark on OrangeSum, or asks about evaluating this task. Reports ROUGE-L.

researchpythongo
0
3
Orb EvalA

Evaluates machine reading comprehension models across diverse linguistic phenomena such as coreference resolution, temporal logic, and causal inference. It tests generalization capabilities by applying synthetic out-of-distribution augmentations to questions, ensuring models rely on reading comprehension rather than external information retrieval. Use when the user wants to benchmark on ORB Benchmark (NewsQA, Quoref, DROP, SQuAD 1.1, SQuAD 2.0, ROPES, DuoRC, NarrativeQA), or asks about evalua...

researchpythongo
0
3
Orbit EvalA

Evaluates recommendation models on candidate item ranking across multiple public sequential recommendation datasets and a large-scale synthetic hidden test (ClueWeb-Reco) to assess generalization to unseen item pools and real-world browsing scenarios. Use when the user wants to benchmark on ML-1M, Amazon Beauty, Amazon Toys, Amazon Sports, Amazon Books, ClueWeb-Reco, or asks about evaluating this task. Reports Recall@10, NDCG@10.

researchpythonperformance
0
3
Orca EvalA

Evaluates Arabic language understanding across seven task clusters, including sentence classification, structured prediction, semantic similarity, NLI, QA, WSD, and topic classification. It probes models' ability to handle diverse Arabic varieties (MSA and dialects) and multiple linguistic levels from tokens to documents. Use when the user wants to benchmark on ORCA, or asks about evaluating this task. Reports ORCA score.

researchpythongo
0
3
Orchid EvalA

Evaluates models on target-independent stance detection (3-way classification) and argumentative dialogue summarization (overall and stance-specific). It probes the ability to classify conflicting viewpoints in Chinese debates and generate concise, faithful summaries aligned with specific stances. Use when the user wants to benchmark on OrChiD, or asks about evaluating this task. Reports Accuracy, ROUGE-1 F1.

researchpythongo
0
3
Orgforge EvalA

This benchmark evaluates RAG and retrieval-augmented agents on synthetic corporate corpora by testing their ability to retrieve artifacts, reason over causal and temporal chains, and detect knowledge gaps. It probes multi-hop reasoning, temporal knowledge-state tracking, and absence-of-evidence detection across structured enterprise artifacts like Slack, JIRA, and emails. Use when the user wants to benchmark on OrgForge Synthetic Corporate Corpus, or asks about evaluating this task. Reports O...

ai-agentspythongo
0
3
Orionbench EvalA

Evaluates unsupervised time series anomaly detection pipelines across diverse real-world and synthetic datasets. It measures detection accuracy for both point and segment anomalies while tracking computational efficiency and model stability over continuous benchmarking cycles. Use when the user wants to benchmark on OrionBench (NASA, NAB, Yahoo S5, UCR), or asks about evaluating this task. Reports F1 score.

researchpythonazure
0
3
Orsi Sod EvalA

This benchmark evaluates optical remote sensing salient object detection models by measuring their ability to accurately segment prominent objects from complex, cluttered backgrounds. It probes structural consistency, boundary precision, and error magnitude across varying object scales and scene complexities. Use when the user wants to benchmark on ORSSD, EORSSD, ORSI-4199, or asks about evaluating this task. Reports maximum F-measure ($F_{\beta}^{max}$).

researchpythonperformance
0
3
Osbad EvalA

Evaluates the ability of statistical and machine learning models to detect anomalies in battery discharge capacity profiles across different chemistries. It probes cross-chemistry generalization and model robustness on imbalanced, rare-anomaly datasets typical of electrochemical systems. Use when the user wants to benchmark on MIT/Stanford (Severson), Tohoku, or asks about evaluating this task. Reports AUROC.

datapythongit
0
3
Osbench EvalA

Evaluates subject-driven image generation and manipulation capabilities, specifically testing identity consistency, prompt adherence, and background preservation across single- and multi-subject scenarios. Use when the user wants to benchmark on OSBench, or asks about evaluating this task. Reports Overall (Generation).

researchpythontesting
0
3
Oscbench EvalA

This benchmark probes a text-to-video model's ability to accurately render and maintain object state transformations (e.g., peeling, slicing) over time. It additionally measures semantic adherence to prompts, scene consistency, and overall perceptual quality to diagnose temporal coherence and physical realism in generated videos. Use when the user wants to benchmark on OSCBench, or asks about evaluating this task. Reports state-change accuracy.

researchpython
0
3
Osmabench EvalA

Evaluates open semantic mapping models' robustness to dynamic indoor lighting and motion conditions. It measures semantic segmentation accuracy and frequency-weighted IoU, alongside LLM-generated visual question answering accuracy on scene graphs. Use when the user wants to benchmark on ReplicaCAD, HM3D, or asks about evaluating this task. Reports mAcc, f-mIoU.

researchpythongo
0
3