All authors
qhjqhj00 avatar

Claude Skills by qhjqhj00

github.com/qhjqhj00
7,636 skillsA× 7,623B× 11C× 1D× 10 installs2,103 views
Doc Key Info Extraction EvalA

Evaluates a model's ability to extract fine-grained key information categories (e.g., total price, date, address) from unstructured, template-agnostic document images. It probes robustness to spatial layout variations, dual-modality feature fusion, and generalization to unseen templates and OCR noise. Use when the user wants to benchmark on SROIE, WildReceipt, or asks about evaluating this task. Reports F1 score.

researchpythongo
0
3
Doc2doc Ir EvalA

Evaluates document-to-document information retrieval systems for regulatory compliance, testing their ability to match long, noisy legislative texts to related legal documents. It probes how well models handle extended query lengths, domain-specific vocabulary, and temporal constraints in legal transposition tasks. Use when the user wants to benchmark on EU2UK, UK2EU, or asks about evaluating this task. Reports R@100.

researchpythongo
0
3
Docbank Layout EvalA

Evaluates a model's ability to identify and classify semantic document structures (e.g., sections, figures, equations) from serialized 2D document pages. It probes multimodal layout understanding by measuring how well token-level predictions align with ground-truth semantic units, even when tokens are discontinuous. Use when the user wants to benchmark on DocBank, or asks about evaluating this task. Reports F1 Score.

researchpythongo
0
3
Docfinqa EvalA

Evaluates long-context financial reasoning and context retrieval capabilities. It probes whether models can accurately retrieve relevant sections from lengthy SEC reports and answer numerical questions grounded in those documents. Use when the user wants to benchmark on DocFinQA, or asks about evaluating this task. Reports HR@k.

ai-agentspythongo
0
3
Docgenome EvalA

This benchmark evaluates multi-modal large language models on their ability to parse and understand complex scientific documents. It probes capabilities across document classification, visual grounding of text elements, open-ended single- and multi-page question answering, layout detection, and modality-to-LaTeX transformation. Use when the user wants to benchmark on DocGenome, or asks about evaluating this task. Reports GPT-acc.

ai-agentspythongo
0
3
Dochplt Docmt EvalA

Evaluates document-level machine translation (DocMT) capabilities of LLMs, probing how context length, fine-tuning strategy, and multilingual training affect translation quality across diverse languages and document structures. Use when the user wants to benchmark on DocHPLT, or asks about evaluating this task. Reports BLEU.

researchpythonapi
0
3
Docile 2023 EvalA

Evaluates a model's ability to localize and extract key information fields and line items from diverse business documents. It specifically probes spatial grounding of text values against predefined field types and tests generalization to previously unseen document layouts. Use when the user wants to benchmark on DocILE, or asks about evaluating this task. Reports Average Precision (PCC-based).

researchpython
0
3
Docile EvalA

Evaluates a model's ability to localize and extract key information (KILE) and recognize line items (LIR) in business documents. It probes multi-modal document understanding, specifically token-level classification and bounding box merging based on OCR and layout features. Use when the user wants to benchmark on DocILE, or asks about evaluating this task. Reports F1.

researchpythongo
0
3
Docker Dl Performance EvalA

Evaluates the performance overhead of Docker containers on deep learning workloads by benchmarking CPU, GPU, I/O, and training speed of representative neural networks (FCN, CNN, RNN) across different frameworks. Use when the user wants to benchmark on MNIST, Cifar10, PTB, or asks about evaluating this task. Reports second per batch.

researchpythongo
0
3
Doclaynet Layout EvalA

Evaluates the ability of object detection models to accurately identify and localize 11 distinct document layout elements (e.g., text, tables, figures, headers) on scanned or digital document pages. It measures robustness across diverse, real-world document types and tests how data splitting strategies and label definitions impact prediction accuracy. Use when the user wants to benchmark on DocLayNet, or asks about evaluating this task. Reports mAP@0.5-0.95.

researchpythongit
0
3
Docred EvalA

Evaluates document-level relation extraction systems on multi-sentence reasoning, entity coreference resolution, and long-range dependency modeling. It measures how well models can predict relational facts between entity pairs across entire documents, including cases requiring evidence from multiple sentences. Use when the user wants to benchmark on DocRED, or asks about evaluating this task. Reports F1.

researchpythongo
0
3
Doctor Rec EvalA

This evaluation probes a model's ability to recommend specialist doctors for patients using implicit interaction data and limited demographic metadata. It specifically tests performance in both warm-start (seen patients) and cold-start (new patients) scenarios, emphasizing the model's capacity to handle popularity bias and recommend less popular specialists. Use when the user wants to benchmark on Doctor Recommendation Dataset, or asks about evaluating this task. Reports PS-nDCG@3.

researchpythongo
0
3
Doctorslimm Bangalore ScoreA

Compute DoctorSlimm/bangalore_score via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of DoctorSlimm/bangalore_score.

developmentpython
0
3
Doctorslimm Kaushiks CriteriaA

Compute DoctorSlimm/kaushiks_criteria via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of DoctorSlimm/kaushiks_criteria.

developmentpython
0
3
Document Haystack EvalA

This benchmark evaluates Vision Language Models' ability to retrieve specific textual or multimodal "needles" embedded within long documents ranging from 5 to 200 pages. It probes visual-text alignment, long-context retrieval capabilities, and performance degradation as document length and token consumption increase. Use when the user wants to benchmark on Document Haystack, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Document Ie EvalA

Evaluates the ability of models to extract key information fields from visually-rich documents (receipts and forms). It probes generalization to unseen document templates by comparing performance on original versus resampled test splits. Use when the user wants to benchmark on SROIE, FUNSD, or asks about evaluating this task. Reports F1.

researchpythongo
0
3
Document Parsing EvalA

Evaluates end-to-end document parsing models on their ability to extract structured content (text, formulas, tables, reading order) from both standardized printed documents and real-world captured images. It measures structural fidelity, multilingual robustness, and decoding stability under visual degradation. Use when the user wants to benchmark on OmniDocBench, XFUND, Wild-OmniDocBench, or asks about evaluating this task. Reports Overall.

researchpythongo
0
3
Document Understanding EvalA

Evaluates a model's capability to perform document understanding tasks, including key information extraction, visual question answering, table question answering, and structural reading comprehension, using textual layout representations derived from OCR and spatial verbalization. Use when the user wants to benchmark on DocVQA, InfographicsVQA, WikiTableQuestions, TabFact, SROIE, or asks about evaluating this task. Reports ANLS, accuracy.

researchpythongo
0
3
Dolphin Arabic Nlg EvalA

Evaluates the natural language generation capabilities of models across 13 diverse Arabic tasks, including machine translation, summarization, question generation, and dialectal normalization. It probes how well models handle linguistic variability across Classical Arabic, Modern Standard Arabic, dialects, and Arabizi, as well as cross-lingual and code-switched scenarios. Use when the user wants to benchmark on Dolphin, or asks about evaluating this task. Reports BLEU.

researchpythongo
0
3
Dom EvalA

Evaluates robotic policies on dynamic object manipulation, measuring their ability to react to moving objects, perceive visual/spatial/motion cues, and generalize across novel objects, scenes, and motion patterns. It specifically probes closed-loop reactivity, dynamic adaptation, long-horizon sequencing, and robustness to disturbances. Use when the user wants to benchmark on DOM, or asks about evaluating this task. Reports Success Rate (SR).

researchpythongo
0
3
Domain Generalization EvalA

This benchmark evaluates a model's out-of-distribution (OOD) generalization capability across multiple domain-shift datasets. It specifically probes whether models rely on true domain-invariant features learned from training domains versus leaking test-domain information through ImageNet pretraining weights or oracle hyperparameter selection. The protocol mandates training from scratch without pretrained weights and evaluating across multiple test domains to ensure a fair comparison of OOD ge...

researchpythongo
0
3
Domain Types EvalA

Evaluates the effectiveness and efficiency of software model checking configurations that use domain types to select abstract domains (BDD vs explicit-value) for variable abstraction. It probes how well different abstraction strategies handle verification tasks across various benchmark suites. Use when the user wants to benchmark on SV-COMP and RERS benchmark sets (SYSTEMC, ECA, LOCK, PRODUCT SIMULATOR, NTDRIVERS, SSH), or asks about evaluating this task. Reports Effectiveness.

researchpythongo
0
3
Domainnet EvalA

Evaluates multi-source domain adaptation methods on image classification tasks across multiple domains with varying visual styles and categories. Use when the user wants to benchmark on DomainNet, or asks about evaluating this task. Reports average accuracy.

researchpythongo
0
3
Domainsum EvalA

Evaluates abstractive text summarization models under varying degrees of domain shift (genre, style, topic) to measure performance degradation and distributional divergence across hierarchical granularity levels. Use when the user wants to benchmark on DomainSum, or asks about evaluating this task. Reports ROUGE.

ai-agentspythongit
0
3
Domino EvalA

Evaluates a robot policy's ability to perform manipulation tasks in environments with moving objects and dynamic spatiotemporal changes. It probes the model's capacity for historical context integration and future state anticipation to maintain control stability and task success under motion. Use when the user wants to benchmark on DOMINO@0.1, or asks about evaluating this task. Reports Success Rate (SR).

researchpythongo
0
3
Dora EvalA

Evaluates RAG-based question answering systems on defense-domain documents, measuring both retrieval effectiveness and end-to-end QA performance including task success, faithfulness, and generation quality. Use when the user wants to benchmark on DoRA, or asks about evaluating this task. Reports task-success.

researchpythongo
0
3
Dori EvalA

Evaluates multimodal large language models' ability to understand and reason about object orientation across four dimensions: frontal alignment, rotational transformations, relative directional relationships, and canonical orientation. It distinguishes between coarse categorical judgments and fine-grained angular estimations to probe 3D spatial reasoning capabilities. Use when the user wants to benchmark on DORI, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Dosrecmc Mammography EvalA

Evaluates cross-domain generalization of mammography classification models under domain shift, specifically testing resilience to variations in pixel intensity distributions across different imaging devices and datasets. Use when the user wants to benchmark on NYU, HCTP, VinDr, CSAW, or asks about evaluating this task. Reports PR-AUC.

researchpythongo
0
3
Dotkaio Competition MathA

Compute dotkaio/competition_math via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of dotkaio/competition_math.

developmentpython
0
3
Dove EvalA

This evaluation probes the robustness and prompt sensitivity of large language models on multiple-choice benchmarks by measuring how performance varies across hundreds of millions of intent-preserving prompt perturbations across multiple dimensions. Use when the user wants to benchmark on DOVE, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Dowis EvalA

This benchmark evaluates instruction-following capabilities of speech-large language models (SLLMs) by comparing performance when prompted with text versus spoken audio across nine diverse tasks. It probes cross-lingual generalization, prompt style robustness, and the model's ability to handle both text and speech modalities for input and output. Use when the user wants to benchmark on DOWIS (Do What I Say), FLEURS, MCIF, YTSeg, or asks about evaluating this task. Reports WER.

researchpythongo
0
3
Downstream EvalA

Evaluates the downstream language and symbolic capabilities of LLMs pretrained on filtered web corpora. It probes general knowledge, reasoning, comprehension, and code/math problem-solving to assess how different data filtering strategies impact model performance. Use when the user wants to benchmark on Dolma (v1.6), Pile-github, or asks about evaluating this task. Reports average normalized accuracy.

researchpythongo
0
3
Downstream Scaling EvalA

This evaluation probes how reliably language model scaling laws predict performance in over-trained regimes, where models are trained with significantly more tokens than parameters. It measures both next-token prediction accuracy on a held-out corpus and generalization across a broad suite of downstream zero-shot and few-shot tasks. Use when the user wants to benchmark on C4 eval, LLM-foundry, or asks about evaluating this task. Reports Validation loss.

researchpythongo
0
3
Dp Gnn Graph Classification EvalA

This evaluation protocol assesses the utility and privacy-utility trade-off of Graph Neural Networks trained with Differentially Private Stochastic Gradient Descent (DP-SGD) on graph-level classification tasks. It probes whether formal privacy guarantees can be maintained across diverse graph structures (molecules, fingerprints, ECG signals, synthetic graphs) without severely degrading predictive performance compared to non-private baselines. Use when the user wants to benchmark on Synthetic,...

researchpythonnode
0
3
Dpflow EvalA

Evaluates optical flow estimation models on standard and high-resolution benchmarks to measure accuracy, generalization across input sizes, and robustness to resolution scaling without tiling. It tests how well adaptive pyramid architectures maintain prediction stability when input dimensions increase from 1K to 8K. Use when the user wants to benchmark on Spring, Middlebury-ST, VIPER, Kubric-NK, MPI-Sintel, KITTI 2015, or asks about evaluating this task. Reports EPE.

researchpythonspring
0
3
Dpg Bench EvalA

Evaluates dense prompt following on multi-requirement prompts by decomposing them into dependency-structured VQA checks spanning entity presence, attributes, relations, and counts. Use when the user wants to benchmark on DPG-Bench, or asks about evaluating this task. Reports Overall.

researchpythongo
0
3
Dph Alignment EvalA

Evaluates language models on natural language understanding, commonsense reasoning, and reading comprehension to measure alignment quality and reasoning preservation. It compares standard log-probability predictions against scores derived from a learned Direct Preference Head (DPH) reward model to assess self-evaluation capabilities. Use when the user wants to benchmark on GLUE, RACE, ARC, OpenBookQA, HellaSwag, WinoGrande, BoolQ, PIQA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Dpo Ppo Multi Bench EvalA

This protocol evaluates language models across factual knowledge, mathematical reasoning, instruction following, code generation, truthfulness, and safety/refusal capabilities. It uses standardized benchmarks to measure how preference optimization methods and data quality impact model performance. Use when the user wants to benchmark on MMLU, GSM8k, Big Bench Hard, TruthfulQA, AlpacaEval, IFEval, HumanEval+, MBPP+, ToxiGen, XSTest, or asks about evaluating this task. Reports average accuracy.

researchpythongo
0
3
Dpo Preference EvalA

This evaluation protocol assesses a language model's ability to align with human preferences across open-ended text generation tasks. It measures how well the model optimizes a reward objective while staying close to a reference policy, and evaluates practical performance via pairwise win rates against baselines. Use when the user wants to benchmark on IMDb, Reddit TL;DR, Anthropic HH, or asks about evaluating this task. Reports win rate.

researchpythongo
0
3
Dpr Retrieval EvalA

Evaluates a model's ability to retrieve relevant passages from a large unstructured corpus for open-domain question answering. It probes semantic matching and dense retrieval capabilities by measuring how often the correct answer span appears in the top-k retrieved passages. Use when the user wants to benchmark on Natural Questions, TriviaQA, WebQuestions, CuratedTREC, SQuAD v1.1, or asks about evaluating this task. Reports top-k retrieval accuracy.

researchpythongo
0
3
Dr Aid EvalA

Evaluates an automated framework's ability to encode natural-language data-governance policies into a formal model, extract data-flow graphs from provenance traces, and correctly trigger compliance obligations across decentralized scientific workflows. Use when the user wants to benchmark on Cyclone tracking workflow, MT3D (Moment Tensor in 3D) workflow, Real-world data-governance policies, or asks about evaluating this task. Reports actioning rules.

researchpythongo
0
3
Dr Spider EvalA

Evaluates the robustness of text-to-SQL models against semantic-preserving and semantic-changing perturbations across database schemas, natural language questions, and SQL queries. It measures how well models maintain execution accuracy when inputs are altered, revealing architectural vulnerabilities in entity linking, decoder design, and value prediction. Use when the user wants to benchmark on Dr.Spider, or asks about evaluating this task. Reports execution accuracy (EX).

researchpythongo
0
3
Dragon EvalA

Evaluates evidence-grounded visual reasoning in diagrams by requiring models to localize bounding boxes of supporting visual elements (e.g., labels, axes, connectors) rather than just predicting the correct answer. It probes a model's ability to visually ground reasoning steps within complex diagrammatic contexts and assesses how well localization quality correlates with reasoning performance. Use when the user wants to benchmark on DRAGON, or asks about evaluating this task. Reports Max Pair...

researchpythongo
0
3
Dragon Rag EvalA

Evaluates retrieval and end-to-end performance of RAG systems on a dynamic, daily-updating news corpus. It probes a model's ability to accurately retrieve relevant document chunks and generate factually consistent responses to knowledge-graph-derived queries. Use when the user wants to benchmark on Public Texts, or asks about evaluating this task. Reports ROUGE-L.

ai-agentspythongo
0
3
Drbench EvalA

Evaluates generative models' ability to perform clinical diagnostic reasoning, including medical knowledge representation, evidence synthesis, and diagnosis generation. The benchmark spans sentence-level to full-note tasks to probe abstractive reasoning and clinical knowledge inference. Use when the user wants to benchmark on DR.BENCH, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Drcd EvalA

Evaluates a model's ability to perform span-based machine reading comprehension in traditional Chinese. It probes factual retrieval and exact answer extraction from provided context paragraphs without requiring complex inference or multiple-choice reasoning. Use when the user wants to benchmark on DRCD, or asks about evaluating this task. Reports F1 score.

researchpythongo
0
3
Dre Bench EvalA

Evaluates large language models' fluid intelligence and abstract rule generalization across four hierarchical cognitive levels (Attribute, Spatial, Sequential, Conceptual). It probes the model's ability to dynamically adapt to varying task complexity and apply learned rules to novel grid-based reasoning problems. Use when the user wants to benchmark on DRE-Bench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Dreambench Customization EvalA

Evaluates a text-to-image model's ability to customize generated images with specific subjects from reference images, both individually and in combination, while maintaining alignment with text prompts. Use when the user wants to benchmark on DreamBench, or asks about evaluating this task. Reports CLIP-I.

researchpythonperformance
0
3
Dreambench EvalA

Evaluates subject-driven image generation by measuring how well the model follows text instructions and preserves the reference subject from the source image. It tests the model's ability to extract and reuse specific objects without fine-tuning. Use when the user wants to benchmark on DreamBench, or asks about evaluating this task. Reports CLIP-T.

researchpython
0
3
Dreambench Subject Control EvalA

Evaluates a diffusion model's ability to generate images that preserve subject identity from reference images while adhering to text prompts, covering both single-subject and multi-subject scenarios. Use when the user wants to benchmark on DreamBench, or asks about evaluating this task. Reports DINO.

researchpython
0
3