All authors
qhjqhj00 avatar

Claude Skills by qhjqhj00

github.com/qhjqhj00
7,636 skillsA× 7,623B× 11C× 1D× 10 installs2,210 views
News Rec EvalA

Evaluates the classification and ranking performance of LLM-based versus deep learning-based news recommendation models. It also measures the diversity of recommended items and how well the recommendations align with individual user history (personalization). Use when the user wants to benchmark on MIND-small, Adressa (one-week), or asks about evaluating this task. Reports AUC.

researchpythongo
0
3
Newsqa EvalA

Probes machine reading comprehension on real-world news articles by requiring models to extract answer spans from context based on natural-language questions. It specifically evaluates the ability to perform reasoning, synthesis, and contextual inference beyond simple keyword matching, while handling multiple valid phrasings for the same answer. Use when the user wants to benchmark on NewsQA, or asks about evaluating this task. Reports F1.

researchpythongo
0
3
Nexar Collision EvalA

Evaluates autonomous ML research agents on their ability to search a mixed categorical-continuous configuration space for optimal model architectures and training setups. It measures convergence speed and final predictive performance on a binary collision prediction task using pre-extracted dashcam video features. Use when the user wants to benchmark on Nexar dashcam collision prediction dataset, or asks about evaluating this task. Reports AP.

researchpythongo
0
3
Next Basket Recommendation EvalA

Evaluates the ability of recommendation algorithms to predict the next basket of items for a user based on their historical purchase sequences. It probes sequential modeling capabilities, handling of item frequency and recency, and robustness across datasets with varying basket lengths and purchase patterns. Use when the user wants to benchmark on TaFeng, Instacart, Dunnhumby, or asks about evaluating this task. Reports Recall@K.

researchpythongo
0
3
Next Transaction Prediction EvalA

Evaluates a model's ability to predict the next user transaction or item interaction based on historical sequential behavior. It probes the model's capacity to capture periodic patterns, user-specific preferences, and generalize across financial and recommendation domains. Use when the user wants to benchmark on WeChat Pay, CCT, MBD-mini, MovieLens-1M, Yelp, or asks about evaluating this task. Reports HR@1.

researchpythongo
0
3
Nfip Flood Loss EvalA

Evaluates machine learning regressors on predicting inter-annual flood economic loss using historical insurance claims and meteorological data. It probes both pointwise prediction accuracy and the fidelity of the predicted loss distribution compared to ground truth, emphasizing temporal generalization over random splits. Use when the user wants to benchmark on NFIP (National Flood Insurance Program), or asks about evaluating this task. Reports R².

researchpythongo
0
3
Nhop L3scoreA

Compute nhop/L3Score via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of nhop/L3Score.

developmentpython
0
3
Nifty Stock Movement EvalA

Evaluates LLMs on predicting short-term stock price movements based on financial news headlines and market context. It probes the model's ability to extract directional sentiment and predictive signals from textual financial data for downstream forecasting tasks. Use when the user wants to benchmark on NIFTY Financial News Headlines Dataset, or asks about evaluating this task. Reports F1 Score.

researchpythongo
0
3
Nih Ap Chest Xray Findings EvalA

Evaluates the ability of deep learning models to perform multi-label classification of 73 fine-grained, sentence-level radiological findings on anterior-posterior (AP) chest X-ray images. It probes whether high-granularity, clinically relevant labels can be effectively learned from limited, semi-automated annotations. Use when the user wants to benchmark on NIH AP Chest X-ray Findings Dataset, or asks about evaluating this task. Reports AUC.

researchpythontesting
0
3
Nih Chest Xray EvalA

Evaluates a model's ability to classify multiple chest X-ray abnormalities and localize them within the image. It probes multi-label disease recognition and spatial localization accuracy under varying strictness thresholds. Use when the user wants to benchmark on NIH Chest X-ray dataset, or asks about evaluating this task. Reports AUC.

researchpythonperformance
0
3
Niletts EvalA

Evaluates the quality of a fine-tuned Text-to-Speech model for Egyptian Arabic dialect synthesis. It measures speech intelligibility, acoustic fidelity, and speaker similarity compared to a baseline model. Use when the user wants to benchmark on NileTTS, or asks about evaluating this task. Reports Word Error Rate (WER).

researchpythongit
0
3
Nim Benchmark EvalA

Evaluates multimodal LLMs' ability to locate and reason about fine-grained details in complex real-world documents. It specifically probes resilience against irrelevant information (distractor images) and measures performance across open- and closed-domain retrieval settings. Use when the user wants to benchmark on ArxiVQA, DUDE, NiM-Benchmark, or asks about evaluating this task. Reports Exact-Match (EM).

researchpythongo
0
3
Nim4 Asr EvalA

Evaluates automatic speech recognition performance across diverse acoustic and linguistic domains, including English, Mandarin, dialects, code-switching, and in-car conversational scenarios. It measures transcription accuracy and hallucination rates to assess model robustness, latency, and customization capabilities. Use when the user wants to benchmark on LibriSpeech, VoxPopuli, MLS-English, AISHELL-1, AISHELL-2, AISHELL-2021-Eval, WeNetSpeech, SpeechIO, WeNetSpeech-Chuan, WeNetSpeech-Yue, K...

businesspythonshell
0
3
Nimaboscarino WeatA

Compute NimaBoscarino/weat via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of NimaBoscarino/weat.

developmentpython
0
3
Ninerrec EvalA

This benchmark evaluates transferable recommendation (TransRec) and modality-only recommendation (MoRec) capabilities. It probes whether models can learn to recommend items using only text and image features without relying on item IDs, and how well they transfer across different domains and modalities. Use when the user wants to benchmark on NineRec (source: Bili_500K, targets: Bili_Food, Bili_Dance, Bili_Movie, Bili_Cartoon, Bili_Music, KU, QB, TN, DY), or asks about evaluating this task. R...

researchpythongit
0
3
Nips 2017 Adv Competition EvalA

Evaluates machine learning models' robustness against adversarial examples across three tracks: generating adversarial perturbations (non-targeted and targeted attacks) and defending against them. It measures both average performance and worst-case vulnerability to imperceptible perturbations. Use when the user wants to benchmark on NIPS 2017 Adversarial Attacks and Defences Competition dataset, or asks about evaluating this task. Reports Score.

researchpythongo
0
3
Nl Code Pair Mining EvalA

Evaluates a machine learning model's ability to automatically extract high-quality, aligned natural language intent and code snippet pairs from Stack Overflow posts. It probes the system's ranking capability, precision, and recall across different programming languages and feature combinations. Use when the user wants to benchmark on Stack Overflow NL-Code Pairs, or asks about evaluating this task. Reports AUC.

researchpythongo
0
3
Nl Object Retrieval EvalA

Evaluates a model's ability to ground natural language queries to specific regions within images by scoring candidate bounding boxes. It probes spatial reasoning, contextual understanding, and cross-modal alignment between text and visual features. Use when the user wants to benchmark on ReferIt, Kitchen, or asks about evaluating this task. Reports P@1.

researchpython
0
3
Nl2gql EvalA

Evaluates a model's capability to translate natural language queries into Graph Query Language (GQL) by measuring syntactic correctness, semantic comprehension, and execution fidelity against a knowledge graph schema. Use when the user wants to benchmark on NL2GQL dataset, or asks about evaluating this task. Reports Execution Accuracy.

researchpythongo
0
3
Nl2sh EvalA

This benchmark evaluates the ability of large language models to translate natural language instructions into executable Bash commands. It probes functional correctness by comparing model-generated commands against ground-truth commands using a functional equivalence heuristic that combines command execution with LLM-based output analysis. Use when the user wants to benchmark on NL2SH, InterCode-ALFA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Nlb21 EvalA

Evaluates latent variable models for their ability to infer neural population dynamics and rates from spiking data. It probes model fidelity in predicting held-out spiking activity, decoding behavioral variables, and capturing autonomous forward dynamics without relying on external labels. Use when the user wants to benchmark on MC_Maze, MC_Maze-L, MC_Maze-M, MC_Maze-S, MC_RTT, Area2_Bump, DMFC_RSG, or asks about evaluating this task. Reports Co-smoothing bps.

researchpythonexpress
0
3
Nlebench Norwegian EvalA

Evaluates generative language models on Norwegian across multiple tasks including conversational dialogue, news summarization, instruction following, document-grounded QA, factual consistency, toxicity, and bias. It probes low-resource language capabilities, cultural understanding, and reasoning via chain-of-thought prompting. Use when the user wants to benchmark on NO-ConvAI2, NO-CNN/DailyMail, NO-Alpaca-Plus, NO-CrowS-Pairs, NO-Multi-QA-Sum, or asks about evaluating this task. Reports BLEU.

researchpythongo
0
3
Nli4pr EvalA

Evaluates whether large language models can correctly determine clinical trial eligibility by performing natural language inference between patient profiles and trial criteria. It probes the model's ability to handle imprecise layman medical terminology compared to precise clinical language in a zero-shot setting. Use when the user wants to benchmark on NLI4PR, or asks about evaluating this task. Reports Macro F1.

researchpythongo
0
3
Nllb Clip Retrieval EvalA

Probes multilingual image-text retrieval capability across low-resource languages. It evaluates how effectively the model aligns visual and textual representations when trained with limited data and frozen encoders. Use when the user wants to benchmark on XTD200, Flickr30k-200, or asks about evaluating this task. Reports R@10.

researchpythongo
0
3
Nllp 2024 Legal Nli EvalA

Probes an LLM's ability to perform Natural Language Inference (NLI) within the legal domain. Specifically, it tests whether the model can correctly classify the logical relationship (entailment, neutral, or contradiction) between a formal legal case summary and an informal social media review. Use when the user wants to benchmark on NLLP 2024 Legal NLI, or asks about evaluating this task. Reports F1.

researchpythongo
0
3
Nlpln TstA

Compute nlpln/tst via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of nlpln/tst.

developmentpython
0
3
Nlt Few Shot Classification EvalA

Evaluates whether few-shot classification methods can adapt to new tasks without using support set labels at test time. It probes the model's ability to cluster or classify query images based solely on support set images and learned representations. Use when the user wants to benchmark on Omniglot, miniImageNet, tieredImageNet, CUB, Meta-Dataset, or asks about evaluating this task. Reports accuracy.

businesspython
0
3
Nlu Service EvalA

Evaluates commercial and open-source NLU platforms on intent classification and named entity recognition across multiple dialogue domains, highlighting limitations in multi-intent support and contextual modeling. Use when the user wants to benchmark on NLU Evaluation Dataset, or asks about evaluating this task. Reports Intent classification accuracy, Entity recognition precision.

researchpythontesting
0
3
Nmt Kd EvalA

Evaluates neural machine translation quality under knowledge distillation settings. It measures how well student models can replicate teacher performance on standard cross-lingual translation benchmarks. Use when the user wants to benchmark on WMT'14 En-De, WMT'14 En-Fr, WMT'16 En-Ro, or asks about evaluating this task. Reports BLEU.

researchpythonperformance
0
3
Nmt Low Resource Indonesian EvalA

Evaluates neural machine translation performance across eight translation directions involving Indonesian and four low-resource Indonesian local languages (Javanese, Sundanese, Minangkabau, Balinese). It probes how different training paradigms (unsupervised, semi-supervised) and data augmentation strategies impact translation quality when parallel data is scarce. Use when the user wants to benchmark on Indonesian Local Language NMT Corpus, or asks about evaluating this task. Reports spm200BLEU.

researchpythonjava
0
3
Nmt Translation EvalA

This evaluation probes a neural machine translation model's ability to translate sentences across multiple language pairs, covering high-resource (English-German, English-French) and low-resource (English-Nepali, English-Sinhala) settings. It measures translation quality using BLEU scores to assess the effectiveness of data diversification strategies without requiring monolingual data or additional parameters. Use when the user wants to benchmark on WMT'14, IWSLT'13/14, Low-resource (Guzmán e...

researchpythontesting
0
3
Nnunet Medical Seg EvalA

Evaluates a self-adapting U-Net framework for medical image segmentation across multiple 3D and 2D tasks. It probes the model's ability to automatically adapt preprocessing, architecture, and training pipelines to achieve robust segmentation performance without manual tuning. Use when the user wants to benchmark on Medical Segmentation Decathlon (Phase 1), or asks about evaluating this task. Reports Dice score.

researchpythontesting
0
3
No2 Prediction EvalA

Evaluates machine learning models' ability to predict ground-level NO2 concentrations in urban areas using multi-source environmental and demographic data. Probes spatial-temporal regression capabilities and model generalization across different cities and time periods. Use when the user wants to benchmark on CityAQVis Urban NO2 Dataset, or asks about evaluating this task. Reports R2 Score.

datapythongit
0
3
Noaa Sst Forecasting EvalA

Evaluates the ability of data-driven models to forecast low-dimensional geophysical dynamics (sea surface temperature and air temperature) from historical time-series observations. It probes long-horizon prediction accuracy, bias-variance trade-offs, and computational efficiency compared to physics-based and deep learning baselines. Use when the user wants to benchmark on NOAA-SST, NOAA-NCEP NAM, or asks about evaluating this task. Reports RMSE.

researchpythongo
0
3
Nobody4 Waf MetricA

Compute nobody4/waf_metric via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of nobody4/waf_metric.

developmentpython
0
3
Node Classification EvalA

This evaluation probes a model's ability to perform node classification on graphs across a spectrum of homophily regimes, from strongly heterophilic to strongly homophilic. It specifically tests the framework's adaptive capability to switch between a combinatorial predictor and a neural refinement stage based on validation performance. Use when the user wants to benchmark on Texas, Cornell, Actor, CiteSeer, Cora, Pubmed, or asks about evaluating this task. Reports classification accuracy.

researchpythongo
0
3
Noisyner EvalA

Evaluates the robustness of noise models and base models under realistic noisy label conditions, measuring how estimation accuracy and base model performance vary with different noise distributions and amounts of clean data. Use when the user wants to benchmark on NoisyNER, or asks about evaluating this task. Reports micro-average F1 score.

researchpythongo
0
3
Noisytoolbench EvalA

This benchmark evaluates how well LLM agents handle ambiguous or unclear user instructions by measuring their ability to ask clarifying questions, execute correct tool calls, and generate accurate final answers. It also assesses interaction efficiency by tracking redundant questions and total action steps. Use when the user wants to benchmark on NoisyToolBench, or asks about evaluating this task. Reports A1.

researchpythongo
0
3
Normalized Mutual Info ScoreA

Compute the normalized_mutual_info_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute normalized_mutual_info_score, or asks how to score with normalized_mutual_info_score.

documentationpython
0
3
NormalizedmutualinfoscoreA

Compute the NormalizedMutualInfoScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute NormalizedMutualInfoScore, or asks how to score with NormalizedMutualInfoScore.

documentationpython
0
3
NormalizedrootmeansquarederrorA

Compute the NormalizedRootMeanSquaredError metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute NormalizedRootMeanSquaredError, or asks how to score with NormalizedRootMeanSquaredError.

documentationpython
0
3
Norne EvalA

Evaluates Named Entity Recognition (NER) performance on Norwegian text, testing the model's ability to identify and classify entity boundaries and types (PER, ORG, LOC, GPE, PROD, EVT, DRV) across Bokmål and Nynorsk variants. Use when the user wants to benchmark on NorNE, or asks about evaluating this task. Reports F1 (strict).

researchpythongo
0
3
Norsumm EvalA

This benchmark evaluates the abstractive summarization capabilities of LLMs on Norwegian news articles. It specifically probes models' ability to generate concise, accurate, and linguistically appropriate summaries in both Bokmål and Nynorsk written variants, while preserving key information and cultural nuance. Use when the user wants to benchmark on NorSumm, or asks about evaluating this task. Reports BERTScore.

researchpythongit
0
3
Norwegian Asr EvalA

Evaluates automatic speech recognition (ASR) models on Norwegian Bokmål and Nynorsk transcriptions, measuring out-of-domain generalization and dialectal robustness across parliamentary and test speech corpora. Use when the user wants to benchmark on NPSC, NST, FLEURS (Norwegian), or asks about evaluating this task. Reports WER.

researchpythonperformance
0
3
Norwegian Nlp Benchmark EvalA

Evaluates contextualized language models on core Norwegian NLP tasks, probing part-of-speech tagging, named entity recognition, sentiment analysis, and negation detection across Bokmål and Nynorsk dialects. Use when the user wants to benchmark on Norwegian Dependency Treebank (NDT), NorNE, NoReC_fine, NoReC_sentence, NoReC_neg, or asks about evaluating this task. Reports Accuracy, strict micro F1.

researchpythongo
0
3
Norwegian Transformer EvalA

Evaluates the cross-lingual transfer and domain adaptation capabilities of a Norwegian BERT model against multilingual and monolingual baselines on token-level (NER/POS) and sequence-level (sentiment/political affiliation) classification tasks. Use when the user wants to benchmark on NorNE, CoNLL-2003, NoReC, Norwegian Parliament Speeches, or asks about evaluating this task. Reports F1 score.

researchpythongo
0
3
Notes Bank EvalA

Evaluates vision-language models on evidence-based visual question answering over unstructured, handwritten scientific notes. The benchmark probes a model's ability to localize relevant visual evidence via bounding boxes, classify content types, and generate natural language answers explicitly grounded in the visual input. Use when the user wants to benchmark on NoTeS-Bank, or asks about evaluating this task. Reports NDCG@5.

researchpythongo
0
3
Noticia EvalA

This benchmark evaluates large language models' ability to interpret misleading clickbait headlines and extract the core information buried in Spanish news articles. It probes the models' capacity for ultra-concise abstractive summarization in a multilingual setting, specifically testing whether they can ignore irrelevant article content and produce brief, accurate summaries. Use when the user wants to benchmark on NoticIA, or asks about evaluating this task. Reports ROUGE-1.

researchpythongo
0
3
Nova EvalA

Evaluates large vision-language models' ability to detect, localize, and reason about rare brain MRI anomalies under extreme clinical and semantic distribution shifts. It probes zero-shot generalization across localization, descriptive captioning, and diagnostic classification without closed-set assumptions. Use when the user wants to benchmark on NOVA, or asks about evaluating this task. Reports Top-1 accuracy.

researchpythongo
0
3
Novel View Extrapolation EvalA

Evaluates a neural radiance field's ability to synthesize high-quality, artifact-free images of solid objects from viewpoints significantly outside the training camera distribution (novel view extrapolation). Use when the user wants to benchmark on Synthetic-NeRF*, MobileObject, or asks about evaluating this task. Reports PSNR.

researchpythontesting
0
3