All categories
Research
Research, evidence gathering, literature, reports, investigation, and synthesis
- 21,376
- 891
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browserBrowse research skills
Showing 9,409–9,432 of 21,376 skills
- Bias Metrics EvalEvaluates whether text classification models exhibit unintended demographic bias by analyzing how model confidence scores are distributed across different identity groups. It probes the model's ability to rank toxic vs. non-toxic content fairly and detect systematic score shifts that threshold-dependent metrics might miss. Use when the user wants to benchmark on Synthetic Bias Test Set, Human-Labeled Online Comments, or asks about evaluating this task. Reports Subgroup AUC, BPSN AUC, BNSP AUC...Votes: 0GitHub stars: 3
- Bias In Bios EvalThis benchmark evaluates occupation prediction models for fairness regarding gender bias. It tests whether classifiers maintain consistent predictions when gender pronouns are swapped (individual fairness) and whether true positive rates are balanced between male and female bios across occupations (group fairness). Use when the user wants to benchmark on Bias in Bios, or asks about evaluating this task. Reports Balanced Accuracy (BA).Votes: 0GitHub stars: 3
- Bias Detection Mitigation EvalEvaluates a pipeline for detecting and mitigating representation bias and explicit stereotypes in text corpora. It measures how well the pipeline generates attribute-specific word lists, quantifies demographic representation imbalances, and identifies stereotypical language compared to human annotations and baselines. Use when the user wants to benchmark on Small Heap, Small Heap Neutral, StereoSet (filtered), Small Heap Annotated, or asks about evaluating this task. Reports DR score.Votes: 0GitHub stars: 3
- Bias Detection EvalProbes a model's ability to detect stereotypical and biased language in text, distinguishing between stereotypical and anti-stereotypical variants, and classifying sentences as biased or unbiased across specific social bias categories. Use when the user wants to benchmark on CrowS-Pairs, BABE, or asks about evaluating this task. Reports Stereotype Score (SS), F1-score.Votes: 0GitHub stars: 3
- Bias Correlation Mitigation EvalThis evaluation protocol assesses the effectiveness of individual and joint bias mitigation strategies across toxicity detection and word embeddings. It probes whether debiasing for one social identity correlates with or affects bias levels in others, and measures the trade-off between bias reduction and model utility. Use when the user wants to benchmark on Jigsaw Toxicity Dataset, CoNLL 2003, or asks about evaluating this task. Reports AUC.Votes: 0GitHub stars: 3
- Bianet Mt EvalEvaluates the impact of the Bianet parallel corpus on Neural Machine Translation performance for English-Turkish and English-Kurdish language pairs in the news domain. It compares baseline models trained on existing corpora against models augmented with Bianet data, and assesses multilingual transfer learning benefits. Use when the user wants to benchmark on WMT2016, Bianet, SETIMES, Ubuntu & GNUME, or asks about evaluating this task. Reports BLEU.Votes: 0GitHub stars: 3
- Bhashaverse Translation EvalEvaluates multilingual machine translation and related sequence-to-sequence tasks across 36 Indian subcontinent languages. It probes the model's ability to handle morphological complexity, script diversity, code-mixing, and domain-specific adaptation through reference-based and reference-free metrics. Use when the user wants to benchmark on FLORES + IN22, Reserved Development Corpora, or asks about evaluating this task. Reports BLEU.Votes: 0GitHub stars: 3
- Bhasa EvalEvaluates large language models on Southeast Asian linguistic and cultural capabilities, probing syntax, semantics, pragmatics, coreference resolution, and scalar implicatures in Indonesian and Tamil. Use when the user wants to benchmark on BHASA LINDSEA (Indonesian & Tamil), or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3
- Beyond Rating EvalEvaluates AI-generated scientific paper reviews across five dimensions: content faithfulness, argumentative alignment, focus consistency, question constructiveness, and AI-likeness. It probes whether models can replicate human-like evaluative reasoning and scoring rather than merely predicting scalar ratings. Use when the user wants to benchmark on AI Review Benchmark, or asks about evaluating this task. Reports MAE.Votes: 0GitHub stars: 3
- Betterbench AssessmentA meta-evaluation framework that scores AI benchmarks across four lifecycle stages to assess their quality, reproducibility, and usability. It evaluates how well benchmarks are designed, implemented, documented, and maintained for both foundation and non-foundation models. Use when the user has predictions and gold and needs to compute lifecycle_score.Votes: 0GitHub stars: 3
- Besstie EvalEvaluates language models' ability to classify sentiment and detect sarcasm across three distinct varieties of English (Australian, Indian, and British). It probes cross-variety generalization and the impact of domain (Google reviews vs. Reddit comments) on model performance. Use when the user wants to benchmark on BESSTIE, or asks about evaluating this task. Reports F-Score.Votes: 0GitHub stars: 3
- Bert4rec EvalEvaluates a model's ability to predict the next item in a user's sequential interaction history using bidirectional context. It probes how well the model captures long-range sequential dependencies and user preferences from implicit feedback. Use when the user wants to benchmark on Amazon Beauty, Steam, MovieLens 1m, MovieLens 20m, or asks about evaluating this task. Reports HR@10.Votes: 0GitHub stars: 3
- Bert Noise Robustness EvalEvaluates BERT's robustness to synthetic character-level noise across sentiment classification and textual similarity tasks. It probes how spelling mistakes and typos disrupt subword tokenization and degrade contextual embeddings under varying noise intensities. Use when the user wants to benchmark on IMDB, SST-2, STS-B, or asks about evaluating this task. Reports F1 score.Votes: 0GitHub stars: 3
- Berst EvalProbes the robustness of Automatic Speech Recognition (ASR) and Speech Emotion Recognition (SER) models under challenging real-world conditions, including varying distances, physical obstructions, and high-intensity vocalizations. It specifically tests whether models can maintain accuracy when linguistic context is removed via nonsense phrases and when acoustic features are degraded by far-field recording and shouting. Use when the user wants to benchmark on BERSt, or asks about evaluating th...Votes: 0GitHub stars: 3
- Ber PerformanceEvaluates the bit error rate (BER) performance of a Reconfigurable Intelligent Surface (RIS) aided spatial media-based modulation system compared to baseline schemes (SM, MBM, QSM) under uncorrelated Rayleigh fading channels. Use when the user has predictions and gold and needs to compute Bit Error Rate (BER).Votes: 0GitHub stars: 3
- Bengalimoralbench EvalEvaluates large language models' ability to perform moral reasoning and align with human ethical judgments within Bengali language and South Asian socio-cultural contexts. It probes cultural grounding, commonsense reasoning, and fairness across five everyday moral domains using native-speaker consensus annotations. Use when the user wants to benchmark on BengaliMoralBench, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3
- Bengali Asr Diarization EvalEvaluates automatic speech recognition (ASR) accuracy and speaker diarization performance on long-form Bengali speech. It measures phonetic robustness and computational efficiency using public and private test splits under strict hardware constraints. Use when the user wants to benchmark on Lipi-Ghor-882, or asks about evaluating this task. Reports WER.Votes: 0GitHub stars: 3
- Bengal Ner El EvalEvaluates the performance of Named Entity Recognition (NER) and Entity Linking (EL) systems on automatically generated corpora. It probes a model's ability to accurately detect entity spans in text and correctly link them to a reference knowledge base (DBpedia) across varying document lengths, entity densities, and languages. Use when the user wants to benchmark on BENGAL (B1-B13, P1-P4, S1-S4), or asks about evaluating this task. Reports micro F1-score.Votes: 0GitHub stars: 3
- Benchx EvalEvaluates Medical Vision-Language Pretraining (MedVLP) models on chest X-ray tasks including multi-label/binary classification, segmentation, report generation, and image-text retrieval. It specifically probes how standardized preprocessing and finetuning strategies affect model performance across heterogeneous architectures. Use when the user wants to benchmark on NIH, VinDr, COVIDx, SIIM, RSNA, Object-CXR, TBX11K, IUXray, MIMIC 5x200, or asks about evaluating this task. Reports AUROC.Votes: 0GitHub stars: 3
- Benchtemp EvalEvaluates the effectiveness and efficiency of Temporal Graph Neural Networks (TGNNs) on link prediction and node classification tasks. It probes model performance across transductive and inductive settings (New-Old/New-New) to ensure fair cross-model comparisons. Use when the user wants to benchmark on BenchTemp (15 datasets), or asks about evaluating this task. Reports AUC.Votes: 0GitHub stars: 3
- Benchread EvalEvaluates retinal anomaly detection models across fundus photography and OCT modalities, testing their ability to distinguish normal from abnormal images and generalize to unseen anomalies under varying supervision levels. Use when the user wants to benchmark on Fundus Benchmark, OCT Benchmark, or asks about evaluating this task. Reports AUC-ROC.Votes: 0GitHub stars: 3
- Benchmd EvalEvaluates modality-agnostic models across 19 real-world medical datasets spanning 1D, 2D, and 3D modalities. Probes performance under data scarcity (few-shot linear evaluation and finetuning) and out-of-distribution generalization across different hospitals and data distributions. Use when the user wants to benchmark on BenchMD, or asks about evaluating this task. Reports AUROC.Votes: 0GitHub stars: 3
- Benchmax EvalBenchMAX evaluates the language-agnostic capabilities of large language models across 17 languages, including non-Latin scripts. It probes instruction following, reasoning, code generation, long-context modeling, tool use, and translation through a rigorously translated and human-post-edited pipeline. Use when the user wants to benchmark on BenchMAX, or asks about evaluating this task. Reports evaluation metrics.Votes: 0GitHub stars: 3
- Benchmark Accuracy EvalEvaluates zero-shot language model performance across a suite of 10 standard NLP benchmarks covering commonsense reasoning, science QA, and language modeling. It measures task accuracy and correlates it with word-level statistical overlap metrics to assess distributional alignment between pre-training data and evaluation sets. Use when the user wants to benchmark on ARC Easy, ARC Challenge, Hellaswag, MMLU, SciQ, OpenBookQA, PIQA, lambada, SocialIQA, SWAG, or asks about evaluating this task. ...Votes: 0GitHub stars: 3