Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 9,361–9,384 of 21,343 skills
Evaluates approximate nearest neighbor (ANN) indexing methods across four realistic, constrained workloads: filtered search, out-of-distribution data, sparse vectors, and streaming updates. It probes the trade-off between search accuracy (recall) and query throughput under strict memory and time constraints. Use when the user wants to benchmark on Big ANN Challenge Datasets, or asks about evaluating this task. Reports 10-recall@10.
Evaluates a graph-based adversarial domain adaptation model's ability to classify Windows malware and benign binaries under concept drift using minimal labeled samples. It probes the model's robustness to real-world malware evolution by comparing performance against baselines on the Big-15 dataset. Use when the user wants to benchmark on Big-15, or asks about evaluating this task. Reports performance metrics.
Evaluates a model's ability to predict novel drug-disease associations by modeling them as a recommendation task using bidirectional behavioral sequences and prototype spaces. It probes cold-start generalization and robustness to highly sparse interaction data. Use when the user wants to benchmark on Gdataset, Cdataset, LRSSL, or asks about evaluating this task. Reports AUPRC.
Evaluates open-source bibliographic reference and citation parsers on their ability to extract structured metadata fields (author, source, year, volume, issue, page, organization) from raw citation text. Compares machine learning-based versus rule-based approaches, and assesses the impact of domain-specific retraining. Use when the user wants to benchmark on Unspecified bibliographic dataset, or asks about evaluating this task. Reports F1.
Evaluates the robustness and sensitivity of multimodal large language models (MLLMs) to perturbations in spoken multiple-choice questions. It probes how models handle variations in language, accent, speaker gender, and answer option ordering, measuring both absolute correctness and prediction stability across conditions. Use when the user wants to benchmark on BiasInEar, or asks about evaluating this task. Reports Question Entropy.
Evaluates text-to-image models for multi-dimensional social biases across demographic attributes (sex, race, age) by measuring implicit distributional divergence, explicit instruction-following accuracy, and whether bias manifests as ignorance or discrimination. Use when the user wants to benchmark on BiasIG, or asks about evaluating this task. Reports Implicit Bias Score ($S_{sum}$).
Evaluates how weight-activation quantization affects model capabilities, stereotypes, fairness, toxicity, and sentiment across demographic subgroups. It probes whether aggressive compression amplifies historical bias, disparate outcomes, and inter-subgroup disparities in generated text. Use when the user wants to benchmark on MMLU, RedditBias, WinoBias, DiscrimEval, DT-Fairness, BOLD, StereoSet, or asks about evaluating this task. Reports MMLU accuracy.
Evaluates whether text classification models exhibit unintended demographic bias by analyzing how model confidence scores are distributed across different identity groups. It probes the model's ability to rank toxic vs. non-toxic content fairly and detect systematic score shifts that threshold-dependent metrics might miss. Use when the user wants to benchmark on Synthetic Bias Test Set, Human-Labeled Online Comments, or asks about evaluating this task. Reports Subgroup AUC, BPSN AUC, BNSP AUC...
This benchmark evaluates occupation prediction models for fairness regarding gender bias. It tests whether classifiers maintain consistent predictions when gender pronouns are swapped (individual fairness) and whether true positive rates are balanced between male and female bios across occupations (group fairness). Use when the user wants to benchmark on Bias in Bios, or asks about evaluating this task. Reports Balanced Accuracy (BA).
Evaluates a pipeline for detecting and mitigating representation bias and explicit stereotypes in text corpora. It measures how well the pipeline generates attribute-specific word lists, quantifies demographic representation imbalances, and identifies stereotypical language compared to human annotations and baselines. Use when the user wants to benchmark on Small Heap, Small Heap Neutral, StereoSet (filtered), Small Heap Annotated, or asks about evaluating this task. Reports DR score.
Probes a model's ability to detect stereotypical and biased language in text, distinguishing between stereotypical and anti-stereotypical variants, and classifying sentences as biased or unbiased across specific social bias categories. Use when the user wants to benchmark on CrowS-Pairs, BABE, or asks about evaluating this task. Reports Stereotype Score (SS), F1-score.
This evaluation protocol assesses the effectiveness of individual and joint bias mitigation strategies across toxicity detection and word embeddings. It probes whether debiasing for one social identity correlates with or affects bias levels in others, and measures the trade-off between bias reduction and model utility. Use when the user wants to benchmark on Jigsaw Toxicity Dataset, CoNLL 2003, or asks about evaluating this task. Reports AUC.
Evaluates the impact of the Bianet parallel corpus on Neural Machine Translation performance for English-Turkish and English-Kurdish language pairs in the news domain. It compares baseline models trained on existing corpora against models augmented with Bianet data, and assesses multilingual transfer learning benefits. Use when the user wants to benchmark on WMT2016, Bianet, SETIMES, Ubuntu & GNUME, or asks about evaluating this task. Reports BLEU.
Evaluates multilingual machine translation and related sequence-to-sequence tasks across 36 Indian subcontinent languages. It probes the model's ability to handle morphological complexity, script diversity, code-mixing, and domain-specific adaptation through reference-based and reference-free metrics. Use when the user wants to benchmark on FLORES + IN22, Reserved Development Corpora, or asks about evaluating this task. Reports BLEU.
Evaluates large language models on Southeast Asian linguistic and cultural capabilities, probing syntax, semantics, pragmatics, coreference resolution, and scalar implicatures in Indonesian and Tamil. Use when the user wants to benchmark on BHASA LINDSEA (Indonesian & Tamil), or asks about evaluating this task. Reports accuracy.
Evaluates AI-generated scientific paper reviews across five dimensions: content faithfulness, argumentative alignment, focus consistency, question constructiveness, and AI-likeness. It probes whether models can replicate human-like evaluative reasoning and scoring rather than merely predicting scalar ratings. Use when the user wants to benchmark on AI Review Benchmark, or asks about evaluating this task. Reports MAE.
A meta-evaluation framework that scores AI benchmarks across four lifecycle stages to assess their quality, reproducibility, and usability. It evaluates how well benchmarks are designed, implemented, documented, and maintained for both foundation and non-foundation models. Use when the user has predictions and gold and needs to compute lifecycle_score.
Evaluates language models' ability to classify sentiment and detect sarcasm across three distinct varieties of English (Australian, Indian, and British). It probes cross-variety generalization and the impact of domain (Google reviews vs. Reddit comments) on model performance. Use when the user wants to benchmark on BESSTIE, or asks about evaluating this task. Reports F-Score.
Evaluates a model's ability to predict the next item in a user's sequential interaction history using bidirectional context. It probes how well the model captures long-range sequential dependencies and user preferences from implicit feedback. Use when the user wants to benchmark on Amazon Beauty, Steam, MovieLens 1m, MovieLens 20m, or asks about evaluating this task. Reports HR@10.
Evaluates BERT's robustness to synthetic character-level noise across sentiment classification and textual similarity tasks. It probes how spelling mistakes and typos disrupt subword tokenization and degrade contextual embeddings under varying noise intensities. Use when the user wants to benchmark on IMDB, SST-2, STS-B, or asks about evaluating this task. Reports F1 score.
Probes the robustness of Automatic Speech Recognition (ASR) and Speech Emotion Recognition (SER) models under challenging real-world conditions, including varying distances, physical obstructions, and high-intensity vocalizations. It specifically tests whether models can maintain accuracy when linguistic context is removed via nonsense phrases and when acoustic features are degraded by far-field recording and shouting. Use when the user wants to benchmark on BERSt, or asks about evaluating th...
Evaluates the bit error rate (BER) performance of a Reconfigurable Intelligent Surface (RIS) aided spatial media-based modulation system compared to baseline schemes (SM, MBM, QSM) under uncorrelated Rayleigh fading channels. Use when the user has predictions and gold and needs to compute Bit Error Rate (BER).
Evaluates large language models' ability to perform moral reasoning and align with human ethical judgments within Bengali language and South Asian socio-cultural contexts. It probes cultural grounding, commonsense reasoning, and fairness across five everyday moral domains using native-speaker consensus annotations. Use when the user wants to benchmark on BengaliMoralBench, or asks about evaluating this task. Reports accuracy.
Evaluates automatic speech recognition (ASR) accuracy and speaker diarization performance on long-form Bengali speech. It measures phonetic robustness and computational efficiency using public and private test splits under strict hardware constraints. Use when the user wants to benchmark on Lipi-Ghor-882, or asks about evaluating this task. Reports WER.