All categories
Research
Research, evidence gathering, literature, reports, investigation, and synthesis
- 21,377
- 891
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browserBrowse research skills
Showing 9,817–9,840 of 21,377 skills
- Adversarial Text Attack EvalEvaluates the robustness of BERT-based text classifiers against word-level adversarial attacks by measuring how well perturbed inputs maintain semantic meaning and syntactic structure while successfully flipping model predictions. It compares three attack methods across three standard classification benchmarks to determine the optimal balance between attack success, semantic preservation, and computational efficiency. Use when the user wants to benchmark on IMDB, AG News, SST2, or asks about ...Votes: 0GitHub stars: 3
- Adversarial Rc EvalThis evaluation probes a model's ability to answer reading comprehension questions under adversarial conditions, specifically testing generalization across datasets constructed by progressively stronger language models. It measures how well models trained on challenging, model-in-the-loop generated questions can handle both adversarial and standard benchmarks. Use when the user wants to benchmark on SQuAD, BiDAF-adversarial, BERT-adversarial, RoBERTa-adversarial, DROP, Natural Questions, or a...Votes: 0GitHub stars: 3
- Adversarial Ood Robustness EvalEvaluates the adversarial and out-of-distribution (OOD) robustness of LLMs across sentiment analysis, natural language inference, and domain-specific classification tasks. It measures how well models maintain performance under adversarial attacks and distribution shifts, and tests the effectiveness of prompt-based enhancement strategies (AHP and ICR). Use when the user wants to benchmark on PromptRobust (SST-2), AdvGlue++, FlipKart, DDXPlus, or asks about evaluating this task. Reports F1.Votes: 0GitHub stars: 3
- Adversarial Nli EvalEvaluates natural language inference models on adversarially crafted examples designed to expose reasoning brittleness and spurious pattern reliance. It probes whether models can generalize to novel, difficult inference cases that specifically target known model weaknesses across iterative rounds of human-and-model-in-the-loop data collection. Use when the user wants to benchmark on ANLI, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3
- Adversarial Nibbler EvalThis benchmark evaluates the robustness of text-to-image models against implicitly adversarial prompts—subtle, non-obvious text inputs that bypass automated safety filters to generate harmful images. It probes the gap between human safety perception and machine safety classification, highlighting long-tail failure modes and context-dependent vulnerabilities in generative AI. Use when the user wants to benchmark on Nibbler, or asks about evaluating this task. Reports false_negative_rate.Votes: 0GitHub stars: 3
- Adversarial Defense EvalEvaluates the robustness, seamlessness, and general utility of LLMs against adversarial inputs (jailbreaks, toxicity, hallucinations, bias) using an inference-time defense framework. Use when the user has predictions and gold and needs to compute robustness score.Votes: 0GitHub stars: 3
- Advbench Asr EvalThis benchmark probes an LLM's susceptibility to jailbreak attacks by measuring how often it generates harmful or policy-violating responses when prompted with malicious objectives. It evaluates both the raw success rate of bypassing safety filters and the relative severity of the generated harmful content through pairwise ranking. Use when the user wants to benchmark on AdvBench, or asks about evaluating this task. Reports Attack Success Rate (ASR).Votes: 0GitHub stars: 3
- Adult EvalThis benchmark evaluates income prediction models for fairness regarding demographic attributes like race and gender. It probes prediction stability under demographic perturbations (individual fairness) and measures equity in true positive rates across protected groups (group fairness). Use when the user wants to benchmark on Adult, or asks about evaluating this task. Reports Balanced Accuracy (BA).Votes: 0GitHub stars: 3
- Adte Tta EvalEvaluates test-time adaptation (TTA) capabilities of vision-language models under distribution shift and class imbalance. It probes how well a model can adapt to out-of-distribution and cross-domain image classification tasks without training, using adaptive entropy-based uncertainty estimation to select confident augmented views. Use when the user wants to benchmark on ImageNet & Cross-Domain Benchmarks, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3
- Ads Violation Cause EvalEvaluates an automated root-cause analysis tool for autonomous driving systems by measuring its ability to correctly identify the faulty component and the specific output message that caused a driving violation in simulation. It also measures the debugging scope reduction and computational efficiency of the tool. Use when the user wants to benchmark on ADS Violation Cause Benchmark, or asks about evaluating this task. Reports component-level success.Votes: 0GitHub stars: 3
- Adrd Bench EvalEvaluates LLMs on domain-specific knowledge and clinical reasoning for Alzheimer's Disease and Related Dementias (ADRD), as well as practical daily caregiving scenarios. It probes both factual recall and error detection capabilities in a medical context. Use when the user wants to benchmark on ADRD-Bench, or asks about evaluating this task. Reports exact match accuracy.Votes: 0GitHub stars: 3
- Adp EvalEvaluates the performance of LLM agents fine-tuned with the Agent Data Protocol (ADP) across software engineering, web browsing, OS/database tool use, and general reasoning tasks. Use when the user wants to benchmark on SWE-Bench Verified, WebArena, AgentBench, GAIA, or asks about evaluating this task. Reports unit test pass rate.Votes: 0GitHub stars: 3
- Adni Fl EvalEvaluates the performance of federated learning algorithms for binary classification of Alzheimer's disease versus normal controls using structural MRI-derived features. It probes how well FL methods handle non-IID data distributions and domain shifts across different scanner parameters (1.5T vs 3.0T) while preserving data privacy. Use when the user wants to benchmark on ADNI, or asks about evaluating this task. Reports ACC.Votes: 0GitHub stars: 3
- Admiere EvalEvaluates a model's ability to understand and represent multimodal idiomaticity by ranking images based on their alignment with a given context sentence containing a nominal compound. It probes vision-language model alignment, figurative language reasoning, and the capacity to distinguish between literal and idiomatic senses. Use when the user wants to benchmark on AdMIRe, or asks about evaluating this task. Reports Top Image Accuracy.Votes: 0GitHub stars: 3
- Admeood EvalEvaluates the out-of-distribution generalization of drug property prediction models under two specific domain shifts: noise-level-based confidence categorization (Noise Shift) and inconsistent labels across sources (Concept Conflict Drift). It probes whether models can maintain predictive performance when trained on in-distribution data and tested on molecular domains with shifted label distributions or conflicting assay results. Use when the user wants to benchmark on ADMEOOD, or asks about ...Votes: 0GitHub stars: 3
- Admedtagger Medical EvalEvaluates the ability of lightweight BERT-based models to classify Polish medical texts into five clinical categories. The benchmark tests knowledge distillation from a large LLM teacher to smaller classifiers, with ground truth curated by medical experts. Use when the user wants to benchmark on ADMEDTAGGER Physician-Validated Test Sets, or asks about evaluating this task. Reports F1 score.Votes: 0GitHub stars: 3
- Adiabatic Quantum Benchmark EvalEvaluates the performance of adiabatic quantum optimization on complex network analysis tasks. It benchmarks quantum annealing against classical methods on Chimera Ising spin glass instances, independent set problems, planted-solution instances, and community detection. Use when the user wants to benchmark on Chimera Ising spin glass instances, Independent set problems, Planted-solution instances, Community detection, or asks about evaluating this task. Reports D-Wave run-time estimation.Votes: 0GitHub stars: 3
- Ader Sr EvalEvaluates continual learning performance for session-based recommendation by measuring how well a model maintains prediction accuracy on historical items while adapting to new sessions over time. It probes stability-plasticity trade-offs by averaging recommendation quality across multiple sequential update cycles. Use when the user wants to benchmark on DIGINETICA, YOOCHOOSE, or asks about evaluating this task. Reports Recall@k.Votes: 0GitHub stars: 3
- Adept Prosody Clone EvalEvaluates a zero-shot multispeaker TTS model's ability to clone both speaker voice and fine-grained prosody from untranscribed reference audio. It measures intelligibility, spectral/prosodic fidelity, and perceptual similarity against human references. Use when the user wants to benchmark on ADEPT, or asks about evaluating this task. Reports Phone Error Rate (PER).Votes: 0GitHub stars: 3
- Ade20k Scene Parse EvalEvaluates a model's ability to perform dense pixel-wise semantic segmentation across 150 common scene categories, including both discrete objects and amorphous 'stuff' classes. It probes fine-grained scene understanding and the model's capacity to handle class imbalance and varying object scales. Use when the user wants to benchmark on SceneParse150, or asks about evaluating this task. Reports Mean IoU.Votes: 0GitHub stars: 3
- Adcraft EvalEvaluates reinforcement learning agents' ability to optimize bidding strategies and budget allocation in a non-stationary, stochastic Search Engine Marketing (SEM) simulation. It probes how well policies handle sparse feedback, shifting reward landscapes, and long-term profitability constraints over a simulated campaign. Use when the user wants to benchmark on AdCraft Environment, or asks about evaluating this task. Reports NCP.Votes: 0GitHub stars: 3
- Adbench EvalEvaluates tabular anomaly detection models on their ability to identify outliers in medium- and high-dimensional datasets by measuring ranking quality and precision-recall trade-offs under a standardized semi-supervised protocol. Use when the user wants to benchmark on ADBench, or asks about evaluating this task. Reports ROC-AUC.Votes: 0GitHub stars: 3
- Adasum Scaling EvalEvaluates the algorithmic and system efficiency of the Adasum distributed gradient combiner compared to naive gradient averaging across different hardware interconnects and model scales. It probes the ability of synchronous SGD to scale to large effective batch sizes while maintaining convergence accuracy and reducing time-to-accuracy. Use when the user wants to benchmark on ImageNet, SQuAD 1.1, MNIST, or asks about evaluating this task. Reports epochs_to_target_accuracy.Votes: 0GitHub stars: 3
- Adaptmmbench EvalEvaluates Vision-Language Models' ability to dynamically select between text-only and tool-augmented reasoning modes, and assesses the quality, efficiency, and final accuracy of their reasoning processes across multimodal domains. Use when the user wants to benchmark on AdaptMMBench, or asks about evaluating this task. Reports Accuracy.Votes: 0GitHub stars: 3