All categories
Research
Research, evidence gathering, literature, reports, investigation, and synthesis
- 21,376
- 891
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browserBrowse research skills
Showing 9,457–9,480 of 21,376 skills
- Beard EvalEvaluates the adversarial robustness of models trained on synthetically distilled datasets. It probes how well different dataset distillation methods preserve model resilience against diverse adversarial attacks across varying image-per-class (IPC) settings. Use when the user wants to benchmark on CIFAR-10, CIFAR-100, TinyImageNet, or asks about evaluating this task. Reports Comprehensive Robustness-Efficiency Index (CREI).Votes: 0GitHub stars: 3
- Beans Zero EvalEvaluates zero-shot generalization of audio-language models on bioacoustic tasks, including species classification, multilabel detection, call-type prediction, lifestage classification, captioning, and individual counting across diverse taxa. Use when the user wants to benchmark on BEANS-Zero, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3
- Beads EvalThis benchmark evaluates language models across multiple tasks to detect, quantify, and mitigate demographic and social biases. It probes classification accuracy for bias/toxicity/sentiment, token-level bias identification, demographic stereotype alignment, and the ability to generate neutral, benign text variants. Use when the user wants to benchmark on BEADs, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3
- Bcs Dbt Classification EvalEvaluates the ability of self-supervised contrastive pre-training and multi-patch fine-tuning to classify imbalanced digital breast tomosynthesis (DBT) slices and volumes as normal or abnormal. It probes the model's robustness to extreme class imbalance and its capacity to preserve spatial resolution through patch-level processing. Use when the user wants to benchmark on BCS-DBT, or asks about evaluating this task. Reports AUC.Votes: 0GitHub stars: 3
- Bbsard EvalEvaluates the ability of retrieval models to match statutory article questions to the correct legal articles in Dutch and French. It benchmarks both zero-shot dense/lexical models and fine-tuned language-specific models on a parallel bilingual dataset. Use when the user wants to benchmark on bBSARD, or asks about evaluating this task. Reports R@k, MAP@k, MRR@k, nDCG@k.Votes: 0GitHub stars: 3
- Bbq EvalEvaluates social bias in question-answering models by measuring accuracy and a bias score across ambiguous and disambiguated contexts. It probes whether models rely on stereotypes when context is under-informative and whether correct answers align with harmful biases. Use when the user wants to benchmark on BBQ, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3
- Bbh Prompting EvalTests the impact of prompt structure and logical validity on language model reasoning performance. Specifically, it compares answer-only, standard chain-of-thought, and logically invalid chain-of-thought prompting strategies on complex reasoning tasks. Use when the user wants to benchmark on BIG-Bench Hard, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3
- Bbh Population Synthesis EvalEvaluates whether simulated gravitational-wave observations (detection rates and chirp mass distributions) can distinguish between different compact binary population synthesis models of binary black hole formation under realistic detector sensitivities and observing durations. Use when the user wants to benchmark on Simulated aLIGO O1/O2 BBH detections, or asks about evaluating this task. Reports posterior probability.Votes: 0GitHub stars: 3
- Bbh Mmlu Predictability EvalThis protocol evaluates how well aggregate and per-task benchmark performance can be predicted from scaled compute using scaling law fits. It probes the monotonicity and predictability of LLM capabilities across compute scaling, distinguishing between stable scaling trends and emergent or non-monotonic behaviors. Use when the user wants to benchmark on BIG-Bench Hard (BBH), MMLU, or asks about evaluating this task. Reports mean absolute error.Votes: 0GitHub stars: 3
- Bbh Host Identification EvalEvaluates the feasibility of identifying host galaxies for binary black hole mergers using next-generation gravitational wave detector networks by comparing estimated localization volumes against theoretical stellar mass and metallicity thresholds. Use when the user wants to benchmark on Grid I: Galaxy Catalogue Injections, Grid II: Maximum & Minimum Sky Sensitivity Injections, or asks about evaluating this task. Reports localization_volume.Votes: 0GitHub stars: 3
- Bbh EvalEvaluates a model's zero-shot in-context learning capability on reasoning-heavy multiple-choice tasks. It compares self-generated demonstrations against direct prompting and chain-of-thought baselines to measure accuracy gains. Use when the user wants to benchmark on BIG-Bench Hard (BBH), or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3
- Bayling2 Multilingual EvalEvaluates multilingual translation quality and cross-lingual reasoning across high-resource and low-resource languages. It probes the model's ability to align languages and transfer capabilities from high-resource to low-resource settings without extensive low-resource instruction data. Use when the user wants to benchmark on Flores-101, WMT22, Belebele, XNLI, GSM8K, or asks about evaluating this task. Reports BLEU (sacrebleu), COMET.Votes: 0GitHub stars: 3
- Bayesian Optimization EvalEvaluates the ability of Bayesian optimization methods to efficiently search discrete spaces (molecules, arithmetic expressions) by maximizing or minimizing a black-box objective function over a limited budget of oracle calls. It probes how well a model aligns its latent representation with the objective landscape to guide search. Use when the user wants to benchmark on Guacamol, TDC DRD3, Arithmetic Expression, or asks about evaluating this task. Reports objective value.Votes: 0GitHub stars: 3
- Bayesian Optical Flow EvalEvaluates a Bayesian statistical inversion method for estimating optical flow fields and quantifying their uncertainty from image pairs, compared against deterministic baselines. Use when the user wants to benchmark on Synthetic benchmark flow fields, Middlebury dataset, or asks about evaluating this task. Reports reconstruction accuracy.Votes: 0GitHub stars: 3
- Bayes Factor Odds RatioEvaluates the likelihood of different binary black hole formation channels (CEE, CHE, SMT) given gravitational wave strain data by comparing Bayesian evidence and prior odds. Use when the user has predictions and gold and needs to compute Bayes factor ($\mathcal{B}$), Odds ratio ($\mathcal{O}$).Votes: 0GitHub stars: 3
- Battleship EvalEvaluates EFCE solvers on a parametric sequential conflict-resolution game where players place ships and fire shots. It probes the solver's ability to construct incentive-compatible correlation plans that maximize social welfare through deterrence and punishment mechanisms. Use when the user wants to benchmark on Battleship, or asks about evaluating this task. Reports Social Welfare (SW).Votes: 0GitHub stars: 3
- Battery Swap Scheduling EvalEvaluates a genetic algorithm enhanced with an LRU strategy for estimating battery swap demand and optimizing 24-hour charging schedules. It probes the algorithm's ability to minimize charging costs while maintaining high user satisfaction and computational efficiency under real-world demand fluctuations. Use when the user wants to benchmark on ST-EVCDP series, UrbanEV series, or asks about evaluating this task. Reports optimization rate (r_opt).Votes: 0GitHub stars: 3
- Batonvoice EvalEvaluates a controllable text-to-speech model's ability to generate intelligible speech and accurately convey specific emotional tones based on text instructions. It probes zero-shot cross-lingual generalization and instruction-following capabilities in speech synthesis. Use when the user wants to benchmark on Seed-TTS, Emotion dataset, or asks about evaluating this task. Reports Emotion Classification Accuracy.Votes: 0GitHub stars: 3
- Baton EvalThis benchmark evaluates a model's ability to understand coarse driving actions and predict bidirectional control transitions between human drivers and automated driving systems. It probes multimodal fusion capabilities by testing whether models can leverage synchronized video, vehicle telemetry, and route context to forecast handovers and takeovers under varying time horizons. Use when the user wants to benchmark on BATON, or asks about evaluating this task. Reports Accuracy, AUPRC.Votes: 0GitHub stars: 3
- Bat Event Flow EvalEvaluates the accuracy and robustness of event-based optical flow estimation models. It probes the model's ability to predict dense 2D motion fields from sparse, asynchronous event streams, handling varying temporal resolutions and occlusions. Use when the user wants to benchmark on DSEC-Flow, MVSEC, or asks about evaluating this task. Reports EPE.Votes: 0GitHub stars: 3
- Bass EvalThis benchmark evaluates audio language models on musical understanding and semantic reasoning across four domains: structural segmentation, lyric transcription, musicological analysis, and artist collaboration. It probes the model's ability to process long-form audio and perform temporal, hierarchical, and musicological reasoning over vocal and structural attributes. Use when the user wants to benchmark on BASS, or asks about evaluating this task. Reports IWER (Normalized Word Error Rate).Votes: 0GitHub stars: 3
- Basqueglue EvalThis benchmark evaluates Basque language models across classical NLP tasks including topic classification, stance detection, coreference detection, and natural language inference. It measures both task-specific performance and overall linguistic competence in a low-resource agglutinative language setting. Use when the user wants to benchmark on BasqueGLUE, or asks about evaluating this task. Reports Avg.Votes: 0GitHub stars: 3
- Basque Multimodal EvalEvaluates multimodal large language models on close-ended visual question answering and open-ended generation tasks in Basque and English. It probes visual reasoning, language proficiency, and the impact of training data composition and backbone LLM choice on low-resource language performance. Use when the user wants to benchmark on VQAv2, A-OKVQA, PixMoCapQA, BertaQA, Wildvision, or asks about evaluating this task. Reports Accuracy.Votes: 0GitHub stars: 3
- Baskethar EvalEvaluates multimodal human activity recognition capabilities in basketball training scenarios by classifying complex dynamic movements from synchronized physiological, inertial, and video sensor data. Use when the user wants to benchmark on BasketHAR, or asks about evaluating this task. Reports F1-score.Votes: 0GitHub stars: 3