All categories
Research
Research, evidence gathering, literature, reports, investigation, and synthesis
- 21,377
- 891
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browserBrowse research skills
Showing 9,577–9,600 of 21,377 skills
- Attestable Audits EvalEvaluates the feasibility and performance of running standard AI safety benchmarks inside Trusted Execution Environments (TEEs) using quantized models. It probes zero-shot reasoning accuracy, toxicity refusal capabilities, and text generation quality under hardware and cryptographic constraints. Use when the user wants to benchmark on MMLU, ToxicChat, Summarization, or asks about evaluating this task. Reports MMLU Accuracy (%).Votes: 0GitHub stars: 3
- Attentionspan EvalEvaluates algorithmic reasoning and out-of-distribution generalization in Transformers by measuring prediction accuracy and attention pattern alignment against ground-truth reference masks on synthetic tasks. Use when the user wants to benchmark on AttentionSpan, or asks about evaluating this task. Reports Accuracy.Votes: 0GitHub stars: 3
- Attentionddi EvalPredicts whether a pair of drugs interacts based on multi-modal similarity features (targets, pathways, side effects, chemical structure, etc.). Evaluates performance on highly imbalanced drug-drug interaction datasets using precision-recall metrics. Use when the user wants to benchmark on DS1, DS2, DS3 (CYP), DS3 (NCYP), or asks about evaluating this task. Reports AUPR.Votes: 0GitHub stars: 3
- Attention Pruning EvalEvaluates the effectiveness of attention head pruning methods in reducing gender bias in large language models while preserving language model utility. It probes the trade-off between fairness mitigation and general language modeling performance across models of varying sizes. Use when the user wants to benchmark on HolisticBias, WikiText-2, or asks about evaluating this task. Reports HolisticBias.Votes: 0GitHub stars: 3
- Atten Mixer EvalEvaluates a session-based recommendation model's ability to predict the next item in a user's browsing session. It probes the model's capacity to capture multi-level user intent and item semantics through attention mechanisms while handling session-specific inductive biases. Use when the user wants to benchmark on Diginetica, Gowalla, Last.fm, or asks about evaluating this task. Reports HR@K, MRR@K.Votes: 0GitHub stars: 3
- Atrw Reid EvalEvaluates Amur tiger re-identification in the wild by measuring how well models can match tiger identities across different camera views and detection/pose conditions. It probes robustness to non-rigid body deformation, extreme pose variation, and domain shifts between controlled (plain) and uncontrolled (wild) environments. Use when the user wants to benchmark on ATRW, or asks about evaluating this task. Reports mAP.Votes: 0GitHub stars: 3
- Atrbench EvalEvaluates SAR automatic target recognition (ATR) capabilities under realistic, wild conditions. It probes fine-grained vehicle classification and detection robustness across varying imaging geometries, scene complexities, and domain shifts (SOC vs EOC settings). Use when the user wants to benchmark on ATRBench, or asks about evaluating this task. Reports overall accuracy (%).Votes: 0GitHub stars: 3
- Atmospheric Subgrid EvalEvaluates neural network parameterizations for predicting subgrid atmospheric processes (e.g., microphysical tendencies, momentum fluxes) using single-column versus non-local (3x3 grid) inputs. It probes the model's ability to capture mesoscale convective dynamics and frontal systems, and assesses performance across different atmospheric stability regimes. Use when the user wants to benchmark on SAM (Simple Atmospheric Model) simulation, or asks about evaluating this task. Reports R^2.Votes: 0GitHub stars: 3
- Atlas EvalThis benchmark probes frontier scientific reasoning across multiple disciplines (e.g., physics, chemistry, biology, computer science, mathematics) using original, multi-step problems. It evaluates a model's ability to generate complex, open-ended, LaTeX-formatted answers and assesses both solution accuracy and inference stability across multiple sampling runs. Use when the user wants to benchmark on ATLAS, or asks about evaluating this task. Reports Accuracy.Votes: 0GitHub stars: 3
- Atlas Chat EvalEvaluates large language models on Moroccan Arabic (Darija) across multiple-choice reasoning, instruction following, translation, summarization, and sentiment analysis. It probes dialect-specific linguistic features, script variability (Arabic vs. Arabizi), and real-world instruction-following capabilities in a low-resource setting. Use when the user wants to benchmark on DarijaMMLU, DarijaHellaSwag, Belebele_Ary, DarijaBench, DarijaAlpacaEval, or asks about evaluating this task. Reports Accu...Votes: 0GitHub stars: 3
- Atlas Benchmark EvalThis benchmark evaluates human motion prediction algorithms by testing their ability to forecast future trajectories given past observations and environmental context. It systematically probes robustness to perception noise, generalization across different social and cultural environments, and performance under varying observation and prediction horizons. Use when the user wants to benchmark on ETH, ATC, THÖR, or asks about evaluating this task. Reports ADE.Votes: 0GitHub stars: 3
- Athletics Anomaly Detection EvalEvaluates performance-based anomaly detection methods to identify athletes with confirmed anti-doping rule violations. It measures how well different algorithms surface sanctioned athletes while balancing precision and recall, accounting for environmental factors like wind and altitude. Use when the user wants to benchmark on 100 m Sprint, or asks about evaluating this task. Reports F1 score.Votes: 0GitHub stars: 3
- Atg Pvd EvalEvaluates a drone-based suspect-and-investigate system for detecting and classifying illegally parked cars, moving cars, and legally parked cars from aerial imagery. Use when the user wants to benchmark on ATG-PVD, or asks about evaluating this task. Reports mAP.Votes: 0GitHub stars: 3
- Atc Asr Domain Shift EvalThis benchmark evaluates the robustness of self-supervised speech recognition models under domain shift in air traffic control communications. It probes few-shot fine-tuning capabilities, sensitivity to audio quality and accents, and potential gender bias in transcription performance. Use when the user wants to benchmark on NATS, ISAVIA, LiveATC-Test, ATCO2-Test, LDC-ATCC, UWB-ATCC, ATCOSIM, or asks about evaluating this task. Reports WER.Votes: 0GitHub stars: 3
- Atc Asr Csd EvalEvaluates automatic speech recognition (ASR) and call sign detection (CSD) capabilities on real-world air traffic control audio. It probes the model's ability to transcribe accented, noisy pilot and controller speech and accurately identify aircraft call signs in domain-specific phraseology. Use when the user wants to benchmark on Airbus ATC Speech Recognition 2018 Challenge Corpus, or asks about evaluating this task. Reports WER.Votes: 0GitHub stars: 3
- Atbench EvalEvaluates vision-language models on five assistive technology tasks for people with visual impairments: panoptic segmentation, depth estimation, optical character recognition, image captioning, and visual question answering. It probes the model's ability to unify multiple visual understanding and generation tasks within a single parameter set using task-specific prompts. Use when the user wants to benchmark on ADE-150, NYU-V2, OCR (IC13, IC15, SVT, IIIT5K, SVTP, CUTE), VizWiz_Cap, VizWiz_VQA,...Votes: 0GitHub stars: 3
- At Add EvalThis evaluation protocol probes the robustness and generalization of audio deepfake detectors under real-world distortions and across heterogeneous audio types. It specifically tests whether models can maintain reliable binary classification performance when facing unseen generation methods, recording condition shifts, and unknown audio categories without relying on type-specific labels. Use when the user wants to benchmark on AT-ADD Challenge Dataset, or asks about evaluating this task. Repo...Votes: 0GitHub stars: 3
- Asvspoof2019 La EvalEvaluates the robustness of audio deepfake detection models against additive noise and measures how speech enhancement algorithms impact spoof detection accuracy. It probes whether improving perceptual speech quality in noisy environments preserves or degrades the discriminative features needed to distinguish real from spoofed audio. Use when the user wants to benchmark on ASVspoof 2019 LA, or asks about evaluating this task. Reports EER.Votes: 0GitHub stars: 3
- Asvspoof2019 EvalEvaluates the robustness of speaker verification systems and anti-spoofing countermeasures against synthesized, voice-converted, and replayed speech attacks. It measures how effectively systems can distinguish genuine speech from spoofed audio and quantifies the real-world impact of spoofing on authentication reliability. Use when the user wants to benchmark on ASVspoof 2019, or asks about evaluating this task. Reports EER, min-tDCF.Votes: 0GitHub stars: 3
- Astrovlbench EvalThis benchmark evaluates the ability of vision-language models to perform multi-modal astronomical reasoning across five distinct observational modalities, including optical imaging, radio interferometry, photometry, light curves, and spectroscopy. It probes whether models can correctly classify celestial objects and interpret physical features, while also testing the impact of prompt guidance and input representation (visual vs. numerical) on classification accuracy and reasoning quality. Us...Votes: 0GitHub stars: 3
- Astrovisbench EvalEvaluates large language models' ability to act as coding assistants for astronomy-specific scientific workflows. It probes domain-specific API usage, data manipulation, and the generation of research-standard visualizations from natural language queries. Use when the user wants to benchmark on AstroVisBench, or asks about evaluating this task. Reports execution-based evaluation.Votes: 0GitHub stars: 3
- Astronomical Semantic Search EvalEvaluates zero-shot semantic retrieval of astronomical images using natural language queries, specifically probing the model's ability to identify rare galactic phenomena (spirals, mergers, gravitational lenses) without curated training labels. It also measures the impact of VLM-based re-ranking on retrieval precision for rare classes. Use when the user wants to benchmark on HSC survey galaxy images, or asks about evaluating this task. Reports nDCG@10.Votes: 0GitHub stars: 3
- Astrochart EvalEvaluates multimodal large language models' ability to comprehend scientific charts and perform knowledge-intensive reasoning in astronomy. It probes visual understanding, data extraction, numerical calculation, and domain-specific inference. Use when the user wants to benchmark on AstroChart, or asks about evaluating this task. Reports Accuracy (%).Votes: 0GitHub stars: 3
- Astro Mcad EvalEvaluates a classifier-based anomaly detection pipeline on simulated astronomical transient light curves. It probes the model's ability to identify rare, out-of-distribution events in real-time without prior exposure to the anomalous classes during training. Use when the user wants to benchmark on Simulated LSST-like transient light curves, or asks about evaluating this task. Reports anomaly score.Votes: 0GitHub stars: 3