All categories
Research
Research, evidence gathering, literature, reports, investigation, and synthesis
- 21,377
- 891
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browserBrowse research skills
Showing 9,529–9,552 of 21,377 skills
- Auv Docking EvalEvaluates the capability of deep reinforcement learning algorithms to perform continuous docking control of an autonomous underwater vehicle (AUV) in a physics-based simulator. It probes the agent's ability to navigate from random initial positions to a target docking station while optimizing a physics-informed reward function that accounts for proximity, orientation, and contact dynamics. Use when the user wants to benchmark on UUV Simulator (DeepLeng AUV model), or asks about evaluating thi...Votes: 0GitHub stars: 3
- Autshumato Nmt EvalEvaluates neural machine translation performance across five Southern African languages (Afrikaans, isiZulu, Northern Sotho, Setswana, Xitsonga) from English. It probes how dataset size and morphological complexity (e.g., agglutinative vs. non-agglutinative) impact translation quality in low-resource settings. Use when the user wants to benchmark on Autshumato, or asks about evaluating this task. Reports BLEU.Votes: 0GitHub stars: 3
- Autovivqa EvalThis benchmark evaluates Vietnamese vision-language models on visual question answering, probing their ability to ground textual queries in images and generate semantically accurate, linguistically fluent responses. It specifically tests reasoning complexity across five levels (recognition, relational, compositional, causal, text-in-image) and measures both exact-match accuracy and generation quality. Use when the user wants to benchmark on AutoViVQA, or asks about evaluating this task. Repor...Votes: 0GitHub stars: 3
- Autored EvalEvaluates the safety alignment and vulnerability of large language models against adversarial red-teaming prompts. It measures how effectively generated or human-crafted harmful instructions can bypass safety filters to elicit unsafe model responses. Use when the user wants to benchmark on AutoRed & Baseline Red-Teaming Datasets, or asks about evaluating this task. Reports Attack Success Rate (ASR).Votes: 0GitHub stars: 3
- Autopet2022 EvalEvaluates the capability of 3D medical image segmentation models to accurately delineate lesions in whole-body FDG-PET/CT scans. It probes the model's ability to handle high-resolution volumetric data and distinguish pathological regions from healthy tissue. Use when the user wants to benchmark on autoPET 2022, or asks about evaluating this task. Reports DSC.Votes: 0GitHub stars: 3
- Automotive Verification EvalEvaluates software model checkers on their ability to verify safety requirements in industrial automotive C code generated from Simulink models. It probes how well tools handle floating-point arithmetic, pointer operations, and complex control logic under bounded and unbounded verification constraints. Use when the user wants to benchmark on DSR, ECC, or asks about evaluating this task. Reports SV-COMP quantile plots.Votes: 0GitHub stars: 3
- Automl Pipeline EvalEvaluates the ability of AutoML systems to automatically discover optimal machine learning pipelines (feature transformers, learners, and hyperparameters) for tabular data. It probes how well meta-learning and graph-based approaches generalize across diverse classification and regression tasks under strict time budgets. Use when the user wants to benchmark on 121-dataset AutoML Benchmark (Open AutoML, PMLB, AL, VolcanoML), or asks about evaluating this task. Reports Macro F1, R².Votes: 0GitHub stars: 3
- Automathtext Math EvalEvaluates the effectiveness of the AutoMathText dataset for continual pretraining by measuring downstream mathematical reasoning performance on the MATH benchmark. It compares models trained on auto-selected high-quality mathematical text versus uniformly sampled filtered text, controlling for token count. Use when the user wants to benchmark on MATH, AutoMathText, or asks about evaluating this task. Reports MATH test accuracy (%).Votes: 0GitHub stars: 3
- Automated Red Teaming EvalEvaluates an LLM's capability to generate effective adversarial prompts (red teaming attacks) for arbitrary safety goals. It measures both the success rate of eliciting targeted behaviors and the diversity of the generated attacks across in-domain and out-of-domain objectives. Use when the user wants to benchmark on garak adversarial goals, or asks about evaluating this task. Reports attack success rate.Votes: 0GitHub stars: 3
- Autograph R1 EvalThis benchmark evaluates the functional utility of reinforcement learning-optimized knowledge graphs in end-to-end retrieval-augmented generation pipelines. It probes whether task-aware RL training improves both graph-based reasoning and text retrieval performance across multiple question-answering benchmarks and model scales. Use when the user wants to benchmark on Natural Questions (NQ), PopQA, HotpotQA, 2WikiMultihopQA, Musique, or asks about evaluating this task. Reports F1 score.Votes: 0GitHub stars: 3
- Autoformalization Compile EvalProbes a model's ability to translate informal natural language mathematical statements into syntactically and semantically valid formal code for theorem provers (Isabelle or Lean4). It measures how well the model captures formal syntax, type-checking rules, and prover-specific conventions without requiring proof generation. Use when the user wants to benchmark on miniF2F, ProofNet, or asks about evaluating this task. Reports Compilation rates (%).Votes: 0GitHub stars: 3
- Autofish EvalFine-grained instance segmentation and length estimation of visually similar fish species under realistic conveyor-belt conditions. The benchmark evaluates model robustness across separated, touching, and occluded fish configurations, using group-based splits to prevent data cross-contamination. Use when the user wants to benchmark on AutoFish, or asks about evaluating this task. Reports mAP.Votes: 0GitHub stars: 3
- Autoeval Video EvalThis benchmark evaluates large vision-language models on open-ended video question answering across nine skill dimensions, including dynamic perception, temporal comprehension, causal reasoning, and response specificity. It probes the model's ability to connect multiple frames, understand temporal dynamics, and generate precise, video-grounded answers rather than generic or hallucinated text. Use when the user wants to benchmark on AutoEval-Video, or asks about evaluating this task. Reports A...Votes: 0GitHub stars: 3
- Autodrive Qa EvalEvaluates vision-language models on urban autonomous driving tasks by testing their ability to answer multiple-choice questions covering perception, prediction, and planning. It probes domain-specific reasoning, hazard detection, speed judgment, and object classification in driving scenarios. Use when the user wants to benchmark on AutoDrive-QA, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3
- Autoctr Ctr EvalEvaluates automated neural architecture search for click-through rate prediction on heterogeneous tabular data. It measures how well discovered architectures predict user clicks compared to human-crafted models, and tests their transferability across different datasets. Use when the user wants to benchmark on Criteo, Avazu, KDD Cup, or asks about evaluating this task. Reports logloss.Votes: 0GitHub stars: 3
- Auto J EvalEvaluates the alignment and preference alignment of LLMs through pairwise comparison, single-response critique generation, and overall rating. It probes the model's ability to consistently identify human-preferred responses, generate structured natural language critiques, and rank outputs according to a reference judge (GPT-4). Use when the user wants to benchmark on Eval-P, Eval-C, Eval-R, AlpacaEval, or asks about evaluating this task. Reports agreement rate.Votes: 0GitHub stars: 3
- Auto Dataset Update EvalEvaluates LLMs on automatically updated benchmark datasets (BIG-bench, MMLU) to measure evaluation stability, data leakage mitigation, and cognitive-level difficulty control via mimicking and extending generation strategies. Use when the user wants to benchmark on BIG-bench, MMLU, or asks about evaluating this task. Reports full-mark rate (%).Votes: 0GitHub stars: 3
- Authorship Attribution EvalEvaluates a model's ability to identify an author from their writing style while suppressing domain-specific (fandom) style leakage. It probes cross-domain generalization and robustness to domain swapping in binary authorship attribution. Use when the user wants to benchmark on Fanfiction Corpus, or asks about evaluating this task. Reports mean macro accuracy.Votes: 0GitHub stars: 3
- Author Centric Review EvalEvaluates an LLM's ability to generate structured, author-centric academic feedback (Summary, Strengths, Weaknesses, Questions) from long research papers. It probes the model's capacity to retrieve salient passages via graph-based retrieval and synthesize constructive pre-submission reviews without relying on full context or multi-agent systems. Use when the user wants to benchmark on ICLR 2024 (ICT), CNT_10, or asks about evaluating this task. Reports human evaluation.Votes: 0GitHub stars: 3
- Authfix Oidc Repair EvalEvaluates an automated program repair system's ability to fix real-world security and logic bugs in OpenID Connect implementations. It measures both the rate of successfully generating correct patches and the semantic quality of those patches compared to human-written developer fixes. Use when the user wants to benchmark on OpenID Connect Bug Dataset, or asks about evaluating this task. Reports fix_accuracy.Votes: 0GitHub stars: 3
- Authenhallu EvalEvaluates an LLM's capability to detect and categorize hallucinations in authentic, real-world human-LLM dialogues. It specifically probes whether models can identify input-conflicting, context-conflicting, and fact-conflicting errors in query-response pairs. Use when the user wants to benchmark on AuthenHallu, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3
- Auslaw Citation EvalEvaluates an LLM's ability to predict the correct legal citation for a given query text. It probes the model's capacity to retrieve or generate accurate references from a large Australian legal corpus, testing both retrieval and generation capabilities in a domain-specific setting. Use when the user wants to benchmark on AusLaw Citation Benchmark, or asks about evaluating this task. Reports ACC@1.Votes: 0GitHub stars: 3
- Aurora Weather Extremes EvalEvaluates the predictability and forecast skill of the Aurora AI weather model across selected extreme weather events, including tropical cyclones, winter freezes, and heatwaves. It probes the model's ability to maintain deterministic track accuracy, temperature amplitude, and spatial pattern fidelity across short-range (1–7 day) to subseasonal (14–21 day) lead times. Use when the user wants to benchmark on Selected Weather Extremes Case Studies, or asks about evaluating this task. Reports Tr...Votes: 0GitHub stars: 3
- Aunp Mechanism EvalEvaluates whether large language models can correctly reason about physicochemical mechanisms in gold nanoparticle synthesis using multiple-choice questions. It probes both factual recall and the depth of mechanistic understanding by measuring prediction accuracy and model confidence derived from output logits. Use when the user wants to benchmark on AuNP Synthesis Mechanism Benchmark, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3