Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 5,017–5,040 of 23,574 skills
Evaluates LLM safety alignment and robustness against adversarial jailbreak attacks, particularly in scientific domains. It also measures the model's ability to maintain general helpfulness, truthfulness, and avoid over-refusal on benign queries. Use when the user wants to benchmark on AdvBench, HarmBench, StrongReject, SciKnowEval (L4), SciSafeEval, LabSafety Bench (Hard), GSM8K, MT-Bench, MMLU, GPQA, SimpleQA, XsTest, or asks about evaluating this task. Reports Attack Success Rate (ASR).
This evaluation probes an agent's ability to learn safe navigation and manipulation policies from human demonstrations in environments with unknown safety constraints. It specifically tests the trade-off between maximizing task reward and minimizing safety violations (cost) under out-of-distribution conditions. Use when the user wants to benchmark on Safety-Gymnasium (SafetyPointGoal1-v0, SafetyPointCircle2-v0, SafetyCarButton1-v0, SafetyCarPush2-v0), or asks about evaluating this task. Repor...
Evaluates the biosafety risks and jailbreak vulnerabilities of protein foundation models by measuring their ability to reconstruct harmful protein sequences and 3D structures from partially masked inputs. It probes whether models can bypass safety filters and generate biologically dangerous proteins when given sequence and structural prompts. Use when the user wants to benchmark on SafeProtein-Bench, or asks about evaluating this task. Reports jailbreak success rate.
Evaluates the ability of Large Vision-Language Models (LVLMs) to detect and refuse harmful visual content without modifying the base model architecture. It measures both safety defense effectiveness on toxic inputs and the preservation of utility on benign inputs. Use when the user wants to benchmark on Toxic Image Categories (Porn, Bloody, Insulting, Alcohol, Cigarette, Gun, Knife, Neutral), or asks about evaluating this task. Reports DSR.
Evaluates the vulnerability of aligned multimodal LLMs to universal adversarial image attacks that bypass safety filters. It measures how often a single optimized image forces the model to generate unsafe or affirmative responses across diverse text prompts. Use when the user wants to benchmark on SafeBench, or asks about evaluating this task. Reports Attack Success Rate (ASR).
Evaluates the safety-aware task planning capabilities of embodied LLM agents in interactive simulation environments. It probes whether agents can proactively reject hazardous instructions, avoid implicit risks in long-horizon planning, and maintain planning performance on safe tasks across varying levels of task abstraction. Use when the user wants to benchmark on SafeAgentBench, or asks about evaluating this task. Reports rejection rate.
Evaluates the capability of classifiers to detect sexist, abusive, offensive, and hate speech in conversational text across multiple granularity levels and established benchmarks. It probes fine-grained toxicity detection, cross-dataset generalization, and performance against strong supervised and LLM baselines. Use when the user wants to benchmark on EDOS (SemEval 2023), OffensEval 2019, AbusEval, HatEval, or asks about evaluating this task. Reports F1.
This benchmark probes the safety judgment and alignment capabilities of professional-level AI agents. It evaluates whether agents can resist executing harmful or risky actions when given complex, domain-specific instructions in fields like finance, law, and healthcare. Use when the user wants to benchmark on SafePro, or asks about evaluating this task. Reports unsafe rate.
Evaluates a hybrid trajectory planning framework that fuses flow matching with model-predictive control. It probes the system's ability to generate adaptive, collision-free motions for a 7-DoF robot manipulator while strictly enforcing safety constraints in real-time across global planning, reactive replanning, and dynamic human-robot handover scenarios. Use when the user wants to benchmark on Custom Robot Manipulation Benchmarks (Exp 1-3), or asks about evaluating this task. Reports adherenc...
Evaluates the ability of reinforcement learning agents to maintain safety constraints and retain knowledge across sequentially changing non-stationary robotic environments, measuring the trade-off between task performance, safety violations, and catastrophic forgetting. Use when the user wants to benchmark on Damaged HalfCheetah Velocity, Damaged Ant Velocity, Safe Continual World, or asks about evaluating this task. Reports Final Task Reward.
Evaluates the safety alignment and refusal capabilities of Video Large Multimodal Models (VLMMs) against everyday adversarial queries and covert, human-red-teamed prompts. It measures whether models can maintain safety guidelines across diverse harmful categories without compromising general utility or falling back to memorized refusals. Use when the user wants to benchmark on SafeVidBench, or asks about evaluating this task. Reports Safety Rate.
Evaluates sparse autoencoder (SAE) architectures across multiple dimensions including reconstruction fidelity, feature disentanglement, concept detection, and practical interpretability tasks. It systematically compares how different SAE designs, dictionary sizes, and sparsity levels impact these capabilities. Use when the user wants to benchmark on Gemma-2-2B, Pythia-160M, or asks about evaluating this task. Reports Loss Recovered.
Evaluates large language models' ability to perform complex, multi-step reasoning and data dependency tracking by executing SQL queries over synthetic, arbitrarily long tables. It probes contextual reasoning, long-context understanding, and structured data manipulation capabilities beyond traditional benchmarks. Use when the user wants to benchmark on S3Eval, or asks about evaluating this task. Reports SQL execution performance.
Evaluates whether multimodal LLMs and agentic systems can reliably generate actionable decision support from operational S2S climate service products. It probes three core capabilities: actionable signal comprehension, uncertainty-conditioned decision-making handoffs, and evidence-grounded planning under dynamic hazards. Use when the user wants to benchmark on S2SServiceBench, or asks about evaluating this task. Reports Rubric Score (CT, ACT, TTH, EG, FC, UC).
Evaluates speech-to-speech models on instruction following, assessing both semantic correctness and paralinguistic/speech quality in a reference-free, head-to-head comparison. Use when the user wants to benchmark on S2S-Arena, or asks about evaluating this task. Reports ELO score.
Evaluates the skill of deep learning post-processing models for global sub-seasonal temperature and precipitation forecasts against climatological baselines and ECMWF recalibrated forecasts. It probes the ability of spatial CNN architectures to correct systematic errors and produce well-calibrated probabilistic tercile predictions over a 2–4 week horizon. Use when the user wants to benchmark on S2S AI Challenge test set (2020), or asks about evaluating this task. Reports RPSS.
Evaluates models' ability to transcribe spoken mathematical equations and sentences into correct LaTeX syntax. It probes audio-to-text conversion, handling of mathematical symbols, and robustness to syntactic variations in LaTeX formatting. Use when the user wants to benchmark on S2L-equations, S2L-sentences, or asks about evaluating this task. Reports CER.
Evaluates whether representation learning models capture biologically meaningful signals in high-content microscopy images. It probes the model's ability to distinguish drug-induced perturbations from controls, predict zero-shot drug-target interactions, and recover known gene-gene relationships from phenotypic embeddings. Use when the user wants to benchmark on RxRx3-core, or asks about evaluating this task. Reports average precision.
Evaluates the ability of embodied agents to follow natural language instructions for navigation in photo-realistic 3D environments. It probes multilingual understanding, spatial reasoning, and dense spatiotemporal grounding by measuring how accurately an agent navigates from a start to a target location. Use when the user wants to benchmark on Room-Across-Room (RxR), or asks about evaluating this task. Reports SR, NDTW.
Evaluates a model's ability to maintain context and generate coherent responses in multi-turn dialogue, comparing stateful event-driven architectures against standard decoder-only LLMs. Use when the user wants to benchmark on MRL Curriculum Datasets (derived from TinyStories), or asks about evaluating this task. Reports MRL Reward Score.
Evaluates object detectors' robustness to real-world spatial domain shifts across hurricane-affected regions in satellite imagery. It measures how well models generalize from in-domain training data to out-of-distribution target domains without fine-tuning. Use when the user wants to benchmark on RWDS-HE, or asks about evaluating this task. Reports mAP.
Evaluates object detectors' robustness to real-world spatial domain shifts across flood-affected regions in satellite imagery. It measures how well models generalize from in-domain training data to out-of-distribution target domains without fine-tuning. Use when the user wants to benchmark on RWDS-FR, or asks about evaluating this task. Reports mAP.
Evaluates object detectors' robustness to real-world spatial domain shifts across different climate zones in satellite imagery. It measures how well models generalize from in-domain training data to out-of-distribution target domains without fine-tuning. Use when the user wants to benchmark on RWDS-CZ, or asks about evaluating this task. Reports mAP.
Evaluates the robustness of modern voice cloning models under realistic deployment conditions, including input variations (accents, text shifts, long context), cross-lingual synthesis, post-processing degradation, and adversarial perturbations. It probes the trade-offs between generation quality, content fidelity, speaker similarity, and deepfake detectability across diverse acoustic and linguistic stressors. Use when the user wants to benchmark on LibriTTS, VCTK, LibriSpeech, RVCBench, or as...