Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 4,993–5,016 of 23,574 skills
Abstractive summarization of naturalistic, messenger-style dialogues. It probes a model's ability to generate concise, human-like summaries from informal, multi-turn conversations containing typos, emoticons, and multi-party interactions. Use when the user wants to benchmark on SAMSum Corpus, or asks about evaluating this task. Reports ROUGE F1 (ROUGE-1, ROUGE-2, ROUGE-L).
Evaluates sequential recommendation models on their ability to predict the next item in a user's chronological interaction history. It specifically probes whether sharpness-aware minimization improves generalization and data efficiency compared to standard Transformers and self-supervised baselines. Use when the user wants to benchmark on Amazon-Beauty, Amazon-Sports, Amazon-Toys, Yelp, or asks about evaluating this task. Reports HR@10.
Evaluates zero-shot singing voice conversion by measuring timbre transfer accuracy and audio quality. It probes the model's ability to disentangle content, pitch, and timbre, and generalize to unseen human and non-human (animal) speakers without fine-tuning. Use when the user wants to benchmark on Custom zero-shot test set, or asks about evaluating this task. Reports MOS-S.
Detects and localizes semantically coordinated multimodal manipulations where visual edits are paired with contextually consistent textual narratives. Probes a model's ability to perform binary classification, multi-label categorization, and fine-grained visual tampering region localization using external celebrity attribute knowledge. Use when the user wants to benchmark on SAMM, or asks about evaluating this task. Reports ACC.
Evaluates the zero-shot segmentation capability of the Segment Anything Model (SAM) across diverse medical imaging modalities and anatomical structures. It probes the model's ability to generalize without task-specific fine-tuning to structured medical targets like organs, lesions, and retinal layers. Use when the user wants to benchmark on Skin Lesion Analysis Toward Melanoma Detection (ISIC), DoFE (Drishiti-GS, RIM-ONE-r3, REFUGE amalgamated), AMOS, MICCAI 2017 Robotic Instrument Segmentati...
Evaluates audio source separation capabilities conditioned on text, visual masks, or temporal spans. It probes open-vocabulary extraction, speaker/music/instrument isolation, and cross-modal grounding in both studio and in-the-wild settings. Use when the user wants to benchmark on SAM Audio Evaluation Set, MUSDB18, or asks about evaluating this task. Reports separation fidelity.
Evaluates the effectiveness and overhead of fine-grained GPU sharing primitives for scheduling deep learning training, hyper-parameter tuning, and inference workloads on a single GPU. Use when the user wants to benchmark on Salus DL Workload Trace & Benchmarks, or asks about evaluating this task. Reports Makespan.
Probes whether tabular models can effectively leverage declarative business knowledge and metadata semantics for prediction tasks, rather than relying solely on statistical correlations in raw features. It evaluates the impact of schema-grounded semantic embeddings on model inductive biases and relative performance across different model families. Use when the user wants to benchmark on SALT-KG, or asks about evaluating this task. Reports ranking metrics.
Evaluates the ability of saliency prediction models to accurately identify visually salient regions and semantic objects in autonomous driving scenarios. It probes whether models can capture critical driving elements like pedestrians and approaching vehicles while mitigating center-bias and peripheral vision neglect. Use when the user wants to benchmark on BDD-A, DR(eye)VE, JAAD, or asks about evaluating this task. Reports D_KL.
Evaluates the quality of synthetically generated Text-to-SQL data by measuring question-SQL alignment and business realism in a sales analytics domain. Use when the user wants to benchmark on Salesforce Sales Analytics Database, or asks about evaluating this task. Reports Question-SQL Alignment (%).
Evaluates the safety, robustness, and helpfulness of Large Language Models across a hierarchical taxonomy of 6 domains, 16 tasks, and 66 categories. It measures performance on benign, adversarial (attack-enhanced), and defense-enhanced prompts, as well as multiple-choice safety questions, while also benchmarking the effectiveness of various attack and defense strategies. Use when the user wants to benchmark on SALAD-Bench, ToxicChat, Beavertails, SafeRLHF, Harmbench, Lifetox, AdvBench-50, or ...
Evaluates large language models on language proficiency, reading comprehension, reasoning, and cultural understanding across three Southeast Asian languages (Indonesian, Vietnamese, Thai). It covers eight diverse tasks including question answering, machine translation, text summarization, multiple-choice exams, commonsense reasoning, machine reading comprehension, natural language inference, and sentiment analysis. Use when the user wants to benchmark on XQuAD, TyDiQA, Flores-200, ThaiSum, In...
Evaluates the scalability and performance trends of scientific AI workloads (3D CNNs) on HPC systems under varying node counts and dataset sizes. It probes how hardware constraints like GPU memory, I/O bandwidth, and network communication affect training efficiency and model convergence. Use when the user wants to benchmark on SAIH-cosmo, or asks about evaluating this task. Reports average_flops.
Evaluates Arabic language models on financial and Shari’ah-compliant reasoning across multiple task types, including multiple-choice questions, extractive summarization, and open-ended question answering. It probes the gap between general Arabic fluency and domain-specific procedural/financial reasoning. Use when the user wants to benchmark on SAHM, or asks about evaluating this task. Reports exact-match accuracy.
This benchmark evaluates speech large language models across five hierarchical levels of understanding, ranging from basic automatic speech recognition and language identification to paralinguistic perception (pitch, volume, emotion), abstract acoustic reasoning (medical cough analysis), and creative/agentic tasks (spoken English coaching). It probes the model's ability to process raw audio, follow instructions, and extract both semantic and non-semantic acoustic features. Use when the user w...
Evaluates embodied vision-and-language navigation (VLN) capabilities within physically executable 3D Gaussian Splatting environments. It probes a model's ability to follow natural language instructions (high- and low-level), navigate to goals without collisions, and exhibit smooth, natural motion continuity rather than mechanical or wall-hugging behaviors. Use when the user wants to benchmark on SAGE-Bench, VLN-CE (R2R Val-Unseen), or asks about evaluating this task. Reports SR.
Evaluates automatic speech recognition (ASR) performance on the Oromo language using real-world, crowd-sourced audio data. It measures how well different model architectures (Conformer trained from scratch, Whisper fine-tuned) transcribe spoken Oromo into text under varying acoustic conditions. Use when the user wants to benchmark on Sagalee, or asks about evaluating this task. Reports WER.
Evaluates the effectiveness of a spectral geometric adversarial attack on 3D mesh autoencoders by measuring how well perturbed meshes deceive a downstream classifier and evade detection. Use when the user wants to benchmark on CoMA, SMAL, or asks about evaluating this task. Reports Targeted classification accuracy.
This evaluation protocol probes the trade-off between safety alignment and reasoning capability in Large Reasoning Models. It measures how post-alignment fine-tuning impacts performance on standard reasoning benchmarks versus the model's propensity to generate harmful responses to malicious prompts. Use when the user wants to benchmark on GPQA, AIME24, MATH500, BeaverTails, or asks about evaluating this task. Reports Reasoning Accuracy.
Quantifies implicit representational harms in pre-trained language models by measuring the disparity in language modeling probabilities between harmful and benign sentences targeting 13 marginalized demographics. It probes whether a model's internal likelihood estimates reflect toxic or stereotypical biases toward specific groups. Use when the user has predictions and gold and needs to compute safety score.
Evaluates the safety alignment of reasoning models by measuring how frequently they comply with harmful or jailbreak prompts across multiple risk categories. It also measures utility retention on standard mathematical and knowledge benchmarks to ensure safety improvements do not degrade general capabilities. Use when the user wants to benchmark on DAN, Wildjailbreak, StrongReject, GSM8K, MMLU, or asks about evaluating this task. Reports attack success rate.
This evaluation probes a model's ability to retain task-specific utility while preserving safety alignment during supervised fine-tuning. It measures how well a method prevents safety degradation when exposed to benign or contaminated fine-tuning data, balancing performance retention against harmful output generation. Use when the user wants to benchmark on SST-2, AGNEWS, GSM8K, PubMedQA, AlpacaEval, JailbreakBench, HarmBench, AdvBench, BeaverTails, or asks about evaluating this task. Reports...
Evaluates the safety alignment of Large Vision Language Models (LVLMs) against safety-awareness benchmarks and multimodal jailbreak attacks. It measures the model's ability to detect and mitigate harmful intents while preserving general multimodal reasoning capabilities. Use when the user wants to benchmark on MSSBench, SIUO, MM-SafetyBench, MML-M, FigStep, or asks about evaluating this task. Reports safety rate.
Evaluates the safety alignment preservation and task utility of LLMs after parameter-efficient fine-tuning under data-poisoning attacks. It measures how well defenses maintain core capabilities while resisting harmful behavior injection. Use when the user wants to benchmark on MMLU, MT-Bench, AdvBench, PolicyEval, or asks about evaluating this task. Reports Attack Success Rate (ASR).