Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 5,449–5,472 of 23,640 skills
This evaluation probes the closed-loop planning robustness and causal reasoning of autonomous vehicle controllers by measuring their ability to handle compounding errors and distribution shifts. It combines real-world driving observations with pseudo-synthetic future scenarios generated via neural rendering to approximate interactive simulation without requiring a full physics engine. Use when the user wants to benchmark on nuPlan (navhard subset), or asks about evaluating this task. Reports ...
Evaluates the ability of masked language models to predict masked amino acids in short peptide sequences, measuring how well the model captures local sequence dependencies without autoregressive assumptions. It specifically probes the model's capacity to generalize to unseen short peptides that were excluded from the reference database during training. Use when the user wants to benchmark on UniRef100-excluded peptides, or asks about evaluating this task. Reports PseudoPPL.
Evaluates sound event localization and detection (SELD) performance on synthetic and real-world audio, measuring classification accuracy, localization precision, and overall detection quality across different network architectures and fine-tuning strategies. Use when the user wants to benchmark on synthetic-test-set, synthetic-training-set, Indoor Recordings, or asks about evaluating this task. Reports SELD.
Evaluates the ability of Estimation of Model Accuracy (EMA) methods to predict the structural quality of protein complex models. It probes global and interface-level accuracy estimation using correlation, ranking, and classification metrics against reference structural scores. Use when the user wants to benchmark on CASP16_inhouse_TOP5_dataset, CASP16_community_dataset, or asks about evaluating this task. Reports Pearson’s correlation (CorrP).
Evaluates a model's ability to learn sequentially from multiple tasks without catastrophic forgetting, measuring both task accuracy and continual learning dynamics like forward/backward transfer and forgetting rates across NLP and vision benchmarks. Use when the user wants to benchmark on Standard & Long, TRACE, ViT Benchmark, or asks about evaluating this task. Reports Accuracy (Acc/AAA).
Evaluates multimodal language models' vision-centric instruction following, multi-image reasoning, and general visual understanding. It probes the model's ability to process single and multiple images, answer factual questions, and perform complex reasoning across a suite of standard benchmarks. Use when the user wants to benchmark on CV-Bench (CVB-2D, CVB-3D), SEED-Bench, MMBench (MMB), MME, QBench2, MMMU, RealWorldQA, MMStar, MMVet, Mantis-Eval, MMT-Bench (MMT), TextVQA, or asks about evalu...
This benchmark evaluates the quality of prototype-based explainable AI (XAI) methods for time series and image data. It probes how well generated prototypes capture model behavior and data structure across nine interpretability properties, including correctness, consistency, continuity, and latent space cohesion. Use when the user wants to benchmark on ECG200, or asks about evaluating this task. Reports Total.
Evaluates protein language models on sequence understanding across multiple downstream tasks (structure, function, interactions, developability) and 3D structure prediction from single amino acid sequences. Use when the user wants to benchmark on CAMEO, CASP15, OOD Protein Sequences (UniProt), or asks about evaluating this task. Reports TM-score.
Evaluates a model's ability to predict the functional or stability impact of amino acid substitutions in proteins without prior experimental data for the specific variant. It probes zero-shot generalization across diverse protein families, taxonomic groups, and mutational depths (single-site vs. deep mutations). Use when the user wants to benchmark on DTm, DDG, ProteinGym, or asks about evaluating this task. Reports TPR@threshold.
Evaluates a model's ability to predict protein-ligand binding using only sequence data. It probes generalization across diverse protein targets, novel chemical scaffolds, and external benchmarks by measuring ranking performance between binders and decoys. Use when the user wants to benchmark on DEL Protein Split, DEL Chemical Library Split, MF-PCBA, Public Binders/Decoys, or asks about evaluating this task. Reports AUROC.
Evaluates the ability of graph neural networks combined with language models to learn structural and sequence representations of proteins. It probes how well the learned embeddings preserve structural similarity via TM-score prediction and generalize to downstream classification tasks across different protein families and out-of-distribution datasets. Use when the user wants to benchmark on Kinase dataset, SCOPe dataset, or asks about evaluating this task. Reports MSE.
This evaluation framework probes the physical validity, functional success, and generalization capability of generative protein models. It emphasizes leakage-aware dataset splits, structural and docking quality metrics, and experimental validation pipelines to ensure designs are biologically plausible and functionally active. Use when the user wants to benchmark on PLINDER, PoseBusters, PDFBench, PoseX, VenusX, FragBench, GeomMotif, or asks about evaluating this task. Reports RMSD.
Evaluates protein language models on classification and regression tasks to assess their capability in predicting protein properties, sub-cellular localization, epitope regions, and mutational effects. Use when the user wants to benchmark on Sub-cellular Localization, Membrane Solubility, Epitope Region Prediction, GB1 Mutational Landscape, or asks about evaluating this task. Reports accuracy, Spearman's rank correlation coefficient.
Evaluates protein language models and geometric deep learning architectures on five realistic downstream biological tasks, including binding affinity prediction, functional annotation, mutation effects, cleavage site detection, and PROTAC interaction modeling. It probes how pretraining objectives, structural information integration, and domain-specific inductive biases affect generalization on limited biological data. Use when the user wants to benchmark on Protap Benchmark, or asks about eva...
Evaluates conversational agents on dialogue safety classification, rule-of-thumb generation, and prosocial response generation. It probes the model's ability to identify unsafe content, generate socially informed guidelines, and produce safe, engaging, and respectful dialogue responses. Use when the user wants to benchmark on PROSOCIALDIALOG, or asks about evaluating this task. Reports accuracy, BLEU-4.
Evaluates a model's ability to decompose sentences into atomic semantic units (propositional segmentation) and determine entailment relationships between text spans. It probes fine-grained compositional semantic alignment and partial entailment recognition beyond sentence-level NLI. Use when the user wants to benchmark on PropSegmEnt, or asks about evaluating this task. Reports Precision/Recall/F1w (macro-averaged).
Evaluates a two-stage malware detection framework that uses a fast ML classifier for initial triage and a deep learning model for borderline cases. It probes the model's ability to accurately classify system call sequences as malicious or benign while balancing detection latency and false positive rates in real-time scenarios. Use when the user wants to benchmark on Propedeutica System Call Dataset, or asks about evaluating this task. Reports accuracy.
Detects propaganda techniques in news articles through a two-stage pipeline: identifying text spans containing propaganda (Span Identification) and classifying the specific rhetorical technique used within those spans (Technique Classification). Use when the user wants to benchmark on SemEval-2020 Task 11 Propaganda Detection, or asks about evaluating this task. Reports F1.
Evaluates a model's ability to translate between natural language mathematics and Lean 3 formal statements (autoformalization and informalization). It measures syntactic validity, semantic correctness, and lexical similarity to assess reasoning over undergraduate-level theory. Use when the user wants to benchmark on ProofNet, or asks about evaluating this task. Reports Accuracy.
This evaluation probes a model's ability to detect prompt injection attacks in realistic deployment settings. It specifically tests whether a detector can distinguish between benign conversational or application-structured inputs and maliciously crafted injections while maintaining a very low false positive rate to avoid costly false alarms. Use when the user wants to benchmark on PromptShield Evaluation Set, or asks about evaluating this task. Reports TPR@0.1%FPR.
Evaluates Chinese large language models on five biomedical NLP tasks: information extraction, text classification, natural language inference, dialogue understanding, and content generation. It probes instruction-following, few-shot in-context learning, and parameter-efficient fine-tuning capabilities in a medical domain context. Use when the user wants to benchmark on PromptCBLUE, or asks about evaluating this task. Reports Instance-level strict micro-F1.
This benchmark evaluates the vulnerability of large language models to prompt injection attacks when generating scientific paper reviews. It probes whether hidden or biased instructions embedded in parsed PDFs can systematically skew the model's review scores and recommendations. Use when the user wants to benchmark on ICLR 2024 Review Dataset, or asks about evaluating this task. Reports Rating.
This benchmark evaluates an LLM's or monitoring system's ability to distinguish between safe user inputs and malicious prompt injection attacks. It measures both false positive rates on legitimate interactions and false negative rates on adversarial prompts to assess overall security robustness. Use when the user wants to benchmark on Gandalf, Tensor-Trust, SPML-Dataset, or asks about evaluating this task. Reports Error Rate (ER).
Evaluates 3D volumetric medical image segmentation capability on prostate MRI scans. It probes the model's ability to accurately delineate organ boundaries under clinical variability and class imbalance using end-to-end fully convolutional networks. Use when the user wants to benchmark on PROMISE2012, or asks about evaluating this task. Reports Dice coefficient.