Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 8,233–8,256 of 20,869 skills
Evaluates cross-lingual speech-to-text retrieval and intent detection capabilities across multiple datasets, testing how well speech queries can retrieve relevant text documents or classify intents without intermediate ASR or translation pipelines. Use when the user wants to benchmark on Kallaama-Retrieval-Eval, Fleurs-Retrieval-Eval, Urban Bus, WolBanking77, or asks about evaluating this task. Reports nDCG@5.
Evaluates zero-shot crosslingual generalization of multilingual LLMs after multitask finetuning. Probes language-agnostic task understanding, robustness to prompt translation, and scaling behavior across NLU, generative, and code tasks. Use when the user wants to benchmark on XNLI, XCOPA, XStoryCloze, XWinograd, HumanEval, or asks about evaluating this task. Reports accuracy.
Evaluates the robustness of multimodal LLMs against explicit and implicit jailbreak attacks while measuring their utility on benign queries. It probes whether a defense model can successfully refuse harmful image-text prompts without over-restricting safe inputs. Use when the user wants to benchmark on JailBreakV, VLGuard, FigStep, MM-SafetyBench, SIUO, MMBench, or asks about evaluating this task. Reports Attack Success Rate (ASR).
Evaluates a model's ability to generate novel, drug-like molecules with high binding affinity for unseen protein pockets in structure-based drug design. It probes the trade-offs between binding energy, molecular properties, and synthesis feasibility. Use when the user wants to benchmark on CrossDocked-100k, or asks about evaluating this task. Reports Vina Dock.
This evaluation probes the ability of multimodal foundation models to generate factual content without hallucination across text, image, and audio-visual modalities. It measures how well reference-free ranking methods correlate with human judgments or gold-standard references to rank model outputs by hallucination severity. Use when the user wants to benchmark on WikiBio, MHaluBench, AVHalluBench, or asks about evaluating this task. Reports System($ ho$).
Evaluates systems' ability to predict target-language pronoun class labels from source-language pronouns using lemmatized, POS-tagged translations and word alignments. It probes cross-lingual anaphora resolution and functional ambiguity handling in machine translation pipelines. Use when the user wants to benchmark on WMT 2016 Cross-lingual Pronoun Prediction Task, or asks about evaluating this task. Reports macro-averaged recall.
Evaluates the intelligibility, speaker similarity, and naturalness of synthesized speech in cross-lingual voice cloning and TTS scenarios. It also measures the accuracy of a language-agnostic speaking rate predictor for duration modeling across multiple languages. Use when the user wants to benchmark on Emilia, Seed-TTS-eval, LibriSpeech-PC test-clean, FLEURS, or asks about evaluating this task. Reports WER.
This evaluation probes a model's ability to perform cross-domain sequential recommendation by leveraging user interaction histories across two domains, even when user overlap is minimal or absent. It measures how well the model captures domain-specific and shared sequential patterns to rank candidate items accurately. Use when the user wants to benchmark on Micro Video, Amazon, or asks about evaluating this task. Reports AUC.
Evaluates an object detection model's ability to generalize across domain shifts (e.g., real-to-artistic, clear-to-foggy, synthetic-to-real) using only labeled source data and unlabeled target data during training. It measures how well the model mitigates domain bias and adapts to unseen target distributions without target annotations. Use when the user wants to benchmark on PASCAL VOC 2007+2012, Clipart1k, Watercolor2k, Cityscapes, Foggy Cityscapes, SIM10K, or asks about evaluating this task...
Probes few-shot image classification generalization across diverse domains and highly variable task regimes (2–20 ways, 1–20 shots) without relying on pre-trained backbones. Use when the user wants to benchmark on Meta-Album, or asks about evaluating this task. Reports accuracy.
Evaluates cross-domain knowledge transfer for click-through rate (CTR) prediction by measuring how well a model trained on a source domain generalizes to a target domain with non-overlapping features. It probes context-aware feature translation and explicit knowledge augmentation in recommendation systems. Use when the user wants to benchmark on Amazon, Taobao, Alibaba Production, or asks about evaluating this task. Reports AUC.
Evaluates continual reinforcement learning capabilities in robotic simulation, specifically measuring how well agents retain performance on previously learned tasks while learning new sequential tasks. It probes catastrophic forgetting, transfer effects, and intrinsic task difficulty across line-following, object-pushing, and reaching benchmarks. Use when the user wants to benchmark on CRoSS, or asks about evaluating this task. Reports average cumulated score.
Evaluates vision-language models' ability to use a crop-and-zoom tool for high-resolution visual question answering, disentangling intrinsic capability improvements from tool-induced gains and harms across multiple benchmarks. Use when the user wants to benchmark on VStar, HR-Bench 4k/8k, VisualProbe Easy/Medium/Harm, or asks about evaluating this task. Reports accuracy.
Evaluates self-supervised remote sensing representations across classification and segmentation tasks using optical and radar-optical inputs. Probes representation quality via finetuning, linear/nonlinear probing, kNN, and clustering. Use when the user wants to benchmark on BigEarthNet, fMoW-Sentinel, EuroSAT, Canadian Cropland, DFC2020, DW-Expert, MARIDA, or asks about evaluating this task. Reports mAP, Top 1 Acc., mIoU.
Evaluates handwritten mathematical expression recognition (HMER) by measuring exact LaTeX sequence matching, tolerant symbol-level error rates, and structural tree prediction accuracy on complex handwritten formulas. Use when the user wants to benchmark on CROHME, HME100K, or asks about evaluating this task. Reports ExpRate.
Evaluates a model's ability to learn generalizable biometric feature representations in a continual learning setting, specifically measuring generalization to unseen identities across sequential learning steps rather than retaining knowledge of previously seen classes. Use when the user wants to benchmark on CRL-face, CRL-person, LFW, Megaface, or asks about evaluating this task. Reports Top 1 accuracy.
Evaluates the ability to detect the onset of systemic instability (criticality) in simulated AI systems by monitoring performance variance across multiple benchmarks. It probes whether a derivative-based threshold can reliably flag phase transitions before functional collapse. Use when the user has predictions and gold and needs to compute percentage of correct classifications.
Evaluates the correctness and computational efficiency of closed-form algorithms for computing critical point probabilities in 2D scalar fields under various parametric and nonparametric noise models. The protocol compares these analytical solutions against Monte Carlo sampling baselines across synthetic and real-world scientific datasets to validate accuracy and speed. Use when the user wants to benchmark on Ackley function (synthetic), Gaussian mixture model (synthetic), E3SM climate data, ...
Evaluates the predictive accuracy and inference efficiency of deep learning recommendation models (DLRM) with compressed embedding tables on large-scale advertising click-through rate datasets. It measures Area Under the ROC Curve (AUC) to assess model quality and samples per second to quantify inference throughput under memory-constrained conditions. Use when the user wants to benchmark on CriteoTB, Criteo Kaggle, or asks about evaluating this task. Reports AUC.
Evaluates the predictive quality and system efficiency of deep learning recommendation models on click-through rate prediction. It measures how well parameter-sharing compression techniques maintain model accuracy while reducing memory footprint and improving training and inference latency. Use when the user wants to benchmark on criteo-kaggle, criteo-tb, or asks about evaluating this task. Reports AUC.
Evaluates an AI system's ability to extract and classify budget allocations for Early Warning System (EWS) investments from heterogeneous financial PDF reports. It probes multi-label classification, numerical budget extraction with tolerance, and evidence retrieval/mapping in climate finance contexts. Use when the user wants to benchmark on MDB Evidence Set, or asks about evaluating this task. Reports Accuracy.
Probes LLM-based multi-agent coordination in dynamic, partially observable wildfire disaster response scenarios. It evaluates capabilities such as spatial reasoning, task designation, plan adaptation, and heterogeneous team collaboration under stochastic dynamics and long-horizon objectives. Use when the user wants to benchmark on CREW-Wildfire, or asks about evaluating this task. Reports task success.
Evaluates a model's ability to predict user creditworthiness based on geographic mobility footprints. It probes whether spatiotemporal visitation patterns and region-level credit signals can reliably distinguish users who pay their mobile phone bills from those who do not. Use when the user wants to benchmark on Hangzhou user mobility dataset, or asks about evaluating this task. Reports AUC.
Evaluates a model's ability to detect fraudulent credit card transactions in a streaming context by learning topological and sequential patterns from transaction graphs without manual feature engineering. Use when the user wants to benchmark on Credit Card Transaction Dataset (Feb-Sep), or asks about evaluating this task. Reports AP.