Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 8,617–8,640 of 21,220 skills
Evaluates an object detection model's ability to generalize across domain shifts (e.g., real-to-artistic, clear-to-foggy, synthetic-to-real) using only labeled source data and unlabeled target data during training. It measures how well the model mitigates domain bias and adapts to unseen target distributions without target annotations. Use when the user wants to benchmark on PASCAL VOC 2007+2012, Clipart1k, Watercolor2k, Cityscapes, Foggy Cityscapes, SIM10K, or asks about evaluating this task...
Probes few-shot image classification generalization across diverse domains and highly variable task regimes (2–20 ways, 1–20 shots) without relying on pre-trained backbones. Use when the user wants to benchmark on Meta-Album, or asks about evaluating this task. Reports accuracy.
Evaluates cross-domain knowledge transfer for click-through rate (CTR) prediction by measuring how well a model trained on a source domain generalizes to a target domain with non-overlapping features. It probes context-aware feature translation and explicit knowledge augmentation in recommendation systems. Use when the user wants to benchmark on Amazon, Taobao, Alibaba Production, or asks about evaluating this task. Reports AUC.
Evaluates continual reinforcement learning capabilities in robotic simulation, specifically measuring how well agents retain performance on previously learned tasks while learning new sequential tasks. It probes catastrophic forgetting, transfer effects, and intrinsic task difficulty across line-following, object-pushing, and reaching benchmarks. Use when the user wants to benchmark on CRoSS, or asks about evaluating this task. Reports average cumulated score.
Evaluates vision-language models' ability to use a crop-and-zoom tool for high-resolution visual question answering, disentangling intrinsic capability improvements from tool-induced gains and harms across multiple benchmarks. Use when the user wants to benchmark on VStar, HR-Bench 4k/8k, VisualProbe Easy/Medium/Harm, or asks about evaluating this task. Reports accuracy.
Evaluates self-supervised remote sensing representations across classification and segmentation tasks using optical and radar-optical inputs. Probes representation quality via finetuning, linear/nonlinear probing, kNN, and clustering. Use when the user wants to benchmark on BigEarthNet, fMoW-Sentinel, EuroSAT, Canadian Cropland, DFC2020, DW-Expert, MARIDA, or asks about evaluating this task. Reports mAP, Top 1 Acc., mIoU.
Evaluates handwritten mathematical expression recognition (HMER) by measuring exact LaTeX sequence matching, tolerant symbol-level error rates, and structural tree prediction accuracy on complex handwritten formulas. Use when the user wants to benchmark on CROHME, HME100K, or asks about evaluating this task. Reports ExpRate.
Evaluates a model's ability to learn generalizable biometric feature representations in a continual learning setting, specifically measuring generalization to unseen identities across sequential learning steps rather than retaining knowledge of previously seen classes. Use when the user wants to benchmark on CRL-face, CRL-person, LFW, Megaface, or asks about evaluating this task. Reports Top 1 accuracy.
Evaluates the ability to detect the onset of systemic instability (criticality) in simulated AI systems by monitoring performance variance across multiple benchmarks. It probes whether a derivative-based threshold can reliably flag phase transitions before functional collapse. Use when the user has predictions and gold and needs to compute percentage of correct classifications.
Evaluates the correctness and computational efficiency of closed-form algorithms for computing critical point probabilities in 2D scalar fields under various parametric and nonparametric noise models. The protocol compares these analytical solutions against Monte Carlo sampling baselines across synthetic and real-world scientific datasets to validate accuracy and speed. Use when the user wants to benchmark on Ackley function (synthetic), Gaussian mixture model (synthetic), E3SM climate data, ...
Evaluates the predictive accuracy and inference efficiency of deep learning recommendation models (DLRM) with compressed embedding tables on large-scale advertising click-through rate datasets. It measures Area Under the ROC Curve (AUC) to assess model quality and samples per second to quantify inference throughput under memory-constrained conditions. Use when the user wants to benchmark on CriteoTB, Criteo Kaggle, or asks about evaluating this task. Reports AUC.
Evaluates the predictive quality and system efficiency of deep learning recommendation models on click-through rate prediction. It measures how well parameter-sharing compression techniques maintain model accuracy while reducing memory footprint and improving training and inference latency. Use when the user wants to benchmark on criteo-kaggle, criteo-tb, or asks about evaluating this task. Reports AUC.
Evaluates an AI system's ability to extract and classify budget allocations for Early Warning System (EWS) investments from heterogeneous financial PDF reports. It probes multi-label classification, numerical budget extraction with tolerance, and evidence retrieval/mapping in climate finance contexts. Use when the user wants to benchmark on MDB Evidence Set, or asks about evaluating this task. Reports Accuracy.
Probes LLM-based multi-agent coordination in dynamic, partially observable wildfire disaster response scenarios. It evaluates capabilities such as spatial reasoning, task designation, plan adaptation, and heterogeneous team collaboration under stochastic dynamics and long-horizon objectives. Use when the user wants to benchmark on CREW-Wildfire, or asks about evaluating this task. Reports task success.
Evaluates a model's ability to predict user creditworthiness based on geographic mobility footprints. It probes whether spatiotemporal visitation patterns and region-level credit signals can reliably distinguish users who pay their mobile phone bills from those who do not. Use when the user wants to benchmark on Hangzhou user mobility dataset, or asks about evaluating this task. Reports AUC.
Evaluates a model's ability to detect fraudulent credit card transactions in a streaming context by learning topological and sequential patterns from transaction graphs without manual feature engineering. Use when the user wants to benchmark on Credit Card Transaction Dataset (Feb-Sep), or asks about evaluating this task. Reports AP.
Evaluates context-aware creative intelligence in multimodal and text-only models by assessing their ability to generate creative, contextually relevant content while maintaining visual factuality across diverse functional and creative writing tasks. Use when the user wants to benchmark on Creation-MMBench, or asks about evaluating this task. Reports VFS, Reward.
Evaluates whether multimodal large language models can perform genuine visual reasoning on charts by inferring values from axes and scales, rather than relying on OCR or pre-existing annotations. It probes the model's ability to interpret complex visual structures and perform multi-step estimation on both synthetic and real-world charts. Use when the user wants to benchmark on CRBench, or asks about evaluating this task. Reports Accuracy.
Evaluates the efficiency and data quality of web crawling strategies for LLM pretraining by measuring downstream model performance after training on crawled or selected documents. It compares graph-connectivity-based, random, and pretraining-influence-based URL scoring methods against an oracle baseline. Use when the user wants to benchmark on ClueWeb22-A (English subset), or asks about evaluating this task. Reports Average performance on 22 core tasks (DCLM evaluation recipe).
CrAIBench probes the robustness of Web3 AI agents against context manipulation attacks, specifically memory injection and prompt injection. It evaluates whether agents can maintain user intent and resist adversarial goals when malicious instructions are embedded in historical memory or active prompts. Use when the user wants to benchmark on CrAIBench, or asks about evaluating this task. Reports Targeted Attack Success Rate (ASR).
Evaluates the factual reliability and hallucination resistance of Retrieval-Augmented Generation (RAG) systems on realistic, dynamic, and long-tail questions. It measures how well models avoid generating incorrect information and appropriately abstain when knowledge is missing. Use when the user wants to benchmark on CRAG, or asks about evaluating this task. Reports truthfulness.
Evaluates the ability of instruction-tuned LLMs to perform domain-specific multiple-choice question answering and text generation tasks. It measures how well models fine-tuned on synthetic, corpus-retrieved data generalize to held-out human-annotated benchmarks in biology, medicine, commonsense, recipe generation, and summarization. Use when the user wants to benchmark on ScienceQA (BioQA), MedMCQA (MedQA), CommonsenseQA 2.0 (CSQA), RecipeNLG (RecipeGen), CNN-DailyMail (Summarization), or ask...
Evaluates few-shot crack segmentation performance under low-light conditions using illumination-invariant features. It tests the model's ability to generalize from well-illuminated support images to unseen low-light query images in both synthetic and real-world scenarios. Use when the user wants to benchmark on ll_CrackSeg9k, LCSD, or asks about evaluating this task. Reports mIOU.
Evaluates a model's ability to answer multiple-choice commonsense reasoning questions. It probes whether providing natural language explanations (human or model-generated) alongside questions improves reasoning performance compared to a baseline without explanations. Use when the user wants to benchmark on CQA, or asks about evaluating this task. Reports Accuracy (%).