Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 4,057–4,080 of 23,477 skills
Evaluates graph neural networks and tabular baselines on detecting fraudulent user rings in travel booking networks. It probes the models' ability to classify individual fraud accounts and recover entire fraud ring structures using heterogeneous graph topology and co-occurrence signals. Use when the user wants to benchmark on TravelFraudBench (TFG), or asks about evaluating this task. Reports AUC-ROC.
This benchmark evaluates large language models' ability to perform event-temporal reasoning by resolving explicit, implicit, and vague temporal references across synthetic household event chains. It systematically probes how model performance degrades with increasing event set length and varying levels of temporal explicitness. Use when the user wants to benchmark on TRAVELER, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates the vulnerability of web agents to prompt injection attacks that redirect their intended tasks. It probes how well agents maintain task fidelity under benign conditions versus how susceptible they are to social-engineering and persuasion-based adversarial injections embedded in web interfaces. Use when the user wants to benchmark on TRAP, or asks about evaluating this task. Reports Attack Success Rate (ASR).
Evaluates spatio-temporal forecasting models on traffic speed, volume, and bike flow prediction tasks. It specifically probes whether simple baselines that account for weekly stationarity (historical average plus linear regression on residuals) can match or outperform complex deep learning architectures across diverse transport datasets. Use when the user wants to benchmark on PeMSD7(M), Urban1, NYC Citi Bike, PeMSD4, SZ-taxi, METR-LA, PEMS-BAY, NYC Bike in- and out-flows, Seattle traffic spe...
Evaluates a hybrid CNN-BERT model's ability to detect and correct errors in English translations. It measures performance across different architectural hyperparameters (kernel size, batch normalization) and linguistic granularity levels. Use when the user wants to benchmark on WMT English-German Parallel Corpus Dataset, Open Subtitles Dataset, or asks about evaluating this task. Reports F1-Score (%).
This evaluation probes a parser's ability to construct accurate projective dependency trees for sentences. It measures how well the model identifies correct syntactic heads and their grammatical relations (labels) under both English and Chinese linguistic conditions. Use when the user wants to benchmark on Penn Treebank (PTB) v5, Chinese Treebank (CTB) v5, or asks about evaluating this task. Reports LAS.
Evaluates the faithfulness and class-specificity of Transformer interpretability methods by measuring how well highlighted input features align with model predictions. It probes explanation quality through pixel/token masking, segmentation overlap, and rationale extraction accuracy. Use when the user wants to benchmark on ImageNet Validation (ILSVRC 2012), ImageNet-Segmentation, Movie Reviews, or asks about evaluating this task. Reports AUC (Positive/Negative Perturbation).
Probes a model's ability to rank machine translation candidates by quality and provide fine-grained, dimensionally structured justifications aligned with MQM standards. It also evaluates the model's robustness to candidate ordering (position bias) when performing comparative translation assessment. Use when the user wants to benchmark on WMT-2024 en-es, WMT-2023 en-de, WMT-2023 zh-en, WMT-2022 en-ru, WMT-2021 en-ja, WMT-2021 ja-en, Hard en-ja, Generic, Haiku 100, Haiku Full, or asks about eva...
Evaluates a boosting-tree kernel transfer learning algorithm for financial risk prediction and fraud detection under domain distribution shifts and data sparsity. It measures how well the model adapts from a source domain to a target domain with limited labeled samples, while maintaining computational efficiency and interpretability. Use when the user wants to benchmark on Tencent Mobile Payment Dataset, LendingClub Dataset, Wine Quality Dataset, or asks about evaluating this task. Reports AUC.
TransBench evaluates machine translation models across three industrial capability levels: basic linguistic quality and robustness, domain-specific proficiency (e-commerce/finance), and cultural adaptation (taboo words and honorifics). It probes whether models maintain translation fidelity under input perturbations, adhere to domain terminology, and correctly handle culturally sensitive expressions without omission or over-translation. Use when the user wants to benchmark on TransBench, or as...
Evaluates the linguistic robustness of LLMs by measuring performance degradation when standard English prompts are transformed into 38 regional dialects and ESL varieties. It probes whether models maintain accuracy and instruction-following capabilities across non-standard linguistic variations. Use when the user wants to benchmark on MMLU, ARC, TruthfulQA, GSM8K, HellaSwag, WinoGrande, IFEval, AlpacaFarm, MT-Bench, or asks about evaluating this task. Reports accuracy.
Evaluates the statistical fidelity and practical utility of synthetic human trajectory generation models by measuring how well generated trajectories perform on downstream mobility tasks compared to real trajectories. It probes whether synthetic data can replace real data without performance degradation across recommendation, prediction, labeling, and simulation tasks. Use when the user wants to benchmark on Foursquare Tokyo (TKY), Foursquare Istanbul (IST), Foursquare New York City (NYC), or...
Evaluates the training efficiency and scalability of AlphaFold-like models on GPU clusters. It measures per-step execution time, overall wall-clock training duration, and convergence speed across different hardware configurations and optimization techniques. Use when the user wants to benchmark on OpenFold dataset, or asks about evaluating this task. Reports step time.
Evaluates a training-free iterative inference method for audio source separation. It probes the model's ability to progressively refine noisy audio mixtures (speech or music) by optimizing blending ratios across multiple inference steps without retraining or architectural changes. Use when the user wants to benchmark on VCTK-DEMAND, DNS Challenge v3, MUSDB18-HQ, or asks about evaluating this task. Reports PESQ, UTMOS, uSDR.
Evaluates the quality of automatically generated multilingual word sense disambiguation (WSD) training corpora by training a supervised WSD system (IMS) on them and measuring performance on standard WSD benchmark datasets. It probes whether synthetic sense-annotated data can match or exceed manually annotated corpora, particularly for low-resource languages. Use when the user wants to benchmark on Senseval-2, Senseval-3, SemEval-2007, SemEval-2013, SemEval-2015, or asks about evaluating this ...
This benchmark evaluates a model's ability to classify the intention of a natural language question relative to a database schema. It probes whether the model can distinguish between answerable queries, questions requiring external knowledge, ambiguous queries, grammatically invalid non-SQL questions, and questions unrelated to the schema. Use when the user wants to benchmark on TRIAGESQL, or asks about evaluating this task. Reports Macro F1.
Evaluates the ability of spatio-temporal graph neural networks to forecast traffic speed under normal and construction work zone disruption conditions. It probes how well models integrate heterogeneous work zone data to capture nonlinear spatio-temporal dependencies and maintain accuracy during significant traffic flow deviations. Use when the user wants to benchmark on Richmond, Tyson’s, or asks about evaluating this task. Reports MAE.
This evaluation protocol assesses a model's ability to forecast future traffic speeds on road networks under varying conditions, including the impact of construction workzones. It probes spatio-temporal dependency modeling by measuring prediction accuracy across multiple forecast horizons (15, 30, and 60 minutes) on real-world highway sensor data. Use when the user wants to benchmark on Tyson's Corner, Los-loop, PEMS-BAY, or asks about evaluating this task. Reports RMSE.
Evaluates spatiotemporal models' ability to localize traffic collision events in time and space, and to forecast network-level congestion and emissions. It probes multi-horizon forecasting accuracy and spatial-temporal coherence under simulated disruption scenarios. Use when the user wants to benchmark on NYC Broadway corridor, or asks about evaluating this task. Reports containment_performance.
Evaluates the ability of spatiotemporal GNN models to forecast future traffic flow (speed or occupancy) based on historical sensor data and road network topology. Use when the user wants to benchmark on METR-LA, PEMS-BAY, PeMS04, or asks about evaluating this task. Reports MAE.
Evaluates the accuracy of multi-modal trajectory forecasting models in predicting the final destination of traffic agents (pedestrians and vehicles) over a future time horizon. Use when the user wants to benchmark on SDD, InD, Argoverse, or asks about evaluating this task. Reports Minimum final displacement error.
Evaluates the ability of deep learning models to classify mobile network traffic into specific application categories (e.g., video, music, games) and background services using extracted packet flow features. It specifically probes the effectiveness of semi-supervised VAE-CNN architectures and model pruning techniques for resource-constrained edge devices. Use when the user wants to benchmark on Private Campus Network Dataset, or asks about evaluating this task. Reports Accuracy.
This evaluation probes a model's ability to detect objective corporate events in financial news and translate those detections into actionable, timely trading signals. It measures how effectively the detected events predict short-term stock price movements and generate excess returns compared to a market benchmark. Use when the user wants to benchmark on EDT, or asks about evaluating this task. Reports Winning Rate.
Evaluates a model's ability to generate aspect-specific summaries from clinical abstracts and accurately cite the supporting source sentences. It probes factual recall, conciseness, and traceability in a medical domain setting. Use when the user wants to benchmark on TracSum, or asks about evaluating this task. Reports Claim Recall.