Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 8,089–8,112 of 20,861 skills
Evaluates deep learning models for binary malware traffic detection using raw network bytestreams. It probes the model's ability to distinguish benign from malicious network flows or packets without relying on handcrafted domain features. Use when the user wants to benchmark on USTCTFC2016, or asks about evaluating this task. Reports accuracy.
This evaluation probes the translation quality and contextual consistency of two commercial machine translation systems (DeepL and Supertext) by having professional raters perform blind pairwise comparisons on full documents. It specifically measures whether LLM-based long-context translation yields superior document-level coherence compared to traditional segment-level systems. Use when the user wants to benchmark on Unspecified source documents, or asks about evaluating this task. Reports p...
Evaluates LLMs' ability to extract and structure information from unstructured text into deep, multi-layer nested JSON formats. It probes format fidelity, field correctness, and structural completeness across varying nesting depths and domains. Use when the user wants to benchmark on DeepJSONEval, or asks about evaluating this task. Reports detailed score.
Evaluates a model's ability to perform counterfactual reasoning in biomedical settings by predicting clinical trial outcomes under perturbations of either the outcome measure or the study arm, using similarity-based retrieval to construct counterfactual pairs. Use when the user wants to benchmark on CT open evaluation sample, or asks about evaluating this task. Reports clinical trial outcome prediction.
Binary classification capability to distinguish background noise from gravitational-wave signals (specifically BBH and SGLF classes) in time-series data. It probes the model's ability to generalize to unseen gravitational wave anomalies using deep latent features. Use when the user wants to benchmark on HDR A3D3 gravitational-wave dataset, or asks about evaluating this task. Reports AUC.
Evaluates a unified multimodal model's capabilities in text-to-image generation, image editing, and world-knowledge reasoning. It probes semantic alignment, long-horizon instruction following, fine-grained attribute binding, and precise text rendering across diverse scenarios. Use when the user wants to benchmark on GenEval, DPG-Bench, UniGenBench, WISE, T2I-CoREBench, ImgEdit, GEdit-EN, UniREditBench, RISE, CVTG-2K, or asks about evaluating this task. Reports GenEval.
Evaluates furniture detection, segmentation, instance retrieval, and set retrieval in indoor scenes. It probes occlusion robustness, fine-grained attribute-based feature learning, and spatial co-occurrence modeling for interior design understanding. Use when the user wants to benchmark on DeepFurniture, or asks about evaluating this task. Reports AP, ACC@K.
Evaluates click-through rate (CTR) prediction models by measuring their ability to correctly rank clicked versus non-clicked instances and output calibrated click probabilities. Use when the user wants to benchmark on Criteo Dataset, Company* Dataset, or asks about evaluating this task. Reports AUC.
Evaluates the robustness of audio-based biometric authentication systems against deepfake speech synthesis attacks. It measures how easily voice cloning models can bypass speaker verification and how effectively anti-spoofing detectors can distinguish genuine from synthetic speech. Use when the user wants to benchmark on AISHELL-3, or asks about evaluating this task. Reports Bypass Rate.
Evaluates large vision-language models on fine-grained visual perception, grounding, hallucination mitigation, and multimodal reasoning. It specifically probes the model's ability to autonomously use image zoom-in tools for interleaved visual-linguistic reasoning (iMCoT) to solve high-resolution and complex visual tasks. Use when the user wants to benchmark on V* Bench, HR-Bench, refCOCO / refCOCO+ / refCOCOg / ReasonSeg, POPE, MathVista, MathVerse, or asks about evaluating this task. Reports...
Binary classification of protein sequences to determine if they are extracellular matrix (ECM) proteins. It evaluates the model's ability to handle class imbalance and generalize across different species and feature extraction methods. Use when the user wants to benchmark on benchmark dataset, independent dataset, ECMPride dataset, or asks about evaluating this task. Reports balanced accuracy.
Evaluates the emotional expressivity and transferability of a generated multi-turn spoken dialogue dataset by training speech emotion recognition models and measuring their classification performance on held-out and zero-shot test sets. Use when the user wants to benchmark on DeepDialogue (SER subset), RAVDESS, or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to predict click-through rates for display advertisements by combining raw image pixels with contextual features. It probes the model's capacity to learn high-level visual semantics and complex nonlinear interactions for ranking and probability calibration in a highly imbalanced, real-world advertising setting. Use when the user wants to benchmark on Commercial Display Ad Dataset (2015), or asks about evaluating this task. Reports relative AUC.
This benchmark evaluates the ability of multi-modal embedding classifiers to distinguish real human motion videos from AI-generated ones. It probes semantic consistency detection, robustness to video laundering (resolution/compression), and generalization to unseen generative models. Use when the user wants to benchmark on DeepAction, or asks about evaluating this task. Reports accuracy.
Evaluates end-to-end speech recognition accuracy across diverse acoustic conditions including clean read speech, accented speech, and noisy speech in English and Mandarin. It benchmarks model performance against both automated baselines and human transcribers to measure real-world applicability. Use when the user wants to benchmark on WSJ eval'92, WSJ eval'93, LibriSpeech test-clean, LibriSpeech test-other, VoxForge Accented Speech, CHiME eval clean, CHiME eval real, CHiME eval sim, Baidu int...
Evaluates an agent's ability to perform long-horizon, multi-step web research to answer complex factual questions. It probes the model's capacity for iterative search, evidence aggregation, and adaptive reasoning under both reproducible offline constraints and live web environments. Use when the user wants to benchmark on BrowseComp-Plus, BrowseComp, GAIA, xbench-DeepSearch, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates a model's ability to infer high-resolution (1km×1km, hourly) PM2.5 concentrations across an urban area using sparse mobile and fixed sensor data combined with multi-scale urban features. It probes spatial-temporal prediction capabilities and measures how well the model integrates local, neighboring, and macro-scale regional transport dynamics to improve air quality estimation accuracy. Use when the user wants to benchmark on Beijing PM2.5 Mobile Sensing Dataset, or as...
Evaluates a model's ability to learn optimal dynamic hedging strategies for financial derivatives under discrete trading and varying risk preferences. The protocol simulates market paths using a Heston stochastic volatility model and trains a neural network to minimize a convex risk measure of the terminal hedging error. Performance is assessed out-of-sample against a theoretical benchmark. Use when the user wants to benchmark on Discretized Heston model, or asks about evaluating this task. R...
Evaluates the effectiveness of a three-stage neural network compression pipeline (pruning, trained quantization, and Huffman coding) in reducing model storage size while preserving classification accuracy on standard computer vision benchmarks. Use when the user wants to benchmark on MNIST, ImageNet (ILSVRC-2012), or asks about evaluating this task. Reports Top-1 Accuracy.
Evaluates the ability of a deep learning model to detect and classify structural bias in heuristic optimization algorithms by analyzing raw performance distributions against a uniform null hypothesis. Use when the user wants to benchmark on BIAS toolbox heuristic pool on $f_0$, or asks about evaluating this task. Reports detection accuracy.
Evaluates a split-inference system's ability to perform real-time semantic segmentation on driving video streams while compensating for simulated network communication delays. It probes temporal prediction capabilities and feature fusion under latency constraints. Use when the user wants to benchmark on BDD100K, or asks about evaluating this task. Reports mIoU.
Evaluates a robot policy's ability to perform bimanual dexterous manipulation tasks under varying levels of tactile dependency. It probes visual-propriocceptive coordination, dynamic object interaction, and contact-rich force control. Use when the user wants to benchmark on DECO-50, or asks about evaluating this task. Reports Success Rate.
This benchmark evaluates a model's ability to decipher undersegmented ancient scripts by aligning unknown character sequences with known language stems. It probes phonological reasoning and unsupervised segmentation capabilities without relying on known language proximity or complete word boundaries. Use when the user wants to benchmark on Gothic, Ugaritic, Iberian, or asks about evaluating this task. Reports P@10.
Evaluates zero-shot question answering robustness against social biases, specifically testing how models adapt to ambiguous versus unambiguous contexts without relying on internal stereotypical knowledge. Use when the user wants to benchmark on BBQ, or asks about evaluating this task. Reports accuracy.