Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 3,529–3,552 of 22,871 skills
Evaluates source-free domain adaptive object detection (SF-DAOD) and unsupervised domain adaptation (UDA) across natural and medical imaging domains. Probes a model's ability to align features and reduce pseudo-label noise when adapting to a target domain without access to target labels during training. Use when the user wants to benchmark on Cityscapes, Foggy Cityscapes, KITTI, SIM10k, BDD100k, RSNA-BSD1K, INBreast, DDSM, or asks about evaluating this task. Reports mAP.
Evaluates semantic segmentation models for day-and-night cloud detection using thermal infrared imagery, specifically testing how auxiliary environmental features improve segmentation accuracy and how well models transfer across different satellite sensors and resolutions. Use when the user wants to benchmark on Landsat Main, or asks about evaluating this task. Reports mIoU.
Evaluates multimodal language models' ability to perform agentic reasoning with images, specifically requiring dynamic visual manipulation and tool-use to solve complex tasks like rotation, jigsaw assembly, and instrument reading. It probes whether models can iteratively process, crop, or transform visual inputs to extract information or solve spatial problems. Use when the user wants to benchmark on TIR-Bench, or asks about evaluating this task. Reports accuracy.
Evaluates the deployment efficiency and accuracy of neural networks on commodity microcontrollers (MCUs) under strict memory and latency constraints. It measures inference speed, memory footprint, and task-specific accuracy across vision, audio, and anomaly detection workloads. Use when the user wants to benchmark on TinyMLPerf (VWW, KWS, AD), or asks about evaluating this task. Reports Accuracy (%), Latency (ms).
Evaluates the efficiency and accuracy of LLM benchmarking by testing how well a small, strategically selected subset of examples predicts overall model performance on standard evaluation scenarios. Use when the user wants to benchmark on HELM, MMLU, AlpacaEval 2.0, Open LLM Leaderboard, or asks about evaluating this task. Reports estimation error.
Evaluates deepfake detectors on distinguishing real from synthetic audio and video, and on attributing synthetic audio to specific TTS generators. It probes robustness to post-processing (DTW alignment, augmentation) and video compression. Use when the user wants to benchmark on TIMIT-TTS, or asks about evaluating this task. Reports AUC.
This benchmark evaluates a model's ability to detect time-dependent and physical mistakes in robotic task executions from video. It specifically probes temporal reasoning, semantic task violation detection, and sim-to-real generalization by comparing frame-level anomaly predictions against ground-truth annotations. Use when the user wants to benchmark on BridgeData V2, Multi-robot dataset, or asks about evaluating this task. Reports F1.
Evaluates Large Language Models' Theory of Mind (ToM) reasoning capabilities across reading comprehension and interactive dialogue scenarios. It specifically probes the model's ability to track character beliefs over time, assess answerability, and determine information access, with a strong focus on first-order and higher-order (up to third-order) belief reasoning. Use when the user wants to benchmark on ToMI, BigToM, FanToM, or asks about evaluating this task. Reports accuracy.
Evaluates the forecasting accuracy of a time series foundation model across diverse real-world and synthetic datasets. It probes the model's ability to capture temporal dynamics, periodicity, and multi-scale patterns over a fixed context window to predict future values. Use when the user wants to benchmark on ETT1, ETT2, Exchange Rate, M1 Monthly, M1 Quarterly, M1 Yearly, M5, Monash M3, NN5, Traffic, Weather, M4 Monthly, Entsoe, Solar with Weather, UK Covid, Sensor Data, or asks about evaluat...
Evaluates the effectiveness of individual architectural modules (e.g., normalization, decomposition, embedding, feedforward types) across diverse time-series forecasting scenarios to identify optimal configurations and predict performance without training. Use when the user wants to benchmark on PEMS03, ETT, Electricity, Social (Unemployment), or asks about evaluating this task. Reports MSE.
Evaluates a large decoder-only Transformer model on standard time series benchmarks for forecasting, imputation, and anomaly detection. It probes the model's few-shot generalization, scalability, and robustness in data-scarce scenarios compared to encoder-only baselines. Use when the user wants to benchmark on ETT, ECL, Traffic, Weather, PEMS, UCR Anomaly Archive, or asks about evaluating this task. Reports MSE.
Evaluates the capability of LLM-based models to forecast multivariate time series across multiple prediction horizons and real-world datasets. It probes pattern-aware temporal modeling and semantic alignment by measuring prediction accuracy under a channel-independent, rolling forecasting setup. Use when the user wants to benchmark on ETTh1, ETTh2, ETTm1, ETTm2, Weather, ECL, Traffic, or asks about evaluating this task. Reports MSE.
Evaluates zero-shot forecasting performance of time series foundation models across diverse datasets and horizons. Probes model capability to capture structural temporal patterns (trend, seasonality, stationarity, complexity) and generalizes to unseen data without leakage. Use when the user wants to benchmark on TIME Benchmark, or asks about evaluating this task. Reports MASE, CRPS.
Evaluates the ability of LLMs and MLLMs to diagnose anomalies in univariate and multivariate time series data. It probes the models' capacity for structured reasoning (generating a 'Thought') and precise action classification ('ActionID') based on raw or visualized temporal data. Use when the user wants to benchmark on RATs40K, or asks about evaluating this task. Reports Label Matching F1.
Evaluates long-term time series forecasting capabilities of foundation models in both zero-shot (unseen datasets) and in-distribution (fine-tuned) settings across multiple prediction horizons. Use when the user wants to benchmark on ETTh1, ETTh2, ETTm1, ETTm2, Weather, Global Temp, or asks about evaluating this task. Reports MSE.
Evaluates how well an automatic multimodal transformer predicts fine-grained human feedback (quality scores, region-level misalignment heatmaps, and text misalignment annotations) on text-to-image generation tasks. Use when the user wants to benchmark on TIFA, or asks about evaluating this task. Reports correlation.
Evaluates the impact of tiered data management (L1–L3) on model performance across general knowledge, reasoning, math, and code domains. It compares models trained on different data quality tiers and different training strategies (mix vs. tiered) to validate data curation and scheduling efficacy. Use when the user wants to benchmark on OpenCompass Benchmarks, or asks about evaluating this task. Reports Average benchmark scores.
Evaluates the ability of traditional and deep learning denoising algorithms to recover sinusoidal dark matter signals from ultra-long, noisy time series data collected by the ABRACADABRA experiment. Use when the user wants to benchmark on TIDMAD, or asks about evaluating this task. Reports mean square error.
Probes a model's ability to learn from inherently subjective or disagreed-upon annotations by treating each annotator's label as a separate example, rather than aggregating them into a single ground truth label. Use when the user wants to benchmark on TID-8, or asks about evaluating this task. Reports exact match accuracy.
Evaluates a robot's ability to perform adversarial interaction and temporal reasoning through a structured pick-and-place board game. It tests the integration of vision-based contour segmentation, game-state reasoning via Minimax, and precise motion planning under hardware constraints. Use when the user wants to benchmark on Yale-CMU-Berkeley objects set, or asks about evaluating this task. Reports t_sub.
Evaluates vision-language models on their ability to generate accurate and linguistically coherent summaries of text-heavy multimodal presentations. It probes cross-modal alignment, long-context understanding, and the impact of different input modalities (raw video, slides, transcripts, interleaved pairs) and token budgets on summarization quality. Use when the user wants to benchmark on TIB-bench, or asks about evaluating this task. Reports $IbR_{overall}$.
Evaluates click models on predicting user click sequences and estimating document relevance from search logs. It also measures how well models recover the underlying distribution of real click data (distributional coverage) and perform when document lists are poorly ranked. Use when the user wants to benchmark on TianGong-ST, or asks about evaluating this task. Reports LL.
Evaluates a model's ability to generate dynamic video scene graphs by predicting inter-object relationships and attributes across multiple temporal frames. It specifically probes temporal consistency, handling of occlusions, and modeling of long-range dependencies in both ground-level and aerial video footage. Use when the user wants to benchmark on ASPIRe, AeroEye-v1.0, or asks about evaluating this task. Reports R@20, mR@20.
Evaluates multimodal large language models on image manipulation, visual perception, mathematical reasoning, and general vision-language tasks. It probes whether autonomous code generation and execution for image processing improves downstream accuracy and reduces hallucination. Use when the user wants to benchmark on MME-RealWorld, HR Bench, MathVista, Hallucination bench, MMStar, or asks about evaluating this task. Reports accuracy.