Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 4,321–4,344 of 23,503 skills
Evaluates the ability of representation learning models to factorize audio into independent semantic factors (timbre, amplitude, frequency). It measures how well the learned latent space aligns with these ground-truth factors using standard disentanglement metrics. Use when the user wants to benchmark on SynTone, or asks about evaluating this task. Reports MIG.
Evaluates the expected cost incurred by a crowdsourcing platform when inferring service states from biased, strategic user reviews under different baseline mechanisms. Use when the user has predictions and gold and needs to compute system loss.
This benchmark evaluates 2D and 3D point tracking capabilities across diverse synthetic domains, including rapid camera motion, articulated objects, and occlusions. It probes a model's ability to maintain spatio-temporal correspondence, handle depth-adaptive spatial errors, and correctly classify occlusion or out-of-frame status under significant distribution shifts. Use when the user wants to benchmark on SynthVerse, or asks about evaluating this task. Reports AJ_3D.
Evaluates the effectiveness of a synthetic referring expression dataset for training language-guided video object segmentation models. It measures segmentation accuracy when models are trained on synthetic versus human annotations and evaluated on standard referring video segmentation benchmarks. Use when the user wants to benchmark on DAVIS-2017, Refer-YouTube-VOS, or asks about evaluating this task. Reports J&F.
Evaluates the ability of LLMs to infer private personal attributes (e.g., occupation, age, location, income) from concatenated user comments. It also assesses the fidelity of synthetic comments compared to real human text via human studies. Use when the user wants to benchmark on SynthPAI, or asks about evaluating this task. Reports 0-1 accuracy.
Probes the ability of an agent to select the optimal tabular data synthesizer for a given dataset and objective (privacy, fidelity, or utility) based on dataset stress profiles and a capability registry. It evaluates whether stress-aware, intent-conditioned matching outperforms heuristics, zero-shot LLMs, and meta-learning baselines in ranking generative models. Use when the user wants to benchmark on OpenML Tabular Benchmark (Abalone, Bean, IndianLiverPatient, Obesity, faults, insurance, wil...
Evaluates a 3D scene representation's capability for novel view synthesis, relighting, and inverse rendering (estimating diffuse albedo and roughness) from posed RGB images. Use when the user wants to benchmark on Synthetic4Relight, or asks about evaluating this task. Reports PSNR.
Evaluates the downstream utility and statistical fidelity of synthetic tabular data generated by various models. It probes whether synthetic data preserves classification accuracy, model selection rankings, feature importance rankings, and distributional similarity compared to real data. Use when the user wants to benchmark on Tabular Classification from Numerical features benchmark suite (filtered), or asks about evaluating this task. Reports AUROC.
Evaluates how synthetic or human-generated column descriptions impact LLM performance on text-to-SQL tasks, and assesses the quality of LLM-generated descriptions across varying semantic difficulty levels. Use when the user wants to benchmark on BIRD-Bench, or asks about evaluating this task. Reports Mean quality scores.
Evaluates the quality and downstream utility of synthetic medical images generated by GANs by measuring how well classifiers trained on synthetic data perform compared to those trained on real data. It probes the trade-offs between image resolution, label complexity, and sample size on both visual fidelity and predictive performance. Use when the user wants to benchmark on Chest radiographs, Brain CT scans, or asks about evaluating this task. Reports AUC_real - AUC_syn.
Evaluates the generalization capability of synthetic image detectors across different generative models, image resolutions, and real-world sources. It probes whether detectors rely on dataset-specific artifacts or scale-dependent biases rather than learning robust forgery signatures. Use when the user wants to benchmark on SuSy Benchmarking Datasets, or asks about evaluating this task. Reports recall.
Evaluates a flow matching generative model's ability to produce realistic 3D subsurface geological models, both unconditionally and conditioned on sparse borehole data. It probes the model's capacity for geological interpolation, structural feature reconstruction, and probabilistic uncertainty estimation. Use when the user wants to benchmark on Synthetic Geology / StructuralGeo Dataset, or asks about evaluating this task. Reports probabilistic confidence intervals.
Evaluates the distributional fidelity of generated pharmacokinetic and drug-target interaction properties against real data, and measures the utility of the synthetic data for downstream regression tasks. Use when the user wants to benchmark on TDCommons/BindingDB PK & DTI Collection, or asks about evaluating this task. Reports Hellinger Distance (HD).
Evaluates whether synthetic data generated by LLMs can effectively substitute real data for benchmarking NLP models, measuring both absolute performance alignment and relative ranking preservation across tasks. Additionally quantifies the self-bias of LLMs when they generate data and subsequently solve the same tasks. Use when the user wants to benchmark on Headlines, Tweet-News, CrossNER-Literature, CrossNER-Politics, SNIPS, ATIS, or asks about evaluating this task. Reports MSPD.
Probes whether music foundation models encode discrete and continuous Western music theory concepts by training linear or MLP classifiers on their internal audio embeddings. Use when the user wants to benchmark on SynTheory, or asks about evaluating this task. Reports accuracy.
Evaluates the effectiveness of an automated framework for synthesizing paralinguistic speech datasets on downstream paralinguistic text-to-speech generation and event detection tasks. It measures how well the generated data improves model performance in producing and recognizing paralinguistic features like laughter, sighs, and gasps compared to real-world annotated datasets. Use when the user wants to benchmark on SynParaSpeech, or asks about evaluating this task. Reports PMOS.
Evaluates unsupervised anomaly detection in medical imaging by training a generative model on healthy images and reconstructing anomalous inputs. It measures how well the model localizes and segments pathological regions by comparing the reconstruction residuals against ground-truth anomaly masks. Use when the user wants to benchmark on BraTS 2023 (Brain MRI), LiTS (Liver CT), Carotid US, or asks about evaluating this task. Reports Dice.
Evaluates a hybrid neural forecasting framework's ability to predict regional particulate matter (PM1, PM2.5, PM10) concentrations. It probes both average forecasting accuracy across spatial grids and the model's capacity to capture rare, high-impact pollution spikes and extreme events. Use when the user wants to benchmark on ERA5 & CAMS, or asks about evaluating this task. Reports Latitude-Weighted RMSE.
Evaluates the impact of curriculum learning and synthetic data generation on legal LLM fine-tuning. It measures performance across legal summarization, classification, and question-answering benchmarks to determine if synthetic data improves model capabilities over real-data-only baselines. Use when the user wants to benchmark on EurLex-Sum, EurLex, LexGLUE, BigLaw-Bench, CUAD, or asks about evaluating this task. Reports training loss.
Evaluates the quality and alignment of a synthetic passage retrieval test collection (SynDL) by measuring how well system rankings on synthetic relevance judgments match those from official human-annotated TREC Deep Learning Track collections. It probes whether LLM-generated judgments can reliably substitute human assessors for deep relevance evaluation and system ranking. Use when the user wants to benchmark on SynDL, or asks about evaluating this task. Reports Kendall rank correlation coeff...
Evaluates a dual-stream text-to-speech model's ability to generate high-quality, speaker-similar speech with low latency and high efficiency under streaming and offline conditions. It probes the model's robustness to complex text, alignment accuracy, and real-time generation speed compared to autoregressive and interleaved baselines. Use when the user wants to benchmark on LibriSpeech test-clean, SeedTTS test-zh, SeedTTS test-hard, or asks about evaluating this task. Reports RTF.
Evaluates the ability of deep learning models to learn and forecast diverse temporal patterns (trends, periodicities, multivariate dependencies) in time series. It also probes model robustness against varying levels of Gaussian and non-Gaussian noise, as well as resilience to point and pulse anomalies. Use when the user wants to benchmark on SynTSBench, or asks about evaluating this task. Reports MSE.
Evaluates an agent's ability to perform open-vocabulary interactive object search in indoor environments using relational semantic reasoning over 3D scene graphs. It probes exploration efficiency, reasoning accuracy, and computational cost compared to embedding-based and LLM-based planners. Use when the user wants to benchmark on SymSearch, OmniGibson, or asks about evaluating this task. Reports Success Rate (SR), Success weighted by Path Length (SPL).
Probes the ability to extract fine-grained mathematical symbols and their textual descriptions from LaTeX-formatted scientific documents. It evaluates both named entity recognition for identifying symbols and descriptions, and relation extraction for linking them according to specific semantic types. Use when the user wants to benchmark on Symlink, or asks about evaluating this task. Reports F-score.