Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 4,369–4,392 of 23,503 skills
Evaluates the acoustic and perceptual quality of singing voice synthesis (SVS) models trained on large-scale multi-singer corpora. It probes direct synthesis capability, cross-domain transfer learning, and data augmentation via joint training. Use when the user wants to benchmark on ACE-Opencpop, ACE-KiSing, or asks about evaluating this task. Reports MOS.
Evaluates a multimodal model's ability to answer visual questions when the query is provided as spoken audio rather than text. It probes speech-vision-language alignment, robustness to synthesized speech variations, and the model's capacity to handle modality-specific prompts. Use when the user wants to benchmark on SEED-Bench, MME, DocVQA, MLS, or asks about evaluating this task. Reports accuracy.
Evaluates how well CLIP aligns text and images by measuring its preference for positive over negative image-text pairs, and analyzes how semantic features (part-of-speech, concreteness, length, frequency, ambiguity) influence this alignment. Use when the user wants to benchmark on SVO-Probes, or asks about evaluating this task. Reports accuracy.
Probes a model's ability to predict social engagement (upvote ratio) from multimodal inputs (images, videos, and text). It evaluates cross-modal fusion and regression capabilities on socially grounded, context-rich data. Use when the user wants to benchmark on SVLD, or asks about evaluating this task. Reports Mean L1point ratio prediction error.
Evaluates a model's ability to classify real-world, cropped street-view house numbers into ten digit classes. It probes robustness to natural scene variations such as background clutter, varying colors, orientations, and focus, which are absent in synthetic datasets like MNIST. Use when the user wants to benchmark on Street View House Numbers (SVHN), or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to refine and correct imperfect SVG code, measuring structural accuracy, visual fidelity, and code efficiency. Use when the user wants to benchmark on SVG-Sophia Code Refinement Benchmark, or asks about evaluating this task. Reports SR.
Evaluates the visual fidelity and text-image alignment of quantized diffusion models by generating images from text prompts and comparing them against reference outputs. It probes whether low-bit quantization preserves distributional similarity, perceptual quality, and human-preferred aesthetics compared to full-precision baselines. Use when the user wants to benchmark on MJHQ-30K, sDCI, or asks about evaluating this task. Reports FID.
This protocol evaluates the ability of singular value decomposition (SVD) and gradient boosting regression trees (GBRT) to decompose unresolved planetary light curves into principal components that physically correspond to specific surface and atmospheric features. It quantifies feature attribution through variance explained, model importance scores, and linear correlations. Use when the user has predictions and gold and needs to compute SVD eigenvalue variance ratio.
This protocol evaluates a singing voice conversion model's ability to transform a source singer's voice into a target singer's timbre while preserving musical naturalness and pitch accuracy. It measures both subjective perceptual quality and objective acoustic fidelity using human ratings and signal processing metrics. Use when the user wants to benchmark on VCTK, NUS-48E, or asks about evaluating this task. Reports MOS (Naturalness), MOS (Similarity).
Evaluates singing voice conversion systems on in-domain (singing-to-singing) and cross-domain (speech-to-singing) speaker conversion. It probes the model's ability to preserve target speaker identity and musical prosody while converting source audio to the target voice. Use when the user wants to benchmark on SVCC 2023, or asks about evaluating this task. Reports perceptual quality (subjective evaluation).
Evaluates large vision-language models' ability to perform sustained temporal reasoning and context tracking across long-form streaming videos. It probes multi-turn dialogue continuity, temporal dependency handling, and complex reasoning skills like counterfactual analysis and spatio-temporal speculation. Use when the user wants to benchmark on SVBench, or asks about evaluating this task. Reports Overall Score (OS).
Evaluates whether NLP models can genuinely solve simple math word problems through arithmetic reasoning versus relying on shallow heuristics like bag-of-words matching or positional cues. It probes model brittleness by testing performance on standard datasets alongside carefully perturbed variants that remove questions or alter operator types. Use when the user wants to benchmark on MAWPS, ASDiv-A, SVAMP, or asks about evaluating this task. Reports accuracy.
Evaluates LLMs' structural and semantic reasoning capabilities in C code vulnerability analysis. It probes whether models rely on memorized patterns or genuinely understand code interdependencies by testing consistency across base, data-flow, control-flow, counterfactual, goal-driven, and predictive scenarios. Use when the user wants to benchmark on SV-TrustEval-C, or asks about evaluating this task. Reports Cons_DFL.
Evaluates Hindi extractive question answering capabilities using multiple-choice questions generated from Wikipedia contexts. It probes a model's ability to comprehend Hindi text, locate relevant information, and select the correct answer from four options under varying context availability settings. Use when the user wants to benchmark on Suvach, or asks about evaluating this task. Reports accuracy.
Evaluates the quality of generated 3D human motions for interactive dialogue, measuring semantic alignment with text/audio, motion quality, audio-motion synchronization, and diversity. Use when the user wants to benchmark on SuSuInterActs, or asks about evaluating this task. Reports R@K.
Evaluates a modular retrieval-augmented generation pipeline for multi-document scientific literature summarization. It probes the model's ability to dynamically generate queries, retrieve relevant papers, and synthesize citation-aware, coherent summaries from multiple sources. Use when the user wants to benchmark on SurveySum, or asks about evaluating this task. Reports Ref-F1.
Evaluates a model's ability to detect and temporally localize anomalous events in long, untrimmed surveillance videos using only video-level labels. It probes the model's robustness to high intra-class variation, ambiguous normal-anomalous boundaries, and varying lighting/occlusion conditions. Use when the user wants to benchmark on Surveillance Anomaly Dataset, or asks about evaluating this task. Reports AUC.
Probes the ability of surrogate-assisted genetic algorithms to optimize continuous 2D functions under tight evaluation budgets, simulating interactive recommendation scenarios where user feedback is sparse and dynamic. Use when the user wants to benchmark on Bohachevsky, Ackley, and Schwefel benchmark functions, or asks about evaluating this task. Reports Best Fitness.
This benchmark evaluates language-guided spatial understanding and reasoning in complex 3D scenes. It probes a model's ability to reason about relative positions, narrative/parametric perspectives, and absolute distances without relying on semantic shortcuts or explicit object names in the prompts. Use when the user wants to benchmark on SURPRISE3D, or asks about evaluating this task. Reports Accuracy (A25/A50).
Evaluates zero-shot surgical video generation models using a four-tiered Surgical Plausibility Pyramid. It probes the model's ability to maintain visual realism while correctly simulating domain-specific surgical causality, instrument handling, tissue feedback, and clinical intent over time. Use when the user wants to benchmark on SurgVeo benchmark, or asks about evaluating this task. Reports Visual Perceptual Plausibility.
Evaluates the correctness and efficiency of gradient-based versus boolean logic-based methods for identifying feature-parameter interactions and transferring trained weights when new features are added to a reinforcement learning model. It measures how well each mapping technique preserves model performance and computational speed during architectural surgery. Use when the user has predictions and gold and needs to compute Interactions Found.
Evaluates a model's ability to predict the next item in a user's interaction sequence by leveraging long-term historical behavior and filtering out noise. It probes the model's capacity to handle varying sequence lengths and efficiently model dynamic user preferences over time. Use when the user wants to benchmark on Taobao, Kuaishou, or asks about evaluating this task. Reports GAUC.
Evaluates multi-modal large language models on hierarchical spatiotemporal reasoning in surgical videos across five clinical dimensions: causal action ordering, cue-action alignment, affordance mapping, micro-transition localization, and anomaly onset tracking. The benchmark tests a model's ability to progressively narrow its focus from global video comprehension to fine-grained frame-level localization while maintaining logical consistency through a chain-of-thought protocol. Use when the us...
This benchmark evaluates fine-grained spatial understanding and reasoning capabilities of vision-language models in real-world driving scenarios. It probes six distinct spatial dimensions: orientation (Yaw), pixel-level localization, depth estimation, pairwise distance, lateral ordering, and front-back relations. Use when the user wants to benchmark on SURDS, or asks about evaluating this task. Reports Score.