Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 4,753–4,776 of 23,552 skills
Evaluates the training efficiency and reconstruction fidelity of a 3D Gaussian Splatting method that uses scale reset and entropy-constrained alpha blending to reduce Gaussian list lengths. Use when the user wants to benchmark on Mip-NeRF 360, Deep Blending, Tanks and Temples, or asks about evaluating this task. Reports PSNR.
Evaluates e-commerce search models on query-product semantic matching, ranking, and multiclass relevance classification. It probes the system's ability to distinguish between Exact matches, Substitutes, Complements, and Irrelevant products across English, Spanish, and Japanese. Use when the user wants to benchmark on Shopping Queries Dataset, or asks about evaluating this task. Reports nDCG.
Evaluates the accuracy and robustness of a neural network-integrated Unscented Kalman Filter for monocular pose tracking of tumbling noncooperative spacecraft. It probes the system's ability to maintain steady-state position and orientation accuracy under domain gaps between synthetic training data and real hardware-in-the-loop test images. Use when the user wants to benchmark on SHIRT, or asks about evaluating this task. Reports e_pose.
Evaluates EFCE solvers on a parametric sequential bargaining game modeling smuggling and inspection. It probes the solver's ability to handle multi-round negotiations, bribery, and deterrence to maximize social welfare. Use when the user wants to benchmark on Sheriff, or asks about evaluating this task. Reports Social Welfare (SW).
Evaluates end-to-end optical music recognition (OMR) systems on their ability to transcribe scanned sheet music images into standardized **kern musical notation. It probes layout analysis, staff/page-level transcription accuracy, and robustness across diverse musical textures such as monophony, pianoform, and quartets. Use when the user wants to benchmark on Sheet Music Benchmark (SMB), or asks about evaluating this task. Reports OMR-NED.
Evaluates the temporal understanding and video-language alignment capabilities of Large Video-Language Models (LVLMs) across three multi-modal video benchmarks. It probes the model's ability to answer questions about video content, track temporal changes, and comprehend complex video sequences without relying on single-frame cues. Use when the user wants to benchmark on VideoBench, MVBench, TempCompass, or asks about evaluating this task. Reports benchmark accuracy (VideoBench, MVBench, TempC...
Evaluates a model's ability to predict the next item in an anonymous user session based on sequential click history. It probes the model's capacity to capture short-term user intent and higher-order item correlations within dynamic session contexts. Use when the user wants to benchmark on YooChoose, Diginetica, or asks about evaluating this task. Reports Hit@20.
Evaluates the computational efficiency (latency and memory) and explanation quality of a Shapley value-based neural network explainer framework against baseline implementations across standard vision models. Use when the user has predictions and gold and needs to compute latency.
Probes how well SHAP-based feature attribution divergence correlates with prediction/output diversity across unsupervised anomaly detection algorithms, and quantifies how this explanation-driven diversity impacts the robustness and accuracy of model ensembles. Use when the user wants to benchmark on ADBench subset (16 datasets: anthyroid, breastw, glass, Hepatitis, Lymphography, mammography, PageBlocks, Pima, Stamps, thyroid, vertebral, vowels, WBC, Wilt, wine, yeast), or asks about evaluatin...
This benchmark evaluates fine-grained hallucination in Large Vision-Language Models (LVLMs) by testing their faithfulness to visual inputs and factuality against external knowledge. It measures model performance under clean conditions and across hierarchical input perturbations (image, instruction, and combination levels) to assess hallucination resistance. Use when the user wants to benchmark on SHALE, or asks about evaluating this task. Reports accuracy, non-hallucination rate.
Evaluates the predictive accuracy and training efficiency of a hierarchical graph neural network that uses Kirchhoff Forest-based stochastic coarsening for graph classification. The benchmark probes whether multi-resolution graph decomposition can maintain competitive performance while significantly reducing computational costs across molecular and social network domains. Use when the user wants to benchmark on MolHIV, MolPPA, COLLAB, DD, REDDIT-MULTI-12K, or asks about evaluating this task. ...
Evaluates the robustness of LLM safety alignment under in-distribution, cross-domain, and noisy supervision settings. It probes whether geometry-aware optimization preserves safety performance while resisting distribution shift and corrupted preference labels. Use when the user wants to benchmark on PKU-SafeRLHF-30K, HH-RLHF-Safety, Do-Not-Answer, HarmBench, SaladBench, or asks about evaluating this task. Reports Win Rate (WR).
Evaluates audio LLMs' ability to comprehend multi-speaker conversations while selectively focusing on a target speaker and ignoring bystanders for privacy. It measures both general audio understanding and selective hearing capability under different instruction modes. Use when the user wants to benchmark on SH-Bench, or asks about evaluating this task. Reports Selective Efficacy (SE).
Evaluates classical and neural-symbolic planners on 3D scene graph environments by measuring their ability to generate valid action sequences for task-driven goals within a strict time limit. Use when the user wants to benchmark on SGPlan, or asks about evaluating this task. Reports task completion.
Evaluates vision-language models on multi-frame spatial reasoning and grounding in volumetric MRI scans. It probes the model's ability to answer clinical questions, generate free-text reasoning, and accurately localize anatomical findings across single slices or full 3D volumes. Use when the user wants to benchmark on SGMRI-VQA, or asks about evaluating this task. Reports A-Score.
Evaluates a model's ability to generate structured scene graphs from images without predefined object boxes. It probes visual relationship reasoning, object detection accuracy, and the model's capacity to produce structurally valid outputs under strict spatial and categorical matching criteria. Use when the user wants to benchmark on VG150, PSG, or asks about evaluating this task. Reports Recall.
Evaluates a model's ability to recognize human emotions (neutral, negative, positive) from dynamic gesture videos. It specifically probes robustness to class imbalance and performance under low-light/high-motion conditions using multimodal inputs. Use when the user wants to benchmark on SGED, or asks about evaluating this task. Reports Accuracy.
Evaluates the ability to align 3D scene graphs by matching semantic entities across scenes with varying spatial overlap and environmental changes. It further tests downstream 3D point cloud registration and mosaicking capabilities using the predicted node alignments to initialize geometric correspondence extraction. Use when the user wants to benchmark on 3RScan (generated sub-scene pairs), or asks about evaluating this task. Reports MRR.
Evaluates models on group activity recognition (GAR) and temporal group activity localization (TGAL) using 3D skeleton sequences from basketball games. It probes spatio-temporal interaction modeling, long-term dependency handling, and the ability to leverage multi-view motion capture data for complex team tactics. Use when the user wants to benchmark on SGA-INTERACT, or asks about evaluating this task. Reports accuracy (mAcc./Top3-mAcc./oAcc.).
Probes whether language models trained via supervised fine-tuning (SFT) truly learn reasoning and planning capabilities or merely memorize instruction templates. It tests generalization across unseen action mappings (instruction variations) and increased grid/card complexity (difficulty variations). Use when the user wants to benchmark on Sokoban, General Points, or asks about evaluating this task. Reports exact-match accuracy.
Evaluates AI-generated image detectors on their ability to generalize across diverse generative models (GANs, diffusion) and resist content bias. It probes robustness using conventional benchmarks, a new content-preserving benchmark (TwinSynths), and low-level vision/perceptual benchmarks to measure how well models rely on texture vs. semantic artifacts. Use when the user wants to benchmark on Conventional benchmark, TwinSynths, Low-level vision and perceptual benchmarks, or asks about evalua...
Evaluates large language models' ability to generate coherent, logically structured financial investment opinions based on company context and questions. It probes reasoning over mere knowledge retrieval by testing performance across varying degrees of familiarity and novelty in companies and questions. Use when the user wants to benchmark on sFIOG, or asks about evaluating this task. Reports ROUGE-L.
Evaluates the generalization and sensitivity of video self-supervised learning models to domain shifts, downstream sample sizes, action similarity, and task shifts beyond action recognition. Use when the user wants to benchmark on UCF-101, NTU-60, FineGym (Gym-99), Something-Something-v2, EPIC-Kitchens-100, Charades, AVA, or asks about evaluating this task. Reports accuracy.
Evaluates the generalization and sensitivity of video self-supervised learning models across four factors: domain shift, sample efficiency, action granularity, and task diversity. It probes how well pre-trained representations transfer to diverse downstream datasets, varying finetuning sample sizes, fine-grained actions, and tasks beyond standard action recognition. Use when the user wants to benchmark on UCF-101, NTU-60, FineGym (Gym-99), Something-Something-v2, EPIC-Kitchens-100, Charades, ...