Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 7,129–7,152 of 20,839 skills
Evaluates the predictive performance of deep neural networks on tabular data across classification and regression tasks, comparing them against tree-based models and other DNNs. It probes how well architectures handle numerical-only versus heterogeneous (numerical + categorical) features at different dataset scales. Use when the user wants to benchmark on Grinsztajn45, or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to generate polygon segmentation masks for generalized referring navigable regions in autonomous driving scenes based on natural language navigation instructions. It specifically probes handling of single-target, multi-target, and no-target scenarios without biasing towards trivial existence predictions. Use when the user wants to benchmark on GRiN-Drive, or asks about evaluating this task. Reports msIoU.
Evaluates aerial-ground cooperative 3D object detection and multi-object tracking in simulated urban environments. Probes cross-view feature alignment, occlusion handling, and communication efficiency under dynamic drone altitudes. Use when the user wants to benchmark on Griffin, or asks about evaluating this task. Reports AP, AMOTA.
Evaluates the ability of embodied agents to learn long-horizon planning and navigation tasks using only terminal rewards, and tests the effectiveness of distilling policies from simplified gridworld experts into visual agents via imitation learning. Use when the user wants to benchmark on PointGoal Navigation, Furniture Moving, 3 vs. 1 Football, or asks about evaluating this task. Reports SPL.
Evaluates 3D semantic segmentation capabilities for power line infrastructure using multi-modal LiDAR and image data. It probes a model's ability to accurately classify geometric and visual features into 11 distinct classes, including critical assets like pylons, cables, and insulators. Use when the user wants to benchmark on GridNet-HD, or asks about evaluating this task. Reports mIoU.
Evaluates a model's ability to segment arbitrary numbers of target objects (including zero) in an image based on a natural language expression. It probes multi-target localization, no-target rejection, and robustness to complex linguistic structures like counting and compound relations. Use when the user wants to benchmark on gRefCOCO, or asks about evaluating this task. Reports generalized IoU (gIoU).
Tests a model's ability to ground natural language expressions that may refer to zero, one, or multiple objects in an image. The model must output a corresponding set of bounding boxes rather than a single box, evaluating its capacity for multi-target and no-target referring expression comprehension. Use when the user wants to benchmark on gRefCOCO, or asks about evaluating this task. Reports set-matching accuracy.
Evaluates the capability of seismic models to detect earthquake events and precisely pick P- and S-wave arrival times from continuous three-component waveform data. It measures both detection accuracy and temporal picking precision under a fixed time-tolerance constraint. Use when the user wants to benchmark on STEAD (Stanford Earthquake Dataset), or asks about evaluating this task. Reports F1.
This benchmark evaluates large language models' ability to answer multiple-choice questions across 45 diverse academic, professional, and governmental subjects in Greek. It specifically probes native language fluency, cultural grounding, and domain-specific knowledge retention under zero-shot and few-shot prompting conditions. Use when the user wants to benchmark on GreekMMLU, or asks about evaluating this task. Reports accuracy.
This benchmark probes free-text legal reasoning and multi-hop statutory citation using Greek Bar exam questions. It requires models to analyze case facts, cite relevant Greek legal articles, and produce open-ended legal analysis. Performance is measured across three dimensions: factual accuracy, correct statutory citation, and quality of legal reasoning. Use when the user wants to benchmark on GreekBarBench, or asks about evaluating this task. Reports Mean.
Evaluates the recommendation accuracy and inference efficiency of GNN-based and matrix factorization models on real-world interaction datasets. It specifically probes how well models perform under standardized evaluation conditions that account for varying negative sampling strategies and representation dimensions. Use when the user wants to benchmark on yelp2018, gowalla, amazon-book, or asks about evaluating this task. Reports NDCG@20.
Evaluates a 3D gravity inversion framework's ability to recover subsurface density structures from gravity data, measuring both computational efficiency (scaling and speedup) and geological accuracy (recovery of synthetic anomalies and fit to field observations). Use when the user wants to benchmark on Synthetic and Field Gravity Inversion Benchmarks, or asks about evaluating this task. Reports speedup.
Evaluates robotic perception and manipulation capabilities in highly cluttered, real-world environments. It benchmarks instance segmentation, 6D object pose estimation, and 6-DoF grasp detection under varying levels of occlusion and scene complexity. Use when the user wants to benchmark on GraspClutter6D, or asks about evaluating this task. Reports Grasp Success Rate (GSR).
This benchmark evaluates the functional impact of 6D object pose estimation and 3D mesh reconstruction methods on robotic grasping performance. It measures how geometric inaccuracies and spatial pose errors propagate to affect the success rate of physics-based grasping attempts in simulation. Use when the user wants to benchmark on YCB-Video (YCB-V), or asks about evaluating this task. Reports grasping success.
This evaluation probes an LLM's ability to generate correct SPARQL queries from natural language questions across diverse knowledge graphs. It measures how well the model can navigate graph structures, handle complex queries, and produce executable results that match ground-truth answers. Use when the user wants to benchmark on WebQuestionsSP (WQSP), ComplexWebQuestions (CWQ), QALD-7, QALD-10, SPINACH, WikiWebQuestions (WWQ), or asks about evaluating this task. Reports F1-score.
Evaluates the test accuracy of single-shot pruning methods at initialization on image classification tasks. It measures how well a pruned sub-network can be trained and generalizes compared to baselines like SNIP and random pruning. Use when the user wants to benchmark on CIFAR-10, CIFAR-100, Tiny-ImageNet, ImageNet, or asks about evaluating this task. Reports test accuracy.
Evaluates multimodal language models' ability to understand language grounding and intuitive physics principles through video-based question answering. It probes capabilities like object detection, feature recognition, and physical plausibility reasoning using simulated Unity environments. Use when the user wants to benchmark on GRASP, or asks about evaluating this task. Reports Accuracy.
Evaluates the predictive accuracy of social recommendation models by forecasting user-item ratings. It jointly leverages user-item interaction graphs and user-user social graphs to learn co-embeddings, testing the model's ability to integrate heterogeneous social tie strengths and opinion signals into rating prediction. Use when the user wants to benchmark on Ciao, Epinions, or asks about evaluating this task. Reports RMSE.
Evaluates the naturalness and prosody quality of synthesized Chinese speech by measuring how closely the generated audio matches human-like pausing and rhythm. It probes the model's ability to capture hierarchical syntactic-semantic dependencies for prosody boundary prediction in text-to-speech systems. Use when the user wants to benchmark on Databaker dataset, or asks about evaluating this task. Reports MOS.
Evaluates graph neural networks' out-of-distribution (OOD) generalization across synthetic, image, molecular, and text graph datasets. It measures how well models maintain performance when tested on domain-shifted splits (e.g., different graph sizes or molecular scaffolds) compared to in-distribution data. Use when the user wants to benchmark on GraphOOD & DrugOOD, or asks about evaluating this task. Reports ROC-AUC, Accuracy.
Evaluates the ability of Graph Neural Networks to induce, compose, and generalize logical rules across synthetic knowledge graphs. It probes relational reasoning, multi-task learning capacity, and catastrophic forgetting in continual learning settings. Use when the user wants to benchmark on GraphLog, or asks about evaluating this task. Reports accuracy.
Evaluates the ability of LLMs to answer knowledge-intensive questions across atomic, aggregated, and multi-hop reasoning scenarios in agricultural, medical, and general domains. It measures how well supervised fine-tuning with synthetic knowledge-graph data improves closed-book QA performance. Use when the user wants to benchmark on SeedEval, PQArefEval, HotpotEval, or asks about evaluating this task. Reports ROUGE-F.
Evaluates session-based recommendation systems by predicting the next item in a user's interaction sequence. It probes the model's ability to capture high-order item relationships and leverage external knowledge graphs for accurate, context-aware ranking. Use when the user wants to benchmark on Tmall, RetailRocket, KKBox, or asks about evaluating this task. Reports P@10.
This evaluation probes a model's ability to perform multihop fact-checking over long-form documents and open-domain QA contexts. It measures how well the system identifies factual inconsistencies or supports claims by reasoning over complex, lengthy grounding texts across general and medical domains. Use when the user wants to benchmark on AggreFact-CNN, AggreFact-Xsum, Summeval, ExpertQA, COVID-Fact, SCIFact, PubHealth, or asks about evaluating this task. Reports balanced accuracy.