Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 7,225–7,248 of 20,839 skills
Evaluates zero-shot time series forecasting models across diverse domains, frequencies, and prediction horizons. It probes a model's ability to generalize to unseen multivariate and univariate series, identifying strengths and weaknesses in short-term versus long-term forecasting. Use when the user wants to benchmark on GIFT-Eval, or asks about evaluating this task. Reports MAPE.
Evaluates the geometric complexity, scene diversity, and sim-to-real transfer capability of the Gibson virtual environment. It benchmarks neural rendering pipelines and embodied agents on tasks like depth estimation, scene classification, and navigation. Use when the user wants to benchmark on Gibson, or asks about evaluating this task. Reports Real-World Transfer Error.
This benchmark evaluates vision-language models' ability to detect and predict hidden interactions in mobile user interfaces. It probes whether models can infer concealed gestures (e.g., long presses, swipes) from before-interaction screenshots and task descriptions, localize the interaction target, and anticipate the resulting UI state changes. Use when the user wants to benchmark on GhostUI, or asks about evaluating this task. Reports Accuracy.
This benchmark probes the susceptibility of vision-language models to tone-induced hallucination under controlled negative-ground-truth conditions. It isolates linguistic prompt intensity as the sole variable to measure both the frequency and severity of unsupported content generation when visual evidence is deliberately absent or illegible. Use when the user wants to benchmark on Ghost-100, or asks about evaluating this task. Reports H-Rate.
Evaluates geometric generative reasoning in unified multimodal models by testing their ability to perform stepwise spatial planning, translate visual constraints into executable code or visual sequences, and verify multi-step geometric construction processes. Use when the user wants to benchmark on GGBench, or asks about evaluating this task. Reports VLM-T.
Evaluates a model's ability to quickly adapt an object detector to novel classes using only a few labeled examples, while preserving performance on previously learned base classes. Use when the user wants to benchmark on MS-COCO (G-FSD benchmark), or asks about evaluating this task. Reports AP.
Evaluates geometric deep learning models' out-of-distribution generalization across scientific domains under conditional, covariate, and concept shifts. It probes how different learning paradigms (ERM, domain adaptation, transfer learning, and OOD generalization) perform when provided with varying amounts of target-domain data. Use when the user wants to benchmark on Track (Particle Tracking Simulation), QMOF (Quantum Metal-organic Frameworks), DrugOOD-3D (3D Conformers of Drug Molecules), or...
Evaluates instruction-tuned decoder-only LLMs on Portuguese language understanding tasks, including natural language inference, paraphrase detection, commonsense reasoning, and multiple-choice question answering. Use when the user wants to benchmark on MRPC, RTE, COPA, ENEM 2022, BLUEX, STS, or asks about evaluating this task. Reports F1 score.
Evaluates German NLP models on aspect-based sentiment analysis (ABSA) tasks, including relevance classification, document-level polarity, aspect/sentiment classification, and opinion target extraction. Use when the user wants to benchmark on GermEval17, or asks about evaluating this task. Reports micro F1.
Evaluates extractive question answering and dense passage retrieval capabilities in German. It probes a model's ability to locate precise answer spans within a given context and retrieve relevant passages from a large corpus. Use when the user wants to benchmark on GermanQuAD, or asks about evaluating this task. Reports Exact Match (EM).
Evaluates the quality of text embeddings for German-language documents by measuring how well they cluster into predefined topical categories. It probes a model's ability to capture semantic similarity and domain-specific nuances across different text lengths (titles vs. full texts) and sources. Use when the user wants to benchmark on BlurbsClusteringS2S/P2P, TenKGnadClusteringS2S/P2P, SubredditClusteringS2S/P2P, or asks about evaluating this task. Reports V-measure.
Evaluates Aspect-Based Sentiment Analysis (ABSA) capabilities on German-language restaurant reviews. It probes models on four subtasks: identifying aspect categories, predicting sentiment polarities for aspects, extracting aspect-sentiment pairs, and end-to-end triplet extraction. Use when the user wants to benchmark on GERestaurant, or asks about evaluating this task. Reports F1 Micro.
Evaluates the quality and linguistic fidelity of an Italian generative language model (GePpeTto) by measuring its perplexity across in-domain and out-of-domain corpora, and profiling its lexical and syntactic complexity against human-written Italian text. Use when the user wants to benchmark on Wikipedia (Italian), ItWac, EUR-Lex Italian Laws, la Repubblica & Il Giornale, Forum Comments, or asks about evaluating this task. Reports Perplexity.
Predicts the geothermal gradient (°C/km) across Colombia using geophysical and geological features. Evaluates model generalization on unseen spatial locations and quantifies prediction error against sparse borehole measurements. Use when the user wants to benchmark on Colombia Geothermal Gradient Dataset, or asks about evaluating this task. Reports R2.
Evaluates the quality of object removal and causal visual artifact removal (shadows, reflections) in images, measuring visual fidelity, structural consistency, and artifact suppression. Use when the user wants to benchmark on RORD-Val, RemovalBench, CausRem, or asks about evaluating this task. Reports FID.
Evaluates a model's ability to perform multimodal numerical reasoning on geometric problems by generating executable symbolic programs from text and diagram inputs. The model must fuse cross-modal information to predict step-by-step reasoning programs. These programs are then executed to select the correct multiple-choice answer from the given options. Use when the user wants to benchmark on GeoQA, or asks about evaluating this task. Reports answer accuracy.
This benchmark evaluates a model's ability to perceive and parse geometric diagrams into a structured formal language. It probes fine-grained visual primitive detection (points, lines, circles, planes) and spatial/semantic relations, testing both syntactic correctness and holistic geometric consistency. Use when the user wants to benchmark on GDP-29K, or asks about evaluating this task. Reports F1-score.
Evaluates expert-level multimodal intelligence in geoscience and remote sensing by testing domain knowledge, perceptual grounding, and spatiotemporal reasoning across diverse sensors, disciplines, and task complexities. Use when the user wants to benchmark on GeoMMBench, or asks about evaluating this task. Reports Micro-averaged accuracy.
Evaluates monocular depth estimation models for their ability to produce geometry-preserving depth maps and accurate 3D point clouds without requiring explicit 3D annotations. It probes scale-and-shift recovery, generalization across indoor and outdoor domains, and consistency under differentiable rendering. Use when the user wants to benchmark on NYU V2, ScanNet, KITTI, ETH3D, 2D3D, or asks about evaluating this task. Reports AbsRel.
This benchmark evaluates the geometric reasoning and problem-solving capabilities of multimodal large language models. It tests whether models can accurately interpret geometric diagrams and accompanying text to produce correct final answers or select the right multiple-choice option. Use when the user wants to benchmark on GeoQA, Geometry3K, PGPS9K, MathVista-mini-GPS, or asks about evaluating this task. Reports Top-1 accuracy.
Evaluates a model's ability to perform geometric matrix completion on multi-network recommendation datasets. It probes how well graph neural networks and low-rank representations can integrate cross-network and within-network features to predict missing user-item ratings. Use when the user wants to benchmark on Douban, Flixster, YahooMusic, ML-100K, ML-1M, or asks about evaluating this task. Reports RMSE.
Evaluates the geometric fidelity and surface reconstruction accuracy of neural 3D scene representations (NeRF and Gaussian Splatting variants) against metric-scale laser scan ground truth. Use when the user wants to benchmark on Robotic Manipulation Scenes, or asks about evaluating this task. Reports CD_{P\rightarrow G}.
Evaluates a CNN's ability to perform semantic segmentation on airborne magnetic data to identify three major lithological groups (dykes, plutons, greywackes). It tests transfer learning from synthetic geostatistical data to real-world geological contexts. Use when the user wants to benchmark on Malartic geological model & synthetic augmentations, or asks about evaluating this task. Reports IOU.
Evaluates a model's ability to predict the geographic coordinates (latitude and longitude) of a query image from worldwide visual data. It probes fine-grained location-aware visual semantics and robustness to geographical heterogeneity across urban, regional, and continental scales. Use when the user wants to benchmark on IM2GPS3k, YFCC4K, or asks about evaluating this task. Reports threshold metric.