Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 9,289–9,312 of 21,341 skills
This evaluation probes a model's ability to detect and classify robotic failures in real-world manipulation tasks. It specifically tests whether a system can distinguish between genuine task-disrupting failures and benign environmental deviations using multimodal observations and nominal demonstrations. Use when the user wants to benchmark on BotFails, Real-π dataset, or asks about evaluating this task. Reports AUROC.
This evaluation probes a model's ability to perform weakly supervised multimodal segmentation of acoustic borehole images by refining threshold-guided pseudo-labels using depth-aligned well logs. It measures how well the predicted segmentation aligns with a provisional target map, testing spatial coherence and multimodal feature fusion rather than absolute geological accuracy. Use when the user wants to benchmark on Antilope25 & Botorosa47 borehole intervals, or asks about evaluating this tas...
Evaluates the ability to refine 6D object poses in cluttered real-world scenes using RGB or RGB-D inputs. It probes generalization to novel objects by measuring pose accuracy against ground truth under symmetry-aware error metrics. Use when the user wants to benchmark on LM-O, T-LESS, TUD-L, IC-BIN, ITODD, HomebrewdDB, YCB-V, or asks about evaluating this task. Reports Average Recall (AR).
Evaluates text-to-multi-view diffusion models on their ability to generate prompt-aligned, high-quality 4-view images and reconstruct consistent 3D objects. It measures image-text alignment and visual fidelity against a synthetic ground-truth distribution. Use when the user wants to benchmark on GPTeval3D, Synthetic GT Distribution, or asks about evaluating this task. Reports FID.
This evaluation probes the mathematical reasoning capability of large language models, specifically focusing on single-step reasoning and the effectiveness of step-aligned in-context learning. It measures how well models can solve challenging math problems across text and multi-modal domains when provided with fine-grained, step-level examples. Use when the user wants to benchmark on MATH500, AQuA, OlympiadBench-TO, MATHBench, AMC-10, AMC-12, MathVision, MathVerse, AIME, or asks about evaluat...
Evaluates stereo and monocular depth/disparity estimation models on images containing specular and transparent surfaces, which violate standard non-Lambertian assumptions and cause significant performance degradation in existing networks. Use when the user wants to benchmark on Booster, or asks about evaluating this task. Reports bad-2.
Evaluates extractive and abstractive summarization models on long-form narrative texts across paragraph, chapter, and book granularities. It probes lexical overlap, semantic similarity, content coverage via question answering, and human-rated fluency, coherence, relevance, and factuality. Use when the user wants to benchmark on BookSum, or asks about evaluating this task. Reports ROUGE-1.
Evaluates vision-language models' ability to perform abstract visual reasoning (AVR) by recognizing fine-grained, abstract visual concepts in Bongard-style matrix problems. It probes capabilities in concept selection, image-to-side classification, and free-form concept description generation. Use when the user wants to benchmark on Bongard-RWR+, or asks about evaluating this task. Reports accuracy.
Evaluates the accuracy and computational efficiency of neural and traditional Shapley value estimators against ground truth attributions across tabular and image datasets. It measures how well different explainers approximate feature importance and how fast they run. Use when the user wants to benchmark on Monks, WBC, Census, Credit, Magic, ImageNette, Pet, or asks about evaluating this task. Reports L1 distance.
Evaluates the classification accuracy, inference latency, and energy efficiency of Spiking Neural Networks (SNNs) trained with Batch Normalization Through Time (BNTT) on standard image and neuromorphic datasets. It probes the model's ability to maintain high accuracy while drastically reducing time-steps and computational cost compared to ANN-SNN conversion and standard surrogate gradient methods. Use when the user wants to benchmark on CIFAR-10, CIFAR-100, Tiny-ImageNet, DVS-CIFAR10, or asks...
Evaluates the sample efficiency and predictive accuracy of batch-mode deep active learning methods for tabular regression tasks. It probes how well different kernel-based selection strategies reduce prediction error over sequential labeling rounds compared to random sampling. Use when the user wants to benchmark on UCI & OpenML Tabular Regression Benchmark, or asks about evaluating this task. Reports RMSE.
Evaluates biomedical language models on a comprehensive suite of downstream NLP tasks, including named entity recognition, relation extraction, sentence similarity, document classification, and question answering. It measures how well domain-specific pretraining transfers to specialized clinical and biomedical text understanding. Use when the user wants to benchmark on BLURB, or asks about evaluating this task. Reports BLURB score.
Evaluates the cross-domain generalization and transfer learning capabilities of pre-trained language models across ten diverse biomedical and clinical NLP tasks. It probes sentence similarity, named entity recognition, relation extraction, document classification, and natural language inference to measure how well domain-specific pre-training captures clinical and biomedical semantics. Use when the user wants to benchmark on MedSTS, BIOSSES, BC5CDR-disease, BC5CDR-chemical, ShARe/CLEFE, DDI, ...
Evaluates BLOOM model variants against BERT-style and GPT-style baselines across diverse NLP tasks including text classification, question answering, zero/few-shot learning, multilingual transfer, and text generation. Use when the user wants to benchmark on GLUE, SQuAD, XNLI, MARC, Zero/FSL Benchmarks, or asks about evaluating this task. Reports accuracy.
Evaluates multimodal emotion recognition systems on detecting blended emotions (presence and salience) across unseen actors. It probes the model's ability to generalize actor-invariant emotional semantics from audio, visual, and combined modalities under a strict threshold-based discretization protocol. Use when the user wants to benchmark on BLEMORE, or asks about evaluating this task. Reports Score.
Evaluates DNN-based grapheme-to-phoneme and duration prediction models for three Indian languages (Hindi, Tamil, Telugu) using crowdsourced ASCII transliterated text. It measures the accuracy of predicted phoneme durations against forced-aligned ground truth to assess component quality for speech synthesis. Use when the user wants to benchmark on Blizzard Challenge 2015 (Hindi, Tamil, Telugu), or asks about evaluating this task. Reports RMSE (frames per phone).
Evaluates large multimodal models on single- and multi-image understanding, covering general VQA, domain knowledge, OCR, hallucination, and interleaved image-text reasoning. Use when the user wants to benchmark on SEED-IMG, MMB(dev), MMStar, MME(norm), RWQA, MMVet, MMMU(val), MathVista, TextVQA, OCRBench, POPE, HalBench, BLINK, QBench, MuirBench, Mantis-Eval, or asks about evaluating this task. Reports benchmark score, average score.
Evaluates the accuracy and robustness of event-based optical flow estimation models on complex scenes with dynamic objects, occlusions, and high-frequency motion. It measures how well models generalize to unseen scenarios and handle fine-grained flow details compared to traditional rigid/static scene benchmarks. Use when the user wants to benchmark on BlinkFlow, DSEC, MVSEC, or asks about evaluating this task. Reports AEE.
This evaluation probes a model's ability to leverage raw visual representations for vision-centric tasks without relying on language priors or domain expertise. It tests pixel-level matching, depth perception, 3D object awareness, and art style recognition across multiple-choice and regression-style tasks. Use when the user wants to benchmark on CV-Bench (Depth Order), SPair-71k, FunKPoint, HPatches, MOCHI, WikiArt (BLINK Art Style), or asks about evaluating this task. Reports multiple-choice...
Evaluates language understanding, linguistic acceptability, sentiment analysis, natural language inference, and factual reasoning under low-resource fine-tuning conditions. Use when the user wants to benchmark on BLiMP, GLUE (subset), SuperGLUE (subset), or asks about evaluating this task. Reports accuracy.
This benchmark probes language models' sensitivity to grammatical acceptability contrasts across 12 linguistic phenomena. It evaluates whether models can reliably distinguish acceptable sentences from minimally ungrammatical ones, revealing strengths in morphological agreement and weaknesses in complex syntactic and semantic constraints. Use when the user wants to benchmark on BLiMP, or asks about evaluating this task. Reports accuracy.
Evaluates the correlation between automatic text generation scores and human quality ratings. Probes a model's ability to accurately rank or score generated translations and data-to-text outputs against human judgments, including robustness to domain/quality drift and few-shot adaptation. Use when the user has predictions and gold and needs to compute Kendall's Tau ($\tau$).
Evaluates the zero-shot forecasting capability of universal time series models across multiple domains and prediction horizons. It probes how well pre-trained models generalize to unseen datasets and measures predictive accuracy using standard error metrics. Use when the user wants to benchmark on ETTh1, ETTh2, ETTm1, ETTm2, Weather, GlobalTemp, GIFT-Eval, or asks about evaluating this task. Reports MSE, MAE.
Evaluates the runtime performance and throughput of a C++ expression template library (SALT) against optimized BLAS implementations (Intel MKL) and other template libraries (Eigen) for standard vector operations. Use when the user has predictions and gold and needs to compute performance (GFLOPS).