Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 9,217–9,240 of 21,273 skills
Evaluates the stability and reliability of global pointwise scores (accuracy, AUC, F1) versus pairwise Bradley-Terry rankings for ordering NLP models across classification and text generation tasks. Use when the user has predictions and gold and needs to compute Bradley-Terry.
Evaluates the ability of audio-language models to align audio with captions and distinguish caption quality across different generation sources (human-human, human-machine, machine-machine). It probes fine-grained semantic and syntactic alignment capabilities under realistic captioning conditions. Use when the user wants to benchmark on BRACE-Main, or asks about evaluating this task. Reports F1-score.
Probes models' robustness in detecting subtle hallucinations in audio captions, specifically those introduced via LLM-driven noun substitution. It measures the ability to identify semantically flawed or factually incorrect descriptions against audio ground truth. Use when the user wants to benchmark on BRACE-Hallucination, or asks about evaluating this task. Reports F1-score.
Evaluates the effectiveness of synthetic 3D MRI tumor ROI generation for data augmentation by measuring downstream binary classification performance on imbalanced brain tumor subtypes. Use when the user wants to benchmark on BraTS 2019, SickKids pLGG, or asks about evaluating this task. Reports AUC.
Evaluates vision-language models' ability to extract structured information (names, types, and connectivity) from Business Process Model and Notation (BPMN) diagrams provided as images. It tests both raw visual understanding and the utility of OCR-enriched inputs for schema-constrained diagram parsing. Use when the user wants to benchmark on BPMN Diagrams (Custom), or asks about evaluating this task. Reports F1 Score.
Evaluates machine translation systems on a contamination-free, multilingual dataset covering diverse domains and registers. It measures translation quality at both sentence and paragraph levels to assess how well models handle linguistic diversity and cultural authenticity across 8 major languages. Use when the user wants to benchmark on BOUQuET, or asks about evaluating this task. Reports CometKiwi.
This evaluation probes a model's ability to forecast and reconstruct the dynamics of partially observed, chaotic geophysical systems. It specifically tests short-term prediction accuracy and long-term topological stability (boundedness) under both in-distribution and out-of-attractor initial conditions. Use when the user wants to benchmark on Lorenz-63, Lorenz-96, or asks about evaluating this task. Reports RMSE.
Evaluates a hierarchical multi-agent reinforcement learning scheduler's ability to optimize task allocation, frequency scaling, and core selection for OpenMP DAG workloads on embedded systems. It probes the trade-off between makespan, energy consumption, and thermal constraints under real-time profiling feedback. Use when the user wants to benchmark on Barcelona OpenMP Tasks Suite (BOTS), or asks about evaluating this task. Reports makespan.
This evaluation probes a model's ability to detect and classify robotic failures in real-world manipulation tasks. It specifically tests whether a system can distinguish between genuine task-disrupting failures and benign environmental deviations using multimodal observations and nominal demonstrations. Use when the user wants to benchmark on BotFails, Real-π dataset, or asks about evaluating this task. Reports AUROC.
This evaluation probes a model's ability to perform weakly supervised multimodal segmentation of acoustic borehole images by refining threshold-guided pseudo-labels using depth-aligned well logs. It measures how well the predicted segmentation aligns with a provisional target map, testing spatial coherence and multimodal feature fusion rather than absolute geological accuracy. Use when the user wants to benchmark on Antilope25 & Botorosa47 borehole intervals, or asks about evaluating this tas...
Evaluates the ability to refine 6D object poses in cluttered real-world scenes using RGB or RGB-D inputs. It probes generalization to novel objects by measuring pose accuracy against ground truth under symmetry-aware error metrics. Use when the user wants to benchmark on LM-O, T-LESS, TUD-L, IC-BIN, ITODD, HomebrewdDB, YCB-V, or asks about evaluating this task. Reports Average Recall (AR).
Evaluates text-to-multi-view diffusion models on their ability to generate prompt-aligned, high-quality 4-view images and reconstruct consistent 3D objects. It measures image-text alignment and visual fidelity against a synthetic ground-truth distribution. Use when the user wants to benchmark on GPTeval3D, Synthetic GT Distribution, or asks about evaluating this task. Reports FID.
This evaluation probes the mathematical reasoning capability of large language models, specifically focusing on single-step reasoning and the effectiveness of step-aligned in-context learning. It measures how well models can solve challenging math problems across text and multi-modal domains when provided with fine-grained, step-level examples. Use when the user wants to benchmark on MATH500, AQuA, OlympiadBench-TO, MATHBench, AMC-10, AMC-12, MathVision, MathVerse, AIME, or asks about evaluat...
Evaluates stereo and monocular depth/disparity estimation models on images containing specular and transparent surfaces, which violate standard non-Lambertian assumptions and cause significant performance degradation in existing networks. Use when the user wants to benchmark on Booster, or asks about evaluating this task. Reports bad-2.
Evaluates extractive and abstractive summarization models on long-form narrative texts across paragraph, chapter, and book granularities. It probes lexical overlap, semantic similarity, content coverage via question answering, and human-rated fluency, coherence, relevance, and factuality. Use when the user wants to benchmark on BookSum, or asks about evaluating this task. Reports ROUGE-1.
Evaluates vision-language models' ability to perform abstract visual reasoning (AVR) by recognizing fine-grained, abstract visual concepts in Bongard-style matrix problems. It probes capabilities in concept selection, image-to-side classification, and free-form concept description generation. Use when the user wants to benchmark on Bongard-RWR+, or asks about evaluating this task. Reports accuracy.
Evaluates the accuracy and computational efficiency of neural and traditional Shapley value estimators against ground truth attributions across tabular and image datasets. It measures how well different explainers approximate feature importance and how fast they run. Use when the user wants to benchmark on Monks, WBC, Census, Credit, Magic, ImageNette, Pet, or asks about evaluating this task. Reports L1 distance.
Evaluates the classification accuracy, inference latency, and energy efficiency of Spiking Neural Networks (SNNs) trained with Batch Normalization Through Time (BNTT) on standard image and neuromorphic datasets. It probes the model's ability to maintain high accuracy while drastically reducing time-steps and computational cost compared to ANN-SNN conversion and standard surrogate gradient methods. Use when the user wants to benchmark on CIFAR-10, CIFAR-100, Tiny-ImageNet, DVS-CIFAR10, or asks...
Evaluates the sample efficiency and predictive accuracy of batch-mode deep active learning methods for tabular regression tasks. It probes how well different kernel-based selection strategies reduce prediction error over sequential labeling rounds compared to random sampling. Use when the user wants to benchmark on UCI & OpenML Tabular Regression Benchmark, or asks about evaluating this task. Reports RMSE.
Evaluates biomedical language models on a comprehensive suite of downstream NLP tasks, including named entity recognition, relation extraction, sentence similarity, document classification, and question answering. It measures how well domain-specific pretraining transfers to specialized clinical and biomedical text understanding. Use when the user wants to benchmark on BLURB, or asks about evaluating this task. Reports BLURB score.
Evaluates the cross-domain generalization and transfer learning capabilities of pre-trained language models across ten diverse biomedical and clinical NLP tasks. It probes sentence similarity, named entity recognition, relation extraction, document classification, and natural language inference to measure how well domain-specific pre-training captures clinical and biomedical semantics. Use when the user wants to benchmark on MedSTS, BIOSSES, BC5CDR-disease, BC5CDR-chemical, ShARe/CLEFE, DDI, ...
Evaluates BLOOM model variants against BERT-style and GPT-style baselines across diverse NLP tasks including text classification, question answering, zero/few-shot learning, multilingual transfer, and text generation. Use when the user wants to benchmark on GLUE, SQuAD, XNLI, MARC, Zero/FSL Benchmarks, or asks about evaluating this task. Reports accuracy.
Evaluates multimodal emotion recognition systems on detecting blended emotions (presence and salience) across unseen actors. It probes the model's ability to generalize actor-invariant emotional semantics from audio, visual, and combined modalities under a strict threshold-based discretization protocol. Use when the user wants to benchmark on BLEMORE, or asks about evaluating this task. Reports Score.
Evaluates DNN-based grapheme-to-phoneme and duration prediction models for three Indian languages (Hindi, Tamil, Telugu) using crowdsourced ASCII transliterated text. It measures the accuracy of predicted phoneme durations against forced-aligned ground truth to assess component quality for speech synthesis. Use when the user wants to benchmark on Blizzard Challenge 2015 (Hindi, Tamil, Telugu), or asks about evaluating this task. Reports RMSE (frames per phone).