Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 5,377–5,400 of 23,640 skills
This benchmark evaluates offline reinforcement learning algorithms on real-world quadrupedal locomotion tasks. It probes the policy's ability to accurately track locomotion commands, maintain energy efficiency, and exhibit stability under real-world environmental stochasticity and terrain variations. Use when the user wants to benchmark on Real-World Quadrupedal Locomotion Dataset, or asks about evaluating this task. Reports Return.
Evaluates a drone's ability to autonomously navigate to a target vehicle using marker-free visual servoing. It measures tracking accuracy, localization precision, and flight efficiency in both simulation and real-world environments. Use when the user wants to benchmark on Custom Simulation & Real-World Flight Dataset, or asks about evaluating this task. Reports NormError.
Evaluates multi-turn, context-dependent question answering where models must track dialog history, resolve coreference, and handle open-ended follow-ups. It specifically probes a system's ability to manage asymmetric knowledge access and correctly identify unanswerable questions within an information-seeking dialogue. Use when the user wants to benchmark on QuAC, or asks about evaluating this task. Reports word-level F1.
Evaluates the calibration of voxel-wise uncertainty estimates in brain tumor segmentation. It measures how effectively uncertainty thresholds filter out incorrect predictions while preserving correct ones, rewarding high confidence in accurate regions and penalizing the loss of correct predictions when filtering uncertain voxels. Use when the user has predictions and gold and needs to compute QU-BraTS unified score.
Evaluates a quantum support vector machine (QSVM) with quantum feature selection for binary fraud detection on real-world card payment data. It probes the model's ability to identify fraudulent transactions using a balanced dataset and compares performance against classical feature selection baselines. Use when the user wants to benchmark on Real-world card payment data (Balanced Data Set), or asks about evaluating this task. Reports Accuracy.
Evaluates the tradeoff between response fluency and factual attribution in retrieval-augmented conversational LLMs. It measures how well models generate coherent, context-aware responses while correctly grounding answers in provided evidence or dialog history. Use when the user wants to benchmark on QReCC, or asks about evaluating this task. Reports Auto-AIS.
Evaluates machine reading comprehension on a low-resource religious domain (Qur'an). It probes a model's ability to extract precise answer spans from Arabic text given a question, testing both exact matching and partial semantic/token overlap. Use when the user wants to benchmark on QRCD, or asks about evaluating this task. Reports pRR.
Evaluates a model's ability to generate semantically equivalent paraphrase sentences from an input question, measuring lexical and semantic overlap with ground truth references. The benchmark probes sentence-level semantic understanding and generative fluency in a question-paraphrase setting. Use when the user wants to benchmark on Quora Question Pairs (QQP), or asks about evaluating this task. Reports BLEU.
Evaluates whether citation growth outpaces publication expansion across leading AI/NLP conferences, measuring the elasticity of scholarly impact relative to scale-driven growth over a decade. Use when the user has predictions and gold and needs to compute QQE.
Measures social bias in medical question-answering systems for pain management by evaluating treatment denial rates across intersectional race-gender profiles. It probes whether AI models exhibit discriminatory prescribing patterns when presented with clinical vignettes containing demographic attributes. Use when the user wants to benchmark on Q-Pain, or asks about evaluating this task. Reports probability_of_no.
Evaluates mathematical reasoning, knowledge-oriented language understanding, and challenging reasoning capabilities of LLMs fine-tuned on domain-specific corpora. Use when the user wants to benchmark on MATH, GSM8K, MMLU, AGIEval, BIG-Bench Hard, or asks about evaluating this task. Reports accuracy.
Evaluates the performance of quantum and classical optimization algorithms on intractable combinatorial problems. It measures solution quality, algorithmic success rates, and computational efficiency to track progress toward quantum advantage. Use when the user wants to benchmark on QOBLIB, or asks about evaluating this task. Reports best_objective_value.
This evaluation protocol tests the scalability and correctness of SMT-based formal verification for quantized neural networks. It measures how well an SMT model-checking framework can prove safety properties or find counterexamples across different quantization levels, network architectures, and SMT solvers. Use when the user wants to benchmark on Iris dataset, Vocalic dataset, AcasXu benchmark, or asks about evaluating this task. Reports verification_time.
Evaluates an agent's ability to distill and synthesize key information from high-noise, multi-turn meeting transcripts based on specific user queries. It probes the model's capacity to maintain global context while filtering irrelevant dialogue turns to produce a query-relevant summary. Use when the user wants to benchmark on QMSUM, or asks about evaluating this task. Reports ROUGE-1.
Assesses a model's ability to predict scalar quantum chemical properties from molecular structures. It evaluates accuracy on five key electronic and vibrational targets using mean absolute error and average ranking. Use when the user wants to benchmark on QM9, or asks about evaluating this task. Reports MAE.
Evaluates the ability of a tokenization framework to generate valid 3D molecular structures and predict quantum mechanical properties. It probes structural validity, geometric plausibility, conditional controllability, and property prediction accuracy on organic molecules. Use when the user wants to benchmark on QM9, or asks about evaluating this task. Reports Mean Absolute Error (MAE).
Evaluates multimodal information retrieval systems across search, recommendation, and deep query answering (DQA) tasks using real-world APP-level user sessions. It probes a model's ability to rank heterogeneous content (text, images, videos) and generate accurate answers augmented by retrieved documents. Use when the user wants to benchmark on Qilin, or asks about evaluating this task. Reports MRR@10.
This evaluation probes a unified vision-language model's ability to perform end-to-end document intelligence, including specialized OCR, general text recognition, document understanding, and key information extraction across diverse document types and multilingual scenarios. Use when the user wants to benchmark on Omni-Doc-Bench v1.5, OLMOCRBench, OCRBench, DocVQA, ChartQA, Nanonets KIE, or asks about evaluating this task. Reports normalized accuracy (0-100).
Evaluates the quality of generated questions across seven dimensions: fluency, clarity, conciseness, relevance, consistency, answerability, and answer consistency. It measures how well automatic metrics and LLMs align with human judgments on these dimensions. Use when the user wants to benchmark on QGEval, or asks about evaluating this task. Reports Pearson correlation.
Evaluates the ability of generative language models to produce paragraph-level questions conditioned on a target answer and a context sentence. It probes domain adaptability and multilingual generalization across diverse extractive QA datasets. Use when the user wants to benchmark on SQuAD v1.1, SQuADShifts, SubjQA, Multilingual QA (JAQuAD, GerQuAD, SberQuAD, KorQuAD, FQuAD, Spanish SQuAD, Italian SQuAD), or asks about evaluating this task. Reports automatic evaluation metrics.
Probes the practical usability and impact of word-level quality estimation highlights on professional translators' post-editing efficiency, accuracy, and workflow. It measures how different highlight modalities (oracle, supervised, unsupervised, none) affect editing effort, productivity, and final translation quality in real-world domain-specific settings. Use when the user wants to benchmark on QE4PE, or asks about evaluating this task. Reports ESA score.
Evaluates the ability of neural re-ranking models to effectively re-order a candidate set of documents based on complex query semantics and entity relationships. It probes fine-grained semantic matching, entity-aware attention, and late aggregation capabilities in information retrieval tasks across news and complex answer domains. Use when the user wants to benchmark on CODEC, TREC Complex Answer Retrieval (CAR) 2017, TREC Robust 2004, TREC News 2021, TREC Core 2018, or asks about evaluating ...
Evaluates a hybrid quantum-classical convolutional neural network on MRI-based brain tumor detection. It probes the model's ability to classify medical images into binary (tumor vs. non-tumor) and multiclass (specific tumor types) categories under class imbalance and limited resolution constraints. Use when the user wants to benchmark on Brain MRI Tumor Dataset, or asks about evaluating this task. Reports accuracy.
Probes vision-language models' ability to interpret quantum calibration plots across diverse visual formats (1D traces, 2D maps, histograms) and perform structured scientific reasoning. It tests capabilities ranging from visual grounding and outcome classification to parameter extraction and operational calibration diagnosis, both in zero-shot and in-context learning settings. Use when the user wants to benchmark on QCalEval, or asks about evaluating this task. Reports accuracy.