Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 5,353–5,376 of 23,601 skills
Probes the practical usability and impact of word-level quality estimation highlights on professional translators' post-editing efficiency, accuracy, and workflow. It measures how different highlight modalities (oracle, supervised, unsupervised, none) affect editing effort, productivity, and final translation quality in real-world domain-specific settings. Use when the user wants to benchmark on QE4PE, or asks about evaluating this task. Reports ESA score.
Evaluates the ability of neural re-ranking models to effectively re-order a candidate set of documents based on complex query semantics and entity relationships. It probes fine-grained semantic matching, entity-aware attention, and late aggregation capabilities in information retrieval tasks across news and complex answer domains. Use when the user wants to benchmark on CODEC, TREC Complex Answer Retrieval (CAR) 2017, TREC Robust 2004, TREC News 2021, TREC Core 2018, or asks about evaluating ...
Evaluates a hybrid quantum-classical convolutional neural network on MRI-based brain tumor detection. It probes the model's ability to classify medical images into binary (tumor vs. non-tumor) and multiclass (specific tumor types) categories under class imbalance and limited resolution constraints. Use when the user wants to benchmark on Brain MRI Tumor Dataset, or asks about evaluating this task. Reports accuracy.
Probes vision-language models' ability to interpret quantum calibration plots across diverse visual formats (1D traces, 2D maps, histograms) and perform structured scientific reasoning. It tests capabilities ranging from visual grounding and outcome classification to parameter extraction and operational calibration diagnosis, both in zero-shot and in-context learning settings. Use when the user wants to benchmark on QCalEval, or asks about evaluating this task. Reports accuracy.
Evaluates the impact of ultra-low precision weight quantization (down to 2 bits) on BERTBASE performance across standard NLP tasks. It compares Hessian-guided mixed and group-wise quantization strategies against direct quantization baselines to measure accuracy retention versus model compression. Use when the user wants to benchmark on SST-2, MNLI, CoNLL-03, SQuAD, or asks about evaluating this task. Reports Acc.
Evaluates sequential recommendation models on their ability to predict the next item in a user's interaction history using multimodal item features. It probes cross-domain generalization, the effectiveness of discrete semantic tokenization, and robustness in sparse interaction scenarios. Use when the user wants to benchmark on Amazon Product Reviews (Instruments, Arts, Games), or asks about evaluating this task. Reports HR@K, NDCG@K.
Evaluates deep learning models for COVID-19 infected region segmentation and binary detection on chest X-ray images. It probes the model's ability to localize pathological regions at the pixel level and classify whole images as positive or negative for infection. Use when the user wants to benchmark on QaTa-COV19, or asks about evaluating this task. Reports F1-Score.
This benchmark evaluates a model's ability to read and reason over full academic research papers to answer information-seeking questions and to identify the specific paragraphs that contain the supporting evidence. It probes document-level comprehension, multi-paragraph reasoning, and handling of diverse answer types including extractive spans, abstractive summaries, yes/no, and unanswerable cases. Use when the user wants to benchmark on QASPER, or asks about evaluating this task. Reports Ans...
Evaluates the factual consistency and information coverage of query-focused summaries for less-resourced languages without reference texts. It probes whether LLMs can preserve source details and align with user intent by comparing answers generated from the source versus answers generated from the candidate summary. Use when the user wants to benchmark on Slovene News Summarization Corpus (MOCHA translation), or asks about evaluating this task. Reports QuestEval F1.
Evaluates document-level information extraction by framing it as a question answering task. It probes a model's ability to extract cross-sentence relation triples from large documents using entity-relation queries and knowledge base alignment. Use when the user wants to benchmark on QA4IE, or asks about evaluating this task. Reports Exact Match (EM), F1-score.
Evaluates the ability of a question-answering system to extract missing objects from partially filled relational tuples by searching through a large document corpus. It probes schema-aware information extraction and the model's capacity to leverage relational coherence across multiple questions. Use when the user wants to benchmark on QA-ZRE, or asks about evaluating this task. Reports Exact Match (EM).
Evaluates attribute and relation hallucination by asking LVLMs to identify object properties and inter-object relationships in images, probing fine-grained visual understanding beyond basic object detection. Use when the user wants to benchmark on QA-VisualGenome, or asks about evaluating this task. Reports Acc.
This benchmark evaluates how well LLM-generated translations preserve the scientific content of original papers. It measures translation fidelity by testing whether a reading model can accurately answer comprehension questions derived from the source text, using only the translated version as context. Use when the user wants to benchmark on Science Across Languages QA Benchmark, or asks about evaluating this task. Reports quiz accuracy.
Evaluates the ability of retrieval models to rank relevant sentences or documents highest for a given question. It probes lexical and semantic matching capabilities in question answering contexts, testing both single-model retrieval and multi-model fusion strategies. Use when the user wants to benchmark on ReQA SQuAD, ReQA NQ, MTEB QA Subset, iapp-wiki-qa-squad, or asks about evaluating this task. Reports MRR.
Evaluates the quality of generated question-answer pairs in an agricultural domain context. It probes a model's ability to produce relevant, accurate, diverse, and fluent Q&A content under varying context conditions. Use when the user wants to benchmark on Agricultural Q&A dataset, or asks about evaluating this task. Reports Relevance.
Evaluates cognition-based hallucination by testing whether LVLMs can leverage world knowledge stored in the LLM to answer entity and relation questions grounded in images. Use when the user wants to benchmark on QA-FB15K, or asks about evaluating this task. Reports Acc.
Evaluates the capability of retrieval-augmented generation systems to answer complex, multi-hop, and long-form questions by iteratively retrieving, structuring, and accumulating evidence from documents. Use when the user wants to benchmark on StrategyQA, ASQA, NQ, 2WikiMultiHopQA, HotpotQA, or asks about evaluating this task. Reports EM, F1, ACC.
Evaluates ranked retrieval lists using graded relevance assessments, balancing precision and cumulative gain while penalizing lower-ranked relevant documents. Use when the user has predictions and gold and needs to compute Q-measure.
Evaluates open-weight multimodal agentic models on visual search, multimodal mathematical reasoning, multi-turn tool use, and video spatial reasoning. It probes the model's ability to dynamically construct context, invoke tools, and perform long-horizon reasoning with high visual token efficiency. Use when the user wants to benchmark on V*, HRBench-4K, HRBench-8K, MathVerse, MathVision, WeMath, DynaMath, TIR-Bench, VSI-Bench, or asks about evaluating this task. Reports accuracy.
Evaluates machine learning models across clinical trial tasks including patient and trial outcome prediction, trial search, and patient simulation, using standardized tabular and sequential data formats. Use when the user wants to benchmark on Tabular Clinical Trial Patient Datasets, TOP Benchmark, Trial Similarity Dataset, Sequential Trial Patient Data, or asks about evaluating this task. Reports AUROC.
Evaluates the numerical correctness, stability, and parallel performance (CPU/GPU strong scaling and distributed weak scaling) of a matrix-free finite difference geodynamic solver. Use when the user wants to benchmark on Pyroclast Stokes & Advection Benchmarks, or asks about evaluating this task. Reports parallel speedup.
Evaluates a model's ability to generate temporal scene graphs where nodes are grounded with pixel-level panoptic segmentation masks instead of bounding boxes, capturing non-rigid objects, backgrounds, and fine-grained interactions in dynamic videos. Use when the user wants to benchmark on PVSG, or asks about evaluating this task. Reports R/mR@20.
Evaluates medium-range global weather forecasting capability using an autoregressive convolutional network. It measures prediction accuracy over 10-day horizons at 6-hour intervals against reanalysis ground truth. Use when the user wants to benchmark on ERA5, or asks about evaluating this task. Reports MAE.
Evaluates video-language models on long-form repetition counting and temporal reasoning. It probes whether models can accurately track state changes and count actions across extended video clips, revealing weaknesses in spatio-temporal tracking compared to supervised baselines. Use when the user wants to benchmark on PushupBench, or asks about evaluating this task. Reports Exact Match.