Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

23,579
skills in category
983
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 5,329–5,352 of 23,579 skills

Qgeval EvalA

Evaluates the quality of generated questions across seven dimensions: fluency, clarity, conciseness, relevance, consistency, answerability, and answer consistency. It measures how well automatic metrics and LLMs align with human judgments on these dimensions. Use when the user wants to benchmark on QGEval, or asks about evaluating this task. Reports Pearson correlation.

researchpythongo
0
3
Qg Bench EvalA

Evaluates the ability of generative language models to produce paragraph-level questions conditioned on a target answer and a context sentence. It probes domain adaptability and multilingual generalization across diverse extractive QA datasets. Use when the user wants to benchmark on SQuAD v1.1, SQuADShifts, SubjQA, Multilingual QA (JAQuAD, GerQuAD, SberQuAD, KorQuAD, FQuAD, Spanish SQuAD, Italian SQuAD), or asks about evaluating this task. Reports automatic evaluation metrics.

researchpythongo
0
3
Qe4pe EvalA

Probes the practical usability and impact of word-level quality estimation highlights on professional translators' post-editing efficiency, accuracy, and workflow. It measures how different highlight modalities (oracle, supervised, unsupervised, none) affect editing effort, productivity, and final translation quality in real-world domain-specific settings. Use when the user wants to benchmark on QE4PE, or asks about evaluating this task. Reports ESA score.

researchpythongo
0
3
Qder Re Ranking EvalA

Evaluates the ability of neural re-ranking models to effectively re-order a candidate set of documents based on complex query semantics and entity relationships. It probes fine-grained semantic matching, entity-aware attention, and late aggregation capabilities in information retrieval tasks across news and complex answer domains. Use when the user wants to benchmark on CODEC, TREC Complex Answer Retrieval (CAR) 2017, TREC Robust 2004, TREC News 2021, TREC Core 2018, or asks about evaluating ...

researchpythongo
0
3
Qcnmri Tumor Classification EvalA

Evaluates a hybrid quantum-classical convolutional neural network on MRI-based brain tumor detection. It probes the model's ability to classify medical images into binary (tumor vs. non-tumor) and multiclass (specific tumor types) categories under class imbalance and limited resolution constraints. Use when the user wants to benchmark on Brain MRI Tumor Dataset, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Qcaleval EvalA

Probes vision-language models' ability to interpret quantum calibration plots across diverse visual formats (1D traces, 2D maps, histograms) and perform structured scientific reasoning. It tests capabilities ranging from visual grounding and outcome classification to parameter extraction and operational calibration diagnosis, both in zero-shot and in-context learning settings. Use when the user wants to benchmark on QCalEval, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Qbert Quantization EvalA

Evaluates the impact of ultra-low precision weight quantization (down to 2 bits) on BERTBASE performance across standard NLP tasks. It compares Hessian-guided mixed and group-wise quantization strategies against direct quantization baselines to measure accuracy retention versus model compression. Use when the user wants to benchmark on SST-2, MNLI, CoNLL-03, SQuAD, or asks about evaluating this task. Reports Acc.

researchpythongo
0
3
Qb4rec EvalA

Evaluates sequential recommendation models on their ability to predict the next item in a user's interaction history using multimodal item features. It probes cross-domain generalization, the effectiveness of discrete semantic tokenization, and robustness in sparse interaction scenarios. Use when the user wants to benchmark on Amazon Product Reviews (Instruments, Arts, Games), or asks about evaluating this task. Reports HR@K, NDCG@K.

researchpythontesting
0
3
Qata Cov19 EvalA

Evaluates deep learning models for COVID-19 infected region segmentation and binary detection on chest X-ray images. It probes the model's ability to localize pathological regions at the pixel level and classify whole images as positive or negative for infection. Use when the user wants to benchmark on QaTa-COV19, or asks about evaluating this task. Reports F1-Score.

researchpythonperformance
0
3
Qasper EvalA

This benchmark evaluates a model's ability to read and reason over full academic research papers to answer information-seeking questions and to identify the specific paragraphs that contain the supporting evidence. It probes document-level comprehension, multi-paragraph reasoning, and handling of diverse answer types including extractive spans, abstractive summaries, yes/no, and unanswerable cases. Use when the user wants to benchmark on QASPER, or asks about evaluating this task. Reports Ans...

researchpythongo
0
3
Qags Questeval EvalA

Evaluates the factual consistency and information coverage of query-focused summaries for less-resourced languages without reference texts. It probes whether LLMs can preserve source details and align with user intent by comparing answers generated from the source versus answers generated from the candidate summary. Use when the user wants to benchmark on Slovene News Summarization Corpus (MOCHA translation), or asks about evaluating this task. Reports QuestEval F1.

researchpythonperformance
0
3
Qa4ie EvalA

Evaluates document-level information extraction by framing it as a question answering task. It probes a model's ability to extract cross-sentence relation triples from large documents using entity-relation queries and knowledge base alignment. Use when the user wants to benchmark on QA4IE, or asks about evaluating this task. Reports Exact Match (EM), F1-score.

researchpythongo
0
3
Qa Zre EvalA

Evaluates the ability of a question-answering system to extract missing objects from partially filled relational tuples by searching through a large document corpus. It probes schema-aware information extraction and the model's capacity to leverage relational coherence across multiple questions. Use when the user wants to benchmark on QA-ZRE, or asks about evaluating this task. Reports Exact Match (EM).

researchpythongo
0
3
Qa Visualgenome EvalA

Evaluates attribute and relation hallucination by asking LVLMs to identify object properties and inter-object relationships in images, probing fine-grained visual understanding beyond basic object detection. Use when the user wants to benchmark on QA-VisualGenome, or asks about evaluating this task. Reports Acc.

researchpythongo
0
3
Qa Translation Fidelity EvalA

This benchmark evaluates how well LLM-generated translations preserve the scientific content of original papers. It measures translation fidelity by testing whether a reading model can accurately answer comprehension questions derived from the source text, using only the translated version as context. Use when the user wants to benchmark on Science Across Languages QA Benchmark, or asks about evaluating this task. Reports quiz accuracy.

researchpythongo
0
3
Qa Retrieval EvalA

Evaluates the ability of retrieval models to rank relevant sentences or documents highest for a given question. It probes lexical and semantic matching capabilities in question answering contexts, testing both single-model retrieval and multi-model fusion strategies. Use when the user wants to benchmark on ReQA SQuAD, ReQA NQ, MTEB QA Subset, iapp-wiki-qa-squad, or asks about evaluating this task. Reports MRR.

researchpythontesting
0
3
Qa Quality EvalA

Evaluates the quality of generated question-answer pairs in an agricultural domain context. It probes a model's ability to produce relevant, accurate, diverse, and fluent Q&A content under varying context conditions. Use when the user wants to benchmark on Agricultural Q&A dataset, or asks about evaluating this task. Reports Relevance.

researchpythongo
0
3
Qa Fb15k EvalA

Evaluates cognition-based hallucination by testing whether LVLMs can leverage world knowledge stored in the LLM to answer entity and relation questions grounded in images. Use when the user wants to benchmark on QA-FB15K, or asks about evaluating this task. Reports Acc.

researchpythongo
0
3
Qa Benchmarks EvalA

Evaluates the capability of retrieval-augmented generation systems to answer complex, multi-hop, and long-form questions by iteratively retrieving, structuring, and accumulating evidence from documents. Use when the user wants to benchmark on StrategyQA, ASQA, NQ, 2WikiMultiHopQA, HotpotQA, or asks about evaluating this task. Reports EM, F1, ACC.

researchpythongo
0
3
Q MeasureA

Evaluates ranked retrieval lists using graded relevance assessments, balancing precision and cumulative gain while penalizing lower-ranked relevant documents. Use when the user has predictions and gold and needs to compute Q-measure.

researchpythongo
0
3
Pyvision Rl EvalA

Evaluates open-weight multimodal agentic models on visual search, multimodal mathematical reasoning, multi-turn tool use, and video spatial reasoning. It probes the model's ability to dynamically construct context, invoke tools, and perform long-horizon reasoning with high visual token efficiency. Use when the user wants to benchmark on V*, HRBench-4K, HRBench-8K, MathVerse, MathVision, WeMath, DynaMath, TIR-Bench, VSI-Bench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Pytrial EvalA

Evaluates machine learning models across clinical trial tasks including patient and trial outcome prediction, trial search, and patient simulation, using standardized tabular and sequential data formats. Use when the user wants to benchmark on Tabular Clinical Trial Patient Datasets, TOP Benchmark, Trial Similarity Dataset, Sequential Trial Patient Data, or asks about evaluating this task. Reports AUROC.

researchpythongo
0
3
Pyroclast Benchmark EvalA

Evaluates the numerical correctness, stability, and parallel performance (CPU/GPU strong scaling and distributed weak scaling) of a matrix-free finite difference geodynamic solver. Use when the user wants to benchmark on Pyroclast Stokes & Advection Benchmarks, or asks about evaluating this task. Reports parallel speedup.

researchpythongo
0
3
Pvsg EvalA

Evaluates a model's ability to generate temporal scene graphs where nodes are grounded with pixel-level panoptic segmentation masks instead of bounding boxes, capturing non-rigid objects, backgrounds, and fine-grained interactions in dynamic videos. Use when the user wants to benchmark on PVSG, or asks about evaluating this task. Reports R/mR@20.

researchpythongo
0
3