Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

23,640
skills in category
985
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 5,401–5,424 of 23,640 skills

Qbert Quantization EvalA

Evaluates the impact of ultra-low precision weight quantization (down to 2 bits) on BERTBASE performance across standard NLP tasks. It compares Hessian-guided mixed and group-wise quantization strategies against direct quantization baselines to measure accuracy retention versus model compression. Use when the user wants to benchmark on SST-2, MNLI, CoNLL-03, SQuAD, or asks about evaluating this task. Reports Acc.

researchpythongo
0
3
Qb4rec EvalA

Evaluates sequential recommendation models on their ability to predict the next item in a user's interaction history using multimodal item features. It probes cross-domain generalization, the effectiveness of discrete semantic tokenization, and robustness in sparse interaction scenarios. Use when the user wants to benchmark on Amazon Product Reviews (Instruments, Arts, Games), or asks about evaluating this task. Reports HR@K, NDCG@K.

researchpythontesting
0
3
Qata Cov19 EvalA

Evaluates deep learning models for COVID-19 infected region segmentation and binary detection on chest X-ray images. It probes the model's ability to localize pathological regions at the pixel level and classify whole images as positive or negative for infection. Use when the user wants to benchmark on QaTa-COV19, or asks about evaluating this task. Reports F1-Score.

researchpythonperformance
0
3
Qasper EvalA

This benchmark evaluates a model's ability to read and reason over full academic research papers to answer information-seeking questions and to identify the specific paragraphs that contain the supporting evidence. It probes document-level comprehension, multi-paragraph reasoning, and handling of diverse answer types including extractive spans, abstractive summaries, yes/no, and unanswerable cases. Use when the user wants to benchmark on QASPER, or asks about evaluating this task. Reports Ans...

researchpythongo
0
3
Qags Questeval EvalA

Evaluates the factual consistency and information coverage of query-focused summaries for less-resourced languages without reference texts. It probes whether LLMs can preserve source details and align with user intent by comparing answers generated from the source versus answers generated from the candidate summary. Use when the user wants to benchmark on Slovene News Summarization Corpus (MOCHA translation), or asks about evaluating this task. Reports QuestEval F1.

researchpythonperformance
0
3
Qa4ie EvalA

Evaluates document-level information extraction by framing it as a question answering task. It probes a model's ability to extract cross-sentence relation triples from large documents using entity-relation queries and knowledge base alignment. Use when the user wants to benchmark on QA4IE, or asks about evaluating this task. Reports Exact Match (EM), F1-score.

researchpythongo
0
3
Qa Zre EvalA

Evaluates the ability of a question-answering system to extract missing objects from partially filled relational tuples by searching through a large document corpus. It probes schema-aware information extraction and the model's capacity to leverage relational coherence across multiple questions. Use when the user wants to benchmark on QA-ZRE, or asks about evaluating this task. Reports Exact Match (EM).

researchpythongo
0
3
Qa Visualgenome EvalA

Evaluates attribute and relation hallucination by asking LVLMs to identify object properties and inter-object relationships in images, probing fine-grained visual understanding beyond basic object detection. Use when the user wants to benchmark on QA-VisualGenome, or asks about evaluating this task. Reports Acc.

researchpythongo
0
3
Qa Translation Fidelity EvalA

This benchmark evaluates how well LLM-generated translations preserve the scientific content of original papers. It measures translation fidelity by testing whether a reading model can accurately answer comprehension questions derived from the source text, using only the translated version as context. Use when the user wants to benchmark on Science Across Languages QA Benchmark, or asks about evaluating this task. Reports quiz accuracy.

researchpythongo
0
3
Qa Retrieval EvalA

Evaluates the ability of retrieval models to rank relevant sentences or documents highest for a given question. It probes lexical and semantic matching capabilities in question answering contexts, testing both single-model retrieval and multi-model fusion strategies. Use when the user wants to benchmark on ReQA SQuAD, ReQA NQ, MTEB QA Subset, iapp-wiki-qa-squad, or asks about evaluating this task. Reports MRR.

researchpythontesting
0
3
Qa Quality EvalA

Evaluates the quality of generated question-answer pairs in an agricultural domain context. It probes a model's ability to produce relevant, accurate, diverse, and fluent Q&A content under varying context conditions. Use when the user wants to benchmark on Agricultural Q&A dataset, or asks about evaluating this task. Reports Relevance.

researchpythongo
0
3
Qa Fb15k EvalA

Evaluates cognition-based hallucination by testing whether LVLMs can leverage world knowledge stored in the LLM to answer entity and relation questions grounded in images. Use when the user wants to benchmark on QA-FB15K, or asks about evaluating this task. Reports Acc.

researchpythongo
0
3
Qa Benchmarks EvalA

Evaluates the capability of retrieval-augmented generation systems to answer complex, multi-hop, and long-form questions by iteratively retrieving, structuring, and accumulating evidence from documents. Use when the user wants to benchmark on StrategyQA, ASQA, NQ, 2WikiMultiHopQA, HotpotQA, or asks about evaluating this task. Reports EM, F1, ACC.

researchpythongo
0
3
Q MeasureA

Evaluates ranked retrieval lists using graded relevance assessments, balancing precision and cumulative gain while penalizing lower-ranked relevant documents. Use when the user has predictions and gold and needs to compute Q-measure.

researchpythongo
0
3
Pyvision Rl EvalA

Evaluates open-weight multimodal agentic models on visual search, multimodal mathematical reasoning, multi-turn tool use, and video spatial reasoning. It probes the model's ability to dynamically construct context, invoke tools, and perform long-horizon reasoning with high visual token efficiency. Use when the user wants to benchmark on V*, HRBench-4K, HRBench-8K, MathVerse, MathVision, WeMath, DynaMath, TIR-Bench, VSI-Bench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Pytrial EvalA

Evaluates machine learning models across clinical trial tasks including patient and trial outcome prediction, trial search, and patient simulation, using standardized tabular and sequential data formats. Use when the user wants to benchmark on Tabular Clinical Trial Patient Datasets, TOP Benchmark, Trial Similarity Dataset, Sequential Trial Patient Data, or asks about evaluating this task. Reports AUROC.

researchpythongo
0
3
Pyroclast Benchmark EvalA

Evaluates the numerical correctness, stability, and parallel performance (CPU/GPU strong scaling and distributed weak scaling) of a matrix-free finite difference geodynamic solver. Use when the user wants to benchmark on Pyroclast Stokes & Advection Benchmarks, or asks about evaluating this task. Reports parallel speedup.

researchpythongo
0
3
Pvsg EvalA

Evaluates a model's ability to generate temporal scene graphs where nodes are grounded with pixel-level panoptic segmentation masks instead of bounding boxes, capturing non-rigid objects, backgrounds, and fine-grained interactions in dynamic videos. Use when the user wants to benchmark on PVSG, or asks about evaluating this task. Reports R/mR@20.

researchpythongo
0
3
Puyun Weather Forecasting EvalA

Evaluates medium-range global weather forecasting capability using an autoregressive convolutional network. It measures prediction accuracy over 10-day horizons at 6-hour intervals against reanalysis ground truth. Use when the user wants to benchmark on ERA5, or asks about evaluating this task. Reports MAE.

researchpythontesting
0
3
Pushupbench EvalA

Evaluates video-language models on long-form repetition counting and temporal reasoning. It probes whether models can accurately track state changes and count actions across extended video clips, revealing weaknesses in spatio-temporal tracking compared to supervised baselines. Use when the user wants to benchmark on PushupBench, or asks about evaluating this task. Reports Exact Match.

researchpythongo
0
3
Punctuation Restoration Iwslt EvalA

Evaluates a model's ability to restore punctuation marks (commas, periods, question marks) in English text. It specifically probes robustness on both manually transcribed transcripts and ASR-generated transcripts, ignoring non-punctuation tokens during evaluation. Use when the user wants to benchmark on IWSLT2011, or asks about evaluating this task. Reports F1-score.

researchpythongo
0
3
Puma Challenge EvalA

Evaluates pixel-level segmentation capability for distinguishing five histopathological tissue classes (tumour, stroma, necrosis, blood vessels, epidermis) in melanoma H&E images. Use when the user wants to benchmark on PUMA Challenge dataset, or asks about evaluating this task. Reports Dice score.

researchpythonperformance
0
3
Pulsnar Alpha Estimation EvalA

This benchmark evaluates the ability of Positive Unlabeled (PU) learning algorithms to accurately estimate the true proportion of positive examples ($\alpha$) within an unlabeled dataset, particularly when selection bias violates the SCAR assumption. It also probes the robustness of downstream classification performance and probability calibration under varying degrees of class imbalance and structured selection bias. Use when the user wants to benchmark on Synthetic SCAR, Synthetic SNAR, UCI...

researchpythongo
0
3
Pulselm EvalA

This benchmark evaluates multimodal physiological reasoning by testing whether large language models can accurately answer closed-ended questions conditioned on raw photoplethysmography (PPG) waveforms. It probes the model's ability to align continuous biosignal representations with natural language queries across diverse physiological domains and assesses cross-dataset generalization beyond the training distribution. Use when the user wants to benchmark on PulseLM, or asks about evaluating t...

researchpythongo
0
3