Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

23,640
skills in category
985
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 5,377–5,400 of 23,640 skills

Quadrupedal Locomotion EvalA

This benchmark evaluates offline reinforcement learning algorithms on real-world quadrupedal locomotion tasks. It probes the policy's ability to accurately track locomotion commands, maintain energy efficiency, and exhibit stability under real-world environmental stochasticity and terrain variations. Use when the user wants to benchmark on Real-World Quadrupedal Locomotion Dataset, or asks about evaluating this task. Reports Return.

researchpythongo
0
3
Quadrotor Visual Servoing EvalA

Evaluates a drone's ability to autonomously navigate to a target vehicle using marker-free visual servoing. It measures tracking accuracy, localization precision, and flight efficiency in both simulation and real-world environments. Use when the user wants to benchmark on Custom Simulation & Real-World Flight Dataset, or asks about evaluating this task. Reports NormError.

researchpythongo
0
3
Quac EvalA

Evaluates multi-turn, context-dependent question answering where models must track dialog history, resolve coreference, and handle open-ended follow-ups. It specifically probes a system's ability to manage asymmetric knowledge access and correctly identify unanswerable questions within an information-seeking dialogue. Use when the user wants to benchmark on QuAC, or asks about evaluating this task. Reports word-level F1.

researchpythongo
0
3
Qu Brats ScoreA

Evaluates the calibration of voxel-wise uncertainty estimates in brain tumor segmentation. It measures how effectively uncertainty thresholds filter out incorrect predictions while preserving correct ones, rewarding high confidence in accurate regions and penalizing the loss of correct predictions when filtering uncertain voxels. Use when the user has predictions and gold and needs to compute QU-BraTS unified score.

researchpythongo
0
3
Qsvm Fraud Detection EvalA

Evaluates a quantum support vector machine (QSVM) with quantum feature selection for binary fraud detection on real-world card payment data. It probes the model's ability to identify fraudulent transactions using a balanced dataset and compares performance against classical feature selection baselines. Use when the user wants to benchmark on Real-world card payment data (Balanced Data Set), or asks about evaluating this task. Reports Accuracy.

researchpythonbackend
0
3
Qrecc Attribution Fluency EvalA

Evaluates the tradeoff between response fluency and factual attribution in retrieval-augmented conversational LLMs. It measures how well models generate coherent, context-aware responses while correctly grounding answers in provided evidence or dialog history. Use when the user wants to benchmark on QReCC, or asks about evaluating this task. Reports Auto-AIS.

researchpythongo
0
3
Qrcd EvalA

Evaluates machine reading comprehension on a low-resource religious domain (Qur'an). It probes a model's ability to extract precise answer spans from Arabic text given a question, testing both exact matching and partial semantic/token overlap. Use when the user wants to benchmark on QRCD, or asks about evaluating this task. Reports pRR.

researchpythongo
0
3
Qqp Paraphrase Generation EvalA

Evaluates a model's ability to generate semantically equivalent paraphrase sentences from an input question, measuring lexical and semantic overlap with ground truth references. The benchmark probes sentence-level semantic understanding and generative fluency in a question-paraphrase setting. Use when the user wants to benchmark on Quora Question Pairs (QQP), or asks about evaluating this task. Reports BLEU.

researchpython
0
3
QqeA

Evaluates whether citation growth outpaces publication expansion across leading AI/NLP conferences, measuring the elasticity of scholarly impact relative to scale-driven growth over a decade. Use when the user has predictions and gold and needs to compute QQE.

researchpythongo
0
3
Qpain EvalA

Measures social bias in medical question-answering systems for pain management by evaluating treatment denial rates across intersectional race-gender profiles. It probes whether AI models exhibit discriminatory prescribing patterns when presented with clinical vignettes containing demographic attributes. Use when the user wants to benchmark on Q-Pain, or asks about evaluating this task. Reports probability_of_no.

researchpythontesting
0
3
Qoc EvalA

Evaluates mathematical reasoning, knowledge-oriented language understanding, and challenging reasoning capabilities of LLMs fine-tuned on domain-specific corpora. Use when the user wants to benchmark on MATH, GSM8K, MMLU, AGIEval, BIG-Bench Hard, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Qoblib EvalA

Evaluates the performance of quantum and classical optimization algorithms on intractable combinatorial problems. It measures solution quality, algorithmic success rates, and computational efficiency to track progress toward quantum advantage. Use when the user wants to benchmark on QOBLIB, or asks about evaluating this task. Reports best_objective_value.

researchpythongo
0
3
Qnn Smt Verification EvalA

This evaluation protocol tests the scalability and correctness of SMT-based formal verification for quantized neural networks. It measures how well an SMT model-checking framework can prove safety properties or find counterexamples across different quantization levels, network architectures, and SMT solvers. Use when the user wants to benchmark on Iris dataset, Vocalic dataset, AcasXu benchmark, or asks about evaluating this task. Reports verification_time.

researchpythonangular
0
3
Qmsum EvalA

Evaluates an agent's ability to distill and synthesize key information from high-noise, multi-turn meeting transcripts based on specific user queries. It probes the model's capacity to maintain global context while filtering irrelevant dialogue turns to produce a query-relevant summary. Use when the user wants to benchmark on QMSUM, or asks about evaluating this task. Reports ROUGE-1.

researchpythongo
0
3
Qm9 Property Prediction EvalA

Assesses a model's ability to predict scalar quantum chemical properties from molecular structures. It evaluates accuracy on five key electronic and vibrational targets using mean absolute error and average ranking. Use when the user wants to benchmark on QM9, or asks about evaluating this task. Reports MAE.

researchpython
0
3
Qm9 Mol Structtok EvalA

Evaluates the ability of a tokenization framework to generate valid 3D molecular structures and predict quantum mechanical properties. It probes structural validity, geometric plausibility, conditional controllability, and property prediction accuracy on organic molecules. Use when the user wants to benchmark on QM9, or asks about evaluating this task. Reports Mean Absolute Error (MAE).

researchpython
0
3
Qilin EvalA

Evaluates multimodal information retrieval systems across search, recommendation, and deep query answering (DQA) tasks using real-world APP-level user sessions. It probes a model's ability to rank heterogeneous content (text, images, videos) and generate accurate answers augmented by retrieved documents. Use when the user wants to benchmark on Qilin, or asks about evaluating this task. Reports MRR@10.

researchpythongo
0
3
Qianfan Ocr EvalA

This evaluation probes a unified vision-language model's ability to perform end-to-end document intelligence, including specialized OCR, general text recognition, document understanding, and key information extraction across diverse document types and multilingual scenarios. Use when the user wants to benchmark on Omni-Doc-Bench v1.5, OLMOCRBench, OCRBench, DocVQA, ChartQA, Nanonets KIE, or asks about evaluating this task. Reports normalized accuracy (0-100).

researchpythongo
0
3
Qgeval EvalA

Evaluates the quality of generated questions across seven dimensions: fluency, clarity, conciseness, relevance, consistency, answerability, and answer consistency. It measures how well automatic metrics and LLMs align with human judgments on these dimensions. Use when the user wants to benchmark on QGEval, or asks about evaluating this task. Reports Pearson correlation.

researchpythongo
0
3
Qg Bench EvalA

Evaluates the ability of generative language models to produce paragraph-level questions conditioned on a target answer and a context sentence. It probes domain adaptability and multilingual generalization across diverse extractive QA datasets. Use when the user wants to benchmark on SQuAD v1.1, SQuADShifts, SubjQA, Multilingual QA (JAQuAD, GerQuAD, SberQuAD, KorQuAD, FQuAD, Spanish SQuAD, Italian SQuAD), or asks about evaluating this task. Reports automatic evaluation metrics.

researchpythongo
0
3
Qe4pe EvalA

Probes the practical usability and impact of word-level quality estimation highlights on professional translators' post-editing efficiency, accuracy, and workflow. It measures how different highlight modalities (oracle, supervised, unsupervised, none) affect editing effort, productivity, and final translation quality in real-world domain-specific settings. Use when the user wants to benchmark on QE4PE, or asks about evaluating this task. Reports ESA score.

researchpythongo
0
3
Qder Re Ranking EvalA

Evaluates the ability of neural re-ranking models to effectively re-order a candidate set of documents based on complex query semantics and entity relationships. It probes fine-grained semantic matching, entity-aware attention, and late aggregation capabilities in information retrieval tasks across news and complex answer domains. Use when the user wants to benchmark on CODEC, TREC Complex Answer Retrieval (CAR) 2017, TREC Robust 2004, TREC News 2021, TREC Core 2018, or asks about evaluating ...

researchpythongo
0
3
Qcnmri Tumor Classification EvalA

Evaluates a hybrid quantum-classical convolutional neural network on MRI-based brain tumor detection. It probes the model's ability to classify medical images into binary (tumor vs. non-tumor) and multiclass (specific tumor types) categories under class imbalance and limited resolution constraints. Use when the user wants to benchmark on Brain MRI Tumor Dataset, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Qcaleval EvalA

Probes vision-language models' ability to interpret quantum calibration plots across diverse visual formats (1D traces, 2D maps, histograms) and perform structured scientific reasoning. It tests capabilities ranging from visual grounding and outcome classification to parameter extraction and operational calibration diagnosis, both in zero-shot and in-context learning settings. Use when the user wants to benchmark on QCalEval, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3