Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

20,843
skills in category
869
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 7,465–7,488 of 20,843 skills

Financial Exam Qa EvalA

Evaluates whether LLMs can perform domain-level conceptual understanding and precise financial reasoning across multilingual professional certification exams. It probes analytical rigor and regulatory knowledge integration rather than simple factual recall. Use when the user wants to benchmark on EFPA, GRFinQA, CFA, CPA, BBF, SAHM, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Financemath EvalA

This benchmark evaluates large language models' ability to perform knowledge-intensive mathematical reasoning within the finance domain. It requires models to integrate college-level financial knowledge with both textual descriptions and tabular data to solve complex problems. Use when the user wants to benchmark on FinanceMATH, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Financebench EvalA

Evaluates large language models' ability to answer financial questions using various retrieval and context strategies, probing numerical reasoning, factuality, and handling of structured or long documents. Use when the user wants to benchmark on FinanceBench, or asks about evaluating this task. Reports correct answer.

researchpythongo
0
3
Fin Bench V2 EvalA

Evaluates large language models on Finnish language capabilities across reading comprehension, commonsense reasoning, sentiment analysis, world knowledge, truthfulness, and alignment. It probes both multiple-choice and generative capabilities under varying prompt formulations (cloze vs. multiple-choice) and shot configurations. Use when the user wants to benchmark on FIN-bench-v2, or asks about evaluating this task. Reports normalized accuracy.

researchpythongo
0
3
Fin Bench EvalA

Evaluates Finnish large language models on a curated suite of 11 tasks spanning arithmetic, reasoning, general knowledge, emotion classification, and linguistic understanding. It probes few-shot generalization and task-specific capabilities in a low-resource language setting. Use when the user wants to benchmark on FIN-bench, or asks about evaluating this task. Reports mean accuracy.

researchpythonaws
0
3
Filler Word Detection EvalA

Evaluates a model's ability to detect and classify filler words (e.g., 'uh', 'um') in naturalistic speech recordings. It probes temporal localization accuracy and fine-grained acoustic classification under varying evaluation granularities. Use when the user wants to benchmark on PodcastFillers, or asks about evaluating this task. Reports F1.

researchpythongo
0
3
Filipino Text Benchmarks EvalA

Evaluates text classification capability in low-resource settings for the Filipino language, specifically measuring model robustness and performance degradation as training data size is systematically reduced. Use when the user wants to benchmark on Hate Speech, Dengue, or asks about evaluating this task. Reports accuracy, hamming loss.

researchpythongo
0
3
Filegrambench EvalA

Evaluates AI agents' ability to personalize based on file-system behavioral traces across procedural, semantic, and episodic memory channels. It probes attribute recognition, behavioral inference, anomaly detection, and grounding using simulated, multimodal, and real-world settings. Use when the user wants to benchmark on FileGramBench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Filbench EvalA

Evaluates LLMs' ability to understand and generate text in Filipino, Tagalog, and Cebuano across cultural knowledge, classical NLP tasks, reading comprehension, and text generation. It probes cultural alignment, factual recall, linguistic processing, and translation capabilities in low-resource Southeast Asian languages. Use when the user wants to benchmark on FilBench, or asks about evaluating this task. Reports FilBench Score.

researchpythongo
0
3
Figurebench EvalA

Evaluates the ability of text-to-illustration models to generate publication-ready scientific figures that balance structural fidelity, visual aesthetics, and communicative clarity based on long-form scientific text. Use when the user wants to benchmark on FigureBench, or asks about evaluating this task. Reports Overall score.

researchpythongo
0
3
Figedit Chart Editing EvalA

This benchmark evaluates the ability of vision-language and image-editing models to perform semantically correct, structure-aware modifications to scientific charts. It probes whether models can follow precise editing instructions while preserving data-encoding consistency, axis coherence, and legend integrity, rather than merely producing pixel-level visual similarity. Use when the user wants to benchmark on FigEdit, or asks about evaluating this task. Reports Instruction-following score.

researchpythongit
0
3
Fig Qa EvalA

Evaluates language models' ability to interpret nonliteral, creative metaphors by selecting the correct literal meaning from two opposing options, and generating sensible interpretations for novel metaphors. It probes commonsense grounding and contextual understanding beyond literal paraphrase tasks. Use when the user wants to benchmark on Fig-QA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Fife EvalA

Evaluates large language models' ability to follow complex financial instructions, with a strong emphasis on precise adherence to formatting, structural constraints, and conditional styling requirements. It probes whether models can maintain procedural compliance rather than just semantic correctness. Use when the user wants to benchmark on FIFE, or asks about evaluating this task. Reports Strict compliance.

researchpythongo
0
3
FidelityA

Probes whether input attribution methods accurately reflect token importance for model predictions, particularly when inputs are adversarially perturbed or masked out-of-distribution. It evaluates the consistency of fidelity scores across different model architectures and under various adversarial attacks. Use when the user has predictions and gold and needs to compute fidelity.

researchpythongo
0
3
FidA

Measures the distributional similarity between original GAN-generated images and their semantically manipulated counterparts. It evaluates whether a latent space transformation preserves overall image quality and realism while altering specific attributes. Use when the user has predictions and gold and needs to compute FID.

researchpythongo
0
3
Fibinet Ctr EvalA

Evaluates the ability of shallow and deep learning models to predict click-through rates (CTR) on large-scale ad impression datasets. It probes how well architectures can model high-order feature interactions and dynamically weight feature importance using bilinear functions and Squeeze-Excitation networks. Use when the user wants to benchmark on Criteo, Avazu, or asks about evaluating this task. Reports AUC.

researchpythongo
0
3
Fgveribench EvalA

Evaluates an LLM verifier's ability to rank multiple candidate answers by their factual correctness and error severity across single-hop and multi-hop questions, using external knowledge retrieval and fine-grained scoring. Use when the user wants to benchmark on FGVeriBench, or asks about evaluating this task. Reports Kendall-tau.

researchpythongo
0
3
Fgs Dti Prediction EvalA

Evaluates a fine-grained selective similarity integration framework for drug-target interaction prediction. It tests the model's ability to dynamically weight multiple drug and target similarity views based on local interaction consistency to predict binary interaction labels. Use when the user wants to benchmark on Nuclear Receptors (NR), G-protein coupled receptors (GPCR), Ion Channel (IC), Enzyme (E), Luo, or asks about evaluating this task. Reports AUC.

researchpythongit
0
3
Fgnn Session Rec EvalA

Evaluates session-based recommendation models by predicting the next item in a user's clickstream session. It probes the model's ability to capture latent, non-temporal transition patterns and item order within a session graph. Use when the user wants to benchmark on Yoochoose, Diginetica, or asks about evaluating this task. Reports R@20.

researchpythongo
0
3
Fgnn Sbr EvalA

Evaluates session-based recommendation models on e-commerce clickstream data by predicting the next item in a session using graph neural networks and cross-session information. Use when the user wants to benchmark on Yoochoose1/64, Yoochoose1/4, Diginetica, or asks about evaluating this task. Reports R@20, MRR@20.

researchpythongo
0
3
Fgner EvalA

Evaluates fine-grained named entity recognition (FgNER) capabilities across multiple languages by measuring how accurately models identify and classify specific entity types within text sequences. The protocol assesses sequence labeling performance using standard span-based metrics. Use when the user wants to benchmark on Various NER datasets (cited in text), or asks about evaluating this task. Reports F1-score.

researchpythongo
0
3
Fgn Weather Forecasting EvalA

Evaluates a functional generative network for medium-range probabilistic weather forecasting against operational ground truth (HRES-fc0) and a diffusion-based baseline (GenCast). Probes the model's ability to capture joint spatial structures and predict tropical cyclone tracks using deterministic and probabilistic scoring rules. Use when the user wants to benchmark on HRES-fc0, ERA5, or asks about evaluating this task. Reports probabilistic metrics.

researchpythongo
0
3
Fgbench EvalA

Tests functional-group reasoning capability by asking models to predict how molecular properties change when specific functional groups are added, removed, or modified at given positions. Use when the user wants to benchmark on FGBench, or asks about evaluating this task. Reports Accuracy (Acc).

researchpythongo
0
3
Fg Mae Transfer EvalA

Evaluates the transfer learning capability of a feature-guided masked autoencoder on remote sensing imagery. It probes the model's ability to adapt to downstream scene classification and semantic segmentation tasks across multispectral and SAR modalities using linear probing and fine-tuning protocols. Use when the user wants to benchmark on BigEarthNet-MM, BigEarthNet-SAR, EuroSAT, EuroSAT-SAR, DFC2020, or asks about evaluating this task. Reports mAP.

researchpython
0
3