Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

20,843
skills in category
869
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 7,441–7,464 of 20,843 skills

Fingen EvalA

Evaluates forward-looking argument generation in finance across text-to-claim, chart-to-argument, and news-to-argument tasks. Probes a model's ability to generate plausible, structured future scenarios and claims based on financial inputs while maintaining factual consistency and handling financial terminology and numerals. Use when the user wants to benchmark on FinGen, or asks about evaluating this task. Reports ROUGE-1.

researchpython
0
3
Fineweb2 Early Signal EvalA

Evaluates the impact of different pre-training data processing steps on downstream model quality by training small language models and measuring performance on a curated suite of multilingual zero-shot benchmarks. Use when the user wants to benchmark on FineWeb2 Early-Signal Benchmark Suite, or asks about evaluating this task. Reports per-category macro-average score.

researchpythongo
0
3
Finevision EvalA

Evaluates vision-language models on a diverse suite of 11 multimodal benchmarks covering visual question answering, chart understanding, document parsing, and general multimodal reasoning. Additionally probes GUI/agentic capabilities on screen interaction tasks. Use when the user wants to benchmark on AI2D, ChartQA, DocVQA, InfoVQA, MME, MMMU, ScienceQA, MMStar, OCRBench, TextVQA, SEED-Bench, Screenspot-V2, Screenspot-Pro, or asks about evaluating this task. Reports mean normalized performanc...

researchpythonperformance
0
3
Fineval EvalA

Evaluates large language models' knowledge and reasoning capabilities in the Chinese financial domain across multiple academic subjects like Finance, Economy, Accounting, and professional Certificates. It tests performance under zero-shot, few-shot, answer-only, and chain-of-thought prompting settings. Use when the user wants to benchmark on FinEval, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Finetasks EvalA

Evaluates the downstream quality of multilingual pretraining corpora by training language models on them and measuring performance on a standardized suite of fine-tuning tasks across Arabic, Hindi, and Turkish. Use when the user wants to benchmark on FineTasks, or asks about evaluating this task. Reports FineTasks scores.

researchpythonperformance
0
3
Finerumfact EvalA

Evaluates a model's ability to perform sentence-level fact verification on generated summaries and localize specific factuality error types. It measures how well the model's judgments align with human annotations across sentence, summary, and system levels. Use when the user wants to benchmark on FineSumFact, or asks about evaluating this task. Reports balanced accuracy (bAcc).

researchpythongo
0
3
Finedialfact EvalA

Evaluates a model's ability to perform fine-grained fact verification on dialogue responses by checking individual atomic facts against external knowledge. It probes whether models can correctly classify facts as supporting, refuting, or lacking sufficient information, particularly in the presence of hallucinations and imbalanced label distributions. Use when the user wants to benchmark on FineDialFact, or asks about evaluating this task. Reports F1-score.

researchpythongo
0
3
Fine T2i EvalA

Probes the effectiveness of a large-scale text-to-image fine-tuning dataset in improving generation quality, text-image alignment, and instruction following across different model architectures (diffusion and autoregressive). Use when the user wants to benchmark on Artificial Analysis Image Arena (Eval Subset), or asks about evaluating this task. Reports human_win_rate.

researchpythongo
0
3
Fine R1 EvalA

Evaluates multi-modal large language models on fine-grained visual recognition (FGVR) tasks across six standard datasets. It probes the model's ability to distinguish visually similar sub-categories in both closed-world (seen categories) and open-world (unseen categories) settings using chain-of-thought reasoning. Use when the user wants to benchmark on CaltechUCSD Bird-200, Stanford Car-196, Stanford Dog-120, Flower-102, Oxford-IIIT Pet-37, FGVC-Aircraft, or asks about evaluating this task. ...

researchpythongo
0
3
Fine Grained Interpretability EvalA

Evaluates the interpretability and faithfulness of neural NLP models by measuring how well token-level saliency rationales align with human-annotated ground truth rationales and model predictions across sentiment analysis, semantic textual similarity, and machine reading comprehension tasks. Use when the user wants to benchmark on Fine-grained Interpretability Benchmark (SA/STS/MRC), or asks about evaluating this task. Reports Token-F1, MAP.

researchpythongo
0
3
Fine Grained Ie EvalA

This benchmark evaluates large language models' capacity for fine-grained information extraction under augmented instructions. It specifically probes two generalization capabilities: adapting to unseen information types and handling unfamiliar task forms. Performance is measured using standard extraction metrics to compare encoder-decoder and decoder-only architectures. Use when the user wants to benchmark on Fine-grained Information Extraction Benchmark, or asks about evaluating this task. R...

researchpythongo
0
3
Fine Grained Classification EvalA

Probes the ability of vision-language models to distinguish visually similar subcategories within broader classes (e.g., specific bird species or car models) using a multiple-choice format. Use when the user wants to benchmark on ImageNet-1K, Oxford Flowers-102, Oxford-IIIT Pet-37, Food-101, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Findingdory EvalA

This benchmark evaluates long-term memory and spatio-temporal reasoning in embodied agents. It requires agents to recall specific past interactions from a video history to select goal frames and navigate to target entities in dynamic, photorealistic environments over long-horizon tasks. Use when the user wants to benchmark on FindingDory, or asks about evaluating this task. Reports LL-SR.

researchpythongo
0
3
Findata EvalA

Evaluates multi-task and single-task learning models on a curated collection of financial NLP tasks (FinDATA) to assess how task diversity, relatedness, and parameter-efficient architectures impact performance. Use when the user wants to benchmark on FinDATA, or asks about evaluating this task. Reports evaluation metrics in Table 2.

researchpythongo
0
3
Find EvalA

Evaluates automated interpretability methods on their ability to describe black-box functions across numeric, string, and semantic domains. It probes both static language model capabilities and interactive agent reasoning, including handling complexities like composition, noise, bias, and approximation. Use when the user wants to benchmark on FIND, or asks about evaluating this task. Reports adequately described rate.

researchpythongo
0
3
Finchart Bench EvalA

Evaluates vision-language models' ability to comprehend real-world financial charts. It probes spatial reasoning, instruction following, and factual extraction across True/False, Multiple Choice, and open-ended Question Answering tasks. Use when the user wants to benchmark on FinChart-Bench, or asks about evaluating this task. Reports Exact Match (EM).

researchpythongo
0
3
Finchain EvalA

Evaluates multi-step symbolic financial reasoning by measuring how well language models generate verifiable chain-of-thought traces aligned with executable financial templates. It probes both the semantic and numeric consistency of intermediate reasoning steps and the accuracy of the final financial answer. Use when the user wants to benchmark on FinChain, or asks about evaluating this task. Reports ChainEval.

researchpythongo
0
3
Financial Sts EvalA

Evaluates a model's ability to detect subtle semantic shifts between pairs of financial narratives by ranking similar pairs higher than dissimilar ones. It probes nuanced understanding of financial language, including intensified sentiment, elaborated details, plan realization, and emerging situations. Use when the user wants to benchmark on LLM-augmented FinSTS, Human-annotated FinSTS, or asks about evaluating this task. Reports AUC.

researchpythongo
0
3
Financial Retrieval EvalA

Evaluates the ability of text embedding models to retrieve relevant financial document passages given complex, long-form queries. It probes domain-specific retrieval capabilities, including sensitivity to company names, tickers, financial metrics, and date-specific information. Use when the user wants to benchmark on Financial Document Retrieval Dataset, or asks about evaluating this task. Reports Recall@1.

researchpythongo
0
3
Financial Phrase Bank Sentiment EvalA

Probes the ability of LLMs and traditional NLP tools to accurately classify financial sentiment (positive, neutral, or negative) from news headlines and earnings-related text. It specifically evaluates how well models capture nuanced, hedged, or domain-specific financial language compared to baseline sentiment engines. Use when the user wants to benchmark on Financial Phrase Bank, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Financial Nlp EvalA

Evaluates the capability of large language models (ChatGPT and GPT-4) and domain-specific models to solve a variety of financial text analytics tasks. It probes performance across sentiment analysis, classification, information extraction, and question answering, measuring how well models handle domain-specific knowledge and structured prediction. Use when the user wants to benchmark on Financial NLP Tasks (Sentiment, Classification, NER, RE, QA), or asks about evaluating this task. Reports a...

researchpythongo
0
3
Financial Nlp Efficiency EvalA

Evaluates large language models on ten financial NLP tasks to measure classification and generation accuracy, inference speed, and a novel Token Efficiency Score (TES) that quantifies the performance-compute trade-off. Use when the user wants to benchmark on Financial NLP tasks (10 datasets), or asks about evaluating this task. Reports Token Efficiency Score (TES).

researchpythongo
0
3
Financial Misinformation Detection EvalA

This benchmark evaluates a model's ability to detect financial misinformation in a reference-free setting, where models must classify financial narrative paragraphs as true or false without external knowledge or source documents. It probes semantic pattern recognition for financial manipulation, omission detection, and domain-specific generalization. Use when the user wants to benchmark on MisD@ICWSM2026, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Financial Llm Zero Shot EvalA

This benchmark evaluates the zero-shot instruction-following and classification capabilities of large language models on four financial NLP tasks: FOMC communication sentiment, general financial sentiment, numerical claim detection, and named entity recognition. It measures how well generative models can perform these tasks without fine-tuning compared to traditional PLMs. Use when the user wants to benchmark on Financial NLP Tasks (FOMC, Sentiment, Claim Detection, NER), or asks about evalua...

researchpythongo
0
3