Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 7,441–7,464 of 20,843 skills
Evaluates forward-looking argument generation in finance across text-to-claim, chart-to-argument, and news-to-argument tasks. Probes a model's ability to generate plausible, structured future scenarios and claims based on financial inputs while maintaining factual consistency and handling financial terminology and numerals. Use when the user wants to benchmark on FinGen, or asks about evaluating this task. Reports ROUGE-1.
Evaluates the impact of different pre-training data processing steps on downstream model quality by training small language models and measuring performance on a curated suite of multilingual zero-shot benchmarks. Use when the user wants to benchmark on FineWeb2 Early-Signal Benchmark Suite, or asks about evaluating this task. Reports per-category macro-average score.
Evaluates vision-language models on a diverse suite of 11 multimodal benchmarks covering visual question answering, chart understanding, document parsing, and general multimodal reasoning. Additionally probes GUI/agentic capabilities on screen interaction tasks. Use when the user wants to benchmark on AI2D, ChartQA, DocVQA, InfoVQA, MME, MMMU, ScienceQA, MMStar, OCRBench, TextVQA, SEED-Bench, Screenspot-V2, Screenspot-Pro, or asks about evaluating this task. Reports mean normalized performanc...
Evaluates large language models' knowledge and reasoning capabilities in the Chinese financial domain across multiple academic subjects like Finance, Economy, Accounting, and professional Certificates. It tests performance under zero-shot, few-shot, answer-only, and chain-of-thought prompting settings. Use when the user wants to benchmark on FinEval, or asks about evaluating this task. Reports accuracy.
Evaluates the downstream quality of multilingual pretraining corpora by training language models on them and measuring performance on a standardized suite of fine-tuning tasks across Arabic, Hindi, and Turkish. Use when the user wants to benchmark on FineTasks, or asks about evaluating this task. Reports FineTasks scores.
Evaluates a model's ability to perform sentence-level fact verification on generated summaries and localize specific factuality error types. It measures how well the model's judgments align with human annotations across sentence, summary, and system levels. Use when the user wants to benchmark on FineSumFact, or asks about evaluating this task. Reports balanced accuracy (bAcc).
Evaluates a model's ability to perform fine-grained fact verification on dialogue responses by checking individual atomic facts against external knowledge. It probes whether models can correctly classify facts as supporting, refuting, or lacking sufficient information, particularly in the presence of hallucinations and imbalanced label distributions. Use when the user wants to benchmark on FineDialFact, or asks about evaluating this task. Reports F1-score.
Probes the effectiveness of a large-scale text-to-image fine-tuning dataset in improving generation quality, text-image alignment, and instruction following across different model architectures (diffusion and autoregressive). Use when the user wants to benchmark on Artificial Analysis Image Arena (Eval Subset), or asks about evaluating this task. Reports human_win_rate.
Evaluates multi-modal large language models on fine-grained visual recognition (FGVR) tasks across six standard datasets. It probes the model's ability to distinguish visually similar sub-categories in both closed-world (seen categories) and open-world (unseen categories) settings using chain-of-thought reasoning. Use when the user wants to benchmark on CaltechUCSD Bird-200, Stanford Car-196, Stanford Dog-120, Flower-102, Oxford-IIIT Pet-37, FGVC-Aircraft, or asks about evaluating this task. ...
Evaluates the interpretability and faithfulness of neural NLP models by measuring how well token-level saliency rationales align with human-annotated ground truth rationales and model predictions across sentiment analysis, semantic textual similarity, and machine reading comprehension tasks. Use when the user wants to benchmark on Fine-grained Interpretability Benchmark (SA/STS/MRC), or asks about evaluating this task. Reports Token-F1, MAP.
This benchmark evaluates large language models' capacity for fine-grained information extraction under augmented instructions. It specifically probes two generalization capabilities: adapting to unseen information types and handling unfamiliar task forms. Performance is measured using standard extraction metrics to compare encoder-decoder and decoder-only architectures. Use when the user wants to benchmark on Fine-grained Information Extraction Benchmark, or asks about evaluating this task. R...
Probes the ability of vision-language models to distinguish visually similar subcategories within broader classes (e.g., specific bird species or car models) using a multiple-choice format. Use when the user wants to benchmark on ImageNet-1K, Oxford Flowers-102, Oxford-IIIT Pet-37, Food-101, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates long-term memory and spatio-temporal reasoning in embodied agents. It requires agents to recall specific past interactions from a video history to select goal frames and navigate to target entities in dynamic, photorealistic environments over long-horizon tasks. Use when the user wants to benchmark on FindingDory, or asks about evaluating this task. Reports LL-SR.
Evaluates multi-task and single-task learning models on a curated collection of financial NLP tasks (FinDATA) to assess how task diversity, relatedness, and parameter-efficient architectures impact performance. Use when the user wants to benchmark on FinDATA, or asks about evaluating this task. Reports evaluation metrics in Table 2.
Evaluates automated interpretability methods on their ability to describe black-box functions across numeric, string, and semantic domains. It probes both static language model capabilities and interactive agent reasoning, including handling complexities like composition, noise, bias, and approximation. Use when the user wants to benchmark on FIND, or asks about evaluating this task. Reports adequately described rate.
Evaluates vision-language models' ability to comprehend real-world financial charts. It probes spatial reasoning, instruction following, and factual extraction across True/False, Multiple Choice, and open-ended Question Answering tasks. Use when the user wants to benchmark on FinChart-Bench, or asks about evaluating this task. Reports Exact Match (EM).
Evaluates multi-step symbolic financial reasoning by measuring how well language models generate verifiable chain-of-thought traces aligned with executable financial templates. It probes both the semantic and numeric consistency of intermediate reasoning steps and the accuracy of the final financial answer. Use when the user wants to benchmark on FinChain, or asks about evaluating this task. Reports ChainEval.
Evaluates a model's ability to detect subtle semantic shifts between pairs of financial narratives by ranking similar pairs higher than dissimilar ones. It probes nuanced understanding of financial language, including intensified sentiment, elaborated details, plan realization, and emerging situations. Use when the user wants to benchmark on LLM-augmented FinSTS, Human-annotated FinSTS, or asks about evaluating this task. Reports AUC.
Evaluates the ability of text embedding models to retrieve relevant financial document passages given complex, long-form queries. It probes domain-specific retrieval capabilities, including sensitivity to company names, tickers, financial metrics, and date-specific information. Use when the user wants to benchmark on Financial Document Retrieval Dataset, or asks about evaluating this task. Reports Recall@1.
Probes the ability of LLMs and traditional NLP tools to accurately classify financial sentiment (positive, neutral, or negative) from news headlines and earnings-related text. It specifically evaluates how well models capture nuanced, hedged, or domain-specific financial language compared to baseline sentiment engines. Use when the user wants to benchmark on Financial Phrase Bank, or asks about evaluating this task. Reports accuracy.
Evaluates the capability of large language models (ChatGPT and GPT-4) and domain-specific models to solve a variety of financial text analytics tasks. It probes performance across sentiment analysis, classification, information extraction, and question answering, measuring how well models handle domain-specific knowledge and structured prediction. Use when the user wants to benchmark on Financial NLP Tasks (Sentiment, Classification, NER, RE, QA), or asks about evaluating this task. Reports a...
Evaluates large language models on ten financial NLP tasks to measure classification and generation accuracy, inference speed, and a novel Token Efficiency Score (TES) that quantifies the performance-compute trade-off. Use when the user wants to benchmark on Financial NLP tasks (10 datasets), or asks about evaluating this task. Reports Token Efficiency Score (TES).
This benchmark evaluates a model's ability to detect financial misinformation in a reference-free setting, where models must classify financial narrative paragraphs as true or false without external knowledge or source documents. It probes semantic pattern recognition for financial manipulation, omission detection, and domain-specific generalization. Use when the user wants to benchmark on MisD@ICWSM2026, or asks about evaluating this task. Reports Accuracy.
This benchmark evaluates the zero-shot instruction-following and classification capabilities of large language models on four financial NLP tasks: FOMC communication sentiment, general financial sentiment, numerical claim detection, and named entity recognition. It measures how well generative models can perform these tasks without fine-tuning compared to traditional PLMs. Use when the user wants to benchmark on Financial NLP Tasks (FOMC, Sentiment, Claim Detection, NER), or asks about evalua...