Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 7,417–7,440 of 20,843 skills
This benchmark probes a model's ability to perform pixel-wise anomaly detection and uncertainty estimation in complex urban driving scenes. It specifically measures how well a segmentation wrapper identifies out-of-distribution objects (e.g., lost & found items, static blends, web overlays) without degrading the underlying semantic segmentation accuracy. Use when the user wants to benchmark on Fishyscapes benchmark, or asks about evaluating this task. Reports AP.
Evaluates computer vision models on fine-grained species classification, multi-label trait identification, and pixel-level trait segmentation in fish images. Probes capabilities in handling long-tailed distributions, out-of-distribution generalization to unseen species, and localizing small/rare anatomical features. Use when the user wants to benchmark on Fish-Vista, or asks about evaluating this task. Reports macro-averaged F1-score, Mean Average Precision (mAP), mean Intersection over Union...
Evaluates speech synthesis models on intelligibility, speaker similarity, and long-form generation across multiple languages. It also assesses subjective qualities like naturalness, instruction-following, and human-level indistinguishability using automated LLM-as-a-Judge and Audio Turing Test frameworks. Use when the user wants to benchmark on Seed-TTS-Eval, CV3-Eval, Minimax Multilingual Testset, Long-TTS-Eval, Audio Turing Test, Emergent TTS Eval, or asks about evaluating this task. Report...
Evaluates the ranking effectiveness and inference latency of single-token decoding for listwise document reranking. It probes whether using only the first-token logits of alphabetical identifiers can accurately rank candidate documents compared to full sequence generation or traditional language modeling objectives. Use when the user wants to benchmark on TREC DL19-22, BEIR, MS MARCO, or asks about evaluating this task. Reports ranking effectiveness.
Evaluates the performance and scalability of a federated inference scheduling framework under varying request loads, measuring throughput, latency, and auto-scaling capabilities on distributed HPC resources. Use when the user wants to benchmark on ShareGPT, or asks about evaluating this task. Reports Request throughput (req/s).
Evaluates multimodal geospatial models on wildfire risk prediction across in-distribution and out-of-distribution regions. It probes the ability of vision-language models to generate chain-of-thought reasoning traces that condition a vision decoder for accurate, interpretable spatial risk raster generation. Use when the user wants to benchmark on FireScope-Bench, or asks about evaluating this task. Reports ROC AUC, QWK.
Evaluates autonomous coding agents on their ability to rediscover established scientific findings by autonomously planning, implementing, and executing experiments from scratch based only on high-level research questions. It probes end-to-end research workflow capabilities, including experimental design, code generation, and evidence-based conclusion formation. Use when the user wants to benchmark on FIRE-Bench, or asks about evaluating this task. Reports F1.
Evaluates the ability of Large Vision-Language Models (LVLMs) to generate accurate, fluent, and comprehensive captions for long videos. It probes semantic alignment, event coverage, and hallucination resistance by comparing model outputs against multi-annotator human references. Use when the user wants to benchmark on FIOVA, or asks about evaluating this task. Reports FIOVA-DQ F1.
Evaluates a multi-agent financial system's ability to answer real-world investor inquiries across macroeconomic, industry, and company analysis scenarios. Probes accuracy, thoroughness, clarity, and professional financial reasoning through automated LLM-judging and human preference testing. Use when the user wants to benchmark on NGA Grand Era Investor Inquiries, or asks about evaluating this task. Reports Overall Score.
Evaluates large language models on structure-aware XBRL tagging for financial information. It probes two subtasks: numeric entity identification (FinNI) and fine-grained concept linking (FinCL) against the US-GAAP taxonomy, testing the model's ability to extract structured facts and align them with hierarchical financial concepts. Use when the user wants to benchmark on FinTagging, or asks about evaluating this task. Reports macro-F1.
Evaluates the quality of a machine-translated extractive QA dataset (FinSQuAD) by training and testing QA models on it, comparing performance against other translated SQuAD datasets and the original English version. It also assesses translation fidelity through backtranslation and manual error analysis. Use when the user wants to benchmark on Finnish SQuAD2.0, SQuAD2.0, or asks about evaluating this task. Reports exact match (EM).
Evaluates financial LLMs across seven text-based tasks (sentiment analysis, NER, number understanding, summarization, stock movement prediction, credit scoring, firm disclosure) and three multimodal/hallucination tasks (ChartQA, FinVQA, FinTerms). It measures domain-specific reasoning, instruction following, and hallucination mitigation in financial contexts. Use when the user wants to benchmark on FinSet, ChartQA, FinVQA, FinTerms-MCQ, FinTerms-Gen, Finance Bench, or asks about evaluating th...
Evaluates the ability of foundation models and tabular baselines to predict financial risk (default, fraud, churn) by classifying customer profiles generated from tabular data. It probes how well profile-based tuning captures richer customer semantics compared to isolated table-based classification. Use when the user wants to benchmark on FinBench, or asks about evaluating this task. Reports F1-score.
Evaluates computer vision models on forest scene understanding, specifically testing instance segmentation, panoptic segmentation, and depth completion in unstructured, densely populated natural environments. Use when the user wants to benchmark on FinnWoodlands, or asks about evaluating this task. Reports mAP@50.
Evaluates Finnish NLP capabilities across four core tasks: part-of-speech tagging, named entity recognition, dependency parsing, and text classification on domain-specific and out-of-domain corpora. It probes a model's ability to handle morphologically rich language and varying text registers. Use when the user wants to benchmark on Finnish NLP Benchmarks, or asks about evaluating this task. Reports Average.
Evaluates named entity recognition systems on Finnish text, testing their ability to identify and classify entities (person, location, organization, product, event, date) in both in-domain news and out-of-domain Wikipedia corpora. It specifically probes domain generalization and the handling of nested entity spans. Use when the user wants to benchmark on Finnish News Corpus, or asks about evaluating this task. Reports F1-score.
Evaluates the performance, power, and resource efficiency of quantized neural networks deployed on various FPGA platforms using the FINN-R framework. It probes the trade-offs between network precision, hardware resource usage, throughput, and classification accuracy across embedded and datacenter-scale hardware. Use when the user wants to benchmark on MNIST, CIFAR-10, GTSRB, SVHN, VOC 2007, ImageNet, or asks about evaluating this task. Reports Top-1 Accuracy.
Evaluates the inference performance of binarized neural networks (BNNs) accelerated on FPGAs using the FINN framework. It measures classification throughput, latency, and accuracy across standard image datasets to assess hardware efficiency and resource utilization. Use when the user wants to benchmark on MNIST, CIFAR-10, SVHN, or asks about evaluating this task. Reports classification throughput (FPS).
Evaluates embedding models on a finance-specific benchmark across seven standard embedding tasks to measure performance relative to general-domain benchmarks. It probes whether general-purpose models suffer a significant performance drop on domain-specific financial text and investigates if this gap is driven by domain shift or inherent dataset complexity. Use when the user wants to benchmark on FinMTEB, or asks about evaluating this task. Reports FinMTEB Score.
Evaluates the multimodal financial reasoning capabilities of LLMs and MLLMs on expert-level question-answer pairs spanning 15 financial domains. It probes the models' ability to interpret complex visual data (charts, tables), apply domain-specific formulas, and perform multi-step logical calculations. Use when the user wants to benchmark on FinMR, or asks about evaluating this task. Reports accuracy.
Evaluates financial multi-modal reasoning capabilities of models on chart-based analysis and domain-specific knowledge. It probes perception, analysis, and reasoning across 18 financial domains and 6 asset classes using multiple-choice and computational problems. Use when the user wants to benchmark on FinMME, or asks about evaluating this task. Reports FinScore.
Evaluates the financial natural language processing capabilities of encoder-only and decoder-only language models across multiple classification tasks. It probes zero-shot prompting, in-context learning strategies, and the impact of data availability (public vs. proprietary) on model performance. Use when the user wants to benchmark on FinSent, FPB, FiQA SA, ESG, FLS, QA, Headlines-PDU, Headlines-PDC, Headlines-PDD, Headlines-PI, Headlines-AC, Headlines-FI, Headlines-PS, NER, FOMC, or asks ab...
Evaluates AI models' ability to assess the quality of financial information disclosure in Chinese. It probes four capabilities: identifying relevant questions, determining question relevance, evaluating answer readability, and measuring answer relevance to the question. Use when the user wants to benchmark on FinTruthQA, or asks about evaluating this task. Reports Accuracy, Precision, Recall, F1-score, Micro/Macro/Weighted F1, Quadratic Weighted Kappa (QWK).
Evaluates a financial domain-specific LLM (FinGPT) across six core NLP tasks: sentiment analysis, text classification, named entity recognition, financial question answering, stock movement prediction, and text summarization. It probes the model's ability to handle domain-specific terminology, numerical reasoning, and structured output generation under instruction-tuning. Use when the user wants to benchmark on FLARE-FPB, FLARE-FIQASA, FinGPT Headline Classification, FinGPT/fingpt-ner, ConvFi...