Evaluates a financial domain-specific LLM (FinGPT) across six core NLP tasks: sentiment analysis, text classification, named entity recognition, financial question answering, stock movement prediction, and text summarization. It probes the model's ability to handle domain-specific terminology, numerical reasoning, and structured output generation under instruction-tuning. Use when the user wants to benchmark on FLARE-FPB, FLARE-FIQASA, FinGPT Headline Classification, FinGPT/fingpt-ner, ConvFi...
Scanned 9/11/2026
Install to Claude Code
npx -y skills add qhjqhj00/research-skills-pool --skill fingpt-financial-eval --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Fingpt Financial Eval?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/qhjqhj00-fingpt-financial-eval)More formats (shields.io, HTML) on the badges page.
---
name: fingpt-financial-eval
description: Evaluates a financial domain-specific LLM (FinGPT) across six core NLP tasks: sentiment analysis, text classification, named entity recognition, financial question answering, stock movement prediction, and text summarization. It probes the model's ability to handle domain-specific terminology, numerical reasoning, and structured output generation under instruction-tuning. Use when the user wants to benchmark on FLARE-FPB, FLARE-FIQASA, FinGPT Headline Classification, FinGPT/fingpt-ner, ConvFinQA, FLARE-FinQA, CIKM18 (flare-ECTSum), StockNet (flare-SM-ACL), BigData22 (flare-SM-BigData), or asks about evaluating this task. Reports F1-score.
metadata:
skill_kind: dataset_eval
source_arxiv: 2507.08015
bibtex_key: djagba2025assessing
confidence: high
---
# fingpt-financial-eval
> Assessing the Capabilities and Limitations of FinGPT Model in Financial NLP Applications — Djagba et al. (2025) (arXiv:2507.08015, 2025)
## What this evaluates
Evaluates a financial domain-specific LLM (FinGPT) across six core NLP tasks: sentiment analysis, text classification, named entity recognition, financial question answering, stock movement prediction, and text summarization. It probes the model's ability to handle domain-specific terminology, numerical reasoning, and structured output generation under instruction-tuning.
## Datasets
- **FLARE-FPB** — total 970; splits: test (970)
- **FLARE-FIQASA** — total 235; splits: test (235)
- **FinGPT Headline Classification** — total ?; splits: test (-1)
- **FinGPT/fingpt-ner** — total 98; splits: test (98); HF `FinGPT/fingpt-ner`
- **ConvFinQA** — total 200; splits: test (200); HF `FinGPT/fingpt-convfinqa`
- **FLARE-FinQA** — total 50; splits: test (50); HF `ChanceFocus/flare-finqa`
- **CIKM18 (flare-ECTSum)** — total ?; splits: test (-1); HF `ChanceFocus/flare-ectsum`
- **StockNet (flare-SM-ACL)** — total ?; splits: test (-1); HF `ChanceFocus/flare-sm-acl`
- **BigData22 (flare-SM-BigData)** — total ?; splits: test (-1); HF `TheFinAI/flare-sm-bigdata`
## Metrics
- `F1-score` **(primary)** — range: [0, 1]
- Harmonic mean of precision and recall, calculated over the positive class (yes/no for classification, or per-entity type for NER).
- `accuracy` — range: [0, 1]
- Proportion of correctly predicted labels or numerically exact answers out of the total evaluated instances.
- `macro F1` — range: [0, 1]
- Unweighted mean of F1-scores computed independently for each class or entity type.
## Input / output format
**Input**: Instruction-style prompts formatted with explicit templates (e.g., `[INST]Classify the sentiment of the following financial headline:<HEADLINE>[/INST]`, `Instruction: Please extract entities... Input: <sentence> Answer:`). Inputs are lowercased, normalized, and tokenized with padding/truncation to a task-specific max length (32–1012 tokens).
**Output**: Text generation constrained to specific categories or values: sentiment labels ('yes', 'no', 'unknown'), entity type strings, numerical values extracted via regex, or stock movement directions ('up', 'down'). Outputs are post-processed and mapped to standardized labels for evaluation.
## Scoring recipe
```python
def score_classification(pred_text, gold_label):
pred = normalize_output(pred_text) # map to yes/no/unknown
if pred == 'unknown': return None # excluded per protocol
return 1 if pred == gold_label else 0
def score_qa(pred_text, gold_num):
nums = extract_numbers(pred_text) # regex parse
if not nums: return None
best_pred = min(nums, key=lambda x: abs(x - gold_num))
return 1 if abs(best_pred - gold_num) < 1e-3 else 0
# Aggregate precision, recall, F1 over non-None results
```
## Common pitfalls
- Outputs categorized as 'unknown' or ambiguous are explicitly excluded from metric aggregation, which can artificially inflate scores if the exclusion rate is high.
- Numerical QA evaluation requires strict regex parsing and filtering of invalid predictions; including malformed outputs skews accuracy.
- Generation parameters (max tokens, decoding strategy) heavily impact structured tasks like NER; greedy decoding with insufficient token limits causes severe hallucination and drops macro F1.
## Evidence (verbatim from paper)
> Model performance was assessed using standard classification metrics, including precision, recall, and F1-score, calculated over the yes and no classes. Any outputs categorized as unknown—due to lack of recognizable sentiment indicators –were excluded from the score aggregation to maintain the reliability of the evaluation.
## Citation
```bibtex
@misc{djagba2025assessing,
title={Assessing the Capabilities and Limitations of FinGPT Model in Financial NLP Applications},
author={Djagba et al. (2025)},
year={2025},
note={arXiv:2507.08015}
}
```
- arXiv: 2507.08015
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!