This benchmark evaluates a model's ability to detect hallucinations in domain-specific question answering systems that use retrieval-augmented generation. It probes whether models can correctly identify when a generated answer contradicts or goes beyond the provided retrieved context, often due to over-reliance on pre-trained knowledge or incomplete retrieval. Use when the user wants to benchmark on DelucionQA, or asks about evaluating this task. Reports Macro F1.
Scanned 9/11/2026
Install to Claude Code
npx -y skills add qhjqhj00/research-skills-pool --skill delucionqa-eval --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Delucionqa Eval?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/qhjqhj00-delucionqa-eval)More formats (shields.io, HTML) on the badges page.
---
name: delucionqa-eval
description: This benchmark evaluates a model's ability to detect hallucinations in domain-specific question answering systems that use retrieval-augmented generation. It probes whether models can correctly identify when a generated answer contradicts or goes beyond the provided retrieved context, often due to over-reliance on pre-trained knowledge or incomplete retrieval. Use when the user wants to benchmark on DelucionQA, or asks about evaluating this task. Reports Macro F1.
metadata:
skill_kind: dataset_eval
source_arxiv: 2312.05200
bibtex_key: sadat2023delucionqa
confidence: high
---
# delucionqa-eval
> DelucionQA: Detecting Hallucinations in Domain-specific Question Answering — Sadat et al. (2023) (arXiv:2312.05200, 2023)
## What this evaluates
This benchmark evaluates a model's ability to detect hallucinations in domain-specific question answering systems that use retrieval-augmented generation. It probes whether models can correctly identify when a generated answer contradicts or goes beyond the provided retrieved context, often due to over-reliance on pre-trained knowledge or incomplete retrieval.
## Datasets
- **DelucionQA** — total ?; splits: train (-1), dev (-1), test (-1); repo https://github.com/boschresearch/DelucionQA
## Metrics
- `Macro F1` **(primary)** — range: percent
- The unweighted mean of the F1 scores for each class (hallucination and non-hallucination). F1 is calculated as 2 * (precision * recall) / (precision + recall) for each class, then averaged across classes.
## Input / output format
**Input**: A question and the retrieved context (retrieval result).
**Output**: A binary label indicating whether the generated answer contains a hallucination or is faithful to the context.
## Scoring recipe
```python
def compute_macro_f1(predictions, gold_labels):
classes = ['hallucination', 'non_hallucination']
f1_scores = []
for cls in classes:
tp = sum(1 for p, g in zip(predictions, gold_labels) if p == cls and g == cls)
fp = sum(1 for p, g in zip(predictions, gold_labels) if p == cls and g != cls)
fn = sum(1 for p, g in zip(predictions, gold_labels) if p != cls and g == cls)
precision = tp / (tp + fp) if (tp + fp) > 0 else 0
recall = tp / (tp + fn) if (tp + fn) > 0 else 0
f1 = 2 * precision * recall / (precision + recall) if (precision + recall) > 0 else 0
f1_scores.append(f1)
return sum(f1_scores) / len(f1_scores)
```
## Common pitfalls
- Data leakage occurs if the same question appears in multiple splits; the protocol requires each unique question to be assigned to exactly one split.
- Simple keyword matching or sentence-level cosine similarity fails to capture one-to-many or many-to-one semantic mappings between context sentences and answer sentences.
- Reporting accuracy instead of Macro F1, which masks class imbalance or uneven performance across hallucination types.
## Evidence (verbatim from paper)
> The best-performing model SIM-OVERLAP achieves a Macro F1 of only 71.1% on the unseen test set. This indicates that DELUCIONQA presents a challenging new task with substantial room for future improvement. From Table 4, we also notice that there are fluctuations in performance across the three splits (train/dev/test) of the dataset. Recall that while randomly dividing the data into multiple splits, we ensure that each unique question ends up in a single split (to avoid data leakage).
## Citation
```bibtex
@misc{sadat2023delucionqa,
title={DelucionQA: Detecting Hallucinations in Domain-specific Question Answering},
author={Sadat et al. (2023)},
year={2023},
note={arXiv:2312.05200}
}
```
- arXiv: 2312.05200
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!