Skip to content
Back to skills

Rag Evaluation Matrix

ASecurity

Use when comparing basic, enhanced, GraphRAG, or agentic RAG designs, evaluating domain-specific RAG quality, tuning retrieval components, or deciding whether agentic RAG is worth its cost.

  • 10 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added October 10, 2026
ai-agentspythonbash

Security analysis

A100/100

Pro scans all 3 files and shows the line behind each finding

Scanned October 10, 2026

npx -y skills add mouadja02/skills --skill rag-evaluation-matrix --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Rag Evaluation Matrix?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Rag Evaluation Matrix
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/mouadja02-rag-evaluation-matrix/badge)](https://www.skillsdirectory.com/skills/mouadja02-rag-evaluation-matrix)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: rag-evaluation-matrix
description: Use when comparing basic, enhanced, GraphRAG, or agentic RAG designs, evaluating domain-specific RAG quality, tuning retrieval components, or deciding whether agentic RAG is worth its cost.
source: "https://arxiv.org/abs/2601.07711"
version: "1.0.0"
---

# RAG Evaluation Matrix

Do not choose agentic RAG by fashion. Compare retrieval designs against domain questions, answerability, correctness, cost, latency, and failure reasons.

## Use When

- The user asks whether to use basic RAG, enhanced RAG, GraphRAG, or agentic RAG.
- A RAG system works on demos but fails on domain-specific questions.
- You need to compare embedding models, rerankers, chunking, query rewriting, or tool orchestration.
- LLM-as-judge scores need alignment with human review.

## Evaluation Flow

1. Build a domain question set with answerable, unanswerable, ambiguous, multi-hop, and adversarial cases.
2. For each candidate pipeline, log retrieved evidence, final answer, citations, latency, token usage, and cost.
3. Score answer correctness and answerability separately.
4. Add evidence quality checks: citation support, contradiction handling, source freshness, and missing-source diagnosis.
5. Segment results by question type instead of reporting only one aggregate score.
6. Inspect low-correctness failures and map them to retrieval, synthesis, tool orchestration, or corpus gaps.
7. Pick the simplest design that meets quality, latency, and cost constraints.

## Matrix

| Dimension | Basic RAG | Enhanced RAG | Agentic RAG |
| --- | --- | --- | --- |
| Best for | Stable FAQ and narrow corpora | Noisy corpora, query mismatch, reranking | Multi-step, ambiguous, tool-rich tasks |
| Main risk | Weak recall and unsupported answers | Pipeline complexity | Cost, latency, loops, tool misuse |
| Eval focus | Retrieval recall and citation support | Component ablations | Trajectory, action choice, stopping behavior |
| Ship gate | Correctness and answerability meet threshold | Ablation proves each module helps | Agentic gains justify extra cost |

## Script

Use the helper to combine per-run JSON metrics into a decision table:

```bash
python scripts/rag_eval_matrix.py results/*.json
```

Expected JSON fields: `pipeline`, `question_type`, `correct`, `answerable_correct`, `latency_ms`, `cost_usd`.

## Common Mistakes

| Mistake | Fix |
| --- | --- |
| Optimizing average score only | Break down by question type and domain |
| Judging unanswerable questions as wrong by default | Score answerability separately |
| Skipping human calibration | Sample judge disagreements and tune rubrics |
| Choosing agentic RAG without ablation | Compare against enhanced RAG at equal budget |
| Ignoring failure reasons | Classify each miss before tuning |

## References

- arXiv: Is Agentic RAG worth it? - https://arxiv.org/abs/2601.07711
- arXiv: RAGalyst - https://arxiv.org/abs/2511.04502
- Hugging Face Papers: RAGalyst - https://huggingface.co/papers/2511.04502
- Hugging Face dataset: RAGalyst QAC - https://huggingface.co/datasets/hoskerelab/ragalyst-qac

Files in this skill

  • SKILL.md3 KB
  • references/rag-metrics.md990 B
  • scripts/rag_eval_matrix.py1.6 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…