Skip to content
Back to skills

Agrag Llm Judge

ASecurity

Run headless AgRAG prompts in a shared thread and produce an LLM-as-judge evaluation report. Use for evaluating retrieval quality, grounding, graph paths, and tool usage of the agrag agent in headless mode, especially when you need a consistent context window via --thread-id.

  • 2 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 27, 2026
toolspythonbash

Security analysis

A100/100

Pro scans all 4 files and shows the line behind each finding

Scanned September 27, 2026

npx -y skills add David-Li0406/meta-skill-evloving --skill agrag-llm-judge --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Agrag Llm Judge?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Agrag Llm Judge
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/david-li0406-agrag-llm-judge/badge)](https://www.skillsdirectory.com/skills/david-li0406-agrag-llm-judge)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: agrag-llm-judge
description: Run headless AgRAG prompts in a shared thread and produce an LLM-as-judge evaluation report. Use for evaluating retrieval quality, grounding, graph paths, and tool usage of the agrag agent in headless mode, especially when you need a consistent context window via --thread-id.
---

# AgRAG LLM Judge

## Overview
Run a short evaluation session against the AgRAG agent in headless mode, capture a transcript, then judge the responses with a structured rubric.

## Workflow

### 1) Choose session inputs
- Pick a thread id and 5-10 prompts.
- Use `references/query-templates.md` for prompts.
- Include at least one negative query.
- Do not run the synthetic generator; use the existing DB state.

### 2) Run headless prompts in a shared thread
Use a single thread id for all prompts so the context stays consistent.

Single prompt:
```bash
poetry run agrag -p "<prompt>" --thread-id <id> --output-format json
```

Batch helper:
```bash
./.codex/skills/agrag-llm-judge/scripts/run_headless_batch.sh <id> prompts.txt out.jsonl
```

### 3) Optional tool audit
Re-run a subset with stream-json to inspect tool calls and results:
```bash
poetry run agrag -p "<prompt>" --thread-id <id> --output-format stream-json
```
Look for missing `tool_result` events or repeated tool calls.

### 4) Detect issues + propose fixes
Run the diagnostics script to identify issues and recommended fixes:
```bash
python .codex/skills/agrag-llm-judge/scripts/diagnose_headless_output.py <output_file>
```
If issues are detected, summarize them and propose fixes, then **wait for user confirmation** before applying any changes.

### 5) Build the transcript
Create a simple Q/A transcript from the outputs:
- User prompt
- Assistant response
- Any cited evidence snippets
- Any graph paths

### 6) LLM-as-judge evaluation
Score each response and summarize issues with fixes.

## Rubric (1-10 each)
- Accuracy: statements match retrieved evidence.
- Relevance: answers the asked entity type and question.
- Groundedness: evidence cites entity IDs; avoid tool-only citations unless no IDs exist.
- Graph path validity: only uses relationship labels returned by `graph_traverse`.
- Tool efficiency: minimal calls, no redundant searches.

## Common failure modes + fixes
- Wrong entity type in results: use "Ranked Results" with entity type tags.
- Invented relationship labels: ensure `graph_traverse` outputs relationship types and only cite those.
- Evidence cites tool names not IDs: format tool outputs with explicit `Entity ID` lines and cite them.

## Resources
- `references/query-templates.md`
- `scripts/run_headless_batch.sh`
- `scripts/diagnose_headless_output.py`

Files in this skill

  • SKILL.md2.6 KB
  • references/query-templates.md1 KB
  • scripts/diagnose_headless_output.py9.6 KB
  • scripts/run_headless_batch.sh817 B

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…