Build grounded question answering and retrieval-augmented generation (RAG) over your own collection of research papers, with answers that cite the exact paper and passage. Use when a user wants to "chat with" or search a folder of PDFs, synthesize evidence across a literature corpus, find which paper says X, or build a vector/hybrid index with SQLite FTS5, pgvector (Postgres), Chroma, Qdrant or FAISS. Includes a zero-dependency local full-text index with citable hits.
7 stars
0 votes
0 copies
0 views
Added October 4, 2026
researchpythonrustgobashsqlapidatabase
Works with
cli
api
Security analysis
A92/100
mediumInstalls packages at runtime which could introduce malicious dependencies
Installs into .claude/skills of the current project.
Are you the author of Paper Corpus Rag?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/kalarislabs-paper-corpus-rag)
---
name: paper-corpus-rag
description: Build grounded question answering and retrieval-augmented generation (RAG) over your own collection of research papers, with answers that cite the exact paper and passage. Use when a user wants to "chat with" or search a folder of PDFs, synthesize evidence across a literature corpus, find which paper says X, or build a vector/hybrid index with SQLite FTS5, pgvector (Postgres), Chroma, Qdrant or FAISS. Includes a zero-dependency local full-text index with citable hits.
license: MIT
compatibility: Python 3.9+ with SQLite FTS5 (standard in CPython builds). Optional pypdf, sentence-transformers, psycopg/pgvector, chromadb or qdrant-client for the vector tiers.
metadata:
version: "1.0"
category: knowledge-and-rag
maintainer: Kalaris Labs
author: Kalaris Labs (Sayan Chowdhury)
tags: RAG, retrieval, vector database, pgvector, Chroma, Qdrant, FAISS, SQLite, full-text search, literature synthesis
---
# RAG over a Research Paper Corpus
The goal is answers a researcher can check: every claim is traceable to a
specific passage of a specific paper. Retrieval quality and citation discipline
matter more than the choice of vector database.
To discover *new* candidate papers before building your local collection, use
`firecrawl-research-index` for hosted paper search and passage reads. Save only
papers you have a lawful copy of, with their DOI or source ID. This skill indexes
the user's own files; Firecrawl's hosted index is not a local corpus download.
## Tier 0: citable full-text search in one minute (no dependencies)
1. Convert PDFs to Markdown first (layout-aware converters beat raw PDF text): use the
`markitdown` or `liteparse` skill, or `pip install pypdf` for the built-in fallback.
2. Index and query:
```bash
python scripts/paper_index.py index papers_md/ --db corpus.sqlite # incremental: re-run after adding papers
python scripts/paper_index.py query "effect of batch size on calibration" --db corpus.sqlite -k 8
python scripts/paper_index.py query '"label smoothing" calibration' --db corpus.sqlite --json --full
```
Each hit is `file#chunk` + section heading + BM25 score. Quote phrases for exact matches.
BM25 is strong for scientific text full of exact terms (gene names, method names, datasets),
so start here before adding embeddings.
## Tier 1: semantic and hybrid retrieval
Add dense embeddings when queries are conceptual ("methods that avoid labeled data").
Use hybrid scoring: BM25 and vector ranks fused with Reciprocal Rank Fusion (RRF),
`score = Σ 1/(60 + rank)`. It is robust and needs no tuning.
Embedding model choice: a scientific-domain or strong general model from
`sentence-transformers` (see that skill). Embed **chunks**, not whole papers. Store the
model name and dimension with the index, because changing models means re-embedding.
**pgvector (Postgres).** Best when you already run Postgres or need SQL filters (year, venue, author):
```sql
CREATE EXTENSION IF NOT EXISTS vector;
CREATE TABLE chunks (id bigserial PRIMARY KEY, paper_id text, chunk int, section text,
body text, tsv tsvector GENERATED ALWAYS AS (to_tsvector('english', body)) STORED,
embedding vector(768));
CREATE INDEX ON chunks USING hnsw (embedding vector_cosine_ops);
CREATE INDEX ON chunks USING gin (tsv);
-- hybrid: fetch top-k from each, fuse with RRF in the application
SELECT id FROM chunks ORDER BY embedding <=> $1 LIMIT 50;
SELECT id FROM chunks WHERE tsv @@ websearch_to_tsquery('english', $2)
ORDER BY ts_rank(tsv, websearch_to_tsquery('english', $2)) DESC LIMIT 50;
```
**Chroma / Qdrant / FAISS.** Local-first prototyping (Chroma), production filtering and
payloads (Qdrant), or in-memory at scale (FAISS). See the `chroma`, `qdrant-vector-search`
and `faiss` skills for APIs. Keep `paper_id`, `chunk`, `section`, `year` as metadata on every vector.
## Chunking rules for papers
- 800-1500 characters with 10-20% overlap. Split on paragraph boundaries, never mid-sentence.
- Keep the section heading with each chunk ("Methods > Training details"), since it disambiguates.
- Index figure/table captions as their own chunks. Drop reference lists from the main index
(they pollute BM25), but keep them in a separate table for citation lookups.
- Store bibliographic metadata (DOI, title, authors, year) per paper, looked up by DOI, not guessed.
## Answering protocol (grounding)
1. Retrieve 8-20 chunks. Re-rank if a cross-encoder is available.
2. Answer **only** from retrieved text. Cite every factual sentence as `[file#chunk]`
(or author-year once mapped to the bibliography).
3. If the corpus does not contain the answer, say so. Do not fill gaps from memory.
4. When papers disagree, present both with citations rather than picking one silently.
5. For quantitative claims, quote the number verbatim with its context (dataset, metric, split).
## Integrity rules
- Never fabricate a citation handle or quote. Every cited `file#chunk` must exist in the index output.
- Verify numbers against the source passage before reporting them.
## Evaluation
Before trusting a pipeline, write 20-50 question→expected-passage pairs from the corpus
and measure recall@k of retrieval. Most "hallucination" in paper QA is actually a
retrieval miss. Re-check after changing chunking, the embedding model or the corpus.
## Related skills
`research-knowledge-graph` (citation/concept graphs of the same corpus),
`markitdown`, `liteparse`, `sentence-transformers`, `chroma`, `qdrant-vector-search`,
`faiss`, `open-notebook`, `literature-review`, `firecrawl-research-index`.