Skip to content
Back to skills

Document Parser

ASecurity

Building robust document ingestion — chunking, metadata, and multi-format pipelines for search and RAG.

  • 2 stars
  • 0 votes
  • 0 copies
  • 3 views
  • Added September 29, 2026
ai-agentsgo

Security analysis

A100/100

Scanned September 29, 2026

npx -y skills add aicodedecode/awesome-muse-skills --skill document-parser --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Document Parser?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Document Parser
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/aicodedecode-document-parser/badge)](https://www.skillsdirectory.com/skills/aicodedecode-document-parser)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: document-parser
description: Building robust document ingestion — chunking, metadata, and multi-format pipelines for search and RAG.
category: document-processing
---

## Overview

Modern applications ingest documents for search, retrieval-augmented generation
(RAG), and analytics. A document parser pipeline turns heterogeneous files
(PDF, DOCX, HTML, spreadsheets, images) into clean, chunked, metadata-rich
units ready for indexing. This skill covers the architecture and the details
that determine retrieval quality.

## When to use

- Building RAG ingestion over PDFs, Office docs, and web pages
- Designing chunking strategies (size, overlap, semantic boundaries)
- Extracting metadata (titles, dates, authors, page numbers, source URLs)
- Handling tables, figures, and scanned content in parsing
- Evaluating and improving retrieval quality end-to-end

## Core concepts

**Parse → clean → chunk → enrich.** The pipeline stages: extract raw text with
structure (headings, tables, lists); clean (fix hyphenation, normalize
whitespace, drop boilerplate like repeated headers/footers); chunk into
retrieval units; enrich each chunk with metadata (source, section, page,
date). Skipping cleaning poisons everything downstream.

**Chunking strategy drives retrieval quality.** Fixed-size chunks with overlap
are the baseline (e.g. ~500 tokens, 10–20% overlap). Better: respect document
structure — chunk by section, keep tables intact, never split a table across
chunks. A chunk should be self-contained enough to make sense out of context,
which is why prepending section headings to each chunk helps enormously.

**Metadata is a retrieval multiplier.** Filters on metadata (date ranges,
document type, source) often beat pure semantic search for precision. Capture
rich metadata at ingest: title, section path, page numbers, author, date,
URL, document version. Store it alongside embeddings, not just in them.

**Tables need special handling.** Serialize tables as Markdown or HTML within
the chunk, keep them whole, and consider generating a textual summary of key
figures ("Q3 revenue was $4.2M, up 12%") as an additional chunk — dense numeric
tables embed poorly on their own.

**Format routing.** Detect file type and route to the right extractor
(native PDF text vs OCR for scans, structure-aware DOCX parsing, readability
extraction for HTML). One generic path for all formats produces mediocre
results for every format.

## Practical workflow

1. **Inventory your corpus:** formats, sizes, languages, scan-vs-native ratio,
   update frequency — the pipeline design follows from this.
2. **Build per-format extractors** with structure preservation (headings,
   tables, lists); add OCR routing for image-only PDFs.
3. **Implement cleaning:** de-hyphenate, normalize whitespace/unicode, strip
   running headers/footers and page numbers, detect and drop boilerplate.
4. **Chunk by structure** with fallback to sized chunks; prepend breadcrumb
   context (document title › section) to each chunk; keep tables atomic.
5. **Attach metadata** at ingest and expose it as filters in retrieval;
   version chunks so re-ingestion is idempotent (hash content, upsert by ID).
6. **Evaluate retrieval, not just parsing:** build a question set with known
   answers, measure hit rate / recall@k, and iterate on chunking and metadata —
   parser quality is proven by retrieval quality.

## Common pitfalls

- **Chunking by fixed size across table boundaries** — half a table in each of
  two chunks retrieves as nonsense; keep structural units atomic.
- **No dedup / idempotent re-ingestion** — re-running the pipeline duplicates
  chunks and corrupts results; content-hash IDs solve this.
- **Dropping page numbers and section context** — users can't verify answers
  without citations; metadata is a feature, not overhead.
- **Indexing boilerplate** (headers, footers, legal disclaimers on every page)
  — pollutes retrieval with repeated near-identical chunks; strip it.
- **One chunk size for all content** — dense technical prose and sparse slides
  need different strategies; tune per document type.
- **Evaluating the embedder while the parser is broken** — if chunks are dirty,
  no embedding model saves you; fix the pipeline bottom-up.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…