Skip to content
Back to skills

003 Name Skill Dfbf66de

ASecurity

Extract text from PDFs for LLM consumption. Use when processing PDFs for RAG, document analysis, or text extraction. Supports API services (Mistral OCR) and local tools (PyMuPDF, pdfplumber). Handles text-based PDFs, tables, and scanned documents with OCR.

  • 9 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added October 11, 2026
toolsbashapi

Works with

  • api

Security analysis

A92/100
  • mediumInstalls packages at runtime which could introduce malicious dependencies

Pro shows the line behind each finding and how to fix it

Scanned October 11, 2026

npx -y skills add tools-only/X-Skills --skill 003-name-skill_dfbf66de --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of 003 Name Skill Dfbf66de?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for 003 Name Skill Dfbf66de
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/tools-only-003-name-skill-dfbf66de/badge)](https://www.skillsdirectory.com/skills/tools-only-003-name-skill-dfbf66de)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

SKILL.md
---
name: extracting-pdf-text
description: Extract text from PDFs for LLM consumption. Use when processing PDFs for RAG, document analysis, or text extraction. Supports API services (Mistral OCR) and local tools (PyMuPDF, pdfplumber). Handles text-based PDFs, tables, and scanned documents with OCR.
---

# Extracting PDF Text for LLMs

This skill provides tools and guidance for extracting text from PDFs in formats suitable for language model consumption.

## Quick Decision Guide

| PDF Type | Best Approach | Script |
|----------|--------------|--------|
| Simple text PDF | PyMuPDF | `scripts/extract_pymupdf.py` |
| PDF with tables | pdfplumber | `scripts/extract_pdfplumber.py` |
| Scanned/image PDF (local) | pytesseract | `scripts/extract_with_ocr.py` |
| Complex layout, highest accuracy | Mistral OCR API | `scripts/extract_mistral_ocr.py` |
| End-to-end RAG pipeline | marker-pdf | `pip install marker-pdf` |

## Recommended Workflow

1. **Try PyMuPDF first** - fastest, handles most text-based PDFs well
2. **If tables are mangled** - switch to pdfplumber
3. **If scanned/image-based** - use Mistral OCR API (best accuracy) or local OCR (free but slower)

## Local Extraction (No API Required)

### PyMuPDF - Fast General Extraction

Best for: Text-heavy PDFs, speed-critical workflows, basic structure preservation.

```bash
uv run scripts/extract_pymupdf.py input.pdf output.md
```

The script outputs markdown with preserved headings and paragraphs. For LLM-optimized output, it uses `pymupdf4llm` which formats text for RAG systems.

### pdfplumber - Table Extraction

Best for: PDFs with tables, financial documents, structured data.

```bash
uv run scripts/extract_pdfplumber.py input.pdf output.md
```

Tables are converted to markdown format. Note: pdfplumber works best on machine-generated PDFs, not scanned documents.

### Local OCR - Scanned Documents

Best for: Scanned PDFs when API access is unavailable.

```bash
uv run scripts/extract_with_ocr.py input.pdf output.txt
```

Requires: `pytesseract`, `pdf2image`, and Tesseract installed (`brew install tesseract` on macOS).

## API-Based Extraction

### Mistral OCR API

Best for: Complex layouts, scanned documents, highest accuracy, multilingual content, math formulas.

**Pricing**: ~1000 pages per dollar (very cost-effective)

```bash
export MISTRAL_API_KEY="your-key"
uv run scripts/extract_mistral_ocr.py input.pdf output.md
```

Features:
- Outputs clean markdown
- Preserves document structure (headings, lists, tables)
- Handles images, math equations, multilingual text
- 95%+ accuracy on complex documents

For detailed API options and other services, see [references/api-services.md](references/api-services.md).

## Output Format Recommendations

For LLM consumption, markdown is preferred:
- Preserves semantic structure (headings become context boundaries)
- Tables remain readable
- Compatible with most RAG chunking strategies

For detailed comparisons of local tools, see [references/local-tools.md](references/local-tools.md).

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…