Skip to content
Back to skills

Pdf To Clean Text

ASecurity

Reads a PDF as clean text instead of page images, for a fraction of the tokens - search it, read chosen pages, or hand whole-document work (summaries, quizzes, translations) to a cheap reader agent. Use whenever the user shares a PDF and wants to know or do something with its content. Flags pages whose text can't be trusted. Not for editing, merging, splitting or filling PDFs.

  • 2 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added October 4, 2026
ai-agentspythonrustshell

Works with

  • claude code

Security analysis

A92/100
  • mediumInstalls packages at runtime which could introduce malicious dependencies

Pro scans all 4 files and shows the line behind each finding

Scanned October 4, 2026

npx -y skills add habib049/pdf-to-clean-text --skill pdf-to-clean-text --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Pdf To Clean Text?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Pdf To Clean Text
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/habib049-pdf-to-clean-text/badge)](https://www.skillsdirectory.com/skills/habib049-pdf-to-clean-text)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: pdf-to-clean-text
description: Reads a PDF as clean text instead of page images, for a fraction of the tokens - search it, read chosen pages, or hand whole-document work (summaries, quizzes, translations) to a cheap reader agent. Use whenever the user shares a PDF and wants to know or do something with its content. Flags pages whose text can't be trusted. Not for editing, merging, splitting or filling PDFs.
license: MIT
compatibility: Python 3.10+, shell access, and pip install -r requirements.txt. Built for Claude Code.
metadata:
  author: habib049
  version: "0.1.0"
---

# PDF to clean text

The script is `scripts/pdf_to_clean_text.py` in this skill's folder; call it by its absolute path. Every command
converts the PDF if needed (about a second per page, then cached by file content) and prints `info:` (pages,
size in tokens) and `warning:` lines on stderr. For errors, setup and limits, read [reference.md](reference.md).

**A few pages long:** read the whole conversion (`python SCRIPT FILE.pdf`) and skip the rest of this.

**A question about part of it:** `python SCRIPT FILE.pdf --find "term"` prints the sections containing the term.
It's literal, so use words the document uses (a heading, number, name). After two misses, or when the question
shares no words with the text, run `--map` (a short card of the document, or its headings with their pages),
then `--pages 12-14`.

**Whole-document work** (summary, quiz, study guide, a table of everything, translation, proofreading): call the
`pdf-reader` agent straight away. Don't convert, map or read the PDF first, and don't check its result yourself:
keeping all of that out of this conversation is the point. Give it exactly this prompt; it has its own procedure:

```
CMD: python <absolute path of scripts/pdf_to_clean_text.py>
PDF: <absolute path of the PDF>
TASK: <the user's request, word for word>
OUTPUT: <absolute path of the file to write>
```

The agent builds notes from the PDF the first time and reuses them later (they're cached by file content, so
"notes from an earlier run" means this same file), and it checks its own citations. Relay its report,
including any warnings in it. Open OUTPUT only if the user asks you to change it. No `pdf-reader` agent? Read
`--pages` in 10-page chunks yourself.

**Warnings,** for the pages you used:

- `scanned page, text came from OCR`: may be garbled. Say so when quoting it, and check names, amounts and dates
  against that page's image.
- `figure with no caption found`: its content isn't in the text. View that page's image if the answer needs it.
- `removed lines repeated across pages`: watermark-like lines were dropped. Look here if something seems missing.

**Cite pages** as the `<!-- page N -->` markers number them: the PDF's own numbering, which can differ from the
numbers printed on the pages.

Files in this skill

  • SKILL.md2.8 KB
  • reference.md2.2 KB
  • requirements.txt44 B
  • scripts/pdf_to_clean_text.py31.1 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…