Reconstruct editable LaTeX from PDF content using page-aware extraction and visual comparison. Preserve source evidence, flag uncertain math/tables/citations, and distinguish text extraction, OCR, reconstruction, and verified compilation.
Pro shows the line behind each finding and how to fix it
Scanned 10/3/2026
npx -y skills add gabrielmoreira/agent-skills-mirror --skill pdf2tex --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Pdf2tex?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/gabrielmoreira-pdf2tex)More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.
---
name: pdf2tex
description: Reconstruct editable LaTeX from PDF content using page-aware extraction and visual comparison. Preserve source evidence, flag uncertain math/tables/citations, and distinguish text extraction, OCR, reconstruction, and verified compilation.
metadata:
version: "1.3.0"
---
## Establish the reconstruction target
Inspect the PDF, requested pages, available tools, and desired output. Determine
whether the user wants content recovery or close visual reconstruction. Preserve
the original PDF and write new artifacts to a separate destination.
A PDF may expose text, font names, coordinates, images, and metadata. It does
not reliably encode its original document class, packages, macros, bibliography
database, comments, or source-file boundaries. Font/creator metadata is evidence
for a candidate setup, not proof of the original engine or class.
## Extract evidence
When PyMuPDF is available, use the bundled helper from this skill's own directory.
The following paths are relative to the repository root; for an installed skill,
substitute its actual location. Dependency installation is separate from extraction.
```sh
python -m pip install -r pdf2tex/requirements.txt
python pdf2tex/scripts/extract_pdf.py paper.pdf --output extraction --pages 1-3,5 --images --render
```
The helper creates a new directory with `report.html`, `text.txt`, `layout.json`, and optional
embedded images and whole-page PNG previews when requested. It records page numbers, raw text spans/font/position data,
metadata, selected-page coverage, and warnings. It refuses existing output
directories and refuses publication if the input fingerprint changes during
extraction. Open `report.html` for offline page/text review; keep the entire
directory together when sharing. It performs no OCR or conversion.
Read [PDF extraction guide](references/pdf-extraction-guide.md) for API details,
alternative readers, columns, fonts, and OCR. Sorted text is not guaranteed
reading order; inspect page layouts and use coordinates. Images can be repeated
or carry separate soft masks. Vector figures and composite panels often need
a page crop or another export workflow. Use optional `--render` previews to
inspect selected pages, including vector/composite figures; these are visual
evidence, not OCR or segmented assets. `--dpi` accepts 72–300 with a per-page pixel limit.
Use optional `--chars` when inspecting scripts or small notation. It adds
character origins/bounding boxes while retaining span text. Page geometry and
rotation matrices help relate unrotated text coordinates to rendered previews;
positions are evidence for candidate readings, not an automatic math parser.
A page without text may be blank, graphical, or scanned. Check it visually before
choosing OCR. OCR requires separate tools and cannot establish the correctness
of equations or tables. Retain page provenance and flag OCR-derived uncertainty.
For a password-protected PDF, use an authorized readable copy.
## Reconstruct without inventing content
Use [structure detection](references/structure-detection.md) to interpret blocks,
[math reconstruction](references/math-reconstruction.md) for notation, and
[table reconstruction](references/table-reconstruction.md) for cells and merged
regions. These heuristics need comparison with the rendered original.
- Select an available class and engine suitable for the target; state inferred
choices. Use a supplied official author kit when exact publication layout is required.
- Preserve selected-page coverage, section order, prose, equations, table values,
captions, footnotes, and references. Escape LaTeX-special characters in prose
without indiscriminately escaping math or generated commands.
- Associate citation markers with bibliography entries only when the mapping
is supported. Keep unmatched markers and uncertainty visible; do not invent
bibliographic metadata or silently assign the nearest reference.
- Preserve ambiguous glyphs, merged table cells, missing images, and illegible
content as source evidence with `% [UNCERTAIN: ...]` or a visible placeholder.
A comment alone must not hide missing content from the generated document.
- Remove headers/footers or join hyphenated lines only after checking that they
are layout artifacts. Preserve meaningful hyphens and repeated scientific text.
- Do not guess original macros or file splitting. A self-contained source is a
useful default, not a claim that it matches the original organization.
## Build, compare, and deliver
Use the selected engine and actual bibliography backend, with additional passes
for cross-references. `latex-rescue` can help when available. If tools or assets
are unavailable, preserve the source and report compilation as unverified.
Compare the rendered reconstruction with the selected original pages: completeness,
reading order, math symbols, tables, figure appearances, captions, and citations.
Matching page counts does not establish fidelity. Check merged cells and OCR
math manually, and distinguish a visual approximation from content verification.
Deliver the new source/assets, input version and selected pages, extraction and
OCR methods actually used, inferred class/engine, build and visual-check results,
and uncertainty locations. Separate recovered content from placeholders. Do not
promise exact original source, perfect reconstruction, or immediate compilation.
Use `latex-polish` or `latex-fmt` only for a further requested editing task.
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!