Extracts structured text from scanned documents and images using Tesseract OCR with custom LSTM training data. Supports table detection via OpenCV contour analysis and PDF/A output generation.
Scanned 6/8/2026
Install via CLI
openskills install agentskillexchange/skills---
name: "Tesseract OCR Document Extractor"
slug: "tesseract-ocr-document-extractor"
description: "Extracts structured text from scanned documents and images using Tesseract OCR with custom LSTM training data. Supports table detection via OpenCV contour analysis and PDF/A output generation."
github_stars: 73614
verification: "security_reviewed"
source: "https://github.com/tesseract-ocr/tesseract"
author: "Tesseract OCR"
category: "Data Extraction & Transformation"
framework: "ChatGPT Agents"
tool_ecosystem:
github_repo: "tesseract-ocr/tesseract"
github_stars: 73614
---
# Tesseract OCR Document Extractor
Extracts structured text from scanned documents and images using Tesseract OCR with custom LSTM training data. Supports table detection via OpenCV contour analysis and PDF/A output generation.
## Installation
Requirements and caveats from upstream:
- **NOTE**: This software depends on other packages that may be licensed under different open source licenses.
Basic usage or getting-started notes:
- It also needs [traineddata](https://tesseract-ocr.github.io/tessdoc/Data-Files.html) files which support the legacy engine, for example those from the [tessdata](https://github.com/tesseract-ocr/tessdata) repository.
- Basic **[command line usage](https://tesseract-ocr.github.io/tessdoc/Command-Line-Usage.html)**:
- Examples can be found in the [documentation](https://tesseract-ocr.github.io/tessdoc/Command-Line-Usage.html#simplest-invocation-to-ocr-an-image).
- Source: https://github.com/tesseract-ocr/tesseract
- Extracted from upstream docs: https://raw.githubusercontent.com/tesseract-ocr/tesseract/HEAD/README.md
## Source
- [Agent Skill Exchange](https://agentskillexchange.com/skills/tesseract-ocr-document-extractor/)
No comments yet. Be the first to comment!
Ultra-compressed communication mode. Cuts token usage ~75% by speaking like caveman while keeping full technical accuracy. Supports intensity levels: lite, full (default), ultra, wenyan-lite, wenyan-full, wenyan-ultra. Use when user says "caveman mode", "talk like caveman", "use caveman", "less tokens", "be brief", or invokes /caveman. Also auto-triggers when token efficiency is requested.