Extracts structured data from scanned documents using Tesseract OCR engine with LSTM models. Supports table detection via OpenCV contour analysis and outputs to CSV, JSON, or Pandas DataFrames.
Scanned 6/8/2026
Install via CLI
openskills install agentskillexchange/skills---
name: "Tesseract OCR Data Extractor"
slug: "tesseract-ocr-data-extractor"
description: "Extracts structured data from scanned documents using Tesseract OCR engine with LSTM models. Supports table detection via OpenCV contour analysis and outputs to CSV, JSON, or Pandas DataFrames."
github_stars: 73629
verification: "security_reviewed"
source: "https://github.com/tesseract-ocr/tesseract"
author: "Tesseract OCR"
category: "Data Extraction & Transformation"
framework: "Gemini"
tool_ecosystem:
github_repo: "tesseract-ocr/tesseract"
github_stars: 73629
---
# Tesseract OCR Data Extractor
Extracts structured data from scanned documents using Tesseract OCR engine with LSTM models. Supports table detection via OpenCV contour analysis and outputs to CSV, JSON, or Pandas DataFrames.
## Prerequisites
Tesseract OCR, OpenCV
## Installation
Requirements and caveats from upstream:
- **NOTE**: This software depends on other packages that may be licensed under different open source licenses.
Basic usage or getting-started notes:
- It also needs [traineddata](https://tesseract-ocr.github.io/tessdoc/Data-Files.html) files which support the legacy engine, for example those from the [tessdata](https://github.com/tesseract-ocr/tessdata) repository.
- Basic **[command line usage](https://tesseract-ocr.github.io/tessdoc/Command-Line-Usage.html)**:
- Examples can be found in the [documentation](https://tesseract-ocr.github.io/tessdoc/Command-Line-Usage.html#simplest-invocation-to-ocr-an-image).
- Source: https://github.com/tesseract-ocr/tesseract
- Extracted from upstream docs: https://raw.githubusercontent.com/tesseract-ocr/tesseract/HEAD/README.md
## Documentation
- https://tesseract-ocr.github.io/tessdoc/
## Source
- [Agent Skill Exchange](https://agentskillexchange.com/skills/tesseract-ocr-data-extractor/)
No comments yet. Be the first to comment!
Ultra-compressed communication mode. Cuts token usage ~75% by speaking like caveman while keeping full technical accuracy. Supports intensity levels: lite, full (default), ultra, wenyan-lite, wenyan-full, wenyan-ultra. Use when user says "caveman mode", "talk like caveman", "use caveman", "less tokens", "be brief", or invokes /caveman. Also auto-triggers when token efficiency is requested.