Documind is an open-source Node.js tool that uses AI to extract structured JSON data from PDFs and other documents. Define a custom schema for what you need, and Documind returns clean, typed data — supporting OpenAI and local LLM backends like Llama 3.2 Vision.
Scanned 6/8/2026
Install via CLI
openskills install agentskillexchange/skills---
name: "Documind AI-Powered Structured Data Extraction from Documents"
slug: "documind-ai-structured-data-extraction"
description: "Documind is an open-source Node.js tool that uses AI to extract structured JSON data from PDFs and other documents. Define a custom schema for what you need, and Documind returns clean, typed data — supporting OpenAI and local LLM backends like Llama 3.2 Vision."
github_stars: 1468
verification: "security_reviewed"
source: "https://github.com/DocumindHQ/documind"
category: "Data Extraction & Transformation"
framework: "Custom Agents"
tool_ecosystem:
github_repo: "DocumindHQ/documind"
github_stars: 1468
npm_package: "documind"
npm_weekly_downloads: 14
---
# Documind AI-Powered Structured Data Extraction from Documents
Documind is an open-source Node.js tool that uses AI to extract structured JSON data from PDFs and other documents. Define a custom schema for what you need, and Documind returns clean, typed data — supporting OpenAI and local LLM backends like Llama 3.2 Vision.
## Installation
Use the upstream install or setup path that matches your environment:
- brew install ghostscript graphicsmagick
- npm install documind
Requirements and caveats from upstream:
- ### **Node.js & NPM**
- Ensure Node.js (v18+) and NPM are installed on your system.
- **documind** requires an **.env** file to store sensitive information like your OpenAI API key.
Basic usage or getting-started notes:
- bash
- # On macOS
- # On Debian/Ubuntu
- Source: https://github.com/DocumindHQ/documind
- Extracted from upstream docs: https://raw.githubusercontent.com/DocumindHQ/documind/HEAD/README.md
## Source
- [Agent Skill Exchange](https://agentskillexchange.com/skills/documind-ai-structured-data-extraction/)
No comments yet. Be the first to comment!
Ultra-compressed communication mode. Cuts token usage ~75% by speaking like caveman while keeping full technical accuracy. Supports intensity levels: lite, full (default), ultra, wenyan-lite, wenyan-full, wenyan-ultra. Use when user says "caveman mode", "talk like caveman", "use caveman", "less tokens", "be brief", or invokes /caveman. Also auto-triggers when token efficiency is requested.