Extracts client name and tax ID from PDF invoice files using Python, based on specific anchor text 'cliente' and 'N.º de contribuinte'.
Scanned 5/30/2026
Install via CLI
openskills install ECNU-ICALK/AutoSkill---
id: "1044a6bf-77ff-4c8b-954c-d464bb69640e"
name: "Extract Name and Tax ID from PDF Invoices"
description: "Extracts client name and tax ID from PDF invoice files using Python, based on specific anchor text 'cliente' and 'N.º de contribuinte'."
version: "0.1.0"
tags:
- "python"
- "pdf extraction"
- "invoice parsing"
- "regex"
- "data extraction"
triggers:
- "extract name and tax id from pdf invoices"
- "write a program to extract cliente and N.º de contribuinte from pdf"
- "parse pdf invoices for name and tax id"
- "extract data from 100 pdf invoices"
---
# Extract Name and Tax ID from PDF Invoices
Extracts client name and tax ID from PDF invoice files using Python, based on specific anchor text 'cliente' and 'N.º de contribuinte'.
## Prompt
# Role & Objective
You are a Python developer tasked with extracting specific data fields from PDF invoice files.
# Operational Rules & Constraints
1. Use a PDF parsing library (e.g., PyPDF2, PyMuPDF, or pdfminer) to extract text from the PDF files.
2. Implement batch processing to handle multiple files (e.g., 100 invoices).
3. Extract the **Name** by searching for the anchor string "cliente" and capturing the text immediately following it.
4. Extract the **Tax ID** by searching for the anchor string "N.º de contribuinte" and capturing the text immediately following it.
5. Use regular expressions or string manipulation to isolate the data.
6. Output the results clearly, indicating the file name, extracted name, and extracted tax ID.
# Anti-Patterns
- Do not hardcode specific file names; allow for iteration over a directory.
- Do not assume the exact position of the text; rely on the anchor strings.
## Triggers
- extract name and tax id from pdf invoices
- write a program to extract cliente and N.º de contribuinte from pdf
- parse pdf invoices for name and tax id
- extract data from 100 pdf invoices
No comments yet. Be the first to comment!
1. **Strip thinking before verifying** — a verifier that sees the reasoning is biased toward agreement. Fresh context, cleaned proof only. 2. **"Does this prove RH?"** — if your theorem's specialization to ζ is a famous open problem, you have a gap. Most reliable red flag. 3. **Short proof → extract the general lemma** — try 2×2 counterexamples. If general form is false, find what's special about THIS instance. 4. **Same gap twice → step back** — the case split may be obscuring a unifie
Monitor Catchtable for open reservation slots and attempt booking using a logged-in Chrome session.
Interactive walkthrough for new users. Learn by doing — each step creates real content in your vault. Three tracks (researcher, manager, personal) with a universal learning arc. Triggers on "/tutorial", "walk me through", "how do I use this".
当用户明确要求"填充示例内容""生成示例""补充 LaTeX 示例"时使用。AI 增强版 LaTeX 示例智能生成器,实现 AI 与硬编码的有机融合:AI 做"语义理解"(分析章节主题、推理资源相关性、生成连贯叙述),硬编码做"结构保护"(格式验证、哈希校验、访问控制)。
当用户明确要求"写/润色 NSFC 标书摘要""生成中文摘要和英文摘要""把中文摘要翻译成英文摘要"时使用。输出中文、英文两个版本(英文必须是中文的忠实翻译版),同时输出标题建议(1个推荐标题+5个候选标题及理由)。中文摘要默认≤400字符,英文摘要默认≤4000字符。输出方式:将结果写入工作目录下的 `NSFC-ABSTRACTS.md`。⚠️ 不适用:用户只想翻译一段与标书无关的通用文本(应直接翻译);用户只想写立项依据/研究内容/研究基础正文(应使用对应 nsfc 系列 skill)。