Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Web Scraper

ASecurity

Ethical web scraping — selectors, pagination, robots.txt, and rate limits — use when extracting data from websites responsibly.

2 stars
0 votes
0 copies
0 views
Added 9/29/2026
ai-agentsjavascriptgojavagitapi

Works with

cursorapi

Security Analysis

A100/100

Scanned 9/29/2026

$npx -y skills add aicodedecode/awesome-muse-skills --skill web-scraper --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Web Scraper?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Web Scraper
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/aicodedecode-web-scraper/badge)](https://www.skillsdirectory.com/skills/aicodedecode-web-scraper)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: web-scraper
description: Ethical web scraping — selectors, pagination, robots.txt, and rate limits — use when extracting data from websites responsibly.
category: web-data
---

## Overview

Web scraping extracts data from websites programmatically — for research,
monitoring, aggregation, or migration. Done responsibly, it's a legitimate
technique; done carelessly, it's a Terms-of-Service violation or a denial-
of-service attack. This skill covers the technical craft and the ethical/
legal boundaries together.

## When to use

- Extracting structured data from websites (prices, listings, articles)
- Choosing between HTTP scraping and headless browsers
- Handling pagination, infinite scroll, and JavaScript-rendered content
- Respecting robots.txt, rate limits, and terms of service
- Building maintainable scrapers that survive site changes

## Core concepts

**Check permission first.** Read `robots.txt` and the site's Terms of
Service before scraping. Prefer official APIs when they exist — they're
faster, stabler, and explicitly permitted. Some sites prohibit scraping
entirely; "it's technically possible" is not permission. When in doubt,
ask the site owner.

**Be a polite guest.** Rate-limit requests (1–2/sec is a common courtesy
baseline, slower for small sites), identify yourself with a real User-Agent
including contact info, scrape during off-peak hours for heavy jobs, and
cache aggressively to avoid re-fetching. Your scraper should be
indistinguishable from light human traffic.

**HTTP first, browser when needed.** Static/server-rendered pages: plain
HTTP requests + HTML parsing (fast, cheap, polite). JavaScript-rendered
content: headless browser (Playwright/Puppeteer — see those skills), which
is 10–100x heavier — use only when necessary, and prefer waiting for
specific elements over arbitrary sleeps.

**Selectors: resilient over clever.** Prefer semantic anchors (data
attributes, ARIA roles, stable class names) over brittle positional XPaths
(`div[3]/span[2]` breaks on the next redesign). Build a validation layer
that fails loudly when expected elements vanish — silent schema drift is
how scrapers produce garbage for weeks.

**Pagination and state.** Handle numbered pages, "load more" buttons, and
infinite scroll (scroll → wait → extract → repeat until no new content).
Track progress (persist visited URLs/page cursors) so interrupted runs
resume instead of restarting — re-scraping from zero hammers the site.

## Practical workflow

1. **Assess:** API available? robots.txt allows? ToS permits? If any answer
   is no, stop or get permission — document the decision.
2. **Prototype the extraction** on a few pages: fetch, parse, map fields;
   verify against the rendered page (what you see is what you should get).
3. **Build resilient selectors** anchored on stable attributes; add
   assertions per field (present? right type? sane range?) that fail the
   run loudly on site changes.
4. **Implement politeness:** rate limiting with jitter, retries with
   exponential backoff on 429/5xx (and back off harder when asked),
   request caching, and off-peak scheduling for large crawls.
5. **Handle sessions properly:** respect login requirements (use your own
   credentials, never circumvent access controls), rotate nothing
   deceptive, and never scrape personal data beyond what's clearly public
   and permitted.
6. **Monitor and maintain:** alert on extraction failures and schema
   changes; version your parsers; keep a changelog of site-structure
   adaptations.

## Common pitfalls

- **Ignoring robots.txt/ToS** — the fastest path to IP bans, legal letters,
  and being blocked by the entire industry's abuse systems.
- **Hammering the site** — concurrent requests at full speed is a DoS;
  rate-limit always, back off on 429s.
- **Brittle selectors** — positional XPaths that break weekly; anchor on
  semantics and validate output.
- **Scraping personal data** — names, emails, photos of private individuals
  carry privacy obligations (GDPR and similar); minimize collection and
  have a lawful basis.
- **Circumventing access controls** — bypassing logins, paywalls, or
  anti-bot measures crosses from scraping into unauthorized access; don't.
- **No change detection** — the site redesigns, your scraper silently
  extracts wrong data for a month; validate every run.

Attribution

aicodedecodeaicodedecode
View sourceSee grades on GitHubMore from aicodedecode →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Terse caveman voice: answer first, fluff gone, every technical fact kept. Use for /caveman, "caveman mode", "talk like caveman", "be brief", "less tokens". Stays on until "stop caveman" or "normal mode".

1100021 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

698621 votes

Writing Skills

Create and manage Claude Code skills in HASH repository following Anthropic best practices. Use when creating new skills, modifying skill-rules.json, understanding trigger patterns, working with hooks, debugging skill activation, or implementing progressive disclosure. Covers skill structure, YAML frontmatter, trigger types (keywords, intent patterns), UserPromptSubmit hook, and the 500-line rule. Includes validation and debugging with SKILL_DEBUG. Examples include rust-error-stack, cargo-dep...

3931 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3421 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Amp, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Grok Build, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

741 votes
View all in ai-agents →