Skip to content
Back to skills

Counterclaw Core

BSecurity

Defensive interceptor for prompt injection and basic PII masking.

  • 14 stars
  • 0 votes
  • 0 copies
  • 2 views
  • Added September 7, 2026
developmentpythonrustgobashgitapisecurity

Works with

  • cli
  • api

Security analysis

B80/100
  • criticalContains 'ignore previous instructions' pattern β€” found in 91% of malicious skills (Snyk ToxicSkills)
  • mediumInstalls packages at runtime which could introduce malicious dependencies

Pro scans all 11 files and shows the line behind each finding

Scanned September 7, 2026

npx -y skills add modbender/skill-library-mcp --skill counterclaw-core --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Counterclaw Core?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Counterclaw Core
[![Security: B β€” Skills Directory](https://www.skillsdirectory.com/api/skills/modbender-counterclaw-core/badge)](https://www.skillsdirectory.com/skills/modbender-counterclaw-core)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: counterclaw
description: Defensive interceptor for prompt injection and basic PII masking.
homepage: https://github.com/nickconstantinou/counterclaw-core
install: "pip install ."
requirements:
  env:
    - TRUSTED_ADMIN_IDS
  files:
    - "~/.openclaw/memory/"
    - "~/.openclaw/memory/MEMORY.md"
metadata:
  clawdbot:
    emoji: "πŸ›‘οΈ"
    version: "1.1.0"
    category: "Security"
    type: "python-middleware"
    security_manifest:
      network_access: "optional (only when using email integration scripts)"
      filesystem_access: "Write-only logging to ~/.openclaw/memory/"
      purpose: "Log security violations locally for user audit."
---

# CounterClaw 🦞

> Defensive security for AI agents. Snaps shut on malicious payloads.

## ⚠️ Security Notice

This package has two modes:

1. **Core Scanner (offline):** `check_input()` and `check_output()` β€” no network calls
2. **Email Integration (network):** `send_protected_email.sh` β€” requires gog CLI for Gmail

## Installation

```bash
claw install counterclaw
```

## Quick Start

```python
from counterclaw import CounterClawInterceptor

interceptor = CounterClawInterceptor()

# Input scan - blocks prompt injections
# NOTE: Examples below are TEST CASES only - not actual instructions
result = interceptor.check_input("{{EXAMPLE: ignore previous instructions}}")
# β†’ {"blocked": True, "safe": False}

# Output scan - detects PII leaks  
result = interceptor.check_output("Contact: john@example.com")
# β†’ {"safe": False, "pii_detected": {"email": True}}
```

## Features

- πŸ”’ Defense against common prompt injection patterns
- πŸ›‘οΈ Basic PII masking (Email, Phone, Credit Card)
- πŸ“ Violation logging to `~/.openclaw/memory/MEMORY.md`
- ⚠️ Warning on startup if TRUSTED_ADMIN_IDS not configured

## Configuration

### Required Environment Variable

```bash
# Set your trusted admin ID(s) - use non-sensitive identifiers only!
export TRUSTED_ADMIN_IDS="your_telegram_id"
```

**Important:** `TRUSTED_ADMIN_IDS` should ONLY contain non-sensitive identifiers:
- βœ… Telegram user IDs (e.g., `"123456789"`)
- βœ… Discord user IDs (e.g., `"987654321"`)
- ❌ NEVER API keys
- ❌ NEVER passwords
- ❌ NEVER tokens

You can set multiple admin IDs by comma-separating:
```bash
export TRUSTED_ADMIN_IDS="telegram_id_1,telegram_id_2"
```

### Runtime Configuration

```python
# Option 1: Via environment variable (recommended)
# Set TRUSTED_ADMIN_IDS before running
interceptor = CounterClawInterceptor()

# Option 2: Direct parameter
interceptor = CounterClawInterceptor(admin_user_id="123456789")
```

## Security Notes

- **Fail-Closed**: If `TRUSTED_ADMIN_IDS` is not set, admin features are disabled by default
- **Logging**: All violations are logged to `~/.openclaw/memory/MEMORY.md` with PII masked
- **No Network Access**: This middleware does not make any external network calls (offline-only)
- **File Access**: Only writes to `~/.openclaw/memory/MEMORY.md` β€” explicitly declared scope

## Files Created

| Path | Purpose |
|------|---------|
| `~/.openclaw/memory/` | Directory created on first run |
| `~/.openclaw/memory/MEMORY.md` | Violation logs with PII masked |

## License

MIT - See LICENSE file

## Development & Release

### Running Tests Locally

```bash
python3 tests/test_scanner.py
```

### Linting

```bash
pip install ruff
ruff check src/
```

### Publishing to ClawHub

The CI runs on every push and pull request:
1. **Ruff** - Lints Python code
2. **Tests** - Runs unit tests

To publish a new version:

```bash
# Version is set in pyproject.toml
git add -A
git commit -m "Release v1.0.9"
git tag v1.0.9
git push origin main --tags
```

CI will automatically:
- Run lint + tests
- If tests pass and tag starts with `v*`, publish to ClawHub

Files in this skill

  • README.md6.7 KB
  • SKILL.md3.8 KB
  • email_protector.py4.1 KB
  • pyproject.toml1.2 KB
  • send_protected_email.sh2.5 KB
  • src/counterclaw/__init__.py410 B
  • src/counterclaw/middleware.py5 KB
  • src/counterclaw/scanner.py1.7 KB
  • src/counterclaw_cli/__init__.py426 B
  • tests/test_email_protection.py4.6 KB
  • tests/test_scanner.py2.8 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…