Skip to content
Back to skills

061 Architecture Aa8fd035

CSecurity

> Internal architecture documentation for contributors and maintainers. ---

  • 9 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added October 11, 2026
documentationpythonrustgosqlapisecurityperformancedocumentation

Works with

  • cli
  • api

Security analysis

C60/100
  • highPerforms destructive filesystem operations
  • criticalImpersonates system messages to override safety constraints

Pro shows the line behind each finding and how to fix it

Scanned October 11, 2026

npx -y skills add tools-only/X-Skills --skill 061-architecture_aa8fd035 --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of 061 Architecture Aa8fd035?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for 061 Architecture Aa8fd035
[![Security: C β€” Skills Directory](https://www.skillsdirectory.com/api/skills/tools-only-061-architecture-aa8fd035/badge)](https://www.skillsdirectory.com/skills/tools-only-061-architecture-aa8fd035)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

SKILL.md
# πŸ—οΈ Prompt Guard Architecture

> Internal architecture documentation for contributors and maintainers.

---

## Overview

Prompt GuardλŠ” **λ‹€μΈ΅ λ°©μ–΄(Defense in Depth)** μ›μΉ™μœΌλ‘œ 섀계됨. 단일 νŒ¨ν„΄μ΄ μ•„λ‹Œ μ—¬λŸ¬ λ ˆμ΄μ–΄μ˜ 검사λ₯Ό 톡해 false positiveλ₯Ό μ€„μ΄λ©΄μ„œ 곡격을 효과적으둜 탐지.

```
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                        INPUT MESSAGE                            β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                               β”‚
                               β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚  Layer 1: Rate Limiting                                         β”‚
β”‚  β€’ Per-user request tracking                                    β”‚
β”‚  β€’ Sliding window algorithm                                     β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                               β”‚
                               β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚  Layer 2: Text Normalization (v2.8.0 expanded)                  β”‚
β”‚  β€’ Homoglyph detection & replacement                            β”‚
β”‚  β€’ Visible delimiter stripping (I+g+n+o+r+e β†’ Ignore)          β”‚
β”‚  β€’ Character spacing collapse (i g n o r e β†’ ignore)            β”‚
β”‚  β€’ Zero-width character removal                                 β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                               β”‚
                               β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚  Layer 2.5: Decode Pipeline (NEW v2.8.0)                        β”‚
β”‚  β€’ Base64 decode + full pattern re-scan                         β”‚
β”‚  β€’ Hex escape decode (\x41\x42)                                 β”‚
β”‚  β€’ ROT13 decode (full-text + per-word)                          β”‚
β”‚  β€’ URL decode (%69%67%6E)                                       β”‚
β”‚  β€’ HTML entity decode (i β†’ i)                              β”‚
β”‚  β€’ Unicode escape decode (\u0069 β†’ i)                           β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                               β”‚
                               β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚  Layer 3: Pattern Matching Engine                               β”‚
β”‚  β€’ Runs against ORIGINAL + all DECODED variants                 β”‚
β”‚  β€’ Critical patterns (immediate block)                          β”‚
β”‚  β€’ Secret/Token requests                                        β”‚
β”‚  β€’ Multi-language injection patterns (10 languages)             β”‚
β”‚  β€’ Scenario jailbreaks                                          β”‚
β”‚  β€’ Social engineering                                           β”‚
β”‚  β€’ Indirect injection                                           β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                               β”‚
                               β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚  Layer 4: Language Detection (NEW v2.8.0)                       β”‚
β”‚  β€’ Detect input language (optional: langdetect)                 β”‚
β”‚  β€’ Flag unsupported languages at MEDIUM severity                β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                               β”‚
                               β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚  Layer 5: Behavioral Analysis                                   β”‚
β”‚  β€’ Repetition detection (token overflow)                        β”‚
β”‚  β€’ Context hijacking patterns                                   β”‚
β”‚  β€’ Multi-turn manipulation                                      β”‚
β”‚  β€’ Invisible character detection                                β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                               β”‚
                               β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚  Layer 5.5: Canary Token Check (NEW v2.8.0)                     β”‚
β”‚  β€’ Check for user-defined canary tokens in message              β”‚
β”‚  β€’ Detects system prompt extraction                             β”‚
β”‚  β€’ CRITICAL severity if canary found                            β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                               β”‚
                               β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚  Layer 6: Context-Aware Decision                                β”‚
β”‚  β€’ Sensitivity adjustment                                       β”‚
β”‚  β€’ Owner bypass rules                                           β”‚
β”‚  β€’ Group context restrictions                                   β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                               β”‚
                               β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                     DetectionResult                             β”‚
β”‚  β€’ severity: SAFE β†’ LOW β†’ MEDIUM β†’ HIGH β†’ CRITICAL              β”‚
β”‚  β€’ action: ALLOW | LOG | WARN | BLOCK | BLOCK_NOTIFY            β”‚
β”‚  β€’ reasons: [matched pattern categories]                        β”‚
β”‚  β€’ decoded_findings: [encoding details]                         β”‚
β”‚  β€’ canary_matches: [leaked canary tokens]                       β”‚
β”‚  β€’ Logged to Markdown and/or JSONL (with hash chain)            β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚  Layer 7: Output Scanner / DLP (NEW v2.8.0)                     β”‚
β”‚  β€’ scan_output() - separate method for LLM responses            β”‚
β”‚  β€’ Canary token leakage detection                               β”‚
β”‚  β€’ Credential format patterns (15+ key formats)                 β”‚
β”‚  β€’ Secret/sensitive path detection                              β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚  Layer 8: Enterprise DLP Sanitizer (NEW v2.8.1)                 β”‚
β”‚  β€’ sanitize_output() - redact-first, block-as-fallback          β”‚
β”‚  β€’ 17 credential patterns β†’ [REDACTED:type] labels              β”‚
β”‚  β€’ Canary token auto-redaction β†’ [REDACTED:canary]              β”‚
β”‚  β€’ Post-redaction re-scan: block if still HIGH+                 β”‚
β”‚  β€’ Returns SanitizeResult with full audit metadata              β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
```

---

## Core Components

### 1. Severity Levels

| Level | Value | Description | Typical Trigger |
|-------|-------|-------------|-----------------|
| SAFE | 0 | No threat detected | Normal conversation |
| LOW | 1 | Minor suspicious signal | Output manipulation |
| MEDIUM | 2 | Clear manipulation attempt | Role manipulation, urgency |
| HIGH | 3 | Dangerous command | Jailbreaks, system access |
| CRITICAL | 4 | Immediate threat | Secret exfil, code execution |

### 2. Action Types

| Action | Description | When Used |
|--------|-------------|-----------|
| `allow` | No intervention | SAFE severity |
| `log` | Record only | Owner requests, LOW severity |
| `warn` | Notify user | MEDIUM severity |
| `block` | Refuse request | HIGH severity |
| `block_notify` | Block + alert owner | CRITICAL severity |

### 3. Pattern Categories

#### πŸ”΄ Critical (Immediate Block)
- `CRITICAL_PATTERNS` - rm -rf, fork bombs, SQL injection, XSS
- `SECRET_PATTERNS` - API key/token/password requests

#### 🟠 v2.6.0 Social Engineering Defense
- `APPROVAL_EXPANSION` - "μ•„κΉŒ ν—ˆλ½ν–ˆμž–μ•„" scope creep
- `CREDENTIAL_PATH_PATTERNS` - credentials.json, .env 경둜
- `BYPASS_COACHING` - "μž‘λ™λ˜κ²Œ λ§Œλ“€μ–΄" bypass help
- `DM_SOCIAL_ENGINEERING` - DM μ‘°μž‘ νŒ¨ν„΄

#### 🟑 v2.5.x Advanced Patterns
- `INDIRECT_INJECTION` - URL/file/image-based injection
- `CONTEXT_HIJACKING` - Fake memory/history manipulation
- `MULTI_TURN_MANIPULATION` - Gradual trust building
- `TOKEN_SMUGGLING` - Invisible Unicode characters
- `SYSTEM_PROMPT_MIMICRY` - `<claude_*>`, `[INST]` λ“±

#### 🟒 v2.4.0 Red Team Patterns
- `SCENARIO_JAILBREAK` - Dream/story/cinema/academic
- `EMOTIONAL_MANIPULATION` - Moral dilemmas, threats
- `AUTHORITY_RECON` - Fake admin, capability probing
- `COGNITIVE_MANIPULATION` - Hypnosis/trance patterns
- `PHISHING_SOCIAL_ENG` - Password reset templates

#### πŸ”΅ Language-Specific
- `PATTERNS_EN` - English patterns
- `PATTERNS_KO` - ν•œκ΅­μ–΄ νŒ¨ν„΄
- `PATTERNS_JA` - ζ—₯本θͺžγƒ‘ターン
- `PATTERNS_ZH` - 中文樑式

---

## Detection Flow

```python
def analyze(message, context):
    # 1. Rate limit check
    if check_rate_limit(user_id):
        return BLOCK

    # 2. Text normalization (v2.8.0: expanded)
    normalized, has_homoglyphs, was_defragmented = normalize(message)
    # Now handles: homoglyphs, delimiter stripping, character spacing
    
    # 3. Critical patterns (highest priority)
    for pattern in CRITICAL_PATTERNS:
        if match(pattern, normalized):
            return CRITICAL
    
    # 4. Secret request patterns
    for lang, patterns in SECRET_PATTERNS:
        if match(pattern, text):
            return CRITICAL
    
    # 5. Versioned pattern sets (newest first)
    # v2.7.0, v2.6.x, v2.5.x, v2.4.0 patterns
    
    # 6. Language-specific patterns (10 languages)
    for lang in [EN, KO, JA, ZH, RU, ES, DE, FR, PT, VI]:
        check_language_patterns(lang)
    
    # 7. Base64 detection (v2.8.0: expanded 40-word list + full pattern re-scan)
    suspicious = detect_base64(message)
    
    # 8. Decode-then-scan (NEW v2.8.0)
    decoded_variants = decode_all(message)  # Base64, Hex, ROT13, URL, HTML, Unicode
    for variant in decoded_variants:
        _scan_text_for_patterns(variant["decoded"])  # Re-run full pattern engine
    
    # 9. Canary token check (NEW v2.8.0)
    canary_matches = check_canary(message)
    
    # 10. Language detection (NEW v2.8.0)
    if detected_language not in SUPPORTED_LANGUAGES:
        flag as unsupported_language_risk
    
    # 11. Behavioral analysis
    check_repetition()
    check_invisible_chars()
    
    # 12. Context-aware adjustment
    adjust_for_sensitivity()
    apply_owner_rules()
    apply_group_restrictions()
    
    # 13. Auto-log (markdown + JSON)
    log_detection()
    log_detection_json()  # NEW v2.8.0: JSONL with hash chain
    
    return DetectionResult(...)

def scan_output(response_text, context):  # NEW v2.8.0
    """DLP: Scan LLM output for data leakage."""
    check_canary(response_text)
    check_credential_formats(response_text)  # 15+ key formats
    check_secret_patterns(response_text)
    check_sensitive_paths(response_text)
    return DetectionResult(scan_type="output")

def sanitize_output(response_text, context):  # NEW v2.8.1
    """Enterprise DLP: Redact-first, block-as-fallback."""
    # Step 1: Redact 17 credential patterns β†’ [REDACTED:type]
    for pattern in CREDENTIAL_REDACTION_PATTERNS:
        text = re.sub(pattern, replacement, text)
    
    # Step 2: Redact canary tokens β†’ [REDACTED:canary]
    for token in canary_tokens:
        text = text.replace(token, "[REDACTED:canary]")
    
    # Step 3: Re-scan redacted text
    post_scan = scan_output(redacted_text)
    
    # Step 4: Block if re-scan still HIGH+, else return redacted text
    if post_scan.severity >= HIGH:
        return SanitizeResult(blocked=True)
    return SanitizeResult(sanitized_text=redacted_text, blocked=False)
```

---

## File Structure

```
prompt-guard/
β”œβ”€β”€ README.md              # User documentation
β”œβ”€β”€ ARCHITECTURE.md        # This file
β”œβ”€β”€ SKILL.md               # Clawdbot skill interface
β”œβ”€β”€ config.example.yaml    # Configuration template
β”œβ”€β”€ requirements.txt       # Dependencies (pyyaml, optional: langdetect)
β”œβ”€β”€ pyproject.toml         # Build config, entry points, dependencies
β”‚
β”œβ”€β”€ prompt_guard/          # Main package (v3.0)
β”‚   β”œβ”€β”€ __init__.py        # Public API + __version__ (re-exports)
β”‚   β”œβ”€β”€ models.py          # Severity, Action, DetectionResult, SanitizeResult
β”‚   β”œβ”€β”€ patterns.py        # 500+ regex patterns (pure data, ~1200 lines)
β”‚   β”œβ”€β”€ normalizer.py      # HOMOGLYPHS dict + normalize() function
β”‚   β”œβ”€β”€ decoder.py         # decode_all() + detect_base64() (Base64/Hex/ROT13/URL/HTML/Unicode)
β”‚   β”œβ”€β”€ scanner.py         # scan_text_for_patterns() (reusable pattern matcher)
β”‚   β”œβ”€β”€ engine.py          # PromptGuard class (analyze, config, rate_limit, canary, language)
β”‚   β”œβ”€β”€ output.py          # scan_output() + sanitize_output() (enterprise DLP)
β”‚   β”œβ”€β”€ logging_utils.py   # log_detection(), log_detection_json(), report_to_hivefence()
β”‚   β”œβ”€β”€ cli.py             # main() CLI entry point
β”‚   β”œβ”€β”€ hivefence.py       # HiveFence threat intelligence client
β”‚   β”œβ”€β”€ audit.py           # System security audit
β”‚   └── analyze_log.py     # Security log analyzer
β”‚
β”œβ”€β”€ scripts/               # Backward-compat shims (deprecated, emit warnings)
β”‚   β”œβ”€β”€ __init__.py        # DeprecationWarning + re-import from prompt_guard
β”‚   └── detect.py          # DeprecationWarning + re-import from prompt_guard
β”‚
└── tests/
    β”œβ”€β”€ test_detect.py     # 121 regression tests
    └── test_detect_cli.py # CLI integration tests
```

---

## Pattern Organization

### Naming Convention
```
{CATEGORY}_{VERSION?} = [
    r"pattern1",
    r"pattern2",
]
```

### Version Tagging in Matches
νŒ¨ν„΄ λ§€μΉ­ μ‹œ 버전 νƒœκ·Έ μΆ”κ°€:
- `new:{category}:{pattern}` - v2.4.0 red team
- `v25:{category}:{pattern}` - v2.5.0 indirect
- `v252:{category}:{pattern}` - v2.5.2 moltbook
- `{lang}:{category}:{pattern}` - language-specific

---

## Configuration Schema

```yaml
prompt_guard:
  # Detection sensitivity
  sensitivity: medium  # low | medium | high | paranoid
  
  # Owner IDs (bypass most restrictions)
  owner_ids:
    - "USER_ID"
  
  # Action per severity
  actions:
    LOW: log
    MEDIUM: warn
    HIGH: block
    CRITICAL: block_notify
  
  # Rate limiting
  rate_limit:
    enabled: true
    max_requests: 30
    window_seconds: 60
  
  # Logging
  logging:
    enabled: true
    path: memory/security-log.md
```

---

## Key Design Decisions

### 1. Regex over ML
- **Pros**: Deterministic, explainable, no model dependencies
- **Cons**: Manual pattern updates needed
- **Reasoning**: Security requires predictability; ML false negatives unacceptable

### 2. Multi-Language First
- All patterns have EN/KO/JA/ZH variants
- Attack language != user language (multilingual attacks common)

### 3. Severity Graduation
- Not binary block/allow
- Owner context matters (more lenient for owners)
- Group context matters (stricter in groups)

### 4. Versioned Patterns
- Clear provenance for each pattern set
- Credits to contributors (ν™λ―Όν‘œ, Moltbook, etc.)
- Easy to audit and roll back

---

## Extension Points

### Adding New Patterns
```python
# 1. Define pattern list
NEW_ATTACK_CATEGORY = [
    r"pattern1",
    r"pattern2",
]

# 2. Add to analysis loop
new_pattern_sets = [
    ...
    (NEW_ATTACK_CATEGORY, "new_category", Severity.HIGH),
]
```

### Adding New Languages
```python
PATTERNS_XX = {
    "instruction_override": [...],
    "role_manipulation": [...],
    ...
}

# Add to all_patterns
all_patterns.append((PATTERNS_XX, "xx"))
```

---

## Performance Notes

- **Regex compilation**: Patterns are compiled on first use (Python caches)
- **Early exit**: CRITICAL patterns checked first
- **Fingerprinting**: Hash-based dedup for repeated attacks
- **Rate limiting**: O(1) sliding window

---

## Security Considerations

### What We DON'T Do
- ❌ Execute user input
- ❌ Log sensitive data in plaintext
- ❌ Trust any "admin" claims without owner_id verification

### What We DO
- βœ… Fail closed (block on uncertainty)
- βœ… Log all suspicious activity
- βœ… Stricter rules in group contexts

---

## Changelog Location

버전별 변경사항은 `detect.py` 상단 docstring에 기둝:

```python
"""
Prompt Guard v2.6.0 - Advanced Prompt Injection Detection

Changelog v2.6.0 (2026-02-01):
- Added Single Approval Expansion detection
- Added Credential Path Harvesting detection
...
"""
```

---

## Credits

- **Core**: @simonkim_nft (κΉ€μ„œμ€€)
- **v2.4.0 Red Team**: ν™λ―Όν‘œ (@kanfrancisco)
- **v2.4.1 Config Fix**: Junho Yeo (@junhoyeo)
- **v2.5.2 Moltbook Patterns**: Community reports

---

*Last updated: 2026-02-07 | v2.8.0*

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…