Back to skills
SKILL.md
061 Architecture Aa8fd035
CSecurity> Internal architecture documentation for contributors and maintainers. ---
- 9 stars
- 0 votes
- 0 copies
- 0 views
- Added October 11, 2026
Works with
Security analysis
60/100- Performs destructive filesystem operations
- Impersonates system messages to override safety constraints
npx -y skills add tools-only/X-Skills --skill 061-architecture_aa8fd035 --agent claude-codeAre you the author of 061 Architecture Aa8fd035?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/tools-only-061-architecture-aa8fd035)# ποΈ Prompt Guard Architecture
> Internal architecture documentation for contributors and maintainers.
---
## Overview
Prompt Guardλ **λ€μΈ΅ λ°©μ΄(Defense in Depth)** μμΉμΌλ‘ μ€κ³λ¨. λ¨μΌ ν¨ν΄μ΄ μλ μ¬λ¬ λ μ΄μ΄μ κ²μ¬λ₯Ό ν΅ν΄ false positiveλ₯Ό μ€μ΄λ©΄μ 곡격μ ν¨κ³Όμ μΌλ‘ νμ§.
```
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β INPUT MESSAGE β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β
βΌ
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β Layer 1: Rate Limiting β
β β’ Per-user request tracking β
β β’ Sliding window algorithm β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β
βΌ
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β Layer 2: Text Normalization (v2.8.0 expanded) β
β β’ Homoglyph detection & replacement β
β β’ Visible delimiter stripping (I+g+n+o+r+e β Ignore) β
β β’ Character spacing collapse (i g n o r e β ignore) β
β β’ Zero-width character removal β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β
βΌ
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β Layer 2.5: Decode Pipeline (NEW v2.8.0) β
β β’ Base64 decode + full pattern re-scan β
β β’ Hex escape decode (\x41\x42) β
β β’ ROT13 decode (full-text + per-word) β
β β’ URL decode (%69%67%6E) β
β β’ HTML entity decode (i β i) β
β β’ Unicode escape decode (\u0069 β i) β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β
βΌ
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β Layer 3: Pattern Matching Engine β
β β’ Runs against ORIGINAL + all DECODED variants β
β β’ Critical patterns (immediate block) β
β β’ Secret/Token requests β
β β’ Multi-language injection patterns (10 languages) β
β β’ Scenario jailbreaks β
β β’ Social engineering β
β β’ Indirect injection β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β
βΌ
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β Layer 4: Language Detection (NEW v2.8.0) β
β β’ Detect input language (optional: langdetect) β
β β’ Flag unsupported languages at MEDIUM severity β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β
βΌ
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β Layer 5: Behavioral Analysis β
β β’ Repetition detection (token overflow) β
β β’ Context hijacking patterns β
β β’ Multi-turn manipulation β
β β’ Invisible character detection β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β
βΌ
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β Layer 5.5: Canary Token Check (NEW v2.8.0) β
β β’ Check for user-defined canary tokens in message β
β β’ Detects system prompt extraction β
β β’ CRITICAL severity if canary found β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β
βΌ
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β Layer 6: Context-Aware Decision β
β β’ Sensitivity adjustment β
β β’ Owner bypass rules β
β β’ Group context restrictions β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β
βΌ
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β DetectionResult β
β β’ severity: SAFE β LOW β MEDIUM β HIGH β CRITICAL β
β β’ action: ALLOW | LOG | WARN | BLOCK | BLOCK_NOTIFY β
β β’ reasons: [matched pattern categories] β
β β’ decoded_findings: [encoding details] β
β β’ canary_matches: [leaked canary tokens] β
β β’ Logged to Markdown and/or JSONL (with hash chain) β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β Layer 7: Output Scanner / DLP (NEW v2.8.0) β
β β’ scan_output() - separate method for LLM responses β
β β’ Canary token leakage detection β
β β’ Credential format patterns (15+ key formats) β
β β’ Secret/sensitive path detection β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β Layer 8: Enterprise DLP Sanitizer (NEW v2.8.1) β
β β’ sanitize_output() - redact-first, block-as-fallback β
β β’ 17 credential patterns β [REDACTED:type] labels β
β β’ Canary token auto-redaction β [REDACTED:canary] β
β β’ Post-redaction re-scan: block if still HIGH+ β
β β’ Returns SanitizeResult with full audit metadata β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
```
---
## Core Components
### 1. Severity Levels
| Level | Value | Description | Typical Trigger |
|-------|-------|-------------|-----------------|
| SAFE | 0 | No threat detected | Normal conversation |
| LOW | 1 | Minor suspicious signal | Output manipulation |
| MEDIUM | 2 | Clear manipulation attempt | Role manipulation, urgency |
| HIGH | 3 | Dangerous command | Jailbreaks, system access |
| CRITICAL | 4 | Immediate threat | Secret exfil, code execution |
### 2. Action Types
| Action | Description | When Used |
|--------|-------------|-----------|
| `allow` | No intervention | SAFE severity |
| `log` | Record only | Owner requests, LOW severity |
| `warn` | Notify user | MEDIUM severity |
| `block` | Refuse request | HIGH severity |
| `block_notify` | Block + alert owner | CRITICAL severity |
### 3. Pattern Categories
#### π΄ Critical (Immediate Block)
- `CRITICAL_PATTERNS` - rm -rf, fork bombs, SQL injection, XSS
- `SECRET_PATTERNS` - API key/token/password requests
#### π v2.6.0 Social Engineering Defense
- `APPROVAL_EXPANSION` - "μκΉ νλ½νμμ" scope creep
- `CREDENTIAL_PATH_PATTERNS` - credentials.json, .env κ²½λ‘
- `BYPASS_COACHING` - "μλλκ² λ§λ€μ΄" bypass help
- `DM_SOCIAL_ENGINEERING` - DM μ‘°μ ν¨ν΄
#### π‘ v2.5.x Advanced Patterns
- `INDIRECT_INJECTION` - URL/file/image-based injection
- `CONTEXT_HIJACKING` - Fake memory/history manipulation
- `MULTI_TURN_MANIPULATION` - Gradual trust building
- `TOKEN_SMUGGLING` - Invisible Unicode characters
- `SYSTEM_PROMPT_MIMICRY` - `<claude_*>`, `[INST]` λ±
#### π’ v2.4.0 Red Team Patterns
- `SCENARIO_JAILBREAK` - Dream/story/cinema/academic
- `EMOTIONAL_MANIPULATION` - Moral dilemmas, threats
- `AUTHORITY_RECON` - Fake admin, capability probing
- `COGNITIVE_MANIPULATION` - Hypnosis/trance patterns
- `PHISHING_SOCIAL_ENG` - Password reset templates
#### π΅ Language-Specific
- `PATTERNS_EN` - English patterns
- `PATTERNS_KO` - νκ΅μ΄ ν¨ν΄
- `PATTERNS_JA` - ζ₯ζ¬θͺγγΏγΌγ³
- `PATTERNS_ZH` - δΈζ樑εΌ
---
## Detection Flow
```python
def analyze(message, context):
# 1. Rate limit check
if check_rate_limit(user_id):
return BLOCK
# 2. Text normalization (v2.8.0: expanded)
normalized, has_homoglyphs, was_defragmented = normalize(message)
# Now handles: homoglyphs, delimiter stripping, character spacing
# 3. Critical patterns (highest priority)
for pattern in CRITICAL_PATTERNS:
if match(pattern, normalized):
return CRITICAL
# 4. Secret request patterns
for lang, patterns in SECRET_PATTERNS:
if match(pattern, text):
return CRITICAL
# 5. Versioned pattern sets (newest first)
# v2.7.0, v2.6.x, v2.5.x, v2.4.0 patterns
# 6. Language-specific patterns (10 languages)
for lang in [EN, KO, JA, ZH, RU, ES, DE, FR, PT, VI]:
check_language_patterns(lang)
# 7. Base64 detection (v2.8.0: expanded 40-word list + full pattern re-scan)
suspicious = detect_base64(message)
# 8. Decode-then-scan (NEW v2.8.0)
decoded_variants = decode_all(message) # Base64, Hex, ROT13, URL, HTML, Unicode
for variant in decoded_variants:
_scan_text_for_patterns(variant["decoded"]) # Re-run full pattern engine
# 9. Canary token check (NEW v2.8.0)
canary_matches = check_canary(message)
# 10. Language detection (NEW v2.8.0)
if detected_language not in SUPPORTED_LANGUAGES:
flag as unsupported_language_risk
# 11. Behavioral analysis
check_repetition()
check_invisible_chars()
# 12. Context-aware adjustment
adjust_for_sensitivity()
apply_owner_rules()
apply_group_restrictions()
# 13. Auto-log (markdown + JSON)
log_detection()
log_detection_json() # NEW v2.8.0: JSONL with hash chain
return DetectionResult(...)
def scan_output(response_text, context): # NEW v2.8.0
"""DLP: Scan LLM output for data leakage."""
check_canary(response_text)
check_credential_formats(response_text) # 15+ key formats
check_secret_patterns(response_text)
check_sensitive_paths(response_text)
return DetectionResult(scan_type="output")
def sanitize_output(response_text, context): # NEW v2.8.1
"""Enterprise DLP: Redact-first, block-as-fallback."""
# Step 1: Redact 17 credential patterns β [REDACTED:type]
for pattern in CREDENTIAL_REDACTION_PATTERNS:
text = re.sub(pattern, replacement, text)
# Step 2: Redact canary tokens β [REDACTED:canary]
for token in canary_tokens:
text = text.replace(token, "[REDACTED:canary]")
# Step 3: Re-scan redacted text
post_scan = scan_output(redacted_text)
# Step 4: Block if re-scan still HIGH+, else return redacted text
if post_scan.severity >= HIGH:
return SanitizeResult(blocked=True)
return SanitizeResult(sanitized_text=redacted_text, blocked=False)
```
---
## File Structure
```
prompt-guard/
βββ README.md # User documentation
βββ ARCHITECTURE.md # This file
βββ SKILL.md # Clawdbot skill interface
βββ config.example.yaml # Configuration template
βββ requirements.txt # Dependencies (pyyaml, optional: langdetect)
βββ pyproject.toml # Build config, entry points, dependencies
β
βββ prompt_guard/ # Main package (v3.0)
β βββ __init__.py # Public API + __version__ (re-exports)
β βββ models.py # Severity, Action, DetectionResult, SanitizeResult
β βββ patterns.py # 500+ regex patterns (pure data, ~1200 lines)
β βββ normalizer.py # HOMOGLYPHS dict + normalize() function
β βββ decoder.py # decode_all() + detect_base64() (Base64/Hex/ROT13/URL/HTML/Unicode)
β βββ scanner.py # scan_text_for_patterns() (reusable pattern matcher)
β βββ engine.py # PromptGuard class (analyze, config, rate_limit, canary, language)
β βββ output.py # scan_output() + sanitize_output() (enterprise DLP)
β βββ logging_utils.py # log_detection(), log_detection_json(), report_to_hivefence()
β βββ cli.py # main() CLI entry point
β βββ hivefence.py # HiveFence threat intelligence client
β βββ audit.py # System security audit
β βββ analyze_log.py # Security log analyzer
β
βββ scripts/ # Backward-compat shims (deprecated, emit warnings)
β βββ __init__.py # DeprecationWarning + re-import from prompt_guard
β βββ detect.py # DeprecationWarning + re-import from prompt_guard
β
βββ tests/
βββ test_detect.py # 121 regression tests
βββ test_detect_cli.py # CLI integration tests
```
---
## Pattern Organization
### Naming Convention
```
{CATEGORY}_{VERSION?} = [
r"pattern1",
r"pattern2",
]
```
### Version Tagging in Matches
ν¨ν΄ λ§€μΉ μ λ²μ νκ·Έ μΆκ°:
- `new:{category}:{pattern}` - v2.4.0 red team
- `v25:{category}:{pattern}` - v2.5.0 indirect
- `v252:{category}:{pattern}` - v2.5.2 moltbook
- `{lang}:{category}:{pattern}` - language-specific
---
## Configuration Schema
```yaml
prompt_guard:
# Detection sensitivity
sensitivity: medium # low | medium | high | paranoid
# Owner IDs (bypass most restrictions)
owner_ids:
- "USER_ID"
# Action per severity
actions:
LOW: log
MEDIUM: warn
HIGH: block
CRITICAL: block_notify
# Rate limiting
rate_limit:
enabled: true
max_requests: 30
window_seconds: 60
# Logging
logging:
enabled: true
path: memory/security-log.md
```
---
## Key Design Decisions
### 1. Regex over ML
- **Pros**: Deterministic, explainable, no model dependencies
- **Cons**: Manual pattern updates needed
- **Reasoning**: Security requires predictability; ML false negatives unacceptable
### 2. Multi-Language First
- All patterns have EN/KO/JA/ZH variants
- Attack language != user language (multilingual attacks common)
### 3. Severity Graduation
- Not binary block/allow
- Owner context matters (more lenient for owners)
- Group context matters (stricter in groups)
### 4. Versioned Patterns
- Clear provenance for each pattern set
- Credits to contributors (νλ―Όν, Moltbook, etc.)
- Easy to audit and roll back
---
## Extension Points
### Adding New Patterns
```python
# 1. Define pattern list
NEW_ATTACK_CATEGORY = [
r"pattern1",
r"pattern2",
]
# 2. Add to analysis loop
new_pattern_sets = [
...
(NEW_ATTACK_CATEGORY, "new_category", Severity.HIGH),
]
```
### Adding New Languages
```python
PATTERNS_XX = {
"instruction_override": [...],
"role_manipulation": [...],
...
}
# Add to all_patterns
all_patterns.append((PATTERNS_XX, "xx"))
```
---
## Performance Notes
- **Regex compilation**: Patterns are compiled on first use (Python caches)
- **Early exit**: CRITICAL patterns checked first
- **Fingerprinting**: Hash-based dedup for repeated attacks
- **Rate limiting**: O(1) sliding window
---
## Security Considerations
### What We DON'T Do
- β Execute user input
- β Log sensitive data in plaintext
- β Trust any "admin" claims without owner_id verification
### What We DO
- β
Fail closed (block on uncertainty)
- β
Log all suspicious activity
- β
Stricter rules in group contexts
---
## Changelog Location
λ²μ λ³ λ³κ²½μ¬νμ `detect.py` μλ¨ docstringμ κΈ°λ‘:
```python
"""
Prompt Guard v2.6.0 - Advanced Prompt Injection Detection
Changelog v2.6.0 (2026-02-01):
- Added Single Approval Expansion detection
- Added Credential Path Harvesting detection
...
"""
```
---
## Credits
- **Core**: @simonkim_nft (κΉμμ€)
- **v2.4.0 Red Team**: νλ―Όν (@kanfrancisco)
- **v2.4.1 Config Fix**: Junho Yeo (@junhoyeo)
- **v2.5.2 Moltbook Patterns**: Community reports
---
*Last updated: 2026-02-07 | v2.8.0*
Attribution
Comments
Loading commentsβ¦