Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Cloud Architect

BSecurity

Use when configuring, automating, deploying, and debugging cloud architect pipelines, containers, servers, and cloud infrastructure.

5 stars
0 votes
0 copies
0 views
Added 9/27/2026
ai-agentsrustgoshellbashsqlnoderailskubernetesawsterraform

Works with

terminalapi

Security Analysis

B88/100
criticalExfiltrates credentials via HTTP — exact pattern from Snyk ToxicSkills study

Pro shows the line behind each finding and how to fix it

Scanned 9/29/2026

$npx -y skills add Harmitx7/tribunal-kit --skill cloud-architect --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Cloud Architect?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Cloud Architect
[![Security: B — Skills Directory](https://www.skillsdirectory.com/api/skills/harmitx7-cloud-architect/badge)](https://www.skillsdirectory.com/skills/harmitx7-cloud-architect)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: cloud-architect
description: "Use when configuring, automating, deploying, and debugging cloud architect pipelines, containers, servers, and cloud infrastructure."
version: 6.0.0
last-updated: 2026-09-29
skills:
  - cicd-pro
  - devops-engineer
  - platform-engineer
tools: Read, Grep, Glob, Bash, Edit, Write
scripts-binding:
  - .agent/scripts/lint_runner.js
  - .agent/scripts/verify_all.js
---

# Cloud Architect — Production AWS Mastery

## Mandatory Pre-Flight Context Inspection
Before reading, generating, or refactoring code in the `cloud-architect` domain, inspect these 5 critical parameters:
1. **System Boundaries & Dependencies**: Verify that all required dependencies exist in target package manifests and environment paths.
2. **Runtime Context & Platform Invariants**: Confirm target platform constraints (Node.js, Browser, Mobile OS, Edge runtime) before applying APIs.
3. **Execution Guardrails**: Identify potential side-effects, state mutations, and unhandled asynchronous exceptions.
4. **Validation & Type Contracts**: Validate input data schemas and strict type constraints across all module interfaces.
5. **Observability & Proof of Execution**: Ensure execution produces tangible verification signals (terminal output, tests, metrics).


## Activation Boundaries
- **Activate when:** Use when configuring, automating, deploying, and debugging cloud architect pipelines, containers, servers, and cloud infrastructure.
- **DO NOT activate when:** The task falls outside the `cloud-architect` domain or is managed by a different dedicated specialist agent.


## 🔁 Multi-Pass Execution Protocol

| Pass | Phase | Core Action | Adaptive Depth |
|:---|:---|:---|:---|
| **Pass 1** | **Understand** | Deconstruct the user's explicit objective, implicit requirements, and platform constraints. | Fast / Standard / Deep |
| **Pass 2** | **Plan** | Decompose task into smallest logical steps; map dependencies, affected files, and tool calls. | Standard / Deep |
| **Pass 3** | **Execute** | Implement solution with production-grade craft, zero placeholders, and strict typing. | All Modes |
| **Pass 4** | **Verify** | Run linters, unit tests, or compiler checks to validate structural correctness. | All Modes |
| **Pass 5** | **Attack & Falsify** | Perform adversarial search for edge-case failures, counterexamples, race conditions, and traps. | Standard / Deep |
| **Pass 6** | **Harden** | Eliminate discovered friction, optimize performance, and harden error boundaries. | Standard / Deep |
| **Pass 7** | **Quality Gate** | Enforce Verification-Before-Completion (VBC) with concrete terminal proof before finalizing. | All Modes |


---

## 🛠️ Technical Architecture & Reference Recipes

## Hallucination Traps (Read First)

- ❌ Confusing availability zones with regions → ✅ `us-east-1` is a region. `us-east-1a` is an AZ within that region. Multi-AZ ≠ multi-region.
- ❌ `s3:*` or `*:*` IAM policies → ✅ IAM policies must follow least privilege. Enumerate exact actions required.
- ❌ Hardcoding ARNs, account IDs, or region strings → ✅ Always use `data.aws_caller_identity.current.account_id`, `var.region`, Terraform variables.
- ❌ Storing secrets in environment variables on EC2/ECS → ✅ Use AWS Secrets Manager. Reference ARN in task definition, not plaintext values.
- ❌ Lambda for all compute → ✅ Lambda cold starts hurt latency-sensitive APIs. ECS Fargate for always-warm containers.
- ❌ Public subnets for application servers → ✅ Application tier lives in private subnets. Only ALB and NAT Gateway in public subnets.

---

## 1. Service Selection Matrix

### Compute

```
ECS Fargate:
  ✅ Long-running HTTP services (APIs, web apps)
  ✅ Container-based workloads (0 server management)
  ✅ Predictable traffic (no cold start concerns)
  ✅ >100ms response time acceptable
  Cost: ~$0.04048/vCPU/hr + $0.004445/GB/hr

Lambda:
  ✅ Event-driven (S3 trigger, SQS consumer, API Gateway)
  ✅ Infrequent, short-duration tasks (<15 min)
  ✅ True scale-to-zero requirements (cost optimization)
  ❌ Cold starts (100ms–1s for first request)
  ❌ Stateful connections (DB connection pooling — use RDS Proxy)
  Cost: $0.20/1M requests + $0.0000166667/GB-second

EC2:
  ✅ Specialized hardware (GPU, high I/O)
  ✅ Legacy apps that can't containerize
  ✅ Reserved capacity for cost optimization at scale
  ❌ Requires patch management and OS maintenance
  Cost: Varies, reserved instances 40-60% cheaper than on-demand
```

### Database

```
RDS PostgreSQL:
  ✅ ACID transactions required
  ✅ Complex JOIN queries
  ✅ Team knows SQL
  ✅ <10TB data
  Multi-AZ: automatic failover in ~60s
  Read replicas: up to 5 (async replication)

Aurora PostgreSQL:
  ✅ RDS PostgreSQL but need: higher availability, faster failover (<30s)
  ✅ >5 read replicas needed (up to 15 Aurora replicas)
  ✅ Global database (multi-region reads)
  ✅ Aurora Serverless v2 for variable/unpredictable load
  Cost: ~2x RDS, worth it for high-availability requirements

DynamoDB:
  ✅ Single-digit ms latency at any scale
  ✅ Key-value or simple document access patterns
  ✅ Global tables (multi-region active-active)
  ✅ Event-driven (DynamoDB Streams → Lambda)
  ❌ Complex queries (no joins, limited filtering)
  ❌ Requires careful partition key design
  On-Demand pricing: $1.25/1M writes, $0.25/1M reads
```

---

## 2. VPC Architecture (Golden Pattern)

```
┌─────────────────── VPC (10.0.0.0/16) ────────────────────┐
│  ┌─── Public Subnet A ──┐  ┌─── Public Subnet B ──┐       │
│  │  10.0.1.0/24         │  │  10.0.2.0/24         │       │
│  │  - ALB               │  │  - ALB (multi-AZ)    │       │
│  │  - NAT Gateway       │  │  - NAT Gateway       │       │
│  └──────────────────────┘  └──────────────────────┘       │
│  ┌─── Private Subnet A ─┐  ┌─── Private Subnet B ─┐       │
│  │  10.0.11.0/24        │  │  10.0.12.0/24        │       │
│  │  - ECS Tasks         │  │  - ECS Tasks         │       │
│  │  - Lambda (VPC)      │  │  - Lambda (VPC)      │       │
│  └──────────────────────┘  └──────────────────────┘       │
│  ┌─── Data Subnet A ────┐  ┌─── Data Subnet B ────┐       │
│  │  10.0.21.0/24        │  │  10.0.22.0/24        │       │
│  │  - RDS Primary       │  │  - RDS Replica       │       │
│  │  - ElastiCache       │  │  - ElastiCache       │       │
│  └──────────────────────┘  └──────────────────────┘       │
└───────────────────────────────────────────────────────────┘
```

### Terraform: VPC Module

```hcl
# modules/networking/main.tf
module "vpc" {
  source  = "terraform-aws-modules/vpc/aws"
  version = "~> 5.0"

  name = "${var.project}-${var.environment}"
  cidr = var.vpc_cidr   # e.g., "10.0.0.0/16"

  azs              = data.aws_availability_zones.available.names
  public_subnets   = var.public_subnet_cidrs    # ["10.0.1.0/24", "10.0.2.0/24"]
  private_subnets  = var.private_subnet_cidrs   # ["10.0.11.0/24", "10.0.12.0/24"]
  database_subnets = var.database_subnet_cidrs  # ["10.0.21.0/24", "10.0.22.0/24"]

  enable_nat_gateway     = true
  single_nat_gateway     = var.environment == "staging"   # save cost in staging
  enable_dns_hostnames   = true
  enable_dns_support     = true

  # Required for ECS/EKS
  public_subnet_tags = {
    "kubernetes.io/role/elb" = 1
  }
  private_subnet_tags = {
    "kubernetes.io/role/internal-elb" = 1
  }

  tags = local.common_tags
}
```

---

## 3. ECS Fargate Service (Complete Terraform)

```hcl
# modules/ecs/main.tf

# Cluster
resource "aws_ecs_cluster" "main" {
  name = "${var.project}-${var.environment}"

  setting {
    name  = "containerInsights"
    value = "enabled"   # CloudWatch Container Insights
  }

  tags = local.common_tags
}

# Task Definition
resource "aws_ecs_task_definition" "api" {
  family                   = "${var.project}-api"
  requires_compatibilities = ["FARGATE"]
  network_mode             = "awsvpc"
  cpu                      = var.task_cpu     # e.g., 512 (0.5 vCPU)
  memory                   = var.task_memory  # e.g., 1024 (1 GB)
  execution_role_arn       = aws_iam_role.ecs_execution.arn
  task_role_arn            = aws_iam_role.ecs_task.arn

  container_definitions = jsonencode([{
    name  = "api"
    image = "${var.ecr_repository_url}:${var.image_tag}"

    portMappings = [{
      containerPort = var.container_port
      protocol      = "tcp"
    }]

    # ✅ Secrets from Secrets Manager (not plaintext env vars)
    secrets = [
      { name = "DATABASE_URL", valueFrom = aws_secretsmanager_secret.db_url.arn },
      { name = "JWT_SECRET",   valueFrom = aws_secretsmanager_secret.jwt.arn }
    ]

    environment = [
      { name = "NODE_ENV", value = var.environment },
      { name = "PORT",     value = tostring(var.container_port) }
    ]

    logConfiguration = {
      logDriver = "awslogs"
      options = {
        "awslogs-group"         = aws_cloudwatch_log_group.api.name
        "awslogs-region"        = var.aws_region
        "awslogs-stream-prefix" = "api"
      }
    }

    healthCheck = {
      command     = ["CMD-SHELL", "wget --quiet --tries=1 --spider http://localhost:${var.container_port}/health || exit 1"]
      interval    = 30
      timeout     = 5
      retries     = 3
      startPeriod = 10
    }
  }])

  tags = local.common_tags
}

# ECS Service
resource "aws_ecs_service" "api" {
  name            = "${var.project}-api"
  cluster         = aws_ecs_cluster.main.id
  task_definition = aws_ecs_task_definition.api.arn
  desired_count   = var.desired_count
  launch_type     = "FARGATE"

  network_configuration {
    subnets          = module.vpc.private_subnets
    security_groups  = [aws_security_group.api.id]
    assign_public_ip = false    # ✅ private subnets never get public IPs
  }

  load_balancer {
    target_group_arn = aws_lb_target_group.api.arn
    container_name   = "api"
    container_port   = var.container_port
  }

  deployment_configuration {
    maximum_percent         = 200
    minimum_healthy_percent = 50

    deployment_circuit_breaker {
      enable   = true
      rollback = true    # auto-rollback on failure
    }
  }

  depends_on = [aws_lb_listener.https]

  tags = local.common_tags
}
```

---

## 4. IAM Least Privilege

```hcl
# ✅ Execution Role — what ECS needs to START your container
resource "aws_iam_role" "ecs_execution" {
  name = "${var.project}-ecs-execution"

  assume_role_policy = jsonencode({
    Version = "2012-10-17"
    Statement = [{
      Effect    = "Allow"
      Principal = { Service = "ecs-tasks.amazonaws.com" }
      Action    = "sts:AssumeRole"
    }]
  })
}

resource "aws_iam_role_policy_attachment" "ecs_execution_managed" {
  role       = aws_iam_role.ecs_execution.name
  policy_arn = "arn:aws:iam::aws:policy/service-role/AmazonECSTaskExecutionRolePolicy"
}

# Allow pulling secrets from Secrets Manager
resource "aws_iam_role_policy" "ecs_execution_secrets" {
  name = "secrets-access"
  role = aws_iam_role.ecs_execution.id

  policy = jsonencode({
    Version = "2012-10-17"
    Statement = [{
      Effect   = "Allow"
      Action   = ["secretsmanager:GetSecretValue"]
      Resource = [
        aws_secretsmanager_secret.db_url.arn,
        aws_secretsmanager_secret.jwt.arn
      ]
      # ✅ Scoped to exact secret ARNs, not secretsmanager:*
    }]
  })
}

# ✅ Task Role — what your APPLICATION code can do at runtime
resource "aws_iam_role_policy" "ecs_task_app" {
  name = "app-permissions"
  role = aws_iam_role.ecs_task.id

  policy = jsonencode({
    Version = "2012-10-17"
    Statement = [
      {
        Effect   = "Allow"
        Action   = ["s3:GetObject", "s3:PutObject"]
        Resource = "${aws_s3_bucket.uploads.arn}/*"
        # ✅ Specific bucket, specific actions — not s3:*
      },
      {
        Effect   = "Allow"
        Action   = ["ses:SendEmail"]
        Resource = "arn:aws:ses:${var.aws_region}:${data.aws_caller_identity.current.account_id}:identity/*"
      }
    ]
  })
}
```

---

## 5. CloudWatch Observability

```hcl
# Structured JSON logging — filter patterns work
resource "aws_cloudwatch_log_group" "api" {
  name              = "/ecs/${var.project}-${var.environment}/api"
  retention_in_days = 30

  tags = local.common_tags
}

# Metric filter — count 5xx errors from structured logs
resource "aws_cloudwatch_log_metric_filter" "errors_5xx" {
  name           = "${var.project}-5xx-errors"
  pattern        = "{ $.status >= 500 }"
  log_group_name = aws_cloudwatch_log_group.api.name

  metric_transformation {
    name      = "5xxErrorCount"
    namespace = "${var.project}/API"
    value     = "1"
    unit      = "Count"
  }
}

# Alarm → SNS → PagerDuty
resource "aws_cloudwatch_metric_alarm" "high_error_rate" {
  alarm_name          = "${var.project}-high-5xx-${var.environment}"
  comparison_operator = "GreaterThanThreshold"
  evaluation_periods  = "2"
  metric_name         = "5xxErrorCount"
  namespace           = "${var.project}/API"
  period              = "60"
  statistic           = "Sum"
  threshold           = "10"    # more than 10 errors in 2 minutes

  alarm_actions = [aws_sns_topic.alerts.arn]
  ok_actions    = [aws_sns_topic.alerts.arn]

  treat_missing_data = "notBreaching"

  tags = local.common_tags
}
```

---

## 6. Multi-Environment Strategy

```hcl
# environments/staging/main.tf
module "app" {
  source = "../../modules/app"

  environment     = "staging"
  desired_count   = 1           # single task in staging
  task_cpu        = 256         # 0.25 vCPU
  task_memory     = 512
  single_nat_gateway = true     # save $32/month vs multi-NAT
}

# environments/production/main.tf
module "app" {
  source = "../../modules/app"

  environment     = "production"
  desired_count   = 3           # 3 tasks across 2 AZs
  task_cpu        = 1024        # 1 vCPU
  task_memory     = 2048
  single_nat_gateway = false    # high availability
}
```

### Terraform State Backend

```hcl
# backend.tf
terraform {
  backend "s3" {
    bucket         = "mycompany-terraform-state"
    key            = "production/app.tfstate"
    region         = "us-east-1"
    encrypt        = true       # AES-256 at rest
    dynamodb_table = "terraform-locks"    # prevents concurrent state writes
  }
}
```

## 🚨 Edge-Case & Failure Mode Matrix

| Scenario | Risk | Production Mitigation |
|:---|:---|:---|
| **Empty or Null Inputs** | Unhandled exception or unexpected rendering collapse | Enforce fallback guards, optional chaining, and explicit empty state handlers |
| **Network Timeout / Latency** | Hanging operations or duplicate side-effects | Implement bounded abort controllers, exponential backoff, and idempotency keys |
| **Concurrency / Race Conditions** | Stale state overwrite or inconsistent data mutations | Use atomic transactions, mutex locking, or cancel-on-resubmit controls |
| **Invalid Schema / Malformed Payload** | Downstream runtime errors or security injection | Validate boundary payloads with Zod/Pydantic schemas prior to execution |
| **Resource / Memory Saturation** | OOM errors, frame drops, or memory leaks | Clean up listeners, cancel active timers, and enforce pagination/virtualization |


## 🏛️ Tribunal Verification & Guardrails

**Active Reviewers:** `pipeline-reviewer` · `devops-engineer` · `resilience-reviewer`
**Slash Command:** `/review` or `/tribunal-full`

### 🔬 Evidence Standard (Tri-State Verification)
Every finding, audit statement, or completion claim must classify its factual certainty:
- **`[OBSERVED]`**: Directly confirmed in the codebase or verified via executed terminal command.
- **`[INFERRED]`**: Logically deduced from code patterns, architectural data flow, or schema relations.
- **`[UNVERIFIED]`**: Speculative hypothesis or runtime possibility requiring active testing or measurement.

### ✅ Pre-Flight Self-Audit Checklist
```
✅ Are strict execution modes (set -euo pipefail) active on all scripts?
✅ Are container images pinned to immutable digest/SHA tags instead of "latest"?
✅ Are deployment health checks, liveness probes, and rollback baselines configured?
✅ Are CI secrets masked and unexposed to untrusted pull requests?
✅ Did I verify environment compatibility across target runtimes?
```

### 🛑 Verification-Before-Completion (VBC) Protocol
**CRITICAL:** You must follow a strict "evidence-based closeout" state machine.
- ❌ **Forbidden:** Declaring a task complete because the output "looks correct."
- ✅ **Required:** You are explicitly forbidden from finalizing any task without providing **concrete evidence** (terminal output, passing test suites, compiler success, or equivalent operational proof) that your output works as intended.

Attribution

Harmitx7Harmitx7
View sourceSee grades on GitHubMore from Harmitx7 →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Terse caveman voice: answer first, fluff gone, every technical fact kept. Use for /caveman, "caveman mode", "talk like caveman", "be brief", "less tokens". Stays on until "stop caveman" or "normal mode".

1100021 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

698461 votes

Writing Skills

Create and manage Claude Code skills in HASH repository following Anthropic best practices. Use when creating new skills, modifying skill-rules.json, understanding trigger patterns, working with hooks, debugging skill activation, or implementing progressive disclosure. Covers skill structure, YAML frontmatter, trigger types (keywords, intent patterns), UserPromptSubmit hook, and the 500-line rule. Includes validation and debugging with SKILL_DEBUG. Examples include rust-error-stack, cargo-dep...

3931 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3421 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Amp, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Grok Build, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

741 votes
View all in ai-agents →