Skip to content
Back to skills

Iac Scan

DSecurity

Reads a codebase's infrastructure as code: the IaC inventory and tooling (Terraform, Pulumi, CloudFormation, CDK, Kubernetes, Helm, Kustomize, Ansible, Compose, serverless), container and dev-environment definitions, declared resources and topology, environment variants, state and drift, and the hardening of what is declared (pinning, root, privileges, exposure, limits). Use for IaC, Terraform, Kubernetes, Dockerfiles, deployment topology, cloud resources in the repo, or an IaC review; report...

  • 3 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added October 2, 2026
developmentpythonrustgobashnodedockerkubernetesawsgcpterraform

Works with

  • cli
  • api

Security analysis

D50/100
  • criticalPipes output to a shell interpreter
  • criticalDownloads and executes remote scripts — classic supply chain attack

Pro scans all 2 files and shows the line behind each finding

Scanned October 3, 2026

npx -y skills add zeljkoobrenovic/sokrates-skills --skill iac-scan --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Iac Scan?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Iac Scan
[![Security: D — Skills Directory](https://www.skillsdirectory.com/api/skills/zeljkoobrenovic-iac-scan/badge)](https://www.skillsdirectory.com/skills/zeljkoobrenovic-iac-scan)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: iac-scan
description: Reads a codebase's infrastructure as code: the IaC inventory and tooling (Terraform, Pulumi, CloudFormation, CDK, Kubernetes, Helm, Kustomize, Ansible, Compose, serverless), container and dev-environment definitions, declared resources and topology, environment variants, state and drift, and the hardening of what is declared (pinning, root, privileges, exposure, limits). Use for IaC, Terraform, Kubernetes, Dockerfiles, deployment topology, cloud resources in the repo, or an IaC review; reports honestly when there is none.
---

# Infrastructure-as-code scan

Application code says what the software does; infrastructure code says what it needs in order to run at all — which machines, which images, which network, which permissions, which limits. That second body of code is often written by different people, reviewed less, and yet decides most of what can go wrong in production. This scanner reads it as code: what it declares, how environments differ, where the truth lives, and where the declared infrastructure is weaker than the application it hosts.

**First read `sokrates-scan-core/SKILL.md`** (sibling skill) — output format, evidence rules, validate/render scripts, `_sokrates` layout. This file adds only what is specific to IaC scanning.

## The one question

*If this repository were the only thing left, what environment could be rebuilt from it, and what would have to be recreated from memory?* Ask it per environment, not per file.

## Scope and boundaries with sibling scanners

- **`tech-stack-scan`** (`infrastructure`) lists the IaC *technologies* — Terraform, Kubernetes, Docker, nix. Read it first; do not repeat the inventory. This scanner reads what those files **declare** (resources, images, users, limits, exposure) and judges it.
- **`cicd-scan`** (`deployment`, `release`) owns *when and by whom* infrastructure is applied: the workflow that runs `terraform apply`, the job that builds and pushes the image, the promotion path between environments. This scanner owns what the applied definitions **contain**. A `terraform apply` step is cicd's; the resources that apply creates are this scanner's; reference cicd's finding for the trigger rather than restating it. When a pipeline applies IaC by hand-rolled steps, that is a `state-and-drift/apply-path` finding here, evidenced by the nearest positive fact. A CI file that *mutates pre-existing infrastructure by name* (`aws s3 cp` to a named bucket, a CloudFront invalidation by distribution id, a deploy to a host the repo never declares) belongs to `cicd-scan` as a pipeline step — but it is this scanner's **best evidence for a coverage gap**, because every resource it names and the repo never declares is a resource nobody could rebuild. Cite those lines under `inventory/coverage`.
- **Containers are this scanner's content** (settled scope): Dockerfiles, `docker-compose*`, `.devcontainer/`, `flake.nix` and other environment definitions are read here for base images and pinning, build stages, the user the process runs as, build-arg and mount surfaces, exposed ports, healthchecks and entrypoints. `tech-stack-scan` keeps the "this project uses Docker" line and `cicd-scan` keeps the build and publish job; this scanner keeps everything inside the file.
- **`network-scan`** owns the application's own connectivity (which endpoints the process opens and calls, timeouts, TLS in client code). This scanner owns *declared* network topology — service definitions, ingress, published ports, security groups, network policies — and references network-scan where the same port appears on both sides.
- **`security-scan`** owns secrets in the tree and application-level trust. Hardening of declared infrastructure (privileged containers, `runAsRoot`, world-open security groups, public buckets, unpinned base images) is **this scanner's** — it is a property of the declaration. A credential *value* found inside an IaC file is security's (`secrets/in-tree`): report the mechanism here (`configuration/secret-injection`-style refs) and cross-reference rather than re-rating the leak.
- **`configuration-scan`** owns how the *application* reads its configuration (sources, precedence, defaults, validation). The values that infrastructure *supplies* — container `env:` blocks, Helm values, Terraform variables feeding the app — are declared here and referenced there; describe the wiring once, in whichever scanner reads the file, and let the other reference it. Environment *variants* of infrastructure (staging vs prod overlays) are this scanner's; environment variants of application settings are configuration's.
- **`reliability-scan`** owns application-level resilience; declared restart policies, replica counts, probes and resource limits are this scanner's, referenced from reliability when it already rated the failure mode.
- **`storage-scan`** owns data at rest in the application; declared volumes, persistent claims and bucket resources are this scanner's topology, referenced to storage's data classes.

**Cross-referencing.** List existing ids first (`grep -h '"id"' _sokrates/reports/ai-insights/*.json --exclude=combined-report.json`) and reference siblings as `sokrates_refs: ["finding:<scanner>/<group>/<slug>"]` — only ids you saw. Never copy another scanner's evidence blocks.

## Workflow

1. **Orient per the core skill.** From `tech-stack-scan` (`infrastructure`, `ci-cd`), `cicd-scan` (`deployment`, `release`), `architecture-scan` (which components exist, so declared services can be matched to them), `network-scan` (endpoints), `configuration-scan` (config sources, if it ran). Sokrates ignore rules skip vendored modules — but note that `.github/`, `deploy/` and `infra/` folders are often excluded from the Sokrates file index, so **always glob the filesystem directly** rather than trusting the index.
2. **Count with the script, then read the definitions.** Run
   ```bash
   python3 <this-skill-path>/scripts/count_iac_sites.py <src-root> --json <scratch>/iac-counts.json
   ```
   It classifies every infrastructure file by tool (Terraform, Kubernetes, Helm, Kustomize, CloudFormation, CDK/Pulumi, Ansible, Docker, Compose, devcontainer, nix, serverless, Vagrant, systemd, Bazel-container rules), counts declared resources per provider and Kubernetes kinds, and counts hardening shapes: unpinned base images and `:latest` tags, `USER root`/absent `USER`, `privileged`, `hostNetwork`/`hostPath`, missing resource limits, `0.0.0.0/0` and public-access properties, plaintext-looking values, build args and mounts, provider and module version pins, backend/state declarations. It also prints a **per-Dockerfile stage record** (each stage's base image, user and entrypoint, plus `final_stage_user` — the one that actually runs) and the **build context**: whether a `.dockerignore` exists, the context size, and the largest files an unfiltered `COPY . .` would pull in. Facts go into `stats`; `*_candidates`/`*_keyword_files` are reading lists only.
3. **Decide the system kind and its live slots.** Not every project has provisioning IaC, and inventing findings for absent tooling is worse than saying so:
   - **No provisioning IaC, containers only** (a CLI tool, a library with a devcontainer): slots `inventory/iac-inventory`, `inventory/coverage`, `resources/container-image` (or one `container-images` finding when the images are variations of one), `environments/dev-environment`, `hardening/*` for what the images do declare, `state-and-drift/apply-path` **only when something is applied by hand or by a pipeline** — when there is no provisioning code at all the absence of an apply path is an entailment, not a finding, and belongs in one line of `coverage` and the posture instead, `iac-posture/posture`. This is the expected shape for most repositories — write it thinly and honestly, do not pad.
   - **Kubernetes application**: `inventory/iac-inventory`, `inventory/coverage`, `resources/workloads`, `resources/services`, `resources/ingress`, `resources/storage`, `environments/environment-matrix`, `hardening/*`, `state-and-drift/apply-path`, `iac-posture/posture`.
   - **Cloud infrastructure repo** (Terraform/CDK/Pulumi): `inventory/iac-inventory`, `inventory/coverage`, `resources/<provider>-<family>` per resource family, `resources/modules-and-reuse`, `environments/environment-matrix` (plus `environments/<env>-overrides` where one environment's deltas deserve a verdict), `state-and-drift/state-backend`, `state-and-drift/apply-path`, `state-and-drift/version-pinning`, `state-and-drift/drift-accommodations`, `hardening/*`, `iac-posture/posture`.
   - **Compose-based service**: `inventory/iac-inventory`, `inventory/coverage`, `resources/services`, `resources/container-image`, `environments/environment-matrix`, `hardening/*`, `iac-posture/posture`.
4. **Read the environment definitions completely.** For each Dockerfile: base images and how pinned (tag, digest, `FROM ... AS` stages), what is installed, the user the final stage runs as, `COPY` of secrets or the whole context, `ARG`s that carry credentials, `EXPOSE`, `HEALTHCHECK`, `ENTRYPOINT`/`CMD`. For compose: services, images, ports published to the host, volumes and bind mounts, `env_file`, restart policies, dependencies. For devcontainer/nix: what the developer environment pins and what it grants (docker socket mounts, privileged flags).
5. **Read the provisioning code by resource family.** Group declarations by what they create — compute, network, storage, identity, secrets, observability — not by file. For each family: which resources, in which environments, with what notable properties. Follow module calls one level (their source and inputs); note where you stopped. For Kubernetes: workloads (replicas, images, probes, limits, security contexts), services and ingress (types, ports, TLS), config and secret objects, RBAC, network policies.
6. **Read environment and variant handling.** How staging differs from production: workspaces, `*.tfvars`, Kustomize overlays, Helm values files per environment, compose overrides, branch-per-environment. The finding worth writing is the *diff*: which properties actually differ (sizes, replica counts, domains, flags) and which are accidentally shared. Note where a variant exists in one dimension but not another (a prod values file with no staging counterpart).
7. **Read state, drift and the apply path.** Where Terraform/Pulumi state lives (backend, locking, encryption); whether it is committed to the repo (a serious finding — state files carry resource attributes and often secrets); what is declared vs what the docs say is created by hand; `lifecycle`/`ignore_changes` blocks and `kubectl apply` scripts that bypass the pipeline; version pinning of providers, modules, charts and actions the IaC pulls.
8. **Judge hardening as declared.** Per the severity table below: image pinning, root and privilege, host access, exposed surfaces, resource limits, public network exposure, encryption-at-rest and logging flags where the provider offers them. Read the file before rating — a `privileged: true` in a devcontainer used by one developer is not a production risk, and the script's counts are candidates, never findings.
9. **Synthesize the posture.** One `iac-posture/posture` finding, `severity: info`, `confidence: likely`: what could be rebuilt from the repository and what could not, which environments are covered, the strongest and weakest declared property, and the three highest-leverage changes — named in the *description* prose with `finding:` refs in `sokrates_refs` (an `info` finding carries no `recommendation`, per the core skill). Evidence cites the root module, the main manifest or the primary Dockerfile.
10. Write findings, validate, render; re-run the merge script if a `combined-report.json` exists. Report per the core workflow. Scanner id: `iac-scan`, version `1.1`.

## Group taxonomy

| group | contents |
|---|---|
| `inventory` | What infrastructure code exists. `iac-inventory` is the **file census** (tools in use, where the definitions live, how many files of each); `coverage` is the **gap statement** (what the runtime environment needs that nothing here declares, and who would rebuild it). Keep the census in one and the judgment in the other — never repeat the file list in `coverage` |
| `resources` | What is declared, by family: container images, workloads and services, network and ingress, storage and volumes, identity and roles, managed services — with the properties that matter |
| `environments` | Environments and variants: how staging/prod/dev differ, workspaces, overlays, values files, compose overrides, the developer environment |
| `state-and-drift` | Where state lives and how it is protected, the apply path (pipeline or by hand), version pinning of providers/modules/charts, `ignore_changes` and other drift accommodations, what is known to be managed outside the code |
| `hardening` | Security and robustness properties of the declaration itself: image pinning, users and privileges, host access, exposed ports, public access, resource limits, probes and restart policies, encryption and logging flags |
| `iac-posture` | The synthesis (one finding, `info`, id `iac-posture/posture`) |

**Precedence**: a property of a *declared* resource → `resources` when it defines the resource, `hardening` when it is a security or robustness judgment on it (an image's tag is `resources/container-image`, the fact that the tag is mutable is `hardening/image-pinning`); a difference *between* environments → `environments`, even when the difference is a hardening property (say which environment is weaker there and reference `hardening`); a version pin of the IaC's own dependencies (providers, modules, charts, base images consumed by name) → `state-and-drift/version-pinning`, except container base images which stay in `hardening/image-pinning`; a secret *value* in an IaC file → reference `security-scan/secrets/in-tree`, only the injection mechanism is described here; an application setting supplied by a manifest → reference `configuration-scan`, describing here only that the manifest supplies it.

## Stable ids

Slugs are **tool, resource-family or mechanism names, never consequences**. Fixed slugs (use only when the subject exists; parametrised slugs take the kebab-case name the tool itself uses — a provider name, a chart name, an environment name):

| group | fixed slugs |
|---|---|
| `inventory` | `iac-inventory` (the one overview: tools, locations, file counts), `coverage` (what the code provisions and what it visibly does not, including infrastructure that CI mutates by name but nothing declares, and whether anything builds or publishes the images the repo defines — the honest absence finding; on a repo with no provisioning IaC this is the load-bearing finding) |
| `resources` | `container-image` (one per image the project *produces or runs*; a multi-stage build is **one** finding, not one per `FROM` — the builder stage is described inside it; `container-images` as a single finding when several produced images are variations of one), `workloads`, `services`, `ingress`, `storage`, `identity`, `managed-services`, `<provider>-<family>` for a cloud family worth its own finding (`aws-networking`, `gcp-iam`), `modules-and-reuse` (module/chart composition and where modules come from) |
| `environments` | `environment-matrix` (the one finding naming the environments and how they are selected), `<env>-overrides` (only when one environment's deltas deserve their own verdict), `dev-environment` (one finding per *distinct* developer-environment mechanism, not one covering all of them: a devcontainer, a Nix flake and a compose-for-development file pin different things and get `dev-environment`, `dev-environment-nix`, `dev-environment-compose`; a single mechanism keeps the bare slug — also used for a container the docs tell developers to build and run locally), `promotion` (only when the IaC itself encodes promotion; the pipeline's promotion is cicd's) |
| `state-and-drift` | `state-backend`, `apply-path` (pipeline, by hand, or absent), `version-pinning` (providers, modules, charts together), `drift-accommodations` (`ignore_changes`, manual-only resources, documented hand-made infrastructure) |
| `hardening` | `image-pinning`, `container-user`, `privileges` (privileged, capabilities, host namespaces, docker-socket mounts), `exposed-surfaces` (published ports, node ports, ingress without TLS), `public-access` (world-open security groups, public buckets, `0.0.0.0/0`), `resource-limits`, `resilience-settings` (replicas, probes, restart policies — referencing reliability), `build-provenance` (what a build pulls in from outside and whether it is verified: `curl … | sh` installers, `ADD https://…`, features or plugins on mutable tags, packages from a rolling channel), `build-secrets` (build args, mounted credentials, whole-context copies — a `COPY . .` is only a finding when nothing filters the context, so pair the script's `copy_whole_context` count with its `build_context` record: no `.dockerignore` plus a large context is the finding, and name the biggest files it drags in), `encryption-and-logging` (provider flags for at-rest encryption, audit/access logs) |
| `iac-posture` | `posture` |

Project-specific findings get a free slug naming the artifact or mechanism (`resources/gpu-node-pool`, `state-and-drift/committed-tfstate`), never the consequence. Several mechanisms sharing a slug are listed in `attributes`.

## What a good finding looks like

Evidence is the `FROM` line, the `USER` line, the `image:` line, the `resource "aws_s3_bucket"` header, the `cidr_blocks` value, the `backend "s3"` block, the values-file key that differs between environments, the `EXPOSE`/`ports:` line. Descriptions speak per resource family or per environment: "the API workload runs three replicas from a mutable `:latest` tag with no CPU or memory limits and no liveness probe; the same manifest sets a read-only root filesystem". One finding per family or mechanism, not per file — twelve manifests that together define one service are one `resources/workloads` finding.

Expect 12–18 findings for a real infrastructure repository, 8–14 for a Kubernetes application, 5–9 for a compose-based service, and **6–9 for a repository with no provisioning IaC** (`iac-inventory`, `coverage`, the container/dev-environment findings, one hardening finding per mechanism actually present — typically one to three — and `posture`) — thin is the correct output there, not a reason to invent. Never merge two hardening mechanisms into one finding to come in under a range: the range follows the slot list, not the other way round. Roughly half `info`. Mandatory slots: `iac-inventory`, `coverage`, `posture`.

Vocabulary to grep once the files are located: Terraform `resource|module|variable|backend|provider|lifecycle|ignore_changes|for_each`; Kubernetes `kind:|replicas:|securityContext|runAsUser|privileged|hostPath|resources:|limits:|livenessProbe|imagePullPolicy|NodePort|LoadBalancer`; Helm `values.yaml|{{ .Values`; Docker `FROM|USER|ARG|EXPOSE|HEALTHCHECK|--mount=type=secret|ENTRYPOINT`; compose `services:|ports:|volumes:|env_file:|privileged:|network_mode:`; cloud openness `0.0.0.0/0|::/0|public-read|PublicAccessBlock|allUsers`.

## Severity calibration

- `critical` — a state file or credential material committed in the repository that grants infrastructure access; a declared resource that is publicly writable (a bucket with `allUsers` write, a database open to `0.0.0.0/0` with a default password in the same file).
- `high` — production compute reachable from the whole internet with no gate declared (management ports open to `0.0.0.0/0`, a database with a public IP and no network restriction); privileged or host-namespace containers in a production workload; a production workload built from an unpinned mutable base image where the pipeline also auto-deploys; no apply path at all for infrastructure **this repository is the owner of** (nothing here can recreate what this team runs). A CLI, a library or a desktop app that merely *consumes* someone else's hosted infrastructure — a download bucket, a registry, an API — is not its owner: that is `info` in `coverage`, naming what could not be rebuilt and who would rebuild it. Ask "would this team be the one to recreate it?" before rating above `info`.
- `medium` — containers running as root in a deployed workload; no resource limits on a shared cluster; ingress without TLS; state backend without locking or encryption; providers, modules or charts pinned to mutable ranges; an environment that silently inherits production values; secrets passed as build args or plain `env:` values (reference security for the value itself); `ignore_changes` on security-relevant attributes.
- `low` — mutable base-image tags in developer or CI-only images; a devcontainer with broad privileges; missing probes or restart policies on non-critical workloads; duplicated manifest blocks that will drift; an environment matrix documented only in prose. Drop to `info` when **nothing publishes or deploys the artifact** — a loose property on an image only ever built locally is a fact about the image, not a risk; say in the description what would raise it ("if this image were published, the mutable tag would become `low`").
- `info` — the inventory, the resource map, the environment matrix, the coverage statement, hardening properties that are correctly set, and the honest "no provisioning IaC exists here" verdict.
- Mitigations lower a finding one rung: the environment is developer-only, the resource is behind a declared network policy, the image is rebuilt and redeployed on every commit, the cluster is single-tenant. Raise nothing on inference alone — an unpinned image in a repository with no deployment path is `low`, not `high`.

## Output

Follow the core workflow: write `_sokrates/reports/ai-insights/iac-scan.json`, validate until OK, render the explorer, re-merge if a combined report exists, report leading with the coverage sentence (what can be rebuilt from the repo and what cannot) and any above-info findings.

`stats` — the script's **facts** are counts of shapes, not verified mechanisms: read the sites behind any key a finding will rest on before copying it. Copy the keys you verified, name every key you dropped or corrected in `count_notes` with the true number, add `count_rule`, and on top the fields below. Zeros are facts and belong in `stats` — *but* when the target has no provisioning IaC, keep only the zero keys that carry the coverage message (`terraform_resources`, `k8s_documents`) and drop the rest of the empty provisioning block into one `count_notes` line ("all Terraform, Kubernetes, Helm and cloud-resource counts are 0"), so the fields that do have content are readable. Name any key you dropped as a false positive in `count_notes`.

- `iac_tools` — list of tools actually in use, e.g. `["Dockerfile", "docker-compose", "devcontainer"]`; `[]` when none
- `provisioning_iac` — `true`/`false`: does the repository contain code that creates infrastructure (as opposed to describing a container)
- `environments` — list of environment names the code distinguishes, `[]` when it does not distinguish any
- `resource_families` — list, e.g. `["container images", "k8s workloads", "ingress", "object storage"]`
- `state_backend` — short string or `none`
- `apply_path` — one of `pipeline`, `manual`, `absent`, `mixed`
- `images` — object per image: `{"<name>": "pinned-digest|pinned-tag|mutable-tag|unpinned"}`
- `runs_as_root` — object per image or workload: `true`/`false`/`unknown`
- `coverage_gaps` — list of runtime elements the repository visibly does not declare (the cloud account, the DNS zone, the database instance)

Files in this skill

  • SKILL.md24 KB
  • scripts/count_iac_sites.py24.6 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…