Defines the features of interest (concerns) a Sokrates analysis tracks across the codebase by path or content regex: debt markers, deprecated code, feature flags, security-sensitive code, unsafe execution, error handling, concurrency, persistence, network, telemetry, configuration access, platform-specific code, integration libraries, domain terms, with a proposal script that measures each candidate with hits and sample lines. Use when Sokrates should track where X is or how much code touches...
Installs into .claude/skills of the current project.
Are you the author of Sokrates Features Of Interest?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/zeljkoobrenovic-sokrates-features-of-interest)
---
name: sokrates-features-of-interest
description: Defines the features of interest (concerns) a Sokrates analysis tracks across the codebase by path or content regex: debt markers, deprecated code, feature flags, security-sensitive code, unsafe execution, error handling, concurrency, persistence, network, telemetry, configuration access, platform-specific code, integration libraries, domain terms, with a proposal script that measures each candidate with hits and sample lines. Use when Sokrates should track where X is or how much code touches Y, or the concern views should mean more than TODOs.
---
# Sokrates features of interest (concerns)
Sokrates "features of interest" are **concerns**: named, regex-defined slices of the main code that cut across components — every file whose path or content matches. Each concern gets its own size, its distribution over components and extensions, its overlap with other concerns, and a trend when analyses are compared over time. They are the mechanism for questions like *how much of the code handles secrets*, *where are the feature flags*, *how big is the legacy compatibility layer*, *which components talk to the network*. The default configuration tracks only `TODOs`; this skill defines the concerns that matter for a given codebase and verifies them on the real tree.
Field semantics (`concernGroups`, `Concern`, `MetaRule`, regex rules) are in `../sokrates-repo-config/references/config-reference.md`. The essentials: a concern classifies **whole files** — a content match anywhere puts the file in the concern, so a concern measures "files that contain X" (and their LOC), never "lines of X"; inline test modules inside main files (Rust `#[cfg(test)]`, Go `_test` in the same package) cannot be excluded — accept it or narrow the pattern to production-only idioms; a concern is a `NamedSourceCodeAspect` (`sourceFileFilters` + optional `files`), matched **against main files only**; `contentPattern` must match an **entire line** (write `.*X.*`), `pathPattern` the entire path (write `.*/foo/.*`); a malformed regex silently matches nothing; `exception: true` vetoes; `metaConcerns` derive concern names from the matched line (`use: "content"`) or the path with `nameOperations` — and **`extract` keeps the whole first regex match, not a capture group**: isolate the value with an `extract` that matches only the value's neighbourhood, then `replace` the surrounding text away.
## Workflow
1. **Measure the candidates:**
```bash
python3 <this-skill-path>/scripts/propose_concerns.py <repo>/_sokrates/config.json -o <scratch>/concerns.json
```
Against the config's main scope it evaluates a catalog of ~30 generic concerns (technical debt, security, robustness, data, integration, observability, configuration, lifecycle, tests-in-main) and the repository's most imported external libraries (imports inside Rust `#[cfg(test)]` regions discounted) as integration candidates, reporting per candidate: files matched, % of main files, LOC of those files (`loc_pct` = LOC of matched files / main LOC — the same "% LOC" the report table uses), matching lines, the share of hits inside test regions, the **most frequent matched tokens** (what the regex actually caught — the fastest way to see false positives and to discover the codebase's own names, e.g. `Feature::UnifiedExec`), spread over depth-2 folders, and stratified `file:line` samples. It also re-measures the existing concerns and proposes a replacement when the default `TODOs` pattern misses this codebase's style (`TODO(name)` is not matched by `(TODO|FIXME)( |:|\t)`). `suggested_concernGroups` is a paste-ready starting point — never paste it whole.
2. **Choose what is *interesting* here.** A concern earns its place when someone would act on its number or its map. Keep 8–15 concerns in 3–5 groups. Selection rules:
- **Drop the obvious**: a candidate touching more than ~60 % of files (serialization in a Rust codebase, logging in a service) is a property of the codebase, not a feature of interest — narrow it (e.g. only `eprintln!`/`console.log` debug prints, only `unwrap()` in non-test code) or drop it.
- **Keep what is actionable**: debt markers, deprecated/legacy/compat code, suppressed warnings, swallowed errors and panics, secrets handling, unsafe/dynamic execution, sandbox/privilege code, external services, platform-specific branches, feature flags, experimental areas. These are the ones people want to see shrink or stay contained.
- **Language idioms are not concerns**: Rust inline `#[cfg(test)]` modules, `unwrap()` in tests, `serde` derives, tokio everywhere — the catalog flags them (`test code inside main scope`, serialization, concurrency) so you can decide; in idiomatic cases drop them rather than "fixing the scope".
- **Add the domain**: the most valuable concerns are repository-specific — a framework being migrated away from (`.*\bjavax\.servlet\b.*`), a protocol version (`.*\bv1_api\b.*`), a business capability's vocabulary (`.*\b(invoice|billing|payment)\w*\b.*`), a client library (`uses tokio`), a compliance topic (PII fields, audit logging). Take the vocabulary from README/docs, and from prior findings in `_sokrates/reports/ai-insights/` when they exist — mine `title`, `group` and `evidence[].snippet` of `security-design-scan` (escape hatches, trust boundaries), `architecture-scan` (migrations, deprecated paths), `domain-language-scan` (glossary, language drift), `evolution-scan` (emerging/dying areas) — and from the library and token lists the script prints. Concerns that make a prior finding *measurable over time* are the best ones.
- **Prefer content over path** for cross-cutting concerns (path-based ones duplicate the component view), but use paths for things that live in folders (`.*/migrations/.*`, `.*/legacy/.*`).
3. **Sharpen each regex** by reading the samples: false positives (the word `secret` in a comment about secrets, `timeout` in a variable name) are cheap to remove with word boundaries, language-specific tokens (`#\[cfg\(`, `@Deprecated`), or an `exception` filter on test-like paths; false negatives are found by grepping the tree for the concept's other spellings. Keep the `.*…*` wrapping. Give every concern a one-line `note` saying what question it answers.
4. **Group and name.** Groups are report sections: `technical debt`, `security`, `robustness`, `integration`, `platform`, `domain: <name>`. Concern names are labels people will read in charts — short noun phrases (`feature flags`, `swallowed errors`, `uses tokio`), not regex descriptions. Stable names keep trends comparable across runs; rename only deliberately.
5. **Verify** with the repo-config preview (it reports every concern's file count and warns on zero hits and regex errors):
```bash
python3 <config-skills-path>/sokrates-repo-config/scripts/preview_config.py <repo>/_sokrates/config.json
```
A concern with 0 files is either a wrong pattern or a fact worth keeping (a "this codebase has no eval" guard) — decide, don't leave it by accident. The preview also flags concerns above 60 % of main files as too broad, and prints each meta concern's derived names. (On macOS use `awk 'NR%k==1'` instead of `shuf` when sampling hits by hand.)
6. **Write and report.** Replace `concernGroups` in `_sokrates/config.json` (keep `TODOs` unless it is genuinely empty; set `analysis.analyzeConcernOverlaps: true` when overlaps between concerns are part of the question, e.g. "how much secret handling is in platform-specific code"). Report the table of concerns (group, name, files, % LOC, what question it answers), the candidates you rejected as too broad or uninteresting, and the next command (`sokrates generateReports`).
## Patterns
- **Meta concerns** for open-ended vocabularies: one rule that turns the matched line into a concern name — e.g. every `#[cfg(target_os = "…")]` value as its own concern:
`{"pathPattern": "", "contentPattern": ".*cfg\\(target_os = \"[a-z]+\"\\).*", "use": "content", "nameOperations": [{"op": "extract", "params": ["target_os = \"[a-z]+\""]}, {"op": "replace", "params": ["target_os = \"", ""]}, {"op": "replace", "params": ["\"", ""]}, {"op": "prepend", "params": ["os: "]}]}`
(`extract` returns the whole match, so it must match only the value's neighbourhood; the `replace` steps strip the rest). The preview evaluates meta concerns and lists the derived names with counts — check that the names are values, not whole lines.
- **Narrowing by scope**: concerns only see `main`; a `test code inside main scope` hit means the *scope* config is wrong — fix it with `sokrates-repo-config`, don't keep it as a concern.
- **Trend use**: concerns are the cheapest way to make a migration or clean-up measurable — define `legacy: <old thing>` and `new: <replacement>` and watch the two curves across analyses.
- **Cost**: every content concern scans every main file line; a dozen is fine, a hundred is slow and unreadable.