Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Clinical Notes Nlp

ASecurity

NLP on clinical text — de-identification, concept extraction, phenotyping from notes, and LLM safety guardrails.

2 stars
0 votes
0 copies
0 views
Added 9/29/2026
ai-agentspythongorailstestingapidocumentation

Works with

cliapi

Security Analysis

A100/100

Scanned 9/29/2026

$npx -y skills add aicodedecode/awesome-muse-skills --skill clinical-notes-nlp --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Clinical Notes Nlp?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Clinical Notes Nlp
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/aicodedecode-clinical-notes-nlp/badge)](https://www.skillsdirectory.com/skills/aicodedecode-clinical-notes-nlp)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: clinical-notes-nlp
description: NLP on clinical text — de-identification, concept extraction, phenotyping from notes, and LLM safety guardrails.
category: scientific
---

## Overview

clinical-notes-nlp covers natural language processing of clinical documentation: discharge
summaries, progress notes, radiology reports, nursing notes. It spans de-identification,
medical concept extraction (UMLS/cTAKES/MedCAT), note-based phenotyping, and the special
caution required when applying large language models to clinical text — where hallucinations
can become chart entries.

Clinical notes are the richest and messiest EHR data: telegraphic, abbreviation-dense,
copy-pasted, and full of implicit context. Standard NLP pipelines underperform until adapted.

## When to use

- De-identifying clinical text for research (HIPAA Safe Harbor, surrogate replacement).
- Extracting concepts: diagnoses, medications, procedures, symptoms with negation/temporality.
- Note-based phenotyping: combining notes with structured data for case definitions.
- Information extraction: tumor staging from pathology reports, ejection fraction from echo
  reports, social determinants from notes.
- Evaluating LLMs for summarization or documentation assistance with safety guardrails.
- Measuring documentation quality: copy-paste detection, note bloat.

## Core concepts

- **De-identification.** HIPAA Safe Harbor: remove 18 identifier categories (names, dates
  except year, locations, contact info, MRNs...). Automated tools (Philter, de-identification
  modules) miss edge cases — validate on samples, especially dates in narrative text and
  clinician names in signatures. Surrogate replacement (realistic fake names/dates) preserves
  readability better than redaction tags.
- **Clinical sublanguage.** Notes use heavy abbreviation ("pt c/o SOB, r/o PE"), telegraphic
  syntax, and institution-specific shorthand. Off-the-shelf NLP models trained on news/wikipedia
  fail; use clinical models (ClinicalBERT, GatorTron) or rule-based clinical pipelines, and
  build an abbreviation dictionary per site.
- **Negation and uncertainty.** "No evidence of pneumonia," "possible PE," "family history of
  MI" — concept extraction without negation/temporality/subject detection (NegEx/ConText
  algorithms) produces garbage phenotypes. Every extracted concept needs: negated? historical?
  about the patient or family? uncertain?
- **Concept normalization.** Map mentions to standard vocabularies (UMLS CUIs, SNOMED CT, RxNorm
  for meds). Tools: cTAKES, MetaMap, MedCAT (which learns site-specific embeddings). Ambiguity
  ("cold" = temperature vs viral illness) needs context-aware disambiguation.
- **Copy-paste and note bloat.** Cloned text propagates errors and inflates concept counts.
  Detect duplication (text similarity across notes); for phenotyping, deduplicate or down-weight
  repeated content — otherwise one copied problem list counts ten times.
- **Phenotyping from notes.** Notes capture what codes miss (symptoms, severity, social
  context). Best practice: combine structured + NLP features, validate against chart review,
  report PPV/sensitivity. Note-only phenotypes inherit documentation bias (what clinicians
  bother to write).
- **LLMs on clinical text.** Useful for summarization and first-draft documentation, dangerous
  as autonomous extractors. Guardrails: human-in-the-loop for anything entering the chart,
  grounded extraction (every claim traceable to source text spans), hallucination testing on
  adversarial cases, and no PHI in prompts to external APIs without a BAA. An LLM summarizing a
  chart must never invent allergies, doses, or code statuses.
- **Evaluation.** Intrinsic (precision/recall/F1 on annotated spans — annotate with dual
  reviewers and adjudication, report inter-annotator agreement) and extrinsic (does the
  phenotype predict what it should?). Report both.

## Practical workflow

1. **Governance.** IRB, de-identification plan, BAA for any external compute. Notes are the most
   sensitive EHR data — treat them accordingly.
2. **Corpus.** Define note types and time windows; sample representatively; de-identify with
   validated tooling.
3. **Annotate.** Guidelines, dual annotation, adjudication, IAA reporting (F1 or kappa on spans).
   Budget more time than expected — clinical annotation is slow.
4. **Extract.** Pipeline: section detection → sentence splitting → concept extraction →
   negation/temporality/subject → normalization. Validate each stage.
5. **Phenotype.** Combine with structured data; validate against chart review; report PPV,
   sensitivity, and failure modes.
6. **LLM use (if any).** Grounded prompts, human review of outputs, hallucination audits, PHI
   controls. Document the human-in-the-loop design.
7. **Monitor.** Documentation practices drift (new templates, scribes, ambient AI); revalidate
   periodically.

Example (Python sketch):
```python
from medcat import CAT
cat = CAT.load_model_pack("umls_model_pack.zip")
doc = cat(doc_text)  # entities with CUIs, negation, temporality metadata
```

## Common pitfalls

- Off-the-shelf NLP without clinical adaptation (abbreviation soup defeats it).
- Concept extraction without negation detection ("no MI" counted as MI).
- Copy-pasted text inflating phenotype counts.
- De-identification validated on easy cases, failing on narrative dates/names.
- LLM hallucinations entering clinical documentation or research phenotypes unchecked.
- PHI sent to external LLM APIs without authorization.
- Single-annotator "gold standards" with no agreement reporting.

Attribution

aicodedecodeaicodedecode
View sourceSee grades on GitHubMore from aicodedecode →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Terse caveman voice: answer first, fluff gone, every technical fact kept. Use for /caveman, "caveman mode", "talk like caveman", "be brief", "less tokens". Stays on until "stop caveman" or "normal mode".

1100021 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

698621 votes

Writing Skills

Create and manage Claude Code skills in HASH repository following Anthropic best practices. Use when creating new skills, modifying skill-rules.json, understanding trigger patterns, working with hooks, debugging skill activation, or implementing progressive disclosure. Covers skill structure, YAML frontmatter, trigger types (keywords, intent patterns), UserPromptSubmit hook, and the 500-line rule. Includes validation and debugging with SKILL_DEBUG. Examples include rust-error-stack, cargo-dep...

3931 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3421 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Amp, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Grok Build, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

741 votes
View all in ai-agents →