Skip to content
Back to skills

Ehr Data Analysis

ASecurity

Electronic health record analysis — phenotyping, missingness, bias, and observational study design with EHR data.

  • 2 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 29, 2026
ai-agentsgosqldocumentation

Works with

  • cli

Security analysis

A100/100

Scanned September 29, 2026

npx -y skills add aicodedecode/awesome-muse-skills --skill ehr-data-analysis --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Ehr Data Analysis?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Ehr Data Analysis
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/aicodedecode-ehr-data-analysis/badge)](https://www.skillsdirectory.com/skills/aicodedecode-ehr-data-analysis)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: ehr-data-analysis
description: Electronic health record analysis — phenotyping, missingness, bias, and observational study design with EHR data.
category: scientific
---

## Overview

ehr-data-analysis covers secondary analysis of electronic health record data: defining computable
phenotypes, handling the systematic missingness and biases of clinical data, and designing
observational studies (cohort, case-control, target-trial emulation) that produce credible
evidence rather than artifacts of documentation practice.

EHR data was collected for billing and clinical care, not research. Every analysis must reckon
with that: who gets tested, who gets coded, and who shows up at all are all non-random.

## When to use

- Defining computable phenotypes from diagnoses, labs, medications, and notes.
- Assessing data quality: completeness, plausibility, temporal coverage.
- Handling missing data that's missing-not-at-random (sicker patients have more data).
- Observational designs: cohort studies, case-control, target trial emulation.
- Confounding control: propensity scores, inverse probability weighting, doubly robust methods.
- Validation: chart review, phenotype PPV estimation.
- Privacy: de-identification, honest-broker workflows, IRB considerations.

## Core concepts

- **Computable phenotyping.** Turning clinical concepts into algorithms: e.g. "type 2 diabetes" =
  (ICD codes AND (A1c ≥6.5% OR diabetes medication)) NOT type 1 codes. Use published phenotype
  libraries (PheKB) as starting points; validate with chart review on a sample and report PPV/
  sensitivity. A phenotype without validation is a guess.
- **Coding is behavior, not biology.** ICD codes reflect billing incentives and documentation
  habits; they under-capture mild disease and over-capture billable conditions. Codes rule in
  poorly and rule out worse. Supplement with labs, meds, and NLP on notes.
- **Informed presence bias.** Patients appear in EHRs because they're sick or engaged with care.
  Comparing EHR patients to "everyone else" bakes in selection bias. Define the denominator
  explicitly (e.g. all patients with ≥1 primary-care visit in the window).
- **Missingness is systematic.** Labs are ordered when clinicians suspect something — missing
  labs usually mean "not suspected," not "unknown." This is MNAR; standard multiple imputation
  assuming MAR can mislead. Sensitivity analyses and explicit missingness indicators are often
  more honest.
- **Immortal time bias.** Defining exposure by a future event (e.g. "received drug X during
  admission") grants exposed patients survival until exposure. Use time-dependent exposures or
  target-trial emulation with proper time zero.
- **Target trial emulation.** Frame the observational question as the RCT you wish you could run:
  eligibility, treatment strategies, time zero, follow-up, outcome. Then emulate each element.
  This discipline eliminates most immortal-time and selection biases by construction.
- **Confounding control.** Propensity scores (matching, weighting, stratification) balance
  measured confounders; report balance (standardized mean differences <0.1). Unmeasured
  confounding remains — use negative controls, E-values, and quantitative bias analysis to bound
  it. Never claim causality from a single observational design.
- **Temporal issues.** Lookback windows for baseline covariates, washout periods for incident
  (not prevalent) user designs, and new-user active-comparator designs to reduce confounding by
  indication.

## Practical workflow

1. **Governance first.** IRB approval, data-use agreements, de-identification plan. Know what
   you may not do before touching the data.
2. **Define the cohort.** Explicit inclusion/exclusion with code lists; denominator definition;
   index date (time zero) rules.
3. **Build phenotypes.** Adapt published algorithms; validate with chart review (n≥50-100);
   report PPV.
4. **Data audit.** Completeness by variable and site, temporal trends (EHR system changes create
   fake trends), plausibility checks, missingness patterns.
5. **Design.** Target-trial emulation framing; new-user active-comparator where possible;
   prespecified analysis plan.
6. **Analyze.** Propensity weighting/matching with balance checks; primary model; negative
   controls; E-value for unmeasured confounding; multiple bias analyses.
7. **Report.** STROBE/RECORD: cohort construction flow, phenotype definitions and validation,
   missingness, balance tables, and explicit discussion of residual bias.

Example (SQL sketch):
```sql
-- incident diabetes phenotype: first qualifying event after 1-year clean window
WITH first_event AS (
  SELECT patient_id, MIN(event_date) AS index_date FROM diagnoses
  WHERE icd_code LIKE 'E11%' GROUP BY patient_id)
SELECT f.* FROM first_event f
WHERE NOT EXISTS (SELECT 1 FROM diagnoses d
  WHERE d.patient_id = f.patient_id AND d.event_date < f.index_date - INTERVAL '1 year'
  AND d.icd_code LIKE 'E11%');
```

## Common pitfalls

- Prevalent-user designs (comparing long-term users to never-users — confounding by indication).
- Immortal time bias from future-defined exposures.
- Treating ICD codes as ground truth without validation.
- MAR-based imputation on systematically missing clinical data.
- Confounding by indication hand-waved away ("we adjusted for age and sex").
- EHR system migrations creating spurious temporal trends.
- Denominator undefined — rates computed over whoever showed up.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…