Skip to content
Back to skills

Reliability Validity Testing

ASecurity

Evaluating measurement quality — test-retest, inter-rater agreement, and construct validity evidence.

  • 2 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 29, 2026
ai-agentsgotesting

Works with

  • cli

Security analysis

A100/100

Scanned September 29, 2026

npx -y skills add aicodedecode/awesome-muse-skills --skill reliability-validity-testing --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Reliability Validity Testing?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Reliability Validity Testing
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/aicodedecode-reliability-validity-testing/badge)](https://www.skillsdirectory.com/skills/aicodedecode-reliability-validity-testing)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: reliability-validity-testing
description: Evaluating measurement quality — test-retest, inter-rater agreement, and construct validity evidence.
category: scientific
---

## Overview

reliability-validity-testing is the applied companion to psychometrics: the concrete procedures for
demonstrating that a measure is consistent (reliability) and measures what it claims (validity).
It covers designing reliability studies, choosing agreement coefficients, gathering validity
evidence, and interpreting the numbers — including what to do when they're disappointing.

Use this when you need to evaluate or defend a measurement instrument; use psychometrics for
building scales and latent-variable modeling.

## When to use

- Planning a test-retest reliability study: interval, sample size, ICC model choice.
- Inter-rater agreement: training raters, choosing kappa vs ICC vs percent agreement.
- Internal consistency: when alpha/omega are appropriate and how to report them.
- Construct validity: convergent/discriminant correlation patterns, known-groups comparisons.
- Criterion validity: validating against a gold standard (sensitivity/specificity framing).
- Evaluating a published instrument's evidence before adopting it.
- Responding to reviewer requests for "reliability and validity" of your measures.

## Core concepts

- **Reliability is consistency, in specific forms.** Internal consistency (items hang together),
  test-retest (stable over time), inter-rater (consistent across judges), parallel forms. A
  measure can have one without the others — report the form relevant to your use. High internal
  consistency does not imply stability over time.
- **ICC, not Pearson, for agreement.** Test-retest and inter-rater reliability need intraclass
  correlation: ICC(2,1) or ICC(3,1) depending on whether raters are random or fixed, absolute
  agreement vs consistency. Pearson correlation ignores systematic bias (rater A always 2 points
  higher still gives r=1.0). Report the ICC model — "ICC = 0.85" without the model is incomplete.
- **Kappa for categorical judgments.** Cohen's kappa for two raters, Fleiss' for more; weighted
  kappa for ordinal categories. Interpret against prevalence — kappa paradoxes occur with very
  high or low base rates, so report raw agreement too.
- **Test-retest interval.** Long enough that memory doesn't inflate agreement, short enough that
  the construct hasn't truly changed. State the rationale; for state-like constructs, expect
  lower values and say so upfront.
- **Validity evidence types.** Content (expert/domain coverage), response processes (cognitive
  interviews — do respondents interpret items as intended?), internal structure (factor analysis,
  invariance), relations to other variables (convergent/discriminant, criterion), and
  consequences (does use of the scores lead to intended outcomes?). Modern standards treat these
  as five sources of evidence for one unified validity argument.
- **Multitrait-multimethod logic.** The classic discriminant-validity design: your measure should
  correlate more strongly with other measures of the same construct (even via different methods)
  than with measures of different constructs via the same method. Same-method correlations are
  inflated by shared method variance — discount them.
- **Attenuation.** Observed correlations are attenuated by unreliability in both measures:
  r_true = r_obs / sqrt(r_xx × r_yy). A "weak" r=0.30 with reliabilities of 0.70 is actually
  ~0.43. Disattenuate when comparing effect sizes across measures with different reliabilities.
- **Standard error of measurement.** SEM = SD × sqrt(1 − reliability). It puts confidence bands
  around individual scores — essential for clinical cutoffs. A screening cutoff is meaningless
  without the SEM around it.

## Practical workflow

1. **Match evidence to use.** Screening decision → criterion validity + SEM at the cutoff;
   research DV → internal consistency + test-retest; observational coding → inter-rater agreement.
2. **Design the study.** Sample size for precise reliability estimates (n≥50 for ICC with narrow
   CIs; more for kappa with rare categories). Train raters with a manual and practice coding.
3. **Collect.** Independent ratings (raters blind to each other and to hypotheses); test-retest
   with the planned interval; no feedback between ratings.
4. **Analyze.** ICC with the correct model + 95% CI; kappa with raw agreement; alpha/omega for
   internal consistency; correlation matrix for convergent/discriminant patterns.
5. **Interpret against standards.** ICC/kappa ≥0.75 good, 0.60-0.74 moderate (context-dependent);
   validity correlations judged against theory, not fixed cutoffs. Report CIs — a point estimate
   of 0.80 with CI [0.55, 0.92] is not "good reliability."
6. **Document.** Full evidence table in the manual/paper: coefficient, model, sample, interval,
   CI. State limits on generalizability (population, context).

## Common pitfalls

- Pearson r reported as "inter-rater reliability" (misses systematic bias).
- Wrong ICC model, or no model specified.
- Kappa interpreted without base rates (paradoxes with rare categories).
- Test-retest interval so short it measures memory, or so long it measures real change.
- Claiming validity from a single correlation ("validated" is not a one-study achievement).
- High alpha from redundant items mistaken for good measurement (bloated specifics).
- Using a measure in a new population without re-establishing reliability there.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…