Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Alterlab Text As Data

ASecurity

Analyzes text as social-science data — topic modeling (BERTopic with embeddings + class-based TF-IDF, LDA/NMF via scikit-learn or gensim), document embeddings (sentence-transformers), dictionary/lexicon methods, and supervised text classification — choosing the method that matches the inferential goal (discovery vs measurement vs prediction) and validating topic reliability rather than trusting one stochastic run. It uses the verified stack (BERTopic, scikit-learn, gensim CoherenceModel, spaC...

158 stars
0 votes
0 copies
0 views
Added 10/6/2026
developmentpythonrustgobashexpressgitapi

Works with

api

Security Analysis

A100/100

Pro scans all 4 files and shows the line behind each finding

Scanned 10/6/2026

$npx -y skills add NVlabs/Skill2Env --skill alterlab-text-as-data --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Alterlab Text As Data?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Alterlab Text As Data
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/nvlabs-alterlab-text-as-data/badge)](https://www.skillsdirectory.com/skills/nvlabs-alterlab-text-as-data)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: alterlab-text-as-data
description: "Analyzes text as social-science data — topic modeling (BERTopic with embeddings + class-based TF-IDF, LDA/NMF via scikit-learn or gensim), document embeddings (sentence-transformers), dictionary/lexicon methods, and supervised text classification — choosing the method that matches the inferential goal (discovery vs measurement vs prediction) and validating topic reliability rather than trusting one stochastic run. It uses the verified stack (BERTopic, scikit-learn, gensim CoherenceModel, spaCy, sentence-transformers) with pinned patterns. Use when the request mentions topic modeling, text as data, computational text analysis, document embeddings, dictionary/sentiment lexicons, or classifying a corpus. For training or fine-tuning transformer models prefer alterlab-transformers; for humanities close-reading corpora prefer alterlab-digital-humanities. Part of the AlterLab Academic Skills suite."
license: MIT
allowed-tools: Read Bash(python:*)
compatibility: "Requires (declare in-session, no runtime install on Anthropic API): bertopic>=0.16, scikit-learn>=1.3, gensim>=4.3, spacy>=3.7 (+ a model like en_core_web_sm), sentence-transformers>=2.2 (pip). Runs locally via `uv run python`; no API key."
metadata:
    skill-author: AlterLab
    version: "1.0.0"
    depends_on: "alterlab-ssci-design-gate, alterlab-transformers (model training), alterlab-digital-humanities; audited by alterlab-ssci-inference-gate"
---

# Text-as-Data — Match the Method to the Inferential Goal

**Skill type: ANALYSIS MODULE.** Turns a corpus into measurements. The discipline is choosing by
*goal* — **discovery** (what themes exist?), **measurement** (how much of concept X?), or
**prediction** (label new documents) — and validating that a topic solution is **reliable**, not a
single lucky stochastic run.

## Core Mission

```
PICK BY GOAL: DISCOVERY vs MEASUREMENT vs PREDICTION. THEN VALIDATE THE TOPICS — ONE RUN IS NOT A RESULT.
```

## When to Use This Skill

- "Run topic modeling on my corpus (BERTopic / LDA)."
- "Measure how much each document expresses concept X (dictionary/lexicon)."
- "Embed my documents and cluster / compare them."
- "Classify these texts into categories."

### Does NOT Trigger

| The request is really about… | Route to | Why not this skill |
|---|---|---|
| Training / fine-tuning a transformer model | `alterlab-transformers` | Model training, not corpus measurement. |
| Humanities close-reading / annotation of texts | `alterlab-digital-humanities` | Interpretive, not quantitative text-as-data. |
| Whether a text method fits the question at all | `alterlab-ssci-design-gate` | Design routing, upstream. |
| Plain tabular statistics | `alterlab-statistical-analysis` | No text. |

## Method by goal (verified stack, pinned)

| Goal | Method | Verified call |
|------|--------|---------------|
| **Discovery** (emergent themes, contextual) | **BERTopic** (v0.17) | `from bertopic import BERTopic; topics, probs = BERTopic().fit_transform(docs)`; `get_topic_info()`, `get_topic(0)`. Plug embeddings via `embedding_model=SentenceTransformer("all-MiniLM-L6-v2")`, a `vectorizer_model=CountVectorizer(min_df=10)`. |
| Discovery (bag-of-words, classic) | **LDA / NMF** | sklearn: `LatentDirichletAllocation(n_components=k).fit(CountVectorizer().fit_transform(docs))`; or gensim `LdaModel(corpus, num_topics=k, id2word=dictionary)`. |
| Topic quality | **coherence** | gensim `CoherenceModel(model=lda, texts=tok, dictionary=d, coherence="c_v").get_coherence()`. |
| **Measurement** (how much of concept X) | **dictionary / lexicon** | count validated lexicon terms; report reliability and validate against hand-coding. |
| Embeddings / similarity | **sentence-transformers** (v5) | `SentenceTransformer("all-MiniLM-L6-v2").encode(texts)` → cluster / cosine-compare. |
| **Prediction** (label documents) | **supervised** | `TfidfVectorizer()` → a sklearn classifier; report held-out F1, not in-sample fit. |
| Linguistic features (POS, entities) | **spaCy** (v3) | `nlp = spacy.load("en_core_web_sm"); doc = nlp(text)`. |

Method-selection helper (stdlib): `scripts/text_method_router.py`. Full patterns, preprocessing,
and the topic-reliability procedure: `references/text_stack.md`.

## The reliability discipline (topic models are stochastic)

LDA (and, to a lesser degree, BERTopic via UMAP) gives different topics on different runs. A single
run is not a result. Required:

1. **Replicate** across seeds/runs; align topics across runs (e.g. Hungarian matching) and score
   stability (top-word overlap, RBO); report a **prototype** (most-representative) solution.
2. **Coherence, not just perplexity** — choose K by `c_v` coherence and human readability, not
   log-likelihood alone.
3. **Determinism where possible** — fix `random_state` (sklearn LDA) and UMAP's `random_state`
   (BERTopic) for reproducibility; still report stability across *different* seeds.
4. **Validate measurement** — a dictionary/topic used as a *measure* of a concept needs validation
   against human coding (precision/recall or correlation), exactly like any other instrument.

## Output Template

```
GOAL:         discovery | measurement | prediction
METHOD:       <BERTopic / LDA / dictionary / embeddings / supervised> + why it matches the goal
MODEL:        <call + K selection by coherence; embedding/vectorizer choices>
RELIABILITY:  <seeds run; topic stability; coherence c_v; determinism settings>
VALIDATION:   <if used as a measure: agreement with human coding>
CLAIM SCOPE:  descriptive/measurement of the corpus; generalization scoped to the corpus/sampling
```

## References

- `references/text_stack.md` — BERTopic/LDA/gensim/spaCy/embeddings patterns, preprocessing, topic-reliability procedure.
- `scripts/text_method_router.py` — stdlib router from goal + corpus features to the right method.

Part of the AlterLab Academic Skills suite.

Attribution

NVlabsNVlabs
View sourceSee grades on GitHubMore from NVlabs →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Clean Code

Pragmatic coding standards - concise, direct, no over-engineering, no unnecessary comments

304955 votes

Browser Extension Developer

Use this skill when developing or maintaining browser extension code in the `browser/` directory, including Chrome/Firefox/Edge compatibility, content scripts, background scripts, or i18n updates.

286712 votes

Seo Optimizer

SEO optimization with keyword analysis, readability assessment, technical validation, content quality. Use for search rankings, blog posts, content audits, or encountering keyword density, readability scores, meta tags, schema markup errors.

2222 votes

Google Official Seo Guide

Official Google SEO guide covering search optimization, best practices, Search Console, crawling, indexing, and improving website search visibility based on official Google documentation

1862 votes

Writing Plans

Use when you have a spec or requirements for a multi-step task, before touching code

2927051 votes
View all in development →