Skip to content
Back to skills

Clean Language Reactor

ASecurity

Clean up Russian vocabulary CSV exports from Language Reactor (YouTube subtitle learning extension). Use when the user asks to process or clean up a Language Reactor CSV file.

  • 2 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 27, 2026
businesspythonbashreact

Security analysis

A100/100

Scanned September 27, 2026

npx -y skills add David-Li0406/meta-skill-evloving --skill clean-language-reactor --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Clean Language Reactor?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Clean Language Reactor
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/david-li0406-clean-language-reactor/badge)](https://www.skillsdirectory.com/skills/david-li0406-clean-language-reactor)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: clean-language-reactor
description: Clean up Russian vocabulary CSV exports from Language Reactor (YouTube subtitle learning extension). Use when the user asks to process or clean up a Language Reactor CSV file.
---

# Clean YouTube Subtitle Vocabulary Exports

Transform vocabulary exports from YouTube subtitle extensions into flashcards for recognition practice.

Supports multiple formats:
- Language Reactor exports (WORD| prefix format)
- Simple CSV exports (word, sentence, timestamp, videoTitle, videoId)

## Workflow

1. **Preprocess** the raw export:
   ```bash
   python3 preprocess_language_reactor.py export.csv -o preprocessed.csv
   ```
   This will:
   - Auto-detect the CSV format
   - Parse the export and extract word data
   - Identify multi-word chunks (collocations) - marked with "Is Chunk=yes"
   - Automatically fetch the YouTube transcript
   - Output `preprocessed.csv` and `preprocessed.transcript.txt`

2. **Clean up** the preprocessed file (with Claude):
   - Read both the preprocessed CSV and the transcript
   - Convert chunks to nominative form (e.g., "собственным успехом" → "со́бственный успе́х")
   - Add translations
   - Create simpler example sentences
   - Add stress marks
   - Output the final 4-column CSV

## Examples

See the `examples/` directory:
- **Preprocessed input**: `examples/language-reactor-preprocessed.csv`
- **Transcript**: `examples/language-reactor-preprocessed.transcript.txt`
- **Cleaned output**: `examples/language-reactor-cleaned.csv`

## Card Purpose

These cards are for **recognition** (passive vocabulary), not production:
- **Front**: Russian word + example sentence
- **Back**: English translation

Use `--mode recognition` when creating the deck:
```bash
python3 create_deck.py --mode recognition cleaned.csv -o passive.apkg
```

## CSV Output Format

4 columns, comma-separated:

| Column | Content |
|--------|---------|
| 1 | Russian word/phrase with stress marks |
| 2 | English translation |
| 3 | Example sentence in Russian |
| 4 | English translation of example |

## Processing Rules

### 1. Read Both Files

When cleaning up, read:
- The preprocessed CSV (word, translation, context, POS)
- The transcript file (full video text for understanding context)

### 2. Handle Chunks and Collocations

**Pre-selected chunks**: Entries marked "Is Chunk=yes" were selected as multi-word phrases by the user. Keep these as collocations but convert to nominative form:
- `собственным успехом` → `со́бственный успе́х` (one's own success)
- `качественной медицине` → `ка́чественная медици́на` (quality healthcare)

**Identify additional collocations**: Some single words should be studied with their common collocations. Use the transcript to identify:
- `птичий` → `на пти́чьих права́х` (on precarious terms)
- `поразить` → `пораже́ны в права́х` (disenfranchised)
- `доступ` → `до́ступ к (+ dat.)` (access to)

### 3. Create Simpler Example Sentences

The subtitle context is often:
- Incomplete (cut off mid-sentence)
- Too long or complex for flashcards
- Contains multiple clauses

Create a **simpler, clearer sentence** that:
- Uses the word with the same meaning as in the video
- Is short enough for a flashcard (under 15 words ideally)
- Is grammatically complete

Use the full transcript to understand the meaning and context.

### 4. Stress Marks (Column 1 Only)

- Add stress mark (´) on the stressed vowel for multisyllabic words
- Skip monosyllables
- Skip words with ё (always stressed)

### 5. Gender for Soft Sign Nouns

- For nouns ending in soft sign (ь), indicate gender with `(m.)` if masculine
- Most soft-sign nouns are feminine, so only mark masculine ones
- Examples: `день (m.)` — day, `гость (m.)` — guest, `дождь (m.)` — rain

### 6. Verb Pairs
- When the word is a verb, include both the perfective and imperfective forms
- When the verb is used with a preposition in the context, include that preposition and the case used with that preposition in this context
Format: `imperfective/perfective` (e.g., `ви́деть/уви́деть`)
- Only include one form if the other isn't commonly used or doesn't make sense
- Include both when learners should know the pair

### 7. CSV Quoting

Wrap any field containing commas in double quotes.

## Example Transformation

**Preprocessed** (from `examples/language-reactor-preprocessed.csv`):
```csv
Word,Translation,Context (RU),Context (EN),POS
птичий,"bird's, avian",Пока у тебя нет гражданства ты всегда на птичьих правах,Until you have citizenship you're always on precarious terms,Adj
```

**Cleaned output** (see `examples/language-reactor-cleaned.csv`):
```
на пти́чьих права́х,on precarious terms (no legal rights),Без гражданства ты на птичьих правах.,Without citizenship you're on precarious terms.
```

Note: The single word `птичий` became the collocation `на пти́чьих права́х` because that's how it was used in context.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…