Make a realistic talking-head "reaction / explainer" UGC ad. An AI creator (A-roll, full-screen or a floating corner cut-out) carries the voiceover while the main screen cuts between B-roll footage and C-roll evidence (real screenshots). Cast an ordinary/relatable creator, generate footage on Seedance 2.0 via Higgsfield with NATIVE audio so the creator speaks his own lines (lip-sync baked in — no TTS, no separate sync pass), grade for "shot-on-a-phone" realism, and composite karaoke-captioned...
Scanned 8/30/2026
Install to Claude Code
npx -y skills add docusphere/ugc-reaction-studio --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of ugc-reaction-studio?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/docusphere-ugc-reaction-studio)More formats (shields.io, HTML) on the badges page.
---
name: ugc-reaction-studio
description: Make a realistic talking-head "reaction / explainer" UGC ad. An AI creator (A-roll, full-screen or a floating corner cut-out) carries the voiceover while the main screen cuts between B-roll footage and C-roll evidence (real screenshots). Cast an ordinary/relatable creator, generate footage on Seedance 2.0 via Higgsfield with NATIVE audio so the creator speaks his own lines (lip-sync baked in — no TTS, no separate sync pass), grade for "shot-on-a-phone" realism, and composite karaoke-captioned in Remotion — all inside Claude Code. Driven by a CSV source of truth (Config + Beats) and an inputs/outputs workspace. Use for a reaction-style product ad, an explainer with on-screen proof, or a talking-head-over-b-roll ad.
---
# UGC Reaction Ad Studio
Make a **mixed-visual reaction ad** that reads as a real person filming on their phone — not an AI ad. An AI creator talks through a tight script; the frame cuts between the creator, supporting footage, and real proof.
**Best for any product where a *claim + on-screen proof* converts** — supplements, skincare, apps, SaaS, courses, gadgets, tools, finance. The proof (C-roll) is whatever's real for that product: reviews, before/after, ratings, DMs, press, a dashboard — **or** a study when the claim is scientific (not a requirement). Weaker fit for pure impulse/aesthetic products with nothing to prove (a candle, a phone case) — those often do better with a plain vibe/UGC clip.
**Companion docs (bundled — read them):** `references/realism-playbook.md` (the exact locked recipes) and `references/lessons.md` (mistakes to skip). This SKILL.md is the backbone; those are the depth.
## The three footage layers (use this vocabulary — not "clips")
- **A-roll — the creator (talking head).** The person speaking to camera; *carries the voiceover.* Shown **full-screen** on hero beats (hook, verdict) or as a **floating corner cut-out** over the main screen on the rest. Always the voice, even when small.
- **B-roll — footage.** Silent supporting cutaways that *illustrate* — the creator working out, product-in-hand shots. Fills the main screen while A-roll narrates from the corner.
- **C-roll — evidence / receipts.** REAL screen-captures that back the claim — **NOT science-only.** Pick whatever proof fits the product: customer **reviews / ratings** (Amazon, Trustpilot, App Store), **before/after** or a results screenshot, real **DMs / comments / Reddit**, **press / "as seen in"**, a **sales/analytics dashboard** or sold-out screenshot, OR a **study/journal** when the claim is scientific. Crisp, never faked, never graded. (Studies are just the strongest proof for an *efficacy* claim — most products use reviews/social proof/before-after instead.)
Each beat = **who's on A-roll** (creator: full / corner / none) + **what's the main visual** (A-roll / B-roll / C-roll).
## Workspace: inputs/ and outputs/
The skill creates these itself. **Never ask the user to make folders/files by hand.**
```
inputs/ ← what the USER feeds in (swap to make a different ad)
├── config.csv locked static style (creator, product, framing, voice, grade, avoid)
├── beats.csv per-beat script + screen plan (source of truth)
├── product.png product image (user-provided OR generated here as a seed)
├── creator-ref.png OPTIONAL — reference photo for the creator's look
└── screenshots/ C-roll evidence, named beat-N-… (captured by the user)
outputs/
├── creator/ locked base character (+ angle refs)
├── footage/ generated A-roll + B-roll, beat-N-[label]-[type].mp4
├── vo/ ONE continuous VO + per-beat splits
├── stage/ per-beat composited stage clips (main + baked corner cut-out)
├── remotion/ the composition
└── final/ rendered ad: reaction-ad-01.mp4
```
The `example-*.csv` shipped with the skill are **templates only** — never edit them per-ad.
## Static vs Dynamic prompts (the "style once" model)
- **STATIC → `config.csv`:** the style layer injected into EVERY generation — locked creator ref, framing (phone selfie, 9:16, 720p), environment (the set/background), wardrobe, lighting, the realism grade, the voice recipe, AVOID.
- **DYNAMIC → `beats.csv` `Gen Prompt`:** per-beat action/motion only. Claude fills this; the user reviews **before** anything generates (prompt-review gate).
---
## Pipeline (the locked order — do NOT reorder)
Order matters because the creator must exist before anything is generated from him. **On the locked native-audio path, Voice + A-roll + Lip-sync all collapse into step 5** — `seedance_2_0` speaks the line as it renders the talking beat, so there is no separate voice step and no sync step. Steps 4 and 6 only expand if you drop to the TTS fallback.
### 0 · Preflight (doctor) + Setup
**Run a quick preflight FIRST and print what's missing — don't let a new user discover deps one failure at a time.** Check and report:
- **Higgsfield MCP** connected? (required — if not, stop and tell them to connect it.)
- **Node + Remotion**, **ffmpeg**, **whisper.cpp** present? (required — offer to set up Remotion via `npx create-video@latest`; name any missing binary.)
- **A browser tool** available (built-in Claude Code browser, or a Playwright/Chrome MCP)? → decides the evidence mode in §7. If none, say "no browser detected — you'll capture the evidence screenshots yourself." (Claude can't install an MCP for them.)
- **pyroomacoustics** (optional) — if absent, note the voice will use the synthetic-room fallback.
Then scaffold `inputs/` + `outputs/`. If `inputs/config.csv` + `inputs/beats.csv` exist, read them; else copy from the `example-*.csv` templates (the templates are never edited per-ad; the working files live in `inputs/`).
### 1 · Product + Brand DNA
Name, what it does, who it's for, the ONE credible claim, and the vibe/positioning. Product image → `inputs/product.png` (user-provided, or generate a clean seed — no on-product text). Keep it original (no third-party IP).
### 2 · Script / Beats (storyboard)
Write the reaction script (hook → problem → switch → evidence → how-I-use → verdict → CTA) into `beats.csv`: per beat the VO line, creator position (A-roll full/corner/none), main visual (A-roll / B-roll:[what] / C-roll:[evidence]), caption. **Define the C-roll evidence needed here** so the user can start capturing it in parallel. Lock the story.
### 3 · Character + Environment (lock ONE base)
Cast an **ordinary, relatable, imperfect** creator (see casting rule) into the set/environment. Lock ONE base reference image (+ a couple angle/half-body-with-product refs). **Seed every future generation from this one base** for consistency.
### 4 · Voice — NATIVE audio on `seedance_2_0` (the locked default)
The creator **speaks his own lines**. Don't make a separate voice at all.
- **⭐ Seedance 2.0 native audio — LOCKED ENGINE.** Generate each talking beat on **`seedance_2_0`** with **`generate_audio: true`** and the beat's line **quoted in the prompt**, so Seedance renders the character *speaking it in its own generated voice* — perfect lip-sync baked in, real phone-voice timbre, no TTS "AI voice" tell. Same face ref on every beat = consistent voice. **This folds Voice + A-roll + Lip-sync into step 5** — there's no VO track and no `sync_so` pass, so **step 6 does not run.** **Gate it cheaply:** generate the HOOK only, get the user's ear on it, then batch the rest. Every native clip needs a **transcribe + trim** pass (it improvises filler/repeats on short lines — `whisper-cli -ml 1 -oj`, trim to the scripted words). See `references/realism-playbook.md` §14. *We reached the best voice this way after preset-TTS, room-treated-TTS, and even the user's own recording all fell short.*
- **Stay on `seedance_2_0`.** Don't switch engines mid-ad — it changes the look *and* the voice. Iterate prompts on `seedance_2_0_mini` (cheaper/faster), then render the keeper on `seedance_2_0`.
- **TTS + lip-sync is a FALLBACK, not an option to offer by default.** Take it only when the voice must be a SPECIFIC or cloned pick the model can't produce. Then: `list_voices` / `create_voice`, the **entire script as ONE continuous read** (never per-line — it drifts), transcribe for word timing, split per beat, ~`atempo=1.13`, plus the **Acoustic Profile** (playbook §13) and the lip-sync pass in §6. Extra chain, lower realism ceiling.
### 5 · Generate A-roll + B-roll (the footage)
Fill `Gen Prompt` per `generate` beat (static Config + dynamic action), show the user, then generate from the locked base:
- **A-roll (talking head):** tight head-and-shoulders, hands out of frame, calm, **engaged with direct eye contact** (a glassy/"lost" gaze is an AI tell — regenerate at the source if it happens) — Seedance, silent, 9:16, 720p.
- **B-roll (footage):** creator action (workout) / product-in-hand — silent. **Vary it to look shot on different days:** different wardrobe AND different locations (home vs a public gym). Keep the hero + corner cam on ONE "today" look; vary only the cutaways.
- **Product-use demo (money shot):** a clip of the creator actually USING the product. Attach `product.png` and **spell out the real mechanics** (e.g. nasal inhaler → tip at one nostril, mouth closed) — regenerate if the model does the wrong action.
- **Two-shot within a beat:** to add a shot without adding a beat (breaks VO/caption timing), cut one beat's main between two b-roll clips while the corner talks over both.
Save to `outputs/footage/beat-N-[label]-[type].mp4`; attach the base ref (and `product.png` on product beats) to EVERY generation.
### 6 · Lip-sync (Sync Lipsync 3) — **LEGACY / TTS-fallback only — normally SKIPPED**
*On the locked native-audio path this step does not run at all* — the talking beats already carry perfectly-synced speech. Only for the TTS fallback: run **every** talking appearance through **`sync_so`** with its per-beat VO segment — full-screen hero beats AND the corner cut-out. **Sync the corner PER BEAT** (upload segment → sync → matte), never loop-mode (loop causes a ~7s head "reset"). See playbook §8. (B-roll/C-roll never get lip-synced.)
### 7 · C-roll / Evidence
The real screenshots planned in step 2, into `inputs/screenshots/` (named `beat-N-…`). Independent — captured any time, just in before compositing. **Never fabricate.** Two capture modes (decided by the §0 preflight):
- **Browser tool present →** Claude navigates to the real source (study/PMC/review) and captures it (frame vertical: portrait viewport, title + key finding). Confirm the source is credible and show the user before using.
- **No browser tool →** give the user a precise capture list (what to screenshot + suggested credible sources) and have them drop the files into `inputs/screenshots/`. This path always works and needs no MCP.
- **Real-vs-fictional product fork (critical — never fabricate):** product-specific proof (reviews, ratings, "sold out", income/results screenshots) is only real for a **REAL** product. For a **fictional/demo/original** product, you MUST use **ingredient/category-level real proof** instead — studies about the active ingredient or category (e.g. niacinamide, menthol, caffeine+protein), which are true regardless of brand. Never invent product reviews/ratings for a product that doesn't have them.
### 8 · Post — cut-out matte + grade
- **Corner cut-out:** `remove_background` (video) → `colorkey` black → clean floating cut-out, **uniform size + a SOLID dark silhouette backing** behind it (LOCKED — use the tight-key `0.02` backing so it keeps his dark hair/shirt and the b-roll never shows *through* him; the old soft halo left see-through holes). Keep corner clips **hands-out-of-frame** (a held product ghosts at the edge). **NEVER green-screen** (AI green won't key), **never a black-shirt source** (key eats the shirt). Playbook §3.
- **Grade (layered):** **clean iPhone grade A** on **A-roll + the corner cut-out** (a real selfie cam is fairly clean); **LIGHT grade on B-roll** (cleaner "higher-quality iPhone"); **C-roll evidence stays crisp and STATIC** (no zoom/pan). Playbook §9 + §6.
### 9 · Karaoke captions
TikTok style from the VO word-timings: **3 words at a time, yellow highlight, the active word flips to black, positioned bottom third.** No badges/labels.
### 10 · Composite + Render (Remotion → outputs/final/)
Per beat: main visual full-screen + floating corner cut-out (when creator=corner) + karaoke captions. **Audio:** Path A = each stage clip carries its **own native audio** (drop the global `<Audio>`, un-mute `OffthreadVideo`, `loudnorm` per beat, **beat-relative captions**); Path B = ONE continuous VO track. Render **9:16, 1080p**. No music (reaction video).
- **Give the user live crop/cut controls** (playbook §15): a `zod` schema + `defaultProps` exposes per-beat `durationSec`/`sourceStartSec` (cut) and `zoom`/`offsetXpct/Y` (crop) in Remotion Studio, with `calculateMetadata` recomputing length. They tweak live and render, or hand off the clean master to **CapCut** for the final cut.
- **`--bundle-cache=false`** after changing only `public/` assets, or you'll silently render the stale bundle.
---
## Locked recipes (summary — full values in references/realism-playbook.md)
- **Casting:** ordinary/imperfect beats attractive (research-backed) — avoid model looks, styled hair, bright eyes, over-smooth skin.
- **Engine (LOCKED):** `seedance_2_0` for all footage; `seedance_2_0_mini` for cheap prompt iteration. One engine per ad — never swap mid-build.
- **Voice (LOCKED):** Seedance **native audio** (`generate_audio:true` + line quoted in prompt), same face ref every beat; transcribe+trim each clip (it improvises). No separate VO, no sync pass. Playbook §14.
- **TTS fallback (rare):** `text2speech_v2` / `elevenlabs`; ONE continuous read; split by word-timing; ~`atempo=1.13`. Mic character = **Acoustic Profile** (Room + Mic-distance + Ambience → EQ + **diffuse** convolution reverb). Never discrete-echo reverb (metallic).
- **Lip-sync (fallback only — normally not run):** `sync_so` (Sync Lipsync 3), `cut_off`; **corner PER BEAT, never loop-mode**. Caps at 3 concurrent (shared with `remove_background`).
- **Cut-out:** `remove_background` (video) → `colorkey=0x000000`, uniform size, **+ SOLID silhouette backing** (tight `0.02` key → no see-through at dark hair/shirt). Hands out of frame. No green screen, no black-shirt source. Bake corner into the main via ffmpeg (no browser-alpha).
- **B-roll:** one **static locked-off** phone angle, no zoom/push-in; one simple literal action; vary shirt+location for "different days". **C-roll static** (no Ken Burns).
- **Grade (layered):** A-roll + corner = **clean iPhone grade A** (no downres, tiny grain); B-roll = light; C-roll crisp. (Old-phone heavy grade is a dial, not the default.)
- **Captions:** karaoke, 3-word, yellow highlight / black active word, bottom.
- **Realism > polish** everywhere; **dial, don't rebuild** (expose knobs: fuzz amount, grade %, speech speed).
## Optional realism upgrades (external — skill runs without them)
- **`pyroomacoustics`** (SHIPPED: `scripts/room_ir.py`) — physically-modelled room IR (real geometry + mic distance) for the Acoustic Profile. `pip install pyroomacoustics numpy scipy soundfile` (CPU, cross-platform). If absent, the skill falls back to the synthetic-noise IR — so this is optional, not required. Playbook §13(A).
- **Cleaner corner matting** (removes the dark-rim workaround; needs a GPU): [MatAnyone](https://github.com/pq-yang/MatAnyone), [BiRefNet HR-matting](https://github.com/ZhengPeng7/BiRefNet), [RobustVideoMatting](https://github.com/PeterL1n/RobustVideoMatting) — swap in for `remove_background`.
- **Fix AI face artifacts** (melty eyes/skin; GPU): [CodeFormer](https://github.com/sczhou/CodeFormer) / [GFPGAN](https://github.com/TencentARC/GFPGAN) — run gently or it over-smooths.
- **Tighter caption timing** (CPU/GPU): [WhisperX](https://github.com/m-bain/whisperx) — wav2vec2 forced alignment, better word timestamps than `whisper.cpp -ml 1`.
- **Open lip-sync alternative** to `sync_so` (GPU): [LatentSync](https://github.com/bytedance/LatentSync).
- **More realistic voice intonation** (off-Higgsfield, GPU): [StyleTTS2](https://github.com/yl4579/StyleTTS2), [F5-TTS](https://github.com/SWivid/F5-TTS), [Chatterbox](https://github.com/resemble-ai/chatterbox) (expressiveness dial), [Kokoro](https://github.com/hexgrad/kokoro); or **prosody transfer** from a real human read via [RVC](https://github.com/RVC-Project/Retrieval-based-Voice-Conversion-WebUI) / [OpenVoice](https://github.com/myshell-ai/OpenVoice). On Higgsfield, drive intonation via script punctuation + `seed_audio`/`qwen` expression params (playbook §Voice).
- **Motion realism:** RIFE interpolation + a touch of handheld shake/film grain (AI clips are eerily stable).
## Rules / guardrails
- One welcome, then straight into it.
- **Prompt-review gate:** the user sees & approves every `Gen Prompt` before generating.
- **Consistency:** lock ONE creator base; seed every clip from it; attach `product.png` to product beats.
- **Validate one sample before batching** (one voice line, one lip-sync, one graded clip) — taste shifts; don't re-do a whole batch.
- **Diagnose by layer** when feedback is "this is off": voice vs treatment vs sync vs motion vs framing.
- **Evidence stays real.** Original product only.
- Generate via **Higgsfield MCP**; composite/render in **Remotion**. If Higgsfield isn't connected → stop and say so; if Remotion isn't set up → `npx create-video@latest` before compositing.
---
## Exact messages (say these, adapt the brackets)
**Opening (the ONLY welcome):**
> 👋 UGC Reaction Ad Studio. Tell me your product and I set everything up — an `inputs/` folder (product, CSVs, screenshots) and `outputs/` (creator, footage, final ad). I cast an ordinary creator, generate his talking beats with a real, natural voice (the creator speaks the lines himself — no robotic TTS), grade it "shot-on-a-phone," and composite karaoke-captioned in Remotion. **Higgsfield connected?** First — **what's the product?** (name, who it's for, the one claim, the vibe; optional product image).
**Step 2 — Script / Beats:**
> Story first. I'll write a ~30–40s reaction script into `beats.csv`: per beat the VO line, where the creator sits (A-roll full / corner), the main visual (A-roll / B-roll footage / C-roll evidence), and the caption. I'll also list the **evidence screenshots you'll need** so you can start grabbing them now. Give me the angle or say "you write it."
**Step 3 — Character:**
> Casting an **ordinary, relatable** creator (not a model — that's the AI tell) into your set, and locking ONE base reference so it's the same person on every shot.
**Step 4 — Voice:**
> For the voice I'll use the most realistic route first: **Seedance generates him actually *speaking* each line** in his own voice (mouth already perfectly synced) — no robotic TTS. I'll make the **hook first** so you can hear it; if you love it, I do the rest. (If you'd rather a specific/cloned voice, I can go the TTS + lip-sync route instead.)
**Step 5–6 — Footage:**
> Generating his talking beats (native voice + tight framing) and the B-roll (footage/product, one steady phone angle) from the locked creator. I transcribe each talking clip to confirm it says the line, and trim any filler.
**Step 7 — Evidence:**
> Drop your captured screenshots into `inputs/screenshots/`, named `beat-N-…`. Real proof only.
**Step 10 — Composite + Render:**
> Compositing in Remotion: main visual per beat + you floating in the corner + karaoke captions (bottom) + the VO. Rendering 9:16 → `outputs/final/`. Here's the ad 👇 want to re-time, swap evidence, or spin a variation?
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!