Skip to content
Back to skills

Paper2video

ASecurity

Turn a research paper or project (just an arXiv link or a paper title is enough; or a repo, PDF, project page, or working directory) into a polished, 3Blue1Brown-style explainer / promo video built with Remotion — script, voice-over (optional TTS key; free or self-recorded alternatives), real-data animations, captions, music, bilingual cuts, vertical cut, covers, and platform copy for YouTube and Bilibili. Use when the user asks for a paper video, project video, research promo, explainer vide...

  • 2 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added October 7, 2026
ai-agentspythongoshellnodegitapibackend

Works with

  • claude code
  • cli
  • api

Security analysis

A92/100
  • mediumInstalls packages at runtime which could introduce malicious dependencies

Pro scans all 20 files and shows the line behind each finding

Scanned October 7, 2026

npx -y skills add Lunamos/paper2video --skill paper2video --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Paper2video?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Paper2video
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/lunamos-paper2video/badge)](https://www.skillsdirectory.com/skills/lunamos-paper2video)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: paper2video
description: Turn a research paper or project (just an arXiv link or a paper title is enough; or a repo, PDF, project page, or working directory) into a polished, 3Blue1Brown-style explainer / promo video built with Remotion — script, voice-over (optional TTS key; free or self-recorded alternatives), real-data animations, captions, music, bilingual cuts, vertical cut, covers, and platform copy for YouTube and Bilibili. Use when the user asks for a paper video, project video, research promo, explainer video, 论文宣传片 / 论文讲解视频, or wants to publish a video about their work.
---

# paper2video

Make a video that a stranger with no context can follow, that stays true to the work, and that looks good enough to share: real data on screen early, one idea per beat, smooth motion that explains rather than decorates.

You are the **lead**: you plan, build the shared pieces, hand out well-scoped work to subagents when useful, integrate, and own correctness. Work autonomously end to end unless the user asks for checkpoints; show the user results (stills, drafts, renders), not questions, whenever a sensible default exists.

Supporting material (read when you reach that stage):
- `reference/rigor.md` — claims ledger, data handling, independent fact-check
- `reference/voice.md` — narration backends (ElevenLabs / Fish Audio API / free edge-tts / self-hosted GPT-SoVITS / the user's own recording / none), QA, pronunciation, open-source TTS know-how
- `reference/visual-style.md` — design system, motion rules, data honesty, review loop
- `reference/platforms.md` — YouTube + Bilibili deliverables: specs, covers, titles, descriptions, chapters
- `reference/vertical.md` — the 9:16 phone cut: safe area, type sizes, one focus per beat, hook, building it from the same audio
- `reference/handoff.md` — using the user's own materials, and exporting for editing software
- `reference/third-party.md` — starting from only an arXiv link/title: finding source, code and data; extracting numbers from LaTeX/figures; demo runs for real per-example data; faithfulness when the paper is not the user's
- `scripts/` — the pipeline (see `scripts/README.md`); `template/` — the Remotion starter (see `template/README.md`)

## 0. Intake (keep it short)

You need: **the work** — an arXiv link or a paper title is enough (then follow `reference/third-party.md` to find source, code and data yourself); a repo, path, project page or readable remote machine also works. **Whose work**: the user's own (first person allowed) or someone else's (third person, faithful, credited — `reference/third-party.md`). **Languages and platforms** (default: an English cut for YouTube and a Chinese cut for Bilibili, both 16:9, plus covers in several sizes and titles/descriptions/chapters for both),  and **narration** (see `reference/voice.md`; default: ElevenLabs if a key is available in env/.env, otherwise edge-tts, and say so). Ask only what you cannot default. Also ask once: first person ("our paper") or third person ("this paper"), and anything that must not appear (unpublished side projects, anonymous submissions, private data).

Treat every remote/shared resource as **read-only** unless told otherwise; list files and sizes before pulling anything large (ask above ~500 MB); never print or commit keys.

## 1. Set up the workshop

1. Create a Remotion project in a new folder (e.g. `npx create-video@latest --blank`, or copy the latest Remotion blank template); do not make the user install anything by hand — install Node deps yourself, and check `ffmpeg`, `python3` and `uv` (install `uv` via its official script or pip if missing).
2. Install the official Remotion agent skills if they are not already available (`~/.agents/skills/remotion-*` or listed in your skills): `npx -y skills add remotion-dev/skills` (see remotion.dev/docs/ai for the current command). Follow them for Remotion API details.
3. Copy `template/` into the project (merge `package.additions.json` into `package.json`, `npm install`), copy `scripts/` to `<project>/scripts/`, create `notes/`, `data/` (raw, git-ignored), `public/data/`, `out/` (current deliverables only, fixed names: `reference/platforms.md` → Files), `review/` (disposable stills and contact sheets), and a project `CLAUDE.md` that records: the goal, audiences, languages, voice choice, the "must not appear" list, source-of-truth document, and any user feedback as it arrives (dated). Keep it current — it is how later sessions (and subagents) stay consistent.
4. `npx tsc --noEmit` and render one still of the template to prove the toolchain works.
5. **Making several videos?** Use a workspace: one folder holding a shared `package.json` / `node_modules` (one install, ~400 MB instead of one per video; all `@remotion/*` at one version) and a shared `.env`, plus a workspace `CLAUDE.md` with the standing rules for every video (Claude Code reads parent `CLAUDE.md` files). Each video is its own folder (`src/`, `public/`, `scenes.json`, `notes/`, `data/`, `out/`, `review/`, `audio_cache/`, `scripts/`, its own `CLAUDE.md`) with `node_modules` and `.env` linked to the workspace root; `scripts/new_video.sh <ShortName>` (copy it to the workspace root) creates one as `<NNN>_<ShortName>`, numbering down from 999 so that sorting by name lists the videos first, newest on top. After changing the shared versions, type-check every video and compare a few rendered frames with its final file.

## 2. Understand the work before writing a word

Collect the material first (`reference/third-party.md` §1–2 if you only have a link or title). Read the paper end to end (and appendix), the README / project page, and skim the code for what the figures are made from. Find the **core finding and why it is surprising**, **what makes the work fun to watch** (the most surprising observation, the cleverest experiment, the example that makes people smile), **the authors' intent** (the question each experiment answers, why they set it up that way, what they tried), the 2–4 claims that carry the story, and the visuals that can show them with *real* data (figure-data exports, result JSONs, logged metrics, demo outputs). Prefer exported figure data or raw results over digitising plots. Write `notes/understanding.md`: the one-sentence story, the claims with their numbers and settings, the candidate visuals, and anything confusing or easy to overstate.

## 3. Story, script and claims

- **Length and focus come first.** **Clear first, no padding**: the length follows the content — explain the method and the technical part properly (that is what a viewer came for), and cut only what adds nothing. Typically 3–5 minutes per cut (≈ 450–750 English words or 1,100–1,600 Chinese characters); no fixed cap, but every extra minute must carry new understanding. A video is not a reading of the paper: pick the **one finding** the whole film serves and the **2–3 pieces of evidence** that carry it, and give everything else one sentence or leave it for the description. Faithful means nothing is distorted, not that every section appears.
- **Make it fun, rigor first.** Choose the most interesting experiments and examples as the evidence beats, and show each the way the authors set it up — the question, what they changed, what happened — so the viewer gets the idea and the authors' intent, not just a number. Humour and surprise come from the real material (a striking example, an unexpected result, a clever control), never from stretching a claim; if a fun line would overstate the paper, drop it.
- **Rough proportions**: hook 15–25 s (the paper's headline result is on screen, as real data, within the first 30 s) · primer 30–40 s · the evidence the largest part · method / mechanism as long as it takes to make it clear, shown rather than recited · takeaway ≤ 15 s. For a well-known older paper, an impact beat (≤ 20–30 s, before the takeaway): what it started or changed, told with sourced facts — named follow-up work with years, where the idea is used now; a citation count only with its source and date; no unsourced superlatives. If the work is a tool or platform viewers can use themselves, add a short how-to (sign up, connect, run a first task — real screenshots, steps from the official docs) and invite them to try it. An impact beat may name one or two notable follow-up papers (no "comment if you want an episode on it" teasers). No limitations scene by default: only when the paper itself treats its limitations as a central part (then ≤ 20 s); otherwise the scope lives in the settings you state with each result. Side experiments and lists of implications get a scene only if they add understanding and can be shown with real data.
- **Cut before you pay**: `python3 scripts/vo.py build --lang <l> --dry-run` prints the film length estimated from reading speed (within ~5 % of the voiced length); trim the script until it fits, then generate the voice.
- **Structure** (adapt, don't force): hook with the strongest real visual (≤ ~15 s) → **a full-screen title card** (`template/src/components/TitleCard.tsx`: the paper's title, authors and institutions, venue or arXiv id, its one-sentence claim (for someone else's paper, "unofficial explainer" goes in the descriptions; on screen only if the user wants it, small); plus the first page of the paper's PDF — title, authors, affiliations, abstract; the first screen of its web page only when there is no PDF — easing in as a clear sheet of paper (solid white, shadow, slight tilt; obvious at a glance as "the paper we are reading", not necessarily readable, never over the text) while the title and institutions are read: `node scripts/paper_shot.mjs <arXiv id | PDF | URL>` → `public/shots/paper.png`, `TitleCard shot="shots/paper.png"`; centre-right at 16:9, under the title at 9:16 and then down under the captions; not on the covers) while the narration names the paper and the authors' institutions (not the authors' names: TTS mangles them, and a name guessed into another script is often wrong — on screen and in the copy, write names exactly as the paper or arXiv spells them, never transliterated) — a scene of its own, on screen by ~20 s at the latest → a primer that lets a no-context viewer follow (what is the task/object, in plain words, animated) → the problem, shown with data → the idea → why it works (mechanism) → results, each with its setting → takeaway, ending on the paper's point. No "links are in the description" / 「论文和代码见简介」 line, spoken or on screen: viewers know where links live, and it ends the film on filler (the description and the title card carry the identifiers). Put good visuals early; don't save them for the end.
- **Write for the ear**: short sentences, one number per sentence, concrete nouns, no "Not X but Y" tics, no fake Q&A, even tone across scene boundaries (an abrupt "Now the fun part!" jars). Give formulas and charts a beat of silence.
- Put the script in `scenes.json` (schema: `template/scenes.example.json`): per scene and language, lines with `text` (captions), `tts` (what is spoken: respellings, sparse audio tags), `anchors` (words that time visual beats), optional `pauseAfterMs`; per scene an optional `chapter`. Write each language natively rather than translating word for word; keep technical terms in English where the audience expects them.
- Start `notes/claims.md` now (`reference/rigor.md`): every number and factual sentence → its source. Write `notes/storyboard.md`: per scene the visual, the anchor-driven beats, the data file, the source line.

## 4. Data → `public/data/*.json`

One agent (you or one subagent) owns turning raw material into clean JSON with a `source` field per file, and writes `notes/data_report.md` that cross-checks every headline number against the paper and lists discrepancies and pitfalls (metric variants, renamed terms, rounding). Scenes only read `public/data`.

## 5. Voice first, then picture

Build the narration before animating (`reference/voice.md`): `scripts/vo.py build --lang <l>` produces per-scene audio and `public/data/vo.<l>.json` with word timings and anchor frames. The picture is then keyed to anchors, so both languages stay in sync automatically. Check the report: ASR round-trip errors, odd prosody, missing anchors; fix wording/respellings, regenerate only affected scenes (outputs are cached by content hash).

## 6. Build the scenes

- You build the theme, shared components and anything reused across scenes first (see `reference/visual-style.md`); then scenes can be parallelised across subagents with **explicit file ownership** (one scene set per agent, nobody edits shared files — they report requested changes back to you). Give each agent: `CLAUDE.md`, the storyboard entry, the `scenes.json` lines and anchor names, the data files, the style rules, the chapter number and title for each of its scenes (parallel agents otherwise number their chapter tags independently), and the requirement to render and inspect stills in every language before reporting (`scripts/review.py`). Afterwards re-run `scripts/zh_font_subset.py` once (new Chinese characters fall back to a system font until then) and, after every Chinese voice build, `scripts/zh_breaks.py` (the captions of both cuts break Chinese only between words).
- The title card's paper page: render it once the paper is fixed (`node scripts/paper_shot.mjs <arXiv id>` downloads the PDF to `data/` and renders page 1; a local PDF or a web article's URL also works; never a hand-made screen grab), pass it as `shot`, key `beats.shot` to the title words (at 9:16 the sheet makes way on the authors beat, `beats.shotMove`); the defaults are meant to work as they are (the text block shrinks itself to clear the captions); check both cuts' stills (the sheet clearly visible, never over the text, captions clear).
- Every scene: headline, beats driven by anchors (with fallbacks), a source line on data shots, SCHEMATIC on illustrations, both languages' on-screen strings, content clear of the caption band.
- Review: render stills at anchor-based frames for the whole film in every language and tile them into contact sheets (`scripts/review.py <Composition> <cut> [--scenes ...]`, one bundle for all frames), and look — overlaps, clipping, empty frames, illegible text, anything off-message. Then review the **pacing**, reading the sheets as the story a viewer gets: flag any stretch over ~20 s without a new real visual, any list of three or more items read out, any formula or theory that is not tied to data on screen, and any scene that does not move the one finding forward — compress or cut those. Iterate.

## 7. Independent fact-check

Spawn a fresh agent that did not build anything to check the script and every on-screen string/number against the paper (prompt in `reference/rigor.md`). One light pass is enough by default: the script and copy right after the script is final (before the paid voice, when wording fixes are cheapest) — numbers, settings, overclaims. Check the screens yourself against `notes/claims.md` while reviewing the sheets; add a second independent pass (screens, plotted data, covers) only for a high-stakes film or when the user asks. Fix MUST-FIX items (including voice-over wording, then regenerate those scenes and re-time the music: `music.py --refit` for a generated bed, the same `--file … --keep-ending N` again for a library track) and most SHOULD-FIX items.

## 8. Sound, render, verify

Music (optional) ducked under speech: by default a royalty-free track (`scripts/music.py --file track.mp3 --keep-ending 20` — CC0 / public domain / CC BY with the credit line in every description; keep a shared library of tracks and their credits in a workspace); composing one with `--generate` costs credits and is opt-in; a few procedural SFX on visual beats (`scripts/sfx_synth.py`, cues in `src/timeline/sfx.ts`; subtle, ≤2 "hits" per film). Render and master with `scripts/finalize.sh` (−14 LUFS / −1.5 dBTP), export captions with `scripts/srt.py`. Verify every file with `scripts/verify.py <mp4> --cut <lang>`: streams, duration against the timeline, loudness, a full-mix ASR pass against the script (with the key terms), and full-frame motion runs (zoom/drift/jitter) to look at. Report honestly what was checked.

## 9. Deliverables and handoff

Per `reference/platforms.md`: horizontal cuts per language, a vertical cut for phones (native 9:16 scenes from the template's `src/vertical/`, same audio, `reference/vertical.md`), covers (16:9 and vertical; `scripts/covers.sh` renders the set into `out/covers/`), SRTs, and `out/social_copy.md` with titles/descriptions/chapters per platform (plus any other platform the user asks for — research its current conventions then). Everything goes into `out/` under the fixed names in `reference/platforms.md` (Files); a new version overwrites the old one — no `_v2` / `_old` copies (move a version you must keep to the Trash or `review/`). If the user wants to finish in an editor or bring their own footage, follow `reference/handoff.md`.

## Efficiency without losing quality
Quality first; cut only waste. Most waste is rework and repetition:
- **Decide before you pay.** Before generating any voice, settle structure and length: hook → title card by ~20 s, one story, explained fully without padding; run `vo.py build --dry-run` for the length estimate; apply the voice's known text rules (respellings, no English in a single-language open-source voice, no fragile short clauses); anchors are words of `text`. A script change after the voice exists costs credits, a re-pick and a re-render.
- **Agents only where they save wall time**: independent heavy work (figure-data extraction across many figures, many scenes in parallel). A ~10-scene film is usually faster built by the lead, which already holds the context; every subagent re-reads the material. Give each one a self-contained brief. One light fact-check, not a checker per stage.
- **Read once, then look things up.** Read the paper yourself, in full, once — as clean text (the LaTeX sections or the extracted article text, not raw HTML with scripts and data URIs). Distil the paper into `notes/understanding.md` and `notes/claims.md`; later use grep / small `sed -n` ranges, inspect a JSON's keys instead of dumping it, and don't re-read files you already summarised.
- **Review what changed.** Contact sheets only for the scenes you touched (`review.py --scenes`, `vreview.py --scenes`, scale 0.5); one full-film sheet before the final render. Render full cuts only after the sheets look right; verify with `--asr none` when the audio did not change.
- **Long jobs in the background** (renders, voice builds, ASR): keep working meanwhile and wait for the completion notice; no sleep/poll loops.
- **Keep tool output short.** Everything printed stays in the context and is re-read on every later call (by mid-film that is a few hundred thousand tokens per call): tail logs, grep for PASS/FAIL/WARN, print counts and diffs of a few lines, write long reports to files and read only the part you need.
- **Never at the expense of the picture**: dense real-data visuals are the point of the film (`reference/visual-style.md` "Density"); spend there.
- **Reuse**: components from earlier films (title card, figure panels, colour bars), voice settings that already worked, music from the library.

## Running on a headless machine
Everything works without a display or GPU: Remotion renders with headless Chrome (downloaded on first use; on Linux it needs the usual shared libraries — `npx remotion browser ensure` reports what is missing), the Python/ffmpeg tools are CLI-only, faster-whisper runs on CPU. Preview by rendering stills/MP4s, or run `npx remotion studio` and forward the port (`ssh -L 3000:localhost:3000 host`). Install Node (nvm), `uv` and a static ffmpeg in the user's home directory if there is no root.

## Iterating with the user

Expect feedback on tone, pacing, clarity and specific frames (they quote timestamps — map them to scene + anchor with `scripts/timeline_info.py`). Record durable preferences in the project `CLAUDE.md`. Re-render only what changed; keep chapters/timestamps in the copy in sync after timing changes.

## Hard-won rules (short list — details in the references)

- Fun from the real material, never at the cost of rigor: the paper's own surprises, clever experiments and examples, shown the way the authors meant them.
- Hook → title → body: within 10–20 s the viewer must know what they are watching (which paper, by whom, what it claims). Keep the hook short, then a full-screen title card that the narration reads out; do this in every cut, vertical included.
- Clear, not padded: the goal is the paper's idea, how it works (the technical part included), its main experiments and its conclusion; length follows that (typically 3–5 minutes). Cut repetition and filler, not explanation — a viewer who didn't understand minute two won't watch minute three.
- Truth over hype: no number without a source; don't state claims more strongly than the paper; say the setting behind each number (model, task, metric). Skip the limitations section unless the paper makes it central.
- Real data on real-data charts; illustrations labelled SCHEMATIC; no count-up number animations; bars from zero unless the axis says log.
- No full-frame shake/punch-in/zoom effects — emphasise the specific card; full-frame scale changes also caused visible text jitter in renders.
- The production tool is a credit line, not the pitch (unless the user explicitly wants that angle).
- Never print, log or commit keys; check paid-API quota before every batch and stop if it would be exceeded.
- Keep the disk clean: renders and bundles are big. Use `scripts/stills.mjs` (it deletes its bundle and closes its browser, also on Ctrl-C); never call Remotion's `bundle()` without deleting the result; review stills go to `review/`, not scratch or system temp folders; check free space before a full render. Close what you start when you are done with it (Remotion Studio, background renders, servers, background shells) — a leftover headless browser keeps eating memory. If the user keeps work on an external/NAS volume, point temp files and tool caches there too: set `PAPER_VIDEO_TMPDIR` (the template's `remotion.config.ts`, `stills.mjs` and `common.py` copy it into `TMPDIR`; some launchers, Claude Code included, reset `TMPDIR` itself) and `UV_CACHE_DIR` / `HF_HOME` / `npm_config_cache` in `.claude/settings.json` → `env`.

Files in this skill

  • SKILL.md22.1 KB
  • reference/handoff.md1.8 KB
  • reference/platforms.md7.2 KB
  • reference/rigor.md3.6 KB
  • reference/third-party.md10.5 KB
  • reference/vertical.md10.8 KB
  • reference/visual-style.md5.3 KB
  • reference/voice.md15.8 KB
  • scripts/README.md8.3 KB
  • scripts/common.py6.8 KB
  • scripts/contact_sheet.py1.3 KB
  • scripts/covers.sh2.5 KB
  • scripts/export_handoff.py7.4 KB
  • scripts/figdata.py16.5 KB
  • scripts/finalize.sh2.2 KB
  • scripts/finalize_vertical.sh911 B
  • scripts/music.py15.6 KB
  • scripts/new_video.sh3.7 KB
  • scripts/paper_shot.mjs10 KB
  • scripts/review.py3 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…