Write MiniMax H3 video generation prompts from a VideoProjectSpec. Use when a video project targets MiniMax H3 (Hailuo) and needs a mode-selected, grammar-compliant prompt for text-to-video, first-frame, last-frame, first-last-frame, or Ref2VA generation, including reference asset roles, 4-15 second duration handling, and 32 kHz stereo audio direction. Works in Claude Code, Codex, Cursor, and any agent that can read local files.
Installs into .claude/skills of the current project.
Are you the author of Minimax H3?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/xxjrq-minimax-h3)
---
name: minimax-h3
description: Write MiniMax H3 video generation prompts from a VideoProjectSpec. Use when a video project targets MiniMax H3 (Hailuo) and needs a mode-selected, grammar-compliant prompt for text-to-video, first-frame, last-frame, first-last-frame, or Ref2VA generation, including reference asset roles, 4-15 second duration handling, and 32 kHz stereo audio direction. Works in Claude Code, Codex, Cursor, and any agent that can read local files.
---
# MiniMax H3 Prompt Writing
Turn a `VideoProjectSpec` (see `schemas/video-project.schema.json`) into a
MiniMax H3 generation prompt. The full grammar lives in
`models/minimax-h3/prompt-grammar.md`; read it before writing a prompt. This
skill is written in this repository's own words and does not copy MiniMax's
prompt-skill files.
## When to use
- The project targets MiniMax H3 (`spec.model == "minimax-h3"`).
- You need one of the five official H3 modes: `text-to-video`, `first-frame`,
`last-frame`, `first-last-frame`, or `ref2va`.
- You have reference assets (keyframe images, style images, reference videos,
audio) that must be wired into the prompt with correct roles.
## Official capability facts
All facts below are from the official MiniMax H3 repository
(source_url: https://github.com/MiniMax-AI/MiniMax-H3, checked_at: 2026-08-19).
| Capability | Official value |
| --- | --- |
| Output duration | 4–15 seconds (hard limit) |
| Output audio | Native 32 kHz stereo, generated together with video |
| Output resolution | 768p shorter side by default (H3-Base); 2K via H3-Regenerate-2K |
| Output frame rate | 24 FPS |
| Aspect ratios | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 (and others) |
| Multimodal context | Text, images, video, and audio understood together |
| FL2VA modes | 0 images = text-to-video; 1 image = first-frame or last-frame; 2 images = first-and-last-frame |
| Ref2VA limits | ≤ 9 images, ≤ 3 videos, ≤ 3 audio, total ≤ 12 files; each video/audio clip 2–15 s, total ≤ 15 s |
| H3-Context-IR | Hosted only, NOT open-sourced; API reproduces the official workflow |
| H3-Regenerate-2K | NOT open-sourced; API available |
| Full 2K pipeline | Requires the official API combined with a locally deployed H3-Base |
## Hard rules
1. **Duration MUST be within 4–15 seconds.** A 3-second or 16-second request
is invalid. Never clamp, round, or silently adjust a duration.
2. **H3-Context-IR and H3-Regenerate-2K are not open-sourced.** Never claim
the complete 2K pipeline is locally open-sourced. Local H3-Base output is
768p; 2K requires the official API.
3. **Ref2VA file limits are hard:** ≤ 9 images, ≤ 3 videos, ≤ 3 audio, total
≤ 12 files. Each reference video/audio clip must be 2–15 seconds long with
total reference duration ≤ 15 seconds.
4. **FL2VA modes accept only keyframe images:** `first-frame` = exactly one
first-frame image, `last-frame` = exactly one last-frame image,
`first-last-frame` = exactly two images (one first, one last). No video or
audio references in FL2VA modes.
5. **Every dynamic capability claim carries a source.** Use
`source_url: https://github.com/MiniMax-AI/MiniMax-H3` and
`checked_at: 2026-08-19` for facts from the official repo.
6. **No fabricated endpoints or pricing.** API endpoints and pricing come only
from official MiniMax platform docs; do not invent them.
## How to build a VideoProjectSpec
A valid spec needs at least: `spec_version`, `project_goal`, `model`
(`"minimax-h3"`), `duration_seconds` (4–15), `aspect_ratio`, `creative`
(`subject`, `scene`, `action` required), and `evidence`. Add:
- `references[]` — each entry `{id, type, role, uri}`. Roles: `first-frame`,
`last-frame`, `style`, `subject`, `motion-reference`, `soundtrack`,
`reference`.
- `frames` — `first_frame_asset_id` / `last_frame_asset_id` pin keyframes;
`continuity_requirements[]` lists what must survive cuts.
- `audio` — `dialogue` (with speaker and language), `ambient`, `music`.
- `constraints` — `preserve[]`, `modify[]`, `forbidden[]`.
- `shots[]` — optional inline timeline; each shot has `description`, `camera`,
`duration_seconds`, `order`.
## Mode selection
Pick the mode from the intent and the reference set:
| Intent | Mode |
| --- | --- |
| New scene from text only, no reference media | `text-to-video` |
| Start from a specific first frame and develop forward | `first-frame` |
| Land on a specific final frame at the end | `last-frame` |
| Interpolate a continuous path between a first and a last frame | `first-last-frame` |
| Combine images + videos + audio (style, subject, motion, soundtrack) | `ref2va` |
Mode is inferred from references: no references → `text-to-video`; one
first-frame image → `first-frame`; one last-frame image → `last-frame`; two
keyframe images → `first-last-frame`; any video/audio reference, more than two
images, or non-keyframe image roles → `ref2va`.
## How to compile the final prompt
1. **Validate duration** — reject anything outside 4–15 s.
2. **Collect references** — merge the manifest and `spec.references`; apply
`frames` role overrides.
3. **Resolve the mode** and validate its reference rules.
4. **Build the prompt** in English (dialogue stays in its original language
inside `<d>`):
- Base modes (`text-to-video`, `first-frame`, `last-frame`,
`first-last-frame`): optional alignment instruction line, then exactly
three fields in order:
`integrated_multimodal_description`, `overall_soundscape`,
`non_diegetic_music`.
- `ref2va`: exactly six sections in order: `subject_definitions`,
`summary`, `retention_analysis`, `detailed_description`,
`overall_soundscape`, `non_diegetic_music`.
5. **Emit the generation request** with `model: "minimax-h3"`, `mode`,
`prompt`, `reference_manifest`, `duration_seconds`, `resolution_hints`,
`provider`, and `evidence`.
## Prompt writing rules
- `[Shot 1]` has no timestamp; later shots use `[Shot N] At MM:SS.mmm, ...`
with strictly increasing cut times inside the duration.
- State style and initial composition at the start of `[Shot 1]`
(e.g. `Live-action, cinematic, a medium-wide shot frames ...`).
- Write camera motion as natural English: motion type + optional amplitude +
optional speed (e.g. `The camera pushes in with small amplitude at slow
speed`).
- Give speakers stable IDs `(S1)`, `(S2)`, ...; write dialogue as
`<d>[Language] exact words</d>` and never translate it.
- Put on-screen text in English double quotation marks, verbatim.
- `overall_soundscape`: 1–4 sentences of ambience and physical sounds; use
`N/A` only for complete silence.
- `non_diegetic_music`: 1–3 sentences of instrumentation, tempo, dynamics; use
`N/A` when there is no audience-only score.
- Ref2VA: define `<Subject N>` / `<Picture N>` / `<Video N>` / `<Audio N>`
labels once in `subject_definitions`, keep them consistent across all six
sections, and use `retention_analysis` markers (`fully_preserved`,
`partially_preserved`, `attribute_transfer`, `weak_reference` for visible
content; `fully_copy`, `partially_copy`, `reference`, `weak_reference` for
audio).
- Match the description's timeline to the requested duration; the final shot
must land exactly at the end.
## Do
- Keep the duration inside 4–15 s and state it explicitly.
- Describe sound in every prompt — H3 generates native 32 kHz stereo audio.
- Preserve dialogue and visible text verbatim in their original language.
- Fold continuity requirements and constraints into the description text.
- Mark anything not in the official repo as "community practice, not
official".
## Don't
- Don't write a prompt for a duration outside 4–15 s.
- Don't claim the full 2K pipeline is locally open-sourced, or that
H3-Context-IR / H3-Regenerate-2K are open-sourced.
- Don't exceed Ref2VA limits (9 images / 3 videos / 3 audio / 12 files).
- Don't use video or audio references in FL2VA modes.
- Don't invent API endpoints, pricing, or third-party capabilities.
- Don't copy MiniMax's prompt-skill files verbatim; write original prompts
following this grammar.
- Don't use movie/TV/brand IP, celebrities, or unlicensed assets in examples.