---
name: generating-ugc-videos
description: Generates a video of a person on camera with a product. Use when the user wants to create a video of a person using, wearing, opening or demonstrating a product. This skill should also be used when the user wants to create a UGC video, creator video, influencer ad, tiktok-style review, unboxing, try-on, OOTD, fit check, haul, talking head, creator testimonial or something similar. Not for a video with no person in it, or a person with no product to show.
license: Apache-2.0
metadata:
  version: "0.6.0"
  category: creative
  summary: "Produces UGC videos of any length - like review, unboxing, try-on, tutorial and more. It casts the actor, writes the script, and storyboards multiple clips to lock in actor and product consistency. It lets you direct every step to get a cohesive, ready-to-run ad in minutes."
---

# UGC Video

Turn a product into a video of a creator on camera.

Videos that do not have a person on camera with a product are out of scope — those should be generated by the
`generating-videos` skill.

**Settled for every job — never ask about these, and never re-decide them:** the delivered video is
`9:16` unless the user explicitly asked for `16:9` · clips are `seedance-2.0-fast` with audio on ·
storyboard sheets are `gpt-image-2` at `16:9` · how many clips there are comes from the length · the
clips are always joined, hard cuts, no transitions.

## Workflow

### Step 1: Read what you have

- **A product image or URL** → hand it to `analyzing-products` for what the product is, how a person
  physically uses it, which parts open or move, and what must stay identical wherever it appears.
- **Run `image_analysis` on every image supplied**, not only the product — what each one shows, and
  whatever the product facts leave out.
- **What each image is for comes from the brief, not from its contents.** A person in a frame does
  not make it a casting photo, and a second object does not make it a prop.
- **No product** → don't guess at one. It becomes the first thing Step 2 asks for.

### Step 2: Interview

**Skip this when the brief already settles the video** — a clear product, a clear format, a length.

Otherwise ask once, bundled into a single message, always with a free-text way out.

| Ask | When |
| --- | --- |
| **The product** — a link or a photo | Neither was supplied. Offer to wait for an upload; a photographed product beats a described one. |
| **How long** | The brief doesn't say. Offer 15s, 30s, 45s or 60s, and let them type their own. |
| **Whether the person in a supplied photo should be the creator** | An image with a person in it was supplied and the brief doesn't say who they are. Ask rather than assume either way. |
| **Whether they have a photo of the delivery package** | The brief is about opening a package and none is attached. If they have one, **wait for it to arrive** before building anything — a promised photo is not a photo. If they don't, a plain unbranded box stands in. |

When the user waives questions, take 15s, say in one sentence what you took, and keep going.

### Step 3: Select the mode and read the corresponding reference file

| The ask is about… | Mode | Reference |
| --- | --- | --- |
| Holding the product and talking about it | **review** | `references/mode-review.md` |
| Opening the package it arrived in | **unboxing** | `references/mode-unboxing.md` |
| Wearing it — clothing, shoes, jewellery, accessories | **try-on** | `references/mode-try-on.md` |
| Showing how to use it, one step at a time | **tutorial** | `references/mode-tutorial.md` |

**Where two rows both fit, these win:**

- **A sealed package is opened on camera** → **unboxing**.
- **The product ends up worn** → **try-on**, even when a bag is opened to reach it.
- **Numbered steps, or "how do I use this"** → **tutorial**, whatever the tone.
- **Nothing to open, wear or work through** → **review**.

### Step 4: Write the product description

Write one description and reuse it unchanged in every storyboard and every clip prompt — rephrasing
it between calls reads as a different product.

- **Material and finish**, surface by surface — matte or gloss, brushed, woven, translucent, or
  whatever this one actually is: *"brushed steel barrel, matte soft-touch collar, a woven wrist
  strap"*.
- **Size** — its width and height, and how it sits against a hand: *"sits in a closed palm, roughly
  9 cm tall and 4 cm across"*. Give the real measurements rather than likening it to another object,
  which the model renders inconsistently.
- **Two to five visual anchors** — features you can verify in the product image and nothing else:
  exact colours, the shape of a closure or handle, the gauge of a chain or strap, a surface finish, a
  distinguishing mark. These carry into every sheet and every clip prompt in the same words.
- **The product mechanism** — how it's built and how it works: the parts that move, how they open or
  close, and where anything dispenses: *"hinged lid at one end, folds back flat; the brush sits inside
  the cap"*.
- **What may be done with each part** — which parts are fixed to it, and the whole of what the moving
  ones allow. Nothing later does anything to the product that isn't on this list.

**Don't transcribe what the label says**, even where the product facts spell it out. Spelling out the
printing invites the model to redraw it, and redrawn text comes back warped. The facts may name it;
the description never repeats it.

### Step 5: Write the concept

**What the user asked for wins.** Wherever the brief is specific — a scenario, a shot list, a named
place, a line they want said, a mood — follow it exactly and keep following it for the rest of the
job. Everything it doesn't cover is yours to decide.

**Who the creator is, and where.** A brief that names them — the founder, the owner, a customer, a
specialist, someone the audience would recognise themselves in — settles it. Where it doesn't, pick
whoever this video is most believable coming from, and put them where it would actually happen: a
room, a workshop, a shop floor, a street, outdoors, etc.

**How many clips.** One storyboard becomes one clip, and a clip runs at most 15 seconds. Fill whole
15-second clips and let the leftover be the last one — 30s is two clips of 15, 50s is 15, 15, 15 and
5. Where the leftover would land under 4 seconds, give the last clip 4 seconds anyway and let the
video run slightly over: 18s becomes 15 and 4.

**What the story can hold.** A **primary action** is one continuous motion at one object — picking a
thing up, tilting it, putting it on, pointing at it. Each one needs two to four seconds, so a clip's
length is the number of primary actions it can carry. Write the story to that number.

These never render, so no part of the story asks for them:

- Anything finer than a whole hand movement — working a clasp, a zip, a drawstring, a pump head.
- Two motions treated as one, or a second object handled in the same breath as the first.
- **A change of state.** Where the product needs to be open, worn or empty, the panel arrives with it
  already so; the transition itself is never rendered.
- **No going back.** Once open, worn or emptied, it holds that state — nothing later reseals, removes
  or restores it.

**Write the concept down.** It is the one description of the video, and every step after this works
from it rather than deciding again:

- who the creator is
- how many clips, and how many seconds each
- what happens in each clip, following the mode guide's arc, within what that clip's length can hold
- roughly how many words each clip can carry, from its seconds

The clip lengths and the story are settled here. How each clip divides into panels is not — that is
the storyboard's to work out. Where the story itself has to change, change it here and say so.

### Step 6: Write the script

Trigger the `writing-video-scripts` skill workflow with the concept from Step 5 — the clip durations,
the beats and the word budget — spoken on camera and lip-synced, and only the claims the user
supplied, used as written. None supplied means none made.

### Step 7: Show the concept and wait

Nothing has been generated yet, and everything after this step costs money. Get the concept approved first.

Show, written out, in one brief, to the point message:

- **The creator** — who they are, their age, enthnicity, how they look, what they're wearing, where they are, etc.
- **What's on screen in each clip** — the setting, what happens, where the product is. Keep this brief and focus on the high level concept rather than stating all small technical details. 
- **What's said in each clip** — the script from Step 6, word for word.

Then ask whether it's right, and say plainly that they can change any of it — the person, the
setting, a beat, or a line. **Expect edits.** Rewrite what they change, show it again, and keep going
until they approve.

Wait for an answer. Don't cast, don't build a sheet, don't generate a clip.

Where the user has said they don't want to be asked, say in one line what you're going with and
carry on.

### Step 8: Cast the creator

**Where the user has already settled who that is** — named in the brief, or confirmed in Step 2 —
their supplied photo *is* the creator image. Skip this step.

Otherwise trigger the `generating-ai-actors` skill workflow:

1. The person from the approved concept — same age, wardrobe and setting, not a fresh invention.
2. Their hands are empty. The product arrives later as its own reference.
3. A phone self-portrait in available light, not a studio portrait.
4. `nano-banana-pro` at `3:4`.

Keep the result. It is the identity reference for every storyboard and every clip that follows.

### Step 9: Build the storyboards

Trigger the `generating-storyboards` skill workflow with: one sheet per clip, what happens in that
clip and how many seconds it runs, which sheet this is and how many there are, the product image(s)
and the creator image from Step 8 as references, that clip's script segment from Step 6, and the
Step 4 product description carried in unchanged. How the clip divides into panels is worked out
there, not here. The sheet's own shape is not yours to set — it is always `16:9` with vertical
panels, whatever ratio the video is delivered in.

### Step 10: Write the clip prompts

Trigger the `writing-video-prompts` skill workflow with: one prompt per storyboard, one cut per panel
of that sheet, `seedance-2.0-fast` with audio on at that clip's duration from the concept and the delivery
ratio, that sheet's panels and the clip's script segment from Step 6, the Step 4 product description
carried in unchanged, and the images that go with it — that storyboard, the creator image, and the
product image(s).

### Step 11: Generate

One `video_generate` call holding one request per storyboard, so they all render at once.

- Per request: the `prompt` from Step 10 · `model` `seedance-2.0-fast` · `duration` for that clip from the
  concept · `generate_audio` on.
- `reference_images` per request: that storyboard, the creator image, and the product image(s), in
  the order they were labelled in the prompt.
- `aspect_ratio` the delivery ratio, the same on every clip. The sheet's own `16:9` shape is the plan
  for the clip and never the shape of the clip itself.
- **Everything goes in as a reference. Nothing goes in as a start frame** — a start frame and
  references sit on different routes and cannot be combined in one call.
- Never generate a second version to compare takes.
- Call `list_video_models` for accepted values rather than guessing — a wrong enum fails a billed
  call.

A clip can come back as `{status: "pending", …}` — a job handle, not a failure. Pass it to
`job_status`, and again if it is still pending. **Never re-run a pending clip**; that abandons a job
you are already paying for.

### Step 12: Stitch and return

Once every clip has finished, join them in order with `video_stitch` — hard cuts, no transitions. A
single clip is already the video.

Hand over the finished video. Don't present storyboards, separate clips, or job handles unless asked.

## Edge cases

- **A generation is refused as unsafe** → usually the creator. Regenerate it in more covering
  clothing, then rebuild the storyboards that used it. Identical inputs fail identically, so change
  something. After a second refusal, name the blocked element rather than retrying blind.
- **A clip is refused because a reference image may show a real person** → not the wardrobe and not
  the prompt; neither changes it. Recast the creator as a different person, rebuild every storyboard
  that used the old one, and resubmit. If it happens again, switch the clip model rather than
  recasting a third time — the one case that overrides `seedance-2.0-fast` being settled.
- **A product URL won't load** → ask for a photo instead.
- **A clip comes back wrong** → deliver it and say what you think is off. Judging taste is the
  user's job.
- **`reference_images` rejected on count** → the error states the limit; drop to it.
- **`error: "no_provider_configured"`** → relay the tool's `hint` (the user must set their key).
- **`analyzing-products` or `video_stitch` unavailable** → do the rest and say which step couldn't run.

## Reference

**Mode guides — read the one you routed to in Step 3.** Each one carries the arc: what the video is
doing, in what order, and what the panels hold.

- `references/mode-review.md` — a creator arguing for a product they already have.
- `references/mode-unboxing.md` — a sealed package opened on camera.
- `references/mode-try-on.md` — a product worn on the body.
- `references/mode-tutorial.md` — a product used, one step at a time.
