> ## Documentation Index
> Fetch the complete documentation index at: https://heygen-1fa696a7.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# HeyGen Video

> HeyGen Video, built on MiniMax H3 and post-trained by HeyGen, generates a whole scene from a prompt: subject, setting and sound together, in one call. Text to video, image to video and reference to video.

<img className="w-full h-44 object-cover rounded-xl" src="https://mintcdn.com/heygen-1fa696a7/hfMXXwJzjE7vBSYZ/images/theme/research-3.webp?fit=max&auto=format&n=hfMXXwJzjE7vBSYZ&q=85&s=c8191ba4743fd85050f1af1a93ed4969" alt="" noZoom width="1400" height="788" data-path="images/theme/research-3.webp" />

**HeyGen Video is our general-purpose video model, built on MiniMax H3 and post-trained by HeyGen.** Give it a prompt and it generates the whole scene: subject, setting, light and sound together, in one call. Each clip runs 5 to 15 seconds with its own audio track of dialogue, ambience and sound effects, so there is no separate speech pass, lip-sync step or avatar to pick.

One model, `heygen-video-1`, called in three ways:

| Mode | You give it | You get |
| - | - | - |
| **Text to video** (`text_to_video`) | A prompt | A new scene with its sound |
| **Image to video** (`image_to_video`) | A prompt and one image | A clip that opens on your image as its first frame |
| **Reference to video** (`reference_to_video`) | A prompt and your images, videos or audio | A new scene that keeps the products, people or places you supplied |

It is a different tool from the [avatar rendering engines](/models). Those animate a look you already own from a script you already wrote. This one invents the subject, the setting, the light and the sound at once, from a description. The rest of this page starts with the call itself, then what the model is good at and how to prompt it, then the full API reference.

export const HERO = [{
  "src": "https://resource2.heygen.ai/videos/66c20e44a62c40a58cf8c51210dab761/66c20e44a62c40a58cf8c51210dab761/01a0f043-0a69-7f00-9129-2c3226ffd9f4.mp4",
  "poster": "/images/heygen-video/L10_train_c.jpg",
  "label": "The audio comes from the same prompt",
  "teaches": "Near silence first, then each sound of the arrival, in order",
  "prompt": "An empty suburban railway platform at dawn, completely still, wet concrete and a yellow safety line, overhead wires, a pole-mounted platform display that reads REALISTIC SOUND in amber LED capitals. Static camera on a tripod at platform height, Shot on an ARRI Alexa Mini, 35mm at f/4, fine film grain, 24 fps, muted true-to-life colour, no teal-orange grade, no CG sheen. Soft overcast daylight. For the first three seconds nothing moves and it is almost silent: faint wind, one distant bird. Then a modern commuter train with a plain unmarked body sweeps in from the left and fills the frame. Audio: a horn blast, wheels hammering the rail joints, a rush of air, brakes squealing as it slows. No people. No logos, brand names, printed words or badges anywhere in frame except the display. No music."
}, {
  "src": "https://resource2.heygen.ai/videos/66c20e44a62c40a58cf8c51210dab761/66c20e44a62c40a58cf8c51210dab761/01a0f07c-1ef6-7718-b775-2eaaf39f3252.mp4",
  "poster": "/images/heygen-video/L16_cello_b.jpg",
  "label": "Naming the lens does the work",
  "teaches": "A body, a focal length and a stop set the look",
  "prompt": "A cellist in a rehearsal room draws one long bow stroke, seen from the side at medium distance, face turned away. Shot on an ARRI Alexa Mini LF, 50mm T2, locked off on a tripod. Warm lamp light and a window. Fine film grain, true-to-life muted colour, no teal-orange grade, no retouching, no beauty filter, no glamour lighting, real textures and materials. One clear subject, simple uncluttered frame, natural physical motion. Faces small or turned away. No logos, brand names, printed words or badges anywhere in frame. No music."
}, {
  "src": "https://resource2.heygen.ai/videos/66c20e44a62c40a58cf8c51210dab761/66c20e44a62c40a58cf8c51210dab761/01a0f07b-7dba-74e5-a0f0-62a03a4edd7b.mp4",
  "poster": "/images/heygen-video/L16_potter_a.jpg",
  "label": "Asking for texture, then forbidding the fix",
  "teaches": "Real materials in; retouching and glamour light ruled out",
  "prompt": "Wide shot of a potter at a spinning wheel in a quiet studio, a clay vessel rising between the potter's hands as the wheel turns. Shot on an ARRI Alexa Mini LF, 35mm T2.8, locked off on a tripod. Soft window daylight. Fine film grain, true-to-life muted colour, no teal-orange grade, no retouching, no beauty filter, no glamour lighting, real textures and materials. One clear subject, simple uncluttered frame, natural physical motion. Faces small or turned away. No logos, brand names, printed words or badges anywhere in frame. No music."
}, {
  "src": "https://resource2.heygen.ai/videos/66c20e44a62c40a58cf8c51210dab761/66c20e44a62c40a58cf8c51210dab761/01a0f043-b312-7ac6-bb47-254fd2364f99.mp4",
  "poster": "/images/heygen-video/L10_dish_a.jpg",
  "label": "Stillness reads as production value",
  "teaches": "A static camera, one slow push, nothing handled",
  "prompt": "A plated restaurant dish on a matte ceramic plate on a dark wooden table, steam rising gently from it, warm low restaurant light, the room softly blurred behind. Shot on an ARRI Alexa Mini, 85mm at f/2, fine film grain, 24 fps, muted true-to-life colour, no teal-orange grade, no CG sheen. Static camera, slow very slight push in. No people in focus, no hands handling anything. No logos, brand names, labels, printed words, numbers or signage anywhere in frame. No music."
}, {
  "src": "https://resource2.heygen.ai/videos/66c20e44a62c40a58cf8c51210dab761/66c20e44a62c40a58cf8c51210dab761/01a0f028-1eff-7bc1-bf91-e344678a8ffd.mp4",
  "poster": "/images/heygen-video/L_ttq_d.jpg",
  "label": "Short lettering, spelled out",
  "teaches": "Three capitalised words, spelled exactly, and nothing else legible",
  "prompt": "subject_definitions:\n<Subject 1> is a real unbranded Formula One car in one clean matte deep ink-navy paint with a single thin bright-cyan stripe, bare carbon fibre and plain black tyres, with no decals, numbers or markings anywhere.\n<Subject 2> is a straight of real race-track asphalt beside a red-and-white painted kerb, with the words TOP-TIER VIDEO QUALITY painted across it in huge white road-marking capitals on three lines, TOP-TIER above VIDEO above QUALITY, the hyphen in TOP-TIER clearly painted, the paint slightly worn and crossed by black rubber tyre marks.\n\nsummary:\n[reference generation] The target video is a single locked-off top-down drone shot looking straight down at the words TOP-TIER VIDEO QUALITY on the track, as <Subject 1> blasts straight over the letters from the bottom of frame to the top at racing speed.\n\nretention_analysis:\n<Subject 2> (appears in [Shot 1]): fully_preserved - the painted words stay legible and spelled exactly TOP-TIER VIDEO QUALITY, with the hyphen, every letter correct.\n\ndetailed_description:\nThe target video is a real overhead drone shot on a DJI Inspire 3 with a Zenmuse X9 full-frame camera and a 24mm lens, perfectly static, hard noon sun with a crisp short shadow under the car, fine organic film grain, true-to-life colour with gentle contrast and natural highlight roll-off, no glossy commercial grade, real material textures, visible asphalt grain and rubber on the racing line, the red-and-white kerb along one edge. No readable text anywhere except TOP-TIER VIDEO QUALITY; no logos, no brand marks, no signage, no watermarks, no captions; hands have exactly five fingers; nothing CG-looking, no plastic skin, no warped geometry.\n[Shot 1] The painted words lie still; the car crosses the whole frame over them in well under a second, its shadow sliding over the paint and a faint puff of heat haze behind it; then the words lie still and alone on the track for the rest of the shot.\n\noverall_soundscape:\nAn engine howl rising from nothing, a violent scream passing directly overhead in a split second, then fading away to the faint tick of hot asphalt and wind.\n\nnon_diegetic_music:\nN/A"
}, {
  "src": "https://resource2.heygen.ai/videos/66c20e44a62c40a58cf8c51210dab761/66c20e44a62c40a58cf8c51210dab761/01a0edd0-2a90-7d63-9a7a-92d97b02bb78.mp4",
  "poster": "/images/heygen-video/Q_in_c.jpg",
  "label": "Naming a real medium, not a vibe",
  "teaches": "Name the print process, the marks it leaves and the inks it may use",
  "prompt": "subject_definitions:\n<Subject 1> is the a hand-printed flat illustration in the style of a mid-century screen print, animated.\n<Subject 2> is the scene rendered in it: a side view of a quiet room with a wooden desk on the left and a round wastepaper basket on the floor on the far right; a hand at the desk throws a folded paper plane toward the basket; the plane glides in a clean arc and drops straight into the basket. The camera holds still.\n\nsummary:\n[reference generation] The target video is a single continuous flat screen-printed illustration animation sequence in which <Subject 2> is rendered entirely through <Subject 1>, with the physical artifacts of the medium visible at all times.\n\nretention_analysis:\n<Subject 1> (appears in [Shot 1]): fully_preserved - the medium's surface, marks and limitations are visible in every frame and are never replaced by a smooth digital render.\n<Subject 2> (appears in [Shot 1]): fully_preserved - the action reads clearly within the constraints of the medium.\n\ndetailed_description:\nThe target video is entirely a flat screen-printed illustration animation. Bold simple shapes with crisp edges and no outlines, visible halftone dot screens and fine paper grain in every fill, slight misregistration between colour layers, gentle limited animation on twos. Only these inks: deep navy #0B0F19 as the ground, bright cyan #12D8F5, fresh mint green #35D9A8, soft pink #E890F8 and warm off-white #F2F5F7. No other colours: no red, orange, brown or yellow. No lettering, logos or numbers anywhere.\n[Shot 1] a side view of a quiet room with a wooden desk on the left and a round wastepaper basket on the floor on the far right; a hand at the desk throws a folded paper plane toward the basket; the plane glides in a clean arc and drops straight into the basket. The camera holds still. The framing holds one continuous view for the full duration with no cut.\n\noverall_soundscape:\nA soft flutter and a soft paper landing. No music.\n\nnon_diegetic_music:\nN/A"
}];

export const STYLES = [{
  "src": "https://resource2.heygen.ai/videos/66c20e44a62c40a58cf8c51210dab761/66c20e44a62c40a58cf8c51210dab761/01a0e650-104e-7238-a596-5da8892cefce.mp4",
  "poster": "/images/heygen-video/L_visor_b.jpg",
  "label": "Live-action cinema",
  "note": "Brand films and launches that should look shot on set"
}, {
  "src": "https://resource2.heygen.ai/videos/66c20e44a62c40a58cf8c51210dab761/66c20e44a62c40a58cf8c51210dab761/01a0e5e7-2bf3-78c1-9d97-a77ab17721e8.mp4",
  "poster": "/images/heygen-video/b_visor.jpg",
  "label": "Plasticine stop-motion",
  "note": "Explainers and onboarding that should feel warm and handmade"
}, {
  "src": "https://resource2.heygen.ai/videos/66c20e44a62c40a58cf8c51210dab761/66c20e44a62c40a58cf8c51210dab761/01a0e65b-f151-780d-b469-4b0c577c5a0c.mp4",
  "poster": "/images/heygen-video/A_visor_b.jpg",
  "label": "Hand-painted cel anime",
  "note": "Campaigns for younger audiences and fan communities"
}, {
  "src": "https://resource2.heygen.ai/videos/66c20e44a62c40a58cf8c51210dab761/66c20e44a62c40a58cf8c51210dab761/01a0e65a-cc3e-7c65-bd29-8b1bb2433422.mp4",
  "poster": "/images/heygen-video/D_visor_b.jpg",
  "label": "Pencil, fineliner and marker animation",
  "note": "Storyboards and concept pitches before anything is shot"
}, {
  "src": "https://resource2.heygen.ai/videos/66c20e44a62c40a58cf8c51210dab761/66c20e44a62c40a58cf8c51210dab761/01a0ed72-1372-7e93-8990-880da728f121.mp4",
  "poster": "/images/heygen-video/P_city_a.jpg",
  "label": "Mid-century screen print",
  "note": "Internal comms and culture films with a poster-art voice"
}, {
  "src": "https://resource2.heygen.ai/videos/66c20e44a62c40a58cf8c51210dab761/66c20e44a62c40a58cf8c51210dab761/01a0e66b-070f-7480-b1ee-966d6e846a90.mp4",
  "poster": "/images/heygen-video/A_k1.jpg",
  "label": "Anime key drawing",
  "note": "Character and product design sign-off before animation"
}, {
  "src": "https://resource2.heygen.ai/videos/66c20e44a62c40a58cf8c51210dab761/66c20e44a62c40a58cf8c51210dab761/01a0e5e8-32a4-7686-b95f-0073913c6384.mp4",
  "poster": "/images/heygen-video/b_k1.jpg",
  "label": "Pencil sketch",
  "note": "Design reviews that show an idea at its earliest stage"
}, {
  "src": "https://resource2.heygen.ai/videos/66c20e44a62c40a58cf8c51210dab761/66c20e44a62c40a58cf8c51210dab761/01a0e5e8-5982-7d8b-aa2c-23cece6db46b.mp4",
  "poster": "/images/heygen-video/b_k2.jpg",
  "label": "Marker rendering",
  "note": "Industrial and product design concepts"
}, {
  "src": "https://resource2.heygen.ai/videos/66c20e44a62c40a58cf8c51210dab761/66c20e44a62c40a58cf8c51210dab761/01a0e5e8-9f4e-7e24-8664-d37fd7c091d9.mp4",
  "poster": "/images/heygen-video/b_k3.jpg",
  "label": "Watercolour wash",
  "note": "Travel, hospitality and wellness brands"
}, {
  "src": "https://resource2.heygen.ai/videos/66c20e44a62c40a58cf8c51210dab761/66c20e44a62c40a58cf8c51210dab761/01a0e5e8-c694-7052-baab-4430bac69f1f.mp4",
  "poster": "/images/heygen-video/b_k4.jpg",
  "label": "Impasto oil",
  "note": "Arts, heritage and luxury storytelling"
}, {
  "src": "https://resource2.heygen.ai/videos/66c20e44a62c40a58cf8c51210dab761/66c20e44a62c40a58cf8c51210dab761/01a0e674-a619-7682-84b0-bd4092669bad.mp4",
  "poster": "/images/heygen-video/A2_k5.jpg",
  "label": "Painted anime background",
  "note": "Setting the world for character-led series"
}, {
  "src": "https://resource2.heygen.ai/videos/66c20e44a62c40a58cf8c51210dab761/66c20e44a62c40a58cf8c51210dab761/01a0e5e9-0d00-7a89-895f-a4d85b494ff4.mp4",
  "poster": "/images/heygen-video/b_k5.jpg",
  "label": "Plasticine model",
  "note": "Product concepts shown as a physical object before tooling"
}];

export const BRANDS = [{
  "src": "https://resource2.heygen.ai/videos/66c20e44a62c40a58cf8c51210dab761/66c20e44a62c40a58cf8c51210dab761/01a0e64f-f8e0-7bab-866a-727d2cbd0b20.mp4",
  "poster": "/images/heygen-video/L_qa_b.jpg",
  "label": "WHAT IF YOU COULD...",
  "note": "Fresh paint on a pit wall, live action"
}, {
  "src": "https://resource2.heygen.ai/videos/66c20e44a62c40a58cf8c51210dab761/66c20e44a62c40a58cf8c51210dab761/01a0e64f-fcd0-7499-bacb-517cfc9587a3.mp4",
  "poster": "/images/heygen-video/L_qb_c.jpg",
  "label": "GENERATE VIDEO LIKE THIS?",
  "note": "Fresh paint on a pit wall, live action"
}, {
  "src": "https://resource2.heygen.ai/videos/66c20e44a62c40a58cf8c51210dab761/66c20e44a62c40a58cf8c51210dab761/01a0ee00-e4f9-75a8-a4e8-4a8e864642c2.mp4",
  "poster": "/images/heygen-video/M_words_c.jpg",
  "label": "VIDEO NEVER LET YOU",
  "note": "Stamped over a running machine, from the Iterate film"
}, {
  "src": "https://resource2.heygen.ai/videos/66c20e44a62c40a58cf8c51210dab761/66c20e44a62c40a58cf8c51210dab761/01a0edf0-a849-7020-b22a-df466747e005.mp4",
  "poster": "/images/heygen-video/T_notanymore_b.jpg",
  "label": "NOT ANYMORE",
  "note": "Stamped screen print, from the Iterate film"
}, {
  "src": "https://resource2.heygen.ai/videos/66c20e44a62c40a58cf8c51210dab761/66c20e44a62c40a58cf8c51210dab761/01a0ee04-bdc1-777c-95ad-45d51b49138d.mp4",
  "poster": "/images/heygen-video/T_tentakes_c.jpg",
  "label": "TEN TAKES",
  "note": "Stamped screen print, from the Iterate film"
}, {
  "src": "https://resource2.heygen.ai/videos/66c20e44a62c40a58cf8c51210dab761/66c20e44a62c40a58cf8c51210dab761/01a0ee04-bdfd-7c9c-82a3-ca28acb0f68d.mp4",
  "poster": "/images/heygen-video/T_itsright_b.jpg",
  "label": "IT'S RIGHT!",
  "note": "Stamped screen print, from the Iterate film"
}, {
  "src": "https://resource2.heygen.ai/videos/66c20e44a62c40a58cf8c51210dab761/66c20e44a62c40a58cf8c51210dab761/01a0f064-2596-7c8c-b5af-2df295eafdf0.mp4",
  "poster": "/images/heygen-video/L13_lp_name_d.jpg",
  "label": "HeyGen Video",
  "note": "Letterpress on cotton paper, the closing card"
}, {
  "src": "https://resource2.heygen.ai/videos/66c20e44a62c40a58cf8c51210dab761/66c20e44a62c40a58cf8c51210dab761/01a0f064-29f6-7d48-a1a3-20ec7ffdd764.mp4",
  "poster": "/images/heygen-video/L13_lp_api_b.jpg",
  "label": "Available via API",
  "note": "Letterpress on cotton paper, the closing card"
}];

export const Grid = ({items, cols = 3, showPrompt = false}) => <div style={{
  display: "grid",
  gridTemplateColumns: `repeat(auto-fill, minmax(${cols === 2 ? 300 : 220}px, 1fr))`,
  gap: "1rem",
  marginTop: "1.5rem"
}}>
    {items.map(v => <div key={v.src}>
        <video controls muted playsInline preload="metadata" poster={v.poster} style={{
  width: "100%",
  borderRadius: "12px",
  display: "block",
  background: "#000"
}}>
          <source src={v.src} type="video/mp4" />
        </video>
        <span style={{
  display: "block",
  marginTop: "0.5rem",
  fontSize: "0.9rem",
  fontWeight: 600
}}>{v.label}</span>
        {v.note && <span style={{
  display: "block",
  fontSize: "0.8rem",
  opacity: 0.65,
  lineHeight: 1.4
}}>{v.note}</span>}
        {showPrompt && v.prompt && <span style={{
  display: "block",
  marginTop: "0.5rem",
  fontSize: "0.78rem",
  opacity: 0.75,
  lineHeight: 1.5,
  fontStyle: "italic",
  whiteSpace: "pre-line"
}}>{v.prompt}</span>}
      </div>)}
  </div>;

## Build with it

| | |
| - | - |
| Model identifier | `heygen-video-1` |
| Create | `POST /v3/models/videos` |
| Retrieve | `GET /v3/models/videos/{video_id}` |
| Auth | `x-api-key` header, from [your API settings](https://app.heygen.com/developers/api) |
| Scopes | `videos:write` to create, `videos:read` to retrieve, `assets:write` to upload references |

### Generate a video

```bash theme={null}
curl -X POST https://api.heygen.com/v3/models/videos \
  -H "x-api-key: $HEYGEN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "heygen-video-1",
    "mode": "text_to_video",
    "prompt": "A plant supervisor in a plain grey coat stands beside a conveyor line and speaks to camera.",
    "duration": 5,
    "resolution": "768p",
    "aspect_ratio": "16:9",
    "seed": 42
  }'
```

Submission returns `202`:

```json theme={null}
{ "data": { "status": "pending", "video_id": "VIDEO_ID" } }
```

Send an `Idempotency-Key` header to retry safely: a retry with the same key within 24 hours returns the original response, and a concurrent duplicate returns `409`.

### Retrieve the result

Poll `GET /v3/models/videos/{video_id}`. `pending` means queued and `processing` means running. The terminal states are `completed`, `failed` and `cancelled`.

```bash theme={null}
curl https://api.heygen.com/v3/models/videos/VIDEO_ID \
  -H "x-api-key: $HEYGEN_API_KEY"
```

```json theme={null}
{
  "data": {
    "video_id": "VIDEO_ID",
    "status": "completed",
    "model": "heygen-video-1",
    "created_at": 1789689600,
    "video_url": "https://resource2.heygen.ai/videos/.../VIDEO.mp4",
    "duration": 5,
    "aspect_ratio": "16:9",
    "width": 1344,
    "height": 768,
    "seed": 42
  }
}
```

`video_url` is a signed link: poll again to refresh it. Failed and cancelled jobs carry `failure_code` (`generation_failed` or `generation_cancelled`) and `failure_message`. An unknown ID, or one from another workspace, returns `404`.

The same job is also readable through `GET /v3/videos/{video_id}`, alongside your other videos, with `video_page_url` and a `title` taken from the first 64 characters of the prompt.

## What it is good at

Sound on for everything below. Every clip on this page is raw model output, with the audio it generated.

<CardGroup cols={2}>
  <Card title="Scenes that hold together" icon="cube">
    Objects stay the same size and shape, stay where they were put, and obey physics. Things do not duplicate, vanish or reappear mid-shot.
  </Card>

  <Card title="Doing what you asked" icon="list-check">
    It follows a brief literally rather than improvising. If you did not ask for dialogue, you do not get dialogue.
  </Card>

  <Card title="Picture and sound together" icon="waveform-lines">
    Room tone, effects and speech are generated with the image from the same prompt, so they match the space you described.
  </Card>

  <Card title="Holding an object you supply" icon="image">
    Give it a photograph in `reference_to_video` and the colour, form and markings of that object survive into the shot.
  </Card>
</CardGroup>

It is at its best on **short, contained shots**: one subject, one place, one action, a locked or barely moving camera. It is least reliable on **long on-screen text**, **soft organic motion** such as petals, paper and hair, and **close hand work** like assembling or operating equipment.

## How to prompt it

Long prompts beat short ones. A few hundred to a few thousand characters gives materially better output, and there is no penalty for detail. The examples below are shots from the HeyGen Video launch film, each with the exact prompt that produced it. The last two use the sectioned layout the film was written in for its longer shots; it is plain text in the same `prompt` field.

<Grid items={HERO} cols={2} showPrompt={true} />

### The rules worth knowing

<AccordionGroup>
  <Accordion title="Name the camera, not the mood">
    A body, a lens and an aperture set depth of field, grain and motion together. "Cinematic" and "high quality" do almost nothing. `Shot on an ARRI Alexa Mini LF, 50mm T2, locked off on a tripod` changes the image.
  </Accordion>

  <Accordion title="Ask for texture, then forbid the correction">
    Describing pores and fine lines gets you part way. Adding `no retouching, no beauty filter, no glamour lighting` gets you the rest. Without the second clause, faces drift toward a retouched look. The same applies to colour: ask for muted colour, then rule out the default with `no teal-orange grade`.
  </Accordion>

  <Accordion title="Rule out marks you did not ask for">
    Plain clothing and equipment will otherwise pick up invented branding. Add `no logos, brand names, printed words or badges anywhere in frame`.
  </Accordion>

  <Accordion title="Keep lettering short and spelled out">
    A short capitalised phrase renders reliably when the prompt spells it exactly and names it as the only lettering in frame: `The only lettering anywhere is the words NOT ANYMORE, spelled exactly NOT ANYMORE`. Keep everything else free of text with the marks clause above. Paragraphs, small print and labelled diagrams read best as real type composited afterwards.
  </Accordion>

  <Accordion title="Prefer stillness to manipulation">
    A static camera with the subject wearing or holding an object is the most reliable composition available. Objects worn on the body render best of all. Close hand work degrades fastest: if a procedure needs explaining, let the speech carry the steps and keep the hands still.
  </Accordion>

  <Accordion title="Say where things touch">
    When an action involves one thing meeting another, name the contact point and say it holds. Without that, the gesture often lands near the target rather than on it.
  </Accordion>

  <Accordion title="Describe the sound">
    Audio comes from the same prompt, so direct it. Name the room tone, the specific effects and their distance. Say `no music` when you do not want a bed, because one may otherwise appear. Written dialogue fits at roughly 2.5 words per second.
  </Accordion>

  <Accordion title="Iterate with a fixed seed">
    Set `seed` yourself and the same prompt returns the same clip, so you can change one clause at a time. Omit it and the server picks a random one, so two identical requests give you different videos.
  </Accordion>
</AccordionGroup>

## Twelve styles

Style is a prompt variable. The launch films shot the same driver close-up in four media, and took one drawing from pencil to oil paint to a clay model, each step an [`image_to_video`](#modes) shot that starts from the last frame of the one before. The reliable way to get a real style is to name an actual production process and the physical marks it leaves, restrict the palette, and rule out the smooth digital default.

<Grid items={STYLES} cols={3} />

## Lettering in the shot

Every title in the launch films was generated inside the shot: painted on a wall, stamped in ink or pressed into paper. The painted track and the platform display under [How to prompt it](#how-to-prompt-it) are two more. Short phrases, spelled out exactly, render reliably, which is what makes these work.

<Note>
  Each prompt spelled the phrase exactly and named it as the only lettering in frame. For longer copy, such as a full lockup with a tagline, add real type afterwards.
</Note>

<Grid items={BRANDS} cols={3} />

## Callbacks

Pass `callback_url` on the create request to be notified when a job finishes, and `callback_id` to correlate the delivery.

| | |
| - | - |
| `callback_url` | Must be an HTTPS URL |
| `callback_id` | At most 256 characters, returned with the delivery |
| Events | `avatar_video.success` and `avatar_video.fail` |
| Success payload | `video_id`, `callback_id`, `url`, `video_page_url`, `video_share_page_url` |
| Failure payload | `video_id`, `callback_id`, `video_page_url`, `msg` |

Send `callback_id` alone to deliver to the webhook endpoints registered for your workspace. One event fires per job, at the terminal status, and a callback is attempted **once**. Treat it as a latency optimisation and keep polling as the fallback for anything you cannot afford to miss.

## Modes

`mode` selects how the model is conditioned. Each mode has its own fields; a field from another mode returns `400` when it carries a value (an empty list or `null` passes).

| Mode | Requires | Takes |
| - | - | - |
| `text_to_video` | `prompt` | the shared fields |
| `image_to_video` | `prompt` and `image` | `image` as the first frame |
| `reference_to_video` (default) | `prompt` and at least one reference image or video | `reference_images`, `reference_videos`, `reference_audio` |

The short forms `t2v`, `i2v` and `ref2va` are accepted as aliases. New integrations should use the long names, which is what validation errors return.

`image_to_video` treats the supplied image as the literal first frame, so the clip opens exactly as that still looks. To place a product or piece of equipment into a scene of your own, use `reference_to_video` and describe the scene around it.

```json theme={null}
{
  "model": "heygen-video-1",
  "mode": "reference_to_video",
  "prompt": "A supervisor wearing the harness in <Picture 1> stands still and speaks to camera.",
  "reference_images": [{ "type": "asset_id", "asset_id": "ASSET_ID" }],
  "duration": 10
}
```

## References and prompt labels

In `reference_to_video`, list order becomes the label you address in the prompt. The first entry of `reference_images` is `<Picture 1>`, the second is `<Picture 2>`, the first entry of `reference_videos` is `<Video 1>`, and so on. Images, videos and audio are numbered independently.

The API does not rewrite these labels. [Prompt enhancement](#prompt-enhancement) can reword the rest of the prompt; set it to `disabled` to keep your text as written.

| List | Maximum |
| - | - |
| `reference_images` | 9 |
| `reference_videos` | 3 |
| `reference_audio` | 3 |
| Total across all three | 12 |

Each reference accepts an HTTPS URL, an uploaded asset ID, or inline base64.

```json theme={null}
{ "type": "url",      "url": "https://example.com/photo.jpg" }
{ "type": "asset_id", "asset_id": "ASSET_ID" }
{ "type": "base64",   "media_type": "image/jpeg", "data": "<base64>" }
```

| Input | Maximum size |
| - | - |
| Image by URL | 16 MB |
| Video or audio by URL | 32 MB |
| Image as base64 | 5 MB |
| Video or audio as base64 | 16 MB |

URLs are fetched server side with SSRF checks and staged before dispatch. The fetcher does not follow redirects, so upload anything you do not control through `POST /v3/assets` and pass the returned `asset_id`.

```bash theme={null}
curl -X POST https://api.heygen.com/v3/assets \
  -H "x-api-key: $HEYGEN_API_KEY" \
  -F "file=@harness.jpg"
```

Uploaded assets must belong to the calling workspace and are reusable across requests.

## Prompt enhancement

Before generation, the prompt passes through an enhancement step that expands it for the model. `prompt_enhancement` picks how:

| Value | Behaviour |
| - | - |
| `turbo` (default) | Fast enhancement pass |
| `quality` | More thorough enhancement pass |
| `disabled` | The prompt goes to the model exactly as you wrote it |

```json theme={null}
{
  "model": "heygen-video-1",
  "mode": "text_to_video",
  "prompt": "A pastry chef pipes a single line of cream along a tart shell, then stops.",
  "prompt_enhancement": "quality"
}
```

Use `disabled` when you have already written a long, fully specified prompt and want it followed word for word.

## Parameters

| Field | Type | Default | Notes |
| - | - | - | - |
| `model` | string | required | `heygen-video-1` |
| `mode` | enum | `reference_to_video` | `text_to_video`, `image_to_video`, `reference_to_video` |
| `prompt` | string | required | 1 to 32,000 characters |
| `duration` | integer | `5` | Any integer from 5 to 15 seconds |
| `resolution` | enum | `768p` | `480p` or `768p` |
| `prompt_enhancement` | enum | `turbo` | `turbo`, `quality` or `disabled` |
| `aspect_ratio` | enum | see below | `21:9`, `16:9`, `4:3`, `1:1`, `3:4` or `9:16`, plus `adaptive` for `reference_to_video`. `image_to_video` follows the first frame. |
| `seed` | integer | random | Unsigned 32-bit. Random when omitted; set it yourself to iterate on a shot. |
| `callback_url` | string | | HTTPS only |
| `callback_id` | string | | At most 256 characters |
| `image` | asset | | `image_to_video` only, required there |
| `reference_images` | array | `[]` | `reference_to_video` only |
| `reference_videos` | array | `[]` | `reference_to_video` only |
| `reference_audio` | array | `[]` | `reference_to_video` only |

The schema is strict: unknown fields are rejected rather than ignored.

## Output

| | |
| - | - |
| Container | MP4, H.264 |
| Frame rate | 24 fps, rounded up to the next whole frame: a 7-second request encodes 175 frames |
| Audio | AAC, 32 kHz stereo, generated dialogue and effects |
| Delivery | HTTPS URL on the job record |

`resolution` names a size class and `aspect_ratio` sets the shape. When you set `aspect_ratio` explicitly, the two resolve to a fixed pixel size:

| `aspect_ratio` | `768p` | `480p` |
| - | - | - |
| `21:9` | 1536 × 672 | 960 × 416 |
| `16:9` | 1344 × 768 | 832 × 480 |
| `4:3` | 1024 × 768 | 640 × 480 |
| `1:1` | 768 × 768 | 480 × 480 |
| `3:4` | 768 × 1024 | 480 × 640 |
| `9:16` | 768 × 1344 | 480 × 832 |

### Default aspect ratio

`text_to_video` has nothing to take a shape from and defaults to `16:9`.

`reference_to_video` defaults to `adaptive`: the aspect ratio of the first reference image, or the first reference video when there are no images. Pass one of the six ratios to override it.

`image_to_video` always follows the first frame, including its EXIF orientation. To change the output shape, crop the image before uploading it.

In both adaptive cases the output follows the source ratio, scaled to the short edge of the resolution with each side rounded to a multiple of 32. A 1280 × 852 reference at `480p` returns 736 × 480. The completed job reports the actual `width` and `height`, and `aspect_ratio` as the reduced pixel ratio, for example `23:15`.

## Seeds

Set `seed` yourself to make a shot repeatable: the same prompt and the same seed, submitted in succession, return an identical file. Hold the seed and change one clause at a time to iterate.

When `seed` is omitted the server picks a random one, so two identical requests return different videos. The completed job reports the seed it used as `seed`.

Treat a seed as reproducible within a deployment rather than as a permanent handle on one render.

## Errors

Errors return `{"error": {"code", "message", "param", "doc_url"}}`. Validation errors use code `invalid_parameter` and set `param` to the field.

| Situation | `param` | Message |
| - | - | - |
| Wrong model identifier | `model` | Must be `heygen-video-1` |
| Unknown mode | `mode` | `mode must be one of: text_to_video, image_to_video, reference_to_video.` |
| No references in `reference_to_video` | `reference_images` | `Provide at least one reference image or video.` |
| Reference list outside `reference_to_video` | the field sent | `reference_images is only accepted with mode reference_to_video.` |
| `image_to_video` without `image` | `image` | `Field required` |
| Duration out of range | `duration` | `Input should be greater than or equal to 5` |
| Unsupported aspect ratio | `aspect_ratio` | Lists the ratios the mode accepts |
| Non-HTTPS callback | `callback_url` | `Value error, callback_url must be an HTTPS URL` |
| Unfetchable reference URL | absent | `Invalid URL in reference_image[0]: Could not download the file. Ensure the URL is publicly accessible.` |
| Oversized inline reference | absent | `Base64 file in reference_image[0] is too large (N bytes).` |
| Unknown field | the field | `Extra inputs are not permitted` |
| Insufficient credits | absent | Returned as `400` |

This route runs on paid API keys.

## Specs at a glance

| | |
| - | - |
| Duration | 5 to 15 seconds |
| Resolution | `480p` or `768p` |
| Aspect ratio | Six ratios, plus `adaptive` for `reference_to_video`; `image_to_video` follows the first frame |
| Jobs | Run to completion once submitted |
| Callbacks | Delivered once; poll as the fallback |
| API keys | Paid API keys |

<CardGroup cols={2}>
  <Card title="Get an API key" icon="key" href="https://app.heygen.com/developers/api">
    Create a key and check your usage.
  </Card>

  <Card title="HeyGen Avatar" icon="user" href="/avatar-v">
    Animate a look you own from a script, with Avatar V, IV and III.
  </Card>
</CardGroup>
