HeyGen Video is our general-purpose video model, built on MiniMax H3 and post-trained by HeyGen. Give it a prompt and it generates the whole scene: subject, setting, light and sound together, in one call. Each clip runs 5 to 15 seconds with its own audio track of dialogue, ambience and sound effects, so there is no separate speech pass, lip-sync step or avatar to pick.
One model, heygen-video-1, called in three ways:
It is a different tool from the avatar rendering engines. Those animate a look you already own from a script you already wrote. This one invents the subject, the setting, the light and the sound at once, from a description. The rest of this page starts with the call itself, then what the model is good at and how to prompt it, then the full API reference.
Build with it
Generate a video
202:
Idempotency-Key header to retry safely: a retry with the same key within 24 hours returns the original response, and a concurrent duplicate returns 409.
Retrieve the result
PollGET /v3/models/videos/{video_id}. pending means queued and processing means running. The terminal states are completed, failed and cancelled.
video_url is a signed link: poll again to refresh it. Failed and cancelled jobs carry failure_code (generation_failed or generation_cancelled) and failure_message. An unknown ID, or one from another workspace, returns 404.
The same job is also readable through GET /v3/videos/{video_id}, alongside your other videos, with video_page_url and a title taken from the first 64 characters of the prompt.
What it is good at
Sound on for everything below. Every clip on this page is raw model output, with the audio it generated.Scenes that hold together
Objects stay the same size and shape, stay where they were put, and obey physics. Things do not duplicate, vanish or reappear mid-shot.
Doing what you asked
It follows a brief literally rather than improvising. If you did not ask for dialogue, you do not get dialogue.
Picture and sound together
Room tone, effects and speech are generated with the image from the same prompt, so they match the space you described.
Holding an object you supply
Give it a photograph in
reference_to_video and the colour, form and markings of that object survive into the shot.How to prompt it
Long prompts beat short ones. A few hundred to a few thousand characters gives materially better output, and there is no penalty for detail. The examples below are shots from the HeyGen Video launch film, each with the exact prompt that produced it. The last two use the sectioned layout the film was written in for its longer shots; it is plain text in the sameprompt field.
The rules worth knowing
Name the camera, not the mood
Name the camera, not the mood
A body, a lens and an aperture set depth of field, grain and motion together. “Cinematic” and “high quality” do almost nothing.
Shot on an ARRI Alexa Mini LF, 50mm T2, locked off on a tripod changes the image.Ask for texture, then forbid the correction
Ask for texture, then forbid the correction
Describing pores and fine lines gets you part way. Adding
no retouching, no beauty filter, no glamour lighting gets you the rest. Without the second clause, faces drift toward a retouched look. The same applies to colour: ask for muted colour, then rule out the default with no teal-orange grade.Rule out marks you did not ask for
Rule out marks you did not ask for
Plain clothing and equipment will otherwise pick up invented branding. Add
no logos, brand names, printed words or badges anywhere in frame.Keep lettering short and spelled out
Keep lettering short and spelled out
A short capitalised phrase renders reliably when the prompt spells it exactly and names it as the only lettering in frame:
The only lettering anywhere is the words NOT ANYMORE, spelled exactly NOT ANYMORE. Keep everything else free of text with the marks clause above. Paragraphs, small print and labelled diagrams read best as real type composited afterwards.Prefer stillness to manipulation
Prefer stillness to manipulation
A static camera with the subject wearing or holding an object is the most reliable composition available. Objects worn on the body render best of all. Close hand work degrades fastest: if a procedure needs explaining, let the speech carry the steps and keep the hands still.
Say where things touch
Say where things touch
When an action involves one thing meeting another, name the contact point and say it holds. Without that, the gesture often lands near the target rather than on it.
Describe the sound
Describe the sound
Audio comes from the same prompt, so direct it. Name the room tone, the specific effects and their distance. Say
no music when you do not want a bed, because one may otherwise appear. Written dialogue fits at roughly 2.5 words per second.Iterate with a fixed seed
Iterate with a fixed seed
Set
seed yourself and the same prompt returns the same clip, so you can change one clause at a time. Omit it and the server picks a random one, so two identical requests give you different videos.Twelve styles
Style is a prompt variable. The launch films shot the same driver close-up in four media, and took one drawing from pencil to oil paint to a clay model, each step animage_to_video shot that starts from the last frame of the one before. The reliable way to get a real style is to name an actual production process and the physical marks it leaves, restrict the palette, and rule out the smooth digital default.
Lettering in the shot
Every title in the launch films was generated inside the shot: painted on a wall, stamped in ink or pressed into paper. The painted track and the platform display under How to prompt it are two more. Short phrases, spelled out exactly, render reliably, which is what makes these work.Each prompt spelled the phrase exactly and named it as the only lettering in frame. For longer copy, such as a full lockup with a tagline, add real type afterwards.
Callbacks
Passcallback_url on the create request to be notified when a job finishes, and callback_id to correlate the delivery.
Send
callback_id alone to deliver to the webhook endpoints registered for your workspace. One event fires per job, at the terminal status, and a callback is attempted once. Treat it as a latency optimisation and keep polling as the fallback for anything you cannot afford to miss.
Modes
mode selects how the model is conditioned. Each mode has its own fields; a field from another mode returns 400 when it carries a value (an empty list or null passes).
The short forms
t2v, i2v and ref2va are accepted as aliases. New integrations should use the long names, which is what validation errors return.
image_to_video treats the supplied image as the literal first frame, so the clip opens exactly as that still looks. To place a product or piece of equipment into a scene of your own, use reference_to_video and describe the scene around it.
References and prompt labels
Inreference_to_video, list order becomes the label you address in the prompt. The first entry of reference_images is <Picture 1>, the second is <Picture 2>, the first entry of reference_videos is <Video 1>, and so on. Images, videos and audio are numbered independently.
The API does not rewrite these labels. Prompt enhancement can reword the rest of the prompt; set it to disabled to keep your text as written.
Each reference accepts an HTTPS URL, an uploaded asset ID, or inline base64.
URLs are fetched server side with SSRF checks and staged before dispatch. The fetcher does not follow redirects, so upload anything you do not control through
POST /v3/assets and pass the returned asset_id.
Prompt enhancement
Before generation, the prompt passes through an enhancement step that expands it for the model.prompt_enhancement picks how:
disabled when you have already written a long, fully specified prompt and want it followed word for word.
Parameters
The schema is strict: unknown fields are rejected rather than ignored.
Output
resolution names a size class and aspect_ratio sets the shape. When you set aspect_ratio explicitly, the two resolve to a fixed pixel size:
Default aspect ratio
text_to_video has nothing to take a shape from and defaults to 16:9.
reference_to_video defaults to adaptive: the aspect ratio of the first reference image, or the first reference video when there are no images. Pass one of the six ratios to override it.
image_to_video always follows the first frame, including its EXIF orientation. To change the output shape, crop the image before uploading it.
In both adaptive cases the output follows the source ratio, scaled to the short edge of the resolution with each side rounded to a multiple of 32. A 1280 × 852 reference at 480p returns 736 × 480. The completed job reports the actual width and height, and aspect_ratio as the reduced pixel ratio, for example 23:15.
Seeds
Setseed yourself to make a shot repeatable: the same prompt and the same seed, submitted in succession, return an identical file. Hold the seed and change one clause at a time to iterate.
When seed is omitted the server picks a random one, so two identical requests return different videos. The completed job reports the seed it used as seed.
Treat a seed as reproducible within a deployment rather than as a permanent handle on one render.
Errors
Errors return{"error": {"code", "message", "param", "doc_url"}}. Validation errors use code invalid_parameter and set param to the field.
This route runs on paid API keys.
Specs at a glance
Get an API key
Create a key and check your usage.
HeyGen Avatar
Animate a look you own from a script, with Avatar V, IV and III.

