One recording in, a voice out. POST /v3/models/audio/voices with "mode": "instant" returns a voice_id that is usually ACTIVE within seconds. No training and no voice clone slot.
Set up with an agent
Not technical? Paste this into Claude Code, Codex, or Cursor and it walks you through every step on this page in your own terminal.Prompt for your agent
1. Create the voice
Upload the recording and pass itsasset_id, or send a public HTTPS URL or base64 audio. The call returns 202 Accepted.
Response
Use MP3 or WAV for uploaded assets and inline base64, with
audio/mpeg or audio/wav for inline data. OGG is supported through a public audio URL; audio/ogg inline data and OGG assets are not accepted.
About the recording: only its first 3 minutes are used. A file can be up to 100 MiB, or 16 MB after decoding for inline base64. A URL must point at the audio file itself and download within 30 seconds; HeyGen fetches it after returning 202, and a fetch failure ends the voice FAILED with INVALID_AUDIO. A silent recording fails.
An instant voice cannot be retrained. To change it, create a new one.
Idempotency-Key is optional. Reusing a key within 24 hours returns the first response, and a duplicate that arrives while the first is still being accepted returns 409 request_in_progress. Without the header, each retry creates a new voice.
2. Wait until it is ACTIVE
PollGET /v3/models/audio/voices/{voice_id} until status is ACTIVE or FAILED.
Response
A voice created without
language has no language field until HeyGen detects it. A failed voice stays in the list as FAILED until you delete it.
Wait for ACTIVE before you generate speech. A speech request for a PENDING voice returns 409 voice_not_ready: keep polling, and send it once the voice is ACTIVE.
3. Generate speech
Once the voice isACTIVE, pass it to Text to Speech. POST /v3/models/audio/tts returns one 44.1 kHz WAV; POST /v3/models/audio/tts/stream streams audio parts over Server-Sent Events from the same body.
model, voice_id, text, language, and expressiveness_boost (0.0–1.0, default 1.0), plus with_timestamps on the stream. seed, speed, pitch_shift, pitch_variance, and <break> tags work with professional voices only and return 400 invalid_parameter for an instant voice.
Languages
Supported base language codes for cloning and speech are:ar, be, bg, ca, cs, da, de, el, en, es, fa, fi, fr, he, hi, hr, hu, id, it, ja, ko, mk, ms, nl, pl, pt, ro, ru, sk, sl, sr, sv, ta, th, tl, tr, uk, vi, zh. Speech also accepts regional tags such as en-US and fr-CA, which normalize to en and fr. A regional tag does not select a regional accent. Unsupported base languages, such as xx, return 400 invalid_parameter.
Manage voices
GET /v3/models/audio/voices lists the workspace’s instant and professional voices, newest first; mode tells them apart. Set limit from 1 to 100 (default 10) and pass next_token as token while has_more is true.
DELETE /v3/models/audio/voices/{voice_id} deletes a voice. Deleting an instant voice that is still PENDING cancels it.
Errors
Speech errors are on Text to Speech. The full catalog is in Error Codes.

