HeyGen Voice is HeyGen’s in-house voice model. It trains a dedicated voice adapter per speaker and synthesizes speech directly from that adapter, giving the highest-fidelity voice output HeyGen offers.
What it does
- Professional voice cloning — train a dedicated voice from 1–10 recordings of the same speaker totaling at least 20 minutes. Training runs asynchronously; poll until the voice is
ACTIVE. See HeyGen Professional Clone. For cloning from a single short recording, see HeyGen Instant Clone on the Starfish engine. - Speech generation — synthesize speech from an
ACTIVEvoice as one completed 44.1 kHz WAV, or stream ordered audio parts over Server-Sent Events with optional word-level timestamps. See HeyGen Voice Speech.
Endpoints
HeyGen Voice and third-party voices
The voices catalog (/v3/voices) covers HeyGen’s 300+ stock voices, designed voices, and HeyGen instant clones on the Starfish engine, synthesized through Third Party Speech. HeyGen Voice is HeyGen’s own model surface: its professional voices are created and used through the /v3/models/audio endpoints, scoped to the workspace that trained them.
