Skip to main content
The HeyGen Voice APIs are in private preview. Your account must be enabled before these endpoints are available — see how to request access.
HeyGen Voice is HeyGen’s in-house voice model. It trains a dedicated voice adapter per speaker and synthesizes speech directly from that adapter, giving the highest-fidelity voice output HeyGen offers.

What it does

  • Professional voice cloning — train a dedicated voice from 1–10 recordings of the same speaker totaling at least 20 minutes. Training runs asynchronously; poll until the voice is ACTIVE. See HeyGen Professional Clone. For cloning from a single short recording, see HeyGen Instant Clone on the Starfish engine.
  • Speech generation — synthesize speech from an ACTIVE voice as one completed 44.1 kHz WAV, or stream ordered audio parts over Server-Sent Events with optional word-level timestamps. See HeyGen Voice Speech.

Endpoints

HeyGen Voice and third-party voices

The voices catalog (/v3/voices) covers HeyGen’s 300+ stock voices, designed voices, and HeyGen instant clones on the Starfish engine, synthesized through Third Party Speech. HeyGen Voice is HeyGen’s own model surface: its professional voices are created and used through the /v3/models/audio endpoints, scoped to the workspace that trained them.