Skip to main content
POST

Overview

Creates a voice of your own from a recording, the same way Voice cloning works in the dashboard. Send a public https link to the sample in audioUrl. The answer is the new voice: pass its id as voiceId in any create endpoint, in Generate Speech or in Speech to Speech. It also shows up in List Voices with isCustom: true.
The sample: MP3, WAV, M4A, AAC, OGG, FLAC or WEBM, 5 to 90 seconds, up to 10 MB. Use one speaker and as little background noise as you can (what is left is removed). The type is read from the file itself, not from the URL.
What happens in the request (a few seconds):
  1. The sample is downloaded. Local, private-network and plain http addresses are refused with 400, and every redirect is checked the same way.
  2. The voice is cloned and saved to your team, with the language, gender, age, use case and accent you give (they label the voice in the dashboard; language also sets the accents you can pick).
  3. The voice records a short introduction in its own language. That sample becomes its templateUrl.
Cloning is free, but it uses one of your plan’s custom voices (cloned and designed voices share the quota; Pro includes 1). Delete a Voice frees one. A managed team needs a live plan.
consent must be true. By sending it you confirm you have the rights to upload and clone this voice, as the dashboard asks before cloning.
BYOK teams clone in their own ElevenLabs account, not here (403, code: "byok"). Voices in that account appear in List Voices and can be used as voiceId right away.

Errors

Send an Idempotency-Key header so a retry after a timeout answers with the first voice instead of cloning it twice.

Authorizations

x-api-key
string
header
required

Headers

Idempotency-Key
string

Makes the request safe to retry. 1-255 printable ASCII characters, one per operation (your job id, or a UUID you store), reused on every retry. Within 24 hours the same key with the same body answers with the first response and the header Idempotent-Replayed: true, without running or charging again. Keys are scoped to your team and the endpoint. 5xx answers and refusals before anything ran (401, 402, 403, 409, 429) are not kept. See Idempotency.

Required string length: 1 - 255
Pattern: ^[\x20-\x7E]+$

Body

application/json
name
string
required

The voice's name in your library.

Required string length: 2 - 100
audioUrl
string<uri>
required

Public https link to the sample file itself: MP3, WAV, M4A, AAC, OGG, FLAC or WEBM, 5 to 90 seconds, up to 10 MB.

Maximum string length: 2000

Must be true: you confirm you have the necessary rights to upload and clone this voice.

Available options:
true
description
string

A note about the voice, returned in List Voices.

Maximum string length: 500
language
string
default:English

The sample's language, a name such as English, Spanish or Portuguese (case-insensitive; the dashboard's list of 55 languages).

gender
enum<string>

Case-insensitive. Stored as Unknown when omitted.

Available options:
male,
female
age
enum<string>

Case-insensitive. Stored as unknown when omitted.

Available options:
young,
middle_aged,
old
useCase
enum<string>
default:general
Available options:
narrative_story,
informative_educational,
conversational,
advertisement,
social_media,
entertainment_tv,
characters_animation,
general
accent
string
default:Standard

One of the accents the dashboard offers for language (for English: Standard, Neutral, American, British, Australian, ...). Case-insensitive; another value answers 400 with the list.

Response

Voice cloned

success
boolean
message
string
data
object

A narration voice.