Skip to main content
POST
Create Podcast
Try it out! Use the API playground on the right to test the Podcast endpoint directly.

Overview

Two AI avatars talking, in one of two shapes:
  • Podcast: each person filmed on their own and cut back and forth, like a real podcast.
  • Dualcast: both people in the same shot; in each clip one talks while the other listens and reacts.
You write the dialogue; every turn becomes a clip of the person who says it, with their own voice. Ideal for:
  • Educational and opinion clips
  • Product mentions in a natural conversation
  • Q&A and “one question” formats
Voice, lip movement and gestures are generated natively by the video model. Each person keeps the same voice in every clip, and the whole conversation shares one accent.

Endpoint


Required Fields

string
default:"podcast"
podcast (needs avatarId and guest) or dualcast (needs dualcastImageKey, dualcastAvatarId, or dualcastAvatars to make the image).
string
required
The dialogue, each turn wrapped in who says it: [A] … [/A] (the host, or the person on the left in dualcast) and [B] … [/B] (the guest, or the person on the right). Lines starting with A: / B: are read the same. Both people must speak. A long turn is split between sentences into clips of up to 10 seconds, up to 24 clips per video.Inside a turn:
  • an action in brackets, in your words: [nods] at the start happens before speaking, in the middle while saying the words around it, at the end after the line
  • [camera] … [/camera] marks the words said to the camera, to the audience, instead of to the other one (the closing CTA, for instance)
  • [product] … [/product] marks the words while the product photo (productImageKey) is on screen, and whoever is talking points up to it as it appears (unless you wrote an action right there)
string
Podcast mode: the host (A): an avatar ID from /v1/avatar/list. An avatar on a podcast set works best: a generated guest sits in the same room.
object
Podcast mode: the guest (B). In dualcast mode only guest.voice is read, as the voice of the person on the right.
string
Dualcast mode, instead of dualcastImageKey: an avatar ID from /v1/avatar/list whose image already shows both people (made in the avatar editor with “Both people in one shot”). A is the person on the left, B the one on the right. It must show exactly two people.
object
Dualcast mode, instead of dualcastImageKey: two avatar IDs (left, right) from /v1/avatar/list. They are put side by side on a podcast set, each with their microphone (one more image): left is A, right is B.
string
With dualcastAvatars, optional: the set they sit on, in your words (max 300 characters), e.g. a modern studio with a wooden slat wall.
string
Dualcast mode: the image with both people, as a storage key from your media library or an https URL. A is the person on the left, B the one on the right. It must show exactly two people; the automatic voices follow who it shows.

Optional Fields

string
Podcast mode: move the host to this set first, in your words (max 300 characters), e.g. a cozy podcast studio with a boom microphone. Only the background changes, and the guest sits in the same room. Useful when the avatar is not on a podcast set.
string
A product photo (storage key from your media library or an https URL), shown at the top of the video on the words marked [product] … [/product].
string
A’s voice, from the same list as guest.voice. Unset: the one that fits the host (podcast), or the person on the left of the image (dualcast).
string
One accent for both people, from the English and Spanish accents (e.g. us-neutral-general-american, spanish-mexico). Unset: the language’s neutral accent.
string
The conversation’s language (e.g. en, es). Unset: detected from the script.
string
Video name (max 100 characters)
string
Music ID from /v1/music/list for background music.
object
Caption settings, same as the other video endpoints (preset, alignment, disabled).
string
HTTPS URL to receive completion notification (max 500 characters).
object
Custom metadata object (max 5KB).

Request Example

Dualcast Example

Authorizations

x-api-key
string
header
required

Body

application/json
script
string
required

The dialogue, each turn wrapped in who says it: [A] … [/A] (the host, or the person on the left in dualcast) and [B] … [/B] (the guest, or the person on the right). Lines starting with "A:" / "B:" are read the same. An action in brackets ("[nods]") at the start of a turn happens before speaking, in the middle while saying the words around it, at the end after the line. [camera] … [/camera] marks the words said to the camera, to the audience. [product] … [/product] marks the words while the product photo is on screen; whoever is talking points up to it as it appears. Both people must speak; up to 24 clips.

Required string length: 1 - 10000
mode
enum<string>
default:podcast

podcast: each person filmed on their own, cut back and forth (needs avatarId and guest). dualcast: both people in one image (needs dualcastImageKey, dualcastAvatarId, or dualcastAvatars to make it).

Available options:
podcast,
dualcast
name
string

Video name (max 100 characters)

Maximum string length: 100
avatarId
string

Podcast mode: the host (A), an avatar ID from /v1/avatar/list. A podcast set works best, since a generated guest sits in the same room.

Maximum string length: 30
guest
object

Podcast mode: the guest (B). In dualcast only guest.voice is read: the voice of the person on the right.

dualcastImageKey
string

Dualcast only: the image with both people, as a storage key from your media library or an https URL. A is the person on the left, B the one on the right.

Maximum string length: 1000
dualcastAvatarId
string

Dualcast mode, instead of dualcastImageKey: an avatar ID from /v1/avatar/list whose image already shows both people (made in the avatar editor with "Both people in one shot"). A is the person on the left, B the one on the right. It must show exactly two people.

Maximum string length: 30
dualcastAvatars
object

Dualcast mode, instead of dualcastImageKey: two avatar IDs from /v1/avatar/list, put side by side on a podcast set (one more image). left is A, right is B.

dualcastSet
string

With dualcastAvatars, optional: the set they sit on, in your words (e.g. "a modern studio with a wooden slat wall").

Maximum string length: 300
hostSet
string

Podcast mode, optional: move the host to this set first, in your words (e.g. "a cozy podcast studio with a boom microphone"). Only the background changes; the guest sits in the same room.

Maximum string length: 300
productImageKey
string

Optional: a product photo (storage key from your media library or an https URL), shown at the top of the video on the words marked [product] … [/product].

Maximum string length: 1000
hostVoice
enum<string>

The host's voice. Unset: the one that fits the avatar.

Available options:
young-woman-25-30,
woman-mid-20s,
woman-mid-30s,
older-woman-50,
young-man-25-30,
man-mid-20s,
man-mid-30s,
older-man-50
accent
string

One accent for both, from the English and Spanish accents (e.g. us-neutral-general-american, spanish-mexico). Unset: the language's neutral one.

language
string

The conversation's language (e.g. en, es). Unset: detected from the script.

musicId
string

Music ID from /v1/music/list for background music

Maximum string length: 30
caption
object
webhook
string

HTTPS URL to receive completion notification

Maximum string length: 500
metadata
object

Custom metadata object (max 5KB)

Response

Talking Avatar created successfully

success
boolean
data
object
message
string