Overview
A podcast video is two people talking. You write the dialogue; every turn becomes a clip of the person who says it, and the clips are cut together. Voice, lip movement and gestures come from the video model, so there is novoiceId to send: each person keeps the same voice in every clip, and both share one accent.
It comes in two shapes, picked with mode:
podcast(default): each person filmed on their own, cut back and forth. The host is one of your avatars (avatarId); the guest is generated in the same room from a description, or is another avatar.dualcast: both people in the same shot. In each clip one talks while the other listens.
GET /v1/avatar/list.
Writing the Script
Wrap every turn in who says it:[A] … [/A] for the host (the person on the left in a dualcast) and [B] … [/B] for the guest (on the right). Both people must speak, and the whole dialogue must fit in 24 clips (a long turn is split into several clips between sentences). Inside a turn you can add:
- an action in brackets, in your own words:
[she nods] Totally.(at the start it happens before speaking); [product] … [/product]around the words while your product photo (productImageKey) is on screen;[camera] … [/camera]around the words said to the audience instead of to the other person, such as the call to action.
Podcast with a Generated Guest
projectId. Poll GET /v1/project/{projectId} or pass a webhook to get the finished video.
Podcast with Two Avatars and a Set
The guest can be another avatar, shown in its own photo.hostSet first moves the host to a set you describe. hostVoice, guest.voice and accent pick the voices instead of letting them be chosen from each avatar’s gender and age.
Dualcast from Two Avatars, with a Product
A dualcast can be made from two avatars, placed side by side on the set you describe indualcastSet (left speaks the [A] turns, right the [B] turns). The product photo appears at the top of the screen on the words marked [product].
dualcastImageKey (an HTTPS URL, or the storage key of an image in your library): it must show exactly two people, A on the left and B on the right, or the call answers 400. dualcastAvatarId does the same with one of your avatars whose image already shows both people.
Parameters
A dualcast needs one of
dualcastAvatars, dualcastImageKey or dualcastAvatarId.
Voices for hostVoice and guest.voice are the ids of Get Catalog with name=voice-presets (for example young-woman-25-30, man-mid-30s).
Accents are the ids of Get Catalog with name=accents (for example us-neutral-general-american, spanish-mexico); each says the language it speaks.
Managed teams pay 100 credits per started 15 seconds of clips (every turn is its own clip, so a short answer still takes a whole clip), plus the images the pipeline generates: the host’s still, the guest, the host’s set, or the two avatars side by side. The estimate is charged when the project is created and refunded if the project fails.Teams in bring-your-own-keys mode are not charged credits for generation, but need their own OpenRouter and Gemini keys in Settings → AI keys. Without them the call answers
402 with code: "missing_credentials" and the missingProviders list, before anything is created or charged. To retry a create call after a timeout or 5xx, send it with an Idempotency-Key: the same key never creates or charges twice.Tips for Podcast Videos
Use Cases
- Talking-head podcasts: Short clips of an interview for TikTok, Reels and Shorts
- Product conversations: Two people talking about a product, with its photo on screen
- Educational content: A host asking the questions your audience would ask
- Testimonials: A conversation instead of a monologue
Handling the Webhook Response
Webhooks are signed: check theHooked-Signature header as shown in Webhooks.