> ## Documentation Index
> Fetch the complete documentation index at: https://docs.hooked.so/llms.txt
> Use this file to discover all available pages before exploring further.

# Podcast

> Create a podcast interview with two AI avatars

<Note>
  **Try it out!** Use the API playground on the right to test the Podcast endpoint directly.
</Note>

## Overview

Two AI avatars talking, in one of two shapes:

* **Podcast**: each person filmed on their own and cut back and forth, like a real podcast.
* **Dualcast**: both people in the same shot; in each clip one talks while the other listens and reacts.

You write the dialogue; every turn becomes a clip of the person who says it, with their own voice. Ideal for:

* Educational and opinion clips
* Product mentions in a natural conversation
* Q\&A and "one question" formats

<Info>
  Voice, lip movement and gestures are generated natively by the video model. Each person keeps the same voice in every clip, and the whole conversation shares one accent.
</Info>

***

## Endpoint

```
POST /v1/project/create/podcast-interview
```

***

## Required Fields

<ParamField body="mode" type="string" default="podcast">
  `podcast` (needs `avatarId` and `guest`) or `dualcast` (needs `dualcastImageKey`, `dualcastAvatarId`, or `dualcastAvatars` to make the image).
</ParamField>

<ParamField body="script" type="string" required>
  The dialogue, each turn wrapped in who says it: `[A] … [/A]` (the host, or the person on the left in dualcast) and `[B] … [/B]` (the guest, or the person on the right). Lines starting with `A:` / `B:` are read the same. Both people must speak. A long turn is split between sentences into clips of up to 10 seconds, up to 24 clips per video.

  Inside a turn:

  * an action in brackets, in your words: `[nods]` at the start happens before speaking, in the middle while saying the words around it, at the end after the line
  * `[camera] … [/camera]` marks the words said to the camera, to the audience, instead of to the other one (the closing CTA, for instance)
  * `[product] … [/product]` marks the words while the product photo (`productImageKey`) is on screen, and whoever is talking points up to it as it appears (unless you wrote an action right there)
</ParamField>

<ParamField body="avatarId" type="string">
  Podcast mode: the host (A): an avatar ID from `/v1/avatar/list`. An avatar on a podcast set works best: a generated guest sits in the same room.
</ParamField>

<ParamField body="guest" type="object">
  Podcast mode: the guest (B). In dualcast mode only `guest.voice` is read, as the voice of the person on the right.

  <Expandable title="Guest Object">
    <ParamField body="source" type="string" default="generated">
      * `generated`: a new person in the same podcast room, seen from the other seat
      * `avatar`: another avatar, seated in the same podcast room from the other seat
    </ParamField>

    <ParamField body="description" type="string">
      Required when `source` is `generated` (max 300 characters): who the guest is, e.g. `a man in his 30s, casual, average looking`.
    </ParamField>

    <ParamField body="avatarId" type="string">
      Required when `source` is `avatar`: an avatar ID from `/v1/avatar/list`.
    </ParamField>

    <ParamField body="voice" type="string">
      The guest's voice: `young-woman-25-30`, `woman-mid-20s`, `woman-mid-30s`, `older-woman-50`, `young-man-25-30`, `man-mid-20s`, `man-mid-30s` or `older-man-50`. Unset: the one that fits the guest.
    </ParamField>
  </Expandable>
</ParamField>

<ParamField body="dualcastAvatarId" type="string">
  Dualcast mode, instead of `dualcastImageKey`: an avatar ID from `/v1/avatar/list` whose image already shows both people (made in the avatar editor with "Both people in one shot"). A is the person on the left, B the one on the right. It must show exactly two people.
</ParamField>

<ParamField body="dualcastAvatars" type="object">
  Dualcast mode, instead of `dualcastImageKey`: two avatar IDs (`left`, `right`) from `/v1/avatar/list`. They are put side by side on a podcast set, each with their microphone (one more image): `left` is A, `right` is B.
</ParamField>

<ParamField body="dualcastSet" type="string">
  With `dualcastAvatars`, optional: the set they sit on, in your words (max 300 characters), e.g. `a modern studio with a wooden slat wall`.
</ParamField>

<ParamField body="dualcastImageKey" type="string">
  Dualcast mode: the image with both people, as a storage key from your media library or an https URL. A is the person on the left, B the one on the right. It must show exactly two people; the automatic voices follow who it shows.
</ParamField>

***

## Optional Fields

<ParamField body="hostSet" type="string">
  Podcast mode: move the host to this set first, in your words (max 300 characters), e.g. `a cozy podcast studio with a boom microphone`. Only the background changes, and the guest sits in the same room. Useful when the avatar is not on a podcast set.
</ParamField>

<ParamField body="productImageKey" type="string">
  A product photo (storage key from your media library or an https URL), shown at the top of the video on the words marked `[product] … [/product]`.
</ParamField>

<ParamField body="hostVoice" type="string">
  A's voice, from the same list as `guest.voice`. Unset: the one that fits the host (podcast), or the person on the left of the image (dualcast).
</ParamField>

<ParamField body="accent" type="string">
  One accent for both people, from the English and Spanish accents (e.g. `us-neutral-general-american`, `spanish-mexico`). Unset: the language's neutral accent.
</ParamField>

<ParamField body="language" type="string">
  The conversation's language (e.g. `en`, `es`). Unset: detected from the script.
</ParamField>

<ParamField body="name" type="string">
  Video name (max 100 characters)
</ParamField>

<ParamField body="musicId" type="string">
  Music ID from `/v1/music/list` for background music.
</ParamField>

<ParamField body="caption" type="object">
  Caption settings, same as the other video endpoints (`preset`, `alignment`, `disabled`).
</ParamField>

<ParamField body="webhook" type="string">
  HTTPS URL to receive completion notification (max 500 characters).
</ParamField>

<ParamField body="metadata" type="object">
  Custom metadata object (max 5KB).
</ParamField>

***

## Request Example

```json theme={null}
{
  "script": "[A] 8 out of 10 people say they struggle to make it to the end of the month. [/A]\n[B] Quick question. Do you consider yourself middle class? [/B]\n[A] [pauses, thinking] Hmm, I'd say yes. [/A]\n[B] [camera] If you want to learn to save, comment SAVE below. [/camera] [/B]",
  "avatarId": "avatar_podcast_host_01",
  "guest": {
    "source": "generated",
    "description": "a man in his 30s, casual, average looking"
  },
  "caption": { "preset": "tiktok", "alignment": "bottom", "disabled": false },
  "webhook": "https://yoursite.com/webhook"
}
```

### Dualcast Example

```json theme={null}
{
  "mode": "dualcast",
  "script": "[A] I look at myself in the mirror in these leggings and go: damn, I look good. [she bursts out laughing] [/A]\n[B] They lift you and cinch your waist, it's next level. [/B]\n[A] I'm leaving a photo up here [product] because otherwise you won't get it. [/product] [/A]\n[B] The link is down below, go take a look. [they both wave goodbye] [/B]",
  "dualcastImageKey": "https://yoursite.com/two-hosts.jpg",
  "hostVoice": "woman-mid-20s",
  "guest": { "source": "generated", "voice": "young-woman-25-30" }
}
```


## OpenAPI

````yaml POST /v1/project/create/podcast-interview
openapi: 3.0.0
info:
  title: Hooked API
  version: 1.0.0
  description: AI Video Generation API
servers:
  - url: https://api.hooked.so
security:
  - ApiKeyAuth: []
tags:
  - name: Images
    description: 'Image editing: background removal and expansion'
paths:
  /v1/project/create/podcast-interview:
    post:
      tags:
        - Videos
      summary: Create Podcast
      description: >-
        Create a podcast interview with two AI avatars: two shots cut back and
        forth (podcast), or both people in one shot (dualcast). One clip per
        turn of the dialogue; voice, lip movement and gestures are generated
        natively by the video model.
      operationId: createPodcastInterview
      requestBody:
        required: true
        content:
          application/json:
            schema:
              type: object
              required:
                - script
              properties:
                mode:
                  type: string
                  enum:
                    - podcast
                    - dualcast
                  default: podcast
                  description: >-
                    `podcast`: each person filmed on their own, cut back and
                    forth (needs avatarId and guest). `dualcast`: both people in
                    one image (needs dualcastImageKey, dualcastAvatarId, or
                    dualcastAvatars to make it).
                script:
                  type: string
                  minLength: 1
                  maxLength: 10000
                  description: >-
                    The dialogue, each turn wrapped in who says it: [A] … [/A]
                    (the host, or the person on the left in dualcast) and [B] …
                    [/B] (the guest, or the person on the right). Lines starting
                    with "A:" / "B:" are read the same. An action in brackets
                    ("[nods]") at the start of a turn happens before speaking,
                    in the middle while saying the words around it, at the end
                    after the line. [camera] … [/camera] marks the words said to
                    the camera, to the audience. [product] … [/product] marks
                    the words while the product photo is on screen; whoever is
                    talking points up to it as it appears. Both people must
                    speak; up to 24 clips.
                name:
                  type: string
                  description: Video name (max 100 characters)
                  maxLength: 100
                avatarId:
                  type: string
                  description: >-
                    Podcast mode: the host (A), an avatar ID from
                    /v1/avatar/list. A podcast set works best, since a generated
                    guest sits in the same room.
                  maxLength: 30
                guest:
                  type: object
                  required:
                    - source
                  description: >-
                    Podcast mode: the guest (B). In dualcast only guest.voice is
                    read: the voice of the person on the right.
                  properties:
                    source:
                      type: string
                      enum:
                        - generated
                        - avatar
                      default: generated
                      description: >-
                        `generated`: a new person in the same podcast room, seen
                        from the other seat. `avatar`: another avatar, seated in
                        that same room from the other seat.
                    description:
                      type: string
                      maxLength: 300
                      description: >-
                        Required when source is generated: who the guest is,
                        e.g. "a man in his 30s, casual, average looking".
                    avatarId:
                      type: string
                      maxLength: 30
                      description: >-
                        Required when source is avatar: an avatar ID from
                        /v1/avatar/list.
                    voice:
                      type: string
                      enum:
                        - young-woman-25-30
                        - woman-mid-20s
                        - woman-mid-30s
                        - older-woman-50
                        - young-man-25-30
                        - man-mid-20s
                        - man-mid-30s
                        - older-man-50
                      description: 'The guest''s voice. Unset: the one that fits the guest.'
                dualcastImageKey:
                  type: string
                  maxLength: 1000
                  description: >-
                    Dualcast only: the image with both people, as a storage key
                    from your media library or an https URL. A is the person on
                    the left, B the one on the right.
                dualcastAvatarId:
                  type: string
                  maxLength: 30
                  description: >-
                    Dualcast mode, instead of dualcastImageKey: an avatar ID
                    from /v1/avatar/list whose image already shows both people
                    (made in the avatar editor with "Both people in one shot").
                    A is the person on the left, B the one on the right. It must
                    show exactly two people.
                dualcastAvatars:
                  type: object
                  required:
                    - left
                    - right
                  description: >-
                    Dualcast mode, instead of dualcastImageKey: two avatar IDs
                    from /v1/avatar/list, put side by side on a podcast set (one
                    more image). left is A, right is B.
                  properties:
                    left:
                      type: string
                      maxLength: 30
                    right:
                      type: string
                      maxLength: 30
                dualcastSet:
                  type: string
                  maxLength: 300
                  description: >-
                    With dualcastAvatars, optional: the set they sit on, in your
                    words (e.g. "a modern studio with a wooden slat wall").
                hostSet:
                  type: string
                  maxLength: 300
                  description: >-
                    Podcast mode, optional: move the host to this set first, in
                    your words (e.g. "a cozy podcast studio with a boom
                    microphone"). Only the background changes; the guest sits in
                    the same room.
                productImageKey:
                  type: string
                  maxLength: 1000
                  description: >-
                    Optional: a product photo (storage key from your media
                    library or an https URL), shown at the top of the video on
                    the words marked [product] … [/product].
                hostVoice:
                  type: string
                  enum:
                    - young-woman-25-30
                    - woman-mid-20s
                    - woman-mid-30s
                    - older-woman-50
                    - young-man-25-30
                    - man-mid-20s
                    - man-mid-30s
                    - older-man-50
                  description: 'The host''s voice. Unset: the one that fits the avatar.'
                accent:
                  type: string
                  description: >-
                    One accent for both, from the English and Spanish accents
                    (e.g. us-neutral-general-american, spanish-mexico). Unset:
                    the language's neutral one.
                language:
                  type: string
                  description: >-
                    The conversation's language (e.g. en, es). Unset: detected
                    from the script.
                musicId:
                  type: string
                  description: Music ID from /v1/music/list for background music
                  maxLength: 30
                caption:
                  type: object
                  properties:
                    preset:
                      type: string
                      enum:
                        - default
                        - beast
                        - umi
                        - tiktok
                        - wrap1
                        - wrap2
                        - ariel
                        - hooked
                        - classic
                        - active
                        - bubble
                        - glass
                        - comic
                        - glow
                        - pastel
                        - neon
                        - retroTV
                        - red
                        - marker
                        - modern
                        - blue
                        - vivid
                      description: Caption preset style
                      default: tiktok
                    alignment:
                      type: string
                      enum:
                        - top
                        - middle
                        - bottom
                      description: Caption position on video
                      default: bottom
                    disabled:
                      type: boolean
                      description: Set to false to show captions on the video
                      default: false
                webhook:
                  type: string
                  description: HTTPS URL to receive completion notification
                  maxLength: 500
                metadata:
                  type: object
                  description: Custom metadata object (max 5KB)
            example:
              script: >-
                [A] 8 out of 10 people say they struggle to make it to the end
                of the month. [/A]

                [B] Quick question. Do you consider yourself middle class? [/B]

                [A] [pauses, thinking] Hmm, I'd say yes. [/A]

                [B] If you want to learn to save, comment SAVE below. [/B]
              avatarId: avatar_podcast_host_01
              guest:
                source: generated
                description: a man in his 30s, casual, average looking
              caption:
                preset: tiktok
                alignment: bottom
                disabled: false
              webhook: https://yoursite.com/webhook
      responses:
        '200':
          description: Talking Avatar created successfully
          content:
            application/json:
              schema:
                type: object
                properties:
                  success:
                    type: boolean
                  data:
                    type: object
                    properties:
                      videoId:
                        type: string
                      projectId:
                        type: string
                      status:
                        type: string
                  message:
                    type: string
              example:
                success: true
                data:
                  videoId: vid_talking_avatar_abc123xyz
                  projectId: proj_ugc_abc123xyz
                  status: STARTED
                message: Talking avatar successfully created
        '400':
          description: Validation error
          content:
            application/json:
              schema:
                type: object
                properties:
                  success:
                    type: boolean
                  message:
                    type: string
              example:
                success: false
                message: 'script: Script must be at least 1 character'
        '401':
          description: Unauthorized
          content:
            application/json:
              schema:
                type: object
                properties:
                  success:
                    type: boolean
                  message:
                    type: string
              example:
                success: false
                message: Invalid API key
      security:
        - ApiKeyAuth: []
components:
  securitySchemes:
    ApiKeyAuth:
      type: apiKey
      in: header
      name: x-api-key

````

This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.