Skip to content

Generating voice & video

Generating is the quickest way to get something to work with. Type a script, pick what you want back, and the Studio produces a watermarked development asset you can play, review and edit.

  1. Select a project (and a conversation, if you’re organising by scene).
  2. Choose a type:
    • Voice — spoken audio only, from your Voice model.
    • Talking-head video — your face animated to the generated speech (uses your Video model + your Voice model).
    • Combined — voice and video produced together.
  3. Optionally pick a role so the asset uses that character’s model.
  4. Type your script or prompt and select Generate.

Generation draws on the model behind the project or role — your own model if you’re generating as yourself (see Building your model), or a licensed / founder likeness otherwise. You only need the model for the type you’re generating: a Voice model is enough for voice, while talking-head needs a Video model too.

Your request is accepted immediately and the asset appears in the timeline as soon as it’s ready — you don’t have to wait on the page. Longer video takes a little longer than voice.

Generation runs in the background, so you can keep queuing ideas. If a model is warming up you may see a short “warming up” message — try again in a moment. Clear messages replace the old failures: if voice or video is briefly unavailable, you’ll be told why rather than hitting an error.

Each asset lands in your project timeline and is filed automatically in the Media Vault. From there you can:

  • Play and review it, and leave comments.
  • Rename it or Save as New to branch a variant.
  • Open it in the editor to trim, arrange and combine it with other assets.
  • Promote it to production when it’s final.

Every generation is metered in tokens and charged to your allowance or wallet. Drafts are inexpensive; see Tokens & billing.