Skip to content

Audio Generation

Auteryn agents can produce audio — not just text and images. Turn a script into natural-sounding narration, or create a multi-speaker, podcast-style conversation, all as part of a normal task.


Text-to-speech

Turn any script or passage into natural spoken narration or voiceover.

Multi-speaker audio

Create a podcast-style conversation between multiple distinct voices from a script.

Slide narration

Add spoken voiceover to presentation slides for a narrated deck.


  1. Describe the audio you want — the text to speak, a voice or style, or the speakers in a conversation.
  2. The agent generates the audio and uploads it to secure cloud storage.
  3. The result appears as a playable artifact in the workspace, ready to download or reuse.

For narrated slide decks, audio narration is added directly to your presentation — see the Presentations deep dive.


Text-to-speech costs 50 credits per 1,000 characters in a preset or designed voice, or 60 credits per 1,000 characters in a cloned voice, per the pricing page. Costs are deducted from your org’s credit balance automatically.


Audio isn’t only an output. When you attach a recording to a chat, or send your agent a voice note through a connected channel, the agent works from what you actually said — models that can’t listen to audio directly are given an automatic transcript first, so the choice of model no longer decides whether a voice note is heard. A transcript is billed at 10 credits per minute of audio, and a repeated recording is transcribed only once.


Audio generation produces a file from a script. If instead you want a real-time, back-and-forth spoken conversation with an agent — where you talk and it replies live — that’s the separate Voice capability.

Audio generation Voice conversation
Output A generated audio file Live two-way conversation
Best for Narration, voiceover, podcasts Talking to your agent in real time
Where Any agent task Voice-enabled agents

  • Be specific about tone and pacing in your instructions (“warm and conversational”, “brisk and professional”).
  • Review generated audio before publishing — models can mispronounce names or unusual terms.
  • For narrated presentations, generate the deck first, then add narration in one pass.