Skip to main content
POST
  1. After submission, a task_id will be returned. If you provided a callback_url, when the task status becomes finished or failed, a POST request will be sent to the callback_url.
  2. Regardless of whether callback_url is provided, you can retrieve the result through the unified Query Task Status endpoint.

Gemini 3.1 Flash TTS

gemini-3-1-flash-tts converts text inputs into expressive speech audio using Google’s Gemini 3.1 Flash TTS model. Results are returned as task files, matching the standard image, video, 3D, and audio task response structure.

Available Model

  • gemini-3-1-flash-tts - Gemini 3.1 Flash text-to-speech generation

Required Parameters

  • text: Text to convert to speech
    • Length: 1-50000 characters
    • Supports natural-language style, pace, accent, and emotional direction inline
    • Supports expressive audio tags such as [sigh], [laughing], [whispering], and [short pause]
    • For multi-speaker synthesis, prefix text lines with speaker aliases that match speakers[].speaker_id

Optional Parameters

  • style_instructions: Optional style, delivery, pace, accent, tone, or emotional direction applied to the request
  • voice: Voice preset for single-speaker synthesis. Ignored when speakers is set. Default: Kore
  • language_code: Optional language enum used to steer multilingual synthesis. If omitted, the model auto-detects from text
  • speakers: Exactly two speaker objects for dialogue. When set, top-level voice is ignored. Each item includes:
    • voice: Gemini voice preset
    • speaker_id: Alias containing only letters, numbers, and underscores with no whitespace, unique within the request, used as a prefix in text
  • temperature: Number controlling delivery variation. Default: 1. Range: 0-2
  • output_format: mp3, wav, or ogg_opus. Default: mp3. Use mp3 for compact playback, wav for uncompressed post-production, and ogg_opus for quality-to-size balance

Voice Presets

Supported values: Achernar, Achird, Algenib, Algieba, Alnilam, Aoede, Autonoe, Callirrhoe, Charon, Despina, Enceladus, Erinome, Fenrir, Gacrux, Iapetus, Kore, Laomedeia, Leda, Orus, Pulcherrima, Puck, Rasalgethi, Sadachbia, Sadaltager, Schedar, Sulafat, Umbriel, Vindemiatrix, Zephyr, Zubenelgenubi.

Credit Billing

Billing is calculated by text length plus style instruction length: 24 credits per 1000 characters.

Request Examples

Single speaker with audio tags

Estimated usage: 152 billable characters, 3.648000 credits.

Multilingual store assistant

Estimated usage: 185 billable characters, 4.440000 credits.

Two speaker podcast

Estimated usage: 200 billable characters, 4.800000 credits.

Notes

  • PoYo’s public parameter name is text.
  • Query status through the standard task status API.
  • Successful responses return files, where the generated speech uses file_type: audio.

Authorizations

Authorization
string
header
required

All API endpoints require Bearer Token authentication.

Get your API Key from the API Key Management Page.

Add it to the request header:

Body

application/json
model
enum<string>
required

Model identifier.

Available options:
gemini-3-1-flash-tts
Example:

"gemini-3-1-flash-tts"

input
object
required
callback_url
string<uri>

Webhook callback URL for result notifications

Example:

"https://your-domain.com/callback"

Response

Task submitted successfully

code
integer
required
Example:

200

data
object
required