curl --request POST \
--url https://api.poyo.ai/api/generate/submit \
--header 'Authorization: Bearer <token>' \
--header 'Content-Type: application/json' \
--data '
{
"model": "gemini-3-1-flash-tts",
"callback_url": "https://your-domain.com/callback",
"input": {
"text": "Good morning, team. [short pause] Our release is stable, and the customer demo starts in ten minutes.",
"style_instructions": "Warm, confident product narrator with clear pacing.",
"voice": "Kore",
"temperature": 1,
"output_format": "mp3"
}
}
'{
"code": 200,
"data": {
"task_id": "task-unified-1757165031-uyujaw3d",
"status": "not_started",
"created_time": "2026-06-25T10:30:00"
}
}{
"code": 123,
"error": {
"message": "<string>",
"type": "<string>"
}
}{
"code": 123,
"error": {
"message": "<string>",
"type": "<string>"
}
}Google
Gemini 3.1 Flash TTS
Generate expressive multilingual speech using Gemini 3.1 Flash TTS
POST
/
api
/
generate
/
submit
curl --request POST \
--url https://api.poyo.ai/api/generate/submit \
--header 'Authorization: Bearer <token>' \
--header 'Content-Type: application/json' \
--data '
{
"model": "gemini-3-1-flash-tts",
"callback_url": "https://your-domain.com/callback",
"input": {
"text": "Good morning, team. [short pause] Our release is stable, and the customer demo starts in ten minutes.",
"style_instructions": "Warm, confident product narrator with clear pacing.",
"voice": "Kore",
"temperature": 1,
"output_format": "mp3"
}
}
'{
"code": 200,
"data": {
"task_id": "task-unified-1757165031-uyujaw3d",
"status": "not_started",
"created_time": "2026-06-25T10:30:00"
}
}{
"code": 123,
"error": {
"message": "<string>",
"type": "<string>"
}
}{
"code": 123,
"error": {
"message": "<string>",
"type": "<string>"
}
}- After submission, a
task_idwill be returned. If you provided acallback_url, when the task status becomesfinishedorfailed, a POST request will be sent to thecallback_url. - Regardless of whether
callback_urlis provided, you can retrieve the result through the unified Query Task Status endpoint.
Gemini 3.1 Flash TTS
gemini-3-1-flash-tts converts text inputs into expressive speech audio using Google’s Gemini 3.1 Flash TTS model. Results are returned as task files, matching the standard image, video, 3D, and audio task response structure.
Available Model
- gemini-3-1-flash-tts - Gemini 3.1 Flash text-to-speech generation
Required Parameters
- text: Text to convert to speech
- Length:
1-50000characters - Supports natural-language style, pace, accent, and emotional direction inline
- Supports expressive audio tags such as
[sigh],[laughing],[whispering], and[short pause] - For multi-speaker synthesis, prefix text lines with speaker aliases that match
speakers[].speaker_id
- Length:
Optional Parameters
- style_instructions: Optional style, delivery, pace, accent, tone, or emotional direction applied to the request
- voice: Voice preset for single-speaker synthesis. Ignored when
speakersis set. Default:Kore - language_code: Optional language enum used to steer multilingual synthesis. If omitted, the model auto-detects from
text - speakers: Exactly two speaker objects for dialogue. When set, top-level
voiceis ignored. Each item includes:voice: Gemini voice presetspeaker_id: Alias containing only letters, numbers, and underscores with no whitespace, unique within the request, used as a prefix intext
- temperature: Number controlling delivery variation. Default:
1. Range:0-2 - output_format:
mp3,wav, orogg_opus. Default:mp3. Usemp3for compact playback,wavfor uncompressed post-production, andogg_opusfor quality-to-size balance
Voice Presets
Supported values:Achernar, Achird, Algenib, Algieba, Alnilam, Aoede, Autonoe, Callirrhoe, Charon, Despina, Enceladus, Erinome, Fenrir, Gacrux, Iapetus, Kore, Laomedeia, Leda, Orus, Pulcherrima, Puck, Rasalgethi, Sadachbia, Sadaltager, Schedar, Sulafat, Umbriel, Vindemiatrix, Zephyr, Zubenelgenubi.
Credit Billing
Billing is calculated by text length plus style instruction length:24 credits per 1000 characters.
Request Examples
Single speaker with audio tags
Estimated usage:152 billable characters, 3.648000 credits.
{
"model": "gemini-3-1-flash-tts",
"callback_url": "https://your-domain.com/callback",
"input": {
"text": "Good morning, team. [short pause] Our release is stable, and the customer demo starts in ten minutes.",
"style_instructions": "Warm, confident product narrator with clear pacing.",
"voice": "Kore",
"temperature": 1,
"output_format": "mp3"
}
}
Multilingual store assistant
Estimated usage:185 billable characters, 4.440000 credits.
{
"model": "gemini-3-1-flash-tts",
"callback_url": "https://your-domain.com/callback",
"input": {
"text": "English: Your order is ready for pickup.\nSpanish: Su pedido esta listo para recoger.\nFrench: Votre commande est prete a etre retiree.",
"style_instructions": "Friendly store assistant voice, natural and concise.",
"voice": "Zephyr",
"language_code": "English (US)",
"temperature": 1,
"output_format": "wav"
}
}
Two speaker podcast
Estimated usage:200 billable characters, 4.800000 credits.
{
"model": "gemini-3-1-flash-tts",
"callback_url": "https://your-domain.com/callback",
"input": {
"text": "Host: Welcome back to Launch Notes. [laughing] Today's build finally passed.\nGuest: Great news. Let's keep the update clear and calm for users.",
"style_instructions": "Conversational podcast tone with distinct speaker energy.",
"speakers": [
{
"speaker_id": "Host",
"voice": "Charon"
},
{
"speaker_id": "Guest",
"voice": "Kore"
}
],
"temperature": 1,
"output_format": "ogg_opus"
}
}
Notes
- PoYo’s public parameter name is
text. - Query status through the standard task status API.
- Successful responses return
files, where the generated speech usesfile_type: audio.
Authorizations
All API endpoints require Bearer Token authentication.
Get your API Key from the API Key Management Page.
Add it to the request header:
Authorization: Bearer YOUR_API_KEY
Body
application/json
