curl --request POST \
--url https://api.poyo.ai/api/generate/submit \
--header 'Authorization: Bearer <token>' \
--header 'Content-Type: application/json' \
--data '
{
"model": "elevenlabs-v3-tts",
"callback_url": "https://your-domain.com/callback",
"input": {
"text": "Hello! This is a test of the text to speech system, powered by ElevenLabs.",
"voice": "Rachel",
"stability": 0.5,
"timestamps": false,
"language_code": "en",
"apply_text_normalization": "auto"
}
}
'{
"code": 200,
"data": {
"task_id": "task-unified-1757165031-uyujaw3d",
"status": "not_started",
"created_time": "2026-05-12T10:30:00"
}
}{
"code": 123,
"error": {
"message": "<string>",
"type": "<string>"
}
}{
"code": 123,
"error": {
"message": "<string>",
"type": "<string>"
}
}ElevenLabs
ElevenLabs V3 TTS
Generate speech audio from text using ElevenLabs V3
POST
/
api
/
generate
/
submit
curl --request POST \
--url https://api.poyo.ai/api/generate/submit \
--header 'Authorization: Bearer <token>' \
--header 'Content-Type: application/json' \
--data '
{
"model": "elevenlabs-v3-tts",
"callback_url": "https://your-domain.com/callback",
"input": {
"text": "Hello! This is a test of the text to speech system, powered by ElevenLabs.",
"voice": "Rachel",
"stability": 0.5,
"timestamps": false,
"language_code": "en",
"apply_text_normalization": "auto"
}
}
'{
"code": 200,
"data": {
"task_id": "task-unified-1757165031-uyujaw3d",
"status": "not_started",
"created_time": "2026-05-12T10:30:00"
}
}{
"code": 123,
"error": {
"message": "<string>",
"type": "<string>"
}
}{
"code": 123,
"error": {
"message": "<string>",
"type": "<string>"
}
}- After submission, a
task_idwill be returned. If you provided acallback_url, when the task status becomesfinishedorfailed, a POST request will be sent to thecallback_url. - Regardless of whether
callback_urlis provided, you can retrieve the result through the unified Query Task Status endpoint.
ElevenLabs V3 TTS
elevenlabs-v3-tts converts text into speech audio. Results are returned as task files, matching the standard image, video, and 3D task response structure.
Available Model
- elevenlabs-v3-tts - ElevenLabs V3 text-to-speech generation
Required Parameters
- text: Text to convert to speech
- Length:
1-5000characters
- Length:
Optional Parameters
- voice: Voice name used for speech generation. Supported values:
Aria,Roger,Sarah,Laura,Charlie,George,Callum,River,Liam,Charlotte,Alice,Matilda,Will,Jessica,Eric,Chris,Brian,Daniel,Lily,Bill,Rachel - stability: Voice stability from
0to1 - timestamps: Boolean. When
true, timestamps may be returned as an additional JSON file - language_code: Optional ISO 639-1 language code used to enforce a language for the model
- apply_text_normalization:
auto,on, oroff
Credit Billing
Billing is calculated by text length:16 credits per 1000 characters.
Notes
- Query status through the standard task status API, not the music detail endpoint.
- Successful responses return
files, where the generated speech usesfile_type: audio. - When
timestampsis enabled and returned, timestamps are uploaded astimestamps.jsonand returned as a separate file item withfile_type: other.
Authorizations
All API endpoints require Bearer Token authentication.
Get your API Key from the API Key Management Page.
Add it to the request header:
Authorization: Bearer YOUR_API_KEY
Body
application/json
