Overview
Convert text to natural-sounding speech using a library of built-in voices, or clone a custom voice from a short audio sample for consistent narration.
Primary Endpoint
/api/v1/video/send_ttsConvert text to speech using AI voices with customizable parameters. ### Voice Selection Provide exactly one of `tts_id` (system voice) or `user_voice_id` (custom cloned voice). Do not send both: synthesis would use the `tts_id` voice while locale defaults and billing follow `user_voice_id`. - `tts_id`: Use built-in system voice (400+ voices available) - `user_voice_id`: Use your own custom cloned voice; omitted `country` and `region` default to `en-US` ### Billing System voice lists return the current account's `ttsRate` in credits per 10 seconds. Charges are based on the generated audio duration: `ceil(duration / 10) * ttsRate`. A valid zero rate means no credits are charged. Fetch the selected voice's current rate before submitting; generated duration is unknown before synthesis, so this rate is not a total-price quote. ### Text Limitations - **API token users**: No local 3000-unit text limit is applied by this controller - **Web users**: Maximum 3000 weighted units after pause tags are removed - The server counts ASCII and Latin-1 characters as one unit, Cyrillic characters as one unit, and other non-Latin-1 characters as two units ### Rate Limiting & Captcha For non-API users, captcha verification may be required after frequent usage: - Use `type` parameter to specify captcha type (`turnstile` or `aliyun_captcha`) - Provide corresponding `turnstile_token` or `captchaVerifyParam` when captcha is required - API token users skip captcha verification If generated audio needs a duration probe and that probe fails, msg includes the audio source failure reason. Invalid audio URLs are cached for 180 seconds; cached failures return the same reason without probing again. ### Request - Send a valid bearer token. The server evaluates the operation in the authenticated caller's access context. - Send an `application/json` body. Required fields: `msg`. ### Behavior - Validates the submitted payload and performs the operation described by the request schema. - Conditional fields, mutually exclusive inputs, limits, and defaults are enforced as documented by their schemas. ### Response - On `200`, consume the fields defined by that response schema below. - JSON object responses, including error responses, normally carry a top-level `trace_id` string for this request; include it when contacting support. It is not a task identifier. - Do not infer undocumented fields or statuses; clients should tolerate additional response properties. ### Errors - `400` — Invalid request parameters. - `401` — Unauthorized - Invalid or missing JWT token. - `500` — Generated audio duration probe failed; msg includes the audio source failure reason. Authentication: set header Authorization: Bearer <token> (supports user JWT or sk_ API token).
Request Parameters
| Name | Type | Required | Description |
|---|---|---|---|
| msg | string | Yes | Text content to convert to speech. Web users are limited to 3000 weighted units; API token users are not checked against this local limit.; minLength: 1 |
| tts_id | string | No | System voice ID (MongoDB ObjectId). Must provide either tts_id or user_voice_id. |
| user_voice_id | string | No | Custom cloned voice ID (MongoDB ObjectId). Used only when tts_id is absent. |
| country | string | No | Locale language part (e.g. 'en', 'zh', 'pt', 'ja'). Optional when using user_voice_id, defaults to 'en'. Combine with `region` to form locale like 'en-US'/'zh-CN'. Values should come from POST /api/v1/anchor/language_list (top-level `value`).; default: "en" |
| region | string | No | Locale region part (e.g. 'US', 'CN', 'BR', 'JP'). Optional when using user_voice_id, defaults to 'US'. Combine with `country` to form locale like 'en-US'/'zh-CN'. Values should come from POST /api/v1/anchor/language_list (child `value`).; default: "US" |
| speechRate | number | No | Speech speed multiplier (0.5-2.0); minimum: 0.5; maximum: 2; default: 1 |
| type | enum: turnstile | aliyun_captcha | No | Captcha type (required when captcha verification is needed) |
| turnstile_token | string | No | Captcha token (required when type is 'turnstile') |
| captchaVerifyParam | string | No | Captcha verification parameter (required when type is 'aliyun_captcha') |
Request schema and conditional rules
{
"type": "object",
"required": [
"msg"
],
"anyOf": [
{
"required": [
"tts_id"
]
},
{
"required": [
"user_voice_id"
]
}
],
"properties": {
"msg": {
"type": "string",
"description": "Text content to convert to speech. Web users are limited to 3000 weighted units; API token users are not checked against this local limit.",
"example": "Welcome. Let's start your AI journey.",
"minLength": 1
},
"tts_id": {
"type": "string",
"description": "System voice ID (MongoDB ObjectId). Must provide either tts_id or user_voice_id.",
"example": "66dc3c1b7dc1f1c483cc5ab8"
},
"user_voice_id": {
"type": "string",
"description": "Custom cloned voice ID (MongoDB ObjectId). Used only when tts_id is absent.",
"example": "66f1234567890abcdef12345"
},
"country": {
"type": "string",
"description": "Locale language part (e.g. 'en', 'zh', 'pt', 'ja'). Optional when using user_voice_id, defaults to 'en'. Combine with `region` to form locale like 'en-US'/'zh-CN'. Values should come from POST /api/v1/anchor/language_list (top-level `value`).",
"example": "en",
"default": "en"
},
"region": {
"type": "string",
"description": "Locale region part (e.g. 'US', 'CN', 'BR', 'JP'). Optional when using user_voice_id, defaults to 'US'. Combine with `country` to form locale like 'en-US'/'zh-CN'. Values should come from POST /api/v1/anchor/language_list (child `value`).",
"example": "US",
"default": "US"
},
"speechRate": {
"type": "number",
"description": "Speech speed multiplier (0.5-2.0)",
"example": 1,
"minimum": 0.5,
"maximum": 2,
"default": 1
},
"type": {
"type": "string",
"description": "Captcha type (required when captcha verification is needed)",
"enum": [
"turnstile",
"aliyun_captcha"
],
"example": "turnstile"
},
"turnstile_token": {
"type": "string",
"description": "Captcha token (required when type is 'turnstile')",
"example": "0x4AAAAAAxxxxxxxxxxxxxxxxxx"
},
"captchaVerifyParam": {
"type": "string",
"description": "Captcha verification parameter (required when type is 'aliyun_captcha')",
"example": "xxxxxxxxxxxx"
}
},
"example": {
"msg": "Welcome. Let's start your AI journey.",
"tts_id": "66dc3c1b7dc1f1c483cc5ab8"
}
}Response Fields
- code: integer
- data: string
- Generated audio preview URL
- trace_id: string
- Trace ID of this HTTP request. Include it when contacting support about this request. It is generated per request and is not a task identifier; use the returned task `_id` to query results.
Request Example
curl -X POST "https://headswap.app/api/v1/video/send_tts" \
-H "Authorization: Bearer YOUR_API_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"msg": "Welcome. Let'\''s start your AI journey.",
"tts_id": "66dc3c1b7dc1f1c483cc5ab8",
"speechRate": 1
}'Related Endpoints
Responses
Text-to-speech generation successful
Invalid request parameters
Unauthorized - Invalid or missing bearer token
Generated audio duration probe failed; msg includes the audio source failure reason.