Developer Docs/Text to Speech API

Text to Speech API

Build Text to Speech integrations with HeadSwap. Review authentication, request parameters, task creation, status, and result endpoints.

Overview

Convert text to natural-sounding speech using a library of built-in voices, or clone a custom voice from a short audio sample for consistent narration.

Primary Endpoint

POST/api/v1/video/send_tts

Convert text to speech using AI voices with customizable parameters. ### Voice Selection Provide exactly one of `tts_id` (system voice) or `user_voice_id` (custom cloned voice). Do not send both: synthesis would use the `tts_id` voice while locale defaults and billing follow `user_voice_id`. - `tts_id`: Use built-in system voice (400+ voices available) - `user_voice_id`: Use your own custom cloned voice; omitted `country` and `region` default to `en-US` ### Billing System voice lists return the current account's `ttsRate` in credits per 10 seconds. Charges are based on the generated audio duration: `ceil(duration / 10) * ttsRate`. A valid zero rate means no credits are charged. Fetch the selected voice's current rate before submitting; generated duration is unknown before synthesis, so this rate is not a total-price quote. ### Text Limitations - **API token users**: No local 3000-unit text limit is applied by this controller - **Web users**: Maximum 3000 weighted units after pause tags are removed - The server counts ASCII and Latin-1 characters as one unit, Cyrillic characters as one unit, and other non-Latin-1 characters as two units ### Rate Limiting & Captcha For non-API users, captcha verification may be required after frequent usage: - Use `type` parameter to specify captcha type (`turnstile` or `aliyun_captcha`) - Provide corresponding `turnstile_token` or `captchaVerifyParam` when captcha is required - API token users skip captcha verification If generated audio needs a duration probe and that probe fails, msg includes the audio source failure reason. Invalid audio URLs are cached for 180 seconds; cached failures return the same reason without probing again. ### Request - Send a valid bearer token. The server evaluates the operation in the authenticated caller's access context. - Send an `application/json` body. Required fields: `msg`. ### Behavior - Validates the submitted payload and performs the operation described by the request schema. - Conditional fields, mutually exclusive inputs, limits, and defaults are enforced as documented by their schemas. ### Response - On `200`, consume the fields defined by that response schema below. - JSON object responses, including error responses, normally carry a top-level `trace_id` string for this request; include it when contacting support. It is not a task identifier. - Do not infer undocumented fields or statuses; clients should tolerate additional response properties. ### Errors - `400` — Invalid request parameters. - `401` — Unauthorized - Invalid or missing JWT token. - `500` — Generated audio duration probe failed; msg includes the audio source failure reason. Authentication: set header Authorization: Bearer <token> (supports user JWT or sk_ API token).

Request Parameters

NameTypeRequiredDescription
msgstringYesText content to convert to speech. Web users are limited to 3000 weighted units; API token users are not checked against this local limit.; minLength: 1
tts_idstringNoSystem voice ID (MongoDB ObjectId). Must provide either tts_id or user_voice_id.
user_voice_idstringNoCustom cloned voice ID (MongoDB ObjectId). Used only when tts_id is absent.
countrystringNoLocale language part (e.g. 'en', 'zh', 'pt', 'ja'). Optional when using user_voice_id, defaults to 'en'. Combine with `region` to form locale like 'en-US'/'zh-CN'. Values should come from POST /api/v1/anchor/language_list (top-level `value`).; default: "en"
regionstringNoLocale region part (e.g. 'US', 'CN', 'BR', 'JP'). Optional when using user_voice_id, defaults to 'US'. Combine with `country` to form locale like 'en-US'/'zh-CN'. Values should come from POST /api/v1/anchor/language_list (child `value`).; default: "US"
speechRatenumberNoSpeech speed multiplier (0.5-2.0); minimum: 0.5; maximum: 2; default: 1
typeenum: turnstile | aliyun_captchaNoCaptcha type (required when captcha verification is needed)
turnstile_tokenstringNoCaptcha token (required when type is 'turnstile')
captchaVerifyParamstringNoCaptcha verification parameter (required when type is 'aliyun_captcha')
Request schema and conditional rules
{
  "type": "object",
  "required": [
    "msg"
  ],
  "anyOf": [
    {
      "required": [
        "tts_id"
      ]
    },
    {
      "required": [
        "user_voice_id"
      ]
    }
  ],
  "properties": {
    "msg": {
      "type": "string",
      "description": "Text content to convert to speech. Web users are limited to 3000 weighted units; API token users are not checked against this local limit.",
      "example": "Welcome. Let's start your AI journey.",
      "minLength": 1
    },
    "tts_id": {
      "type": "string",
      "description": "System voice ID (MongoDB ObjectId). Must provide either tts_id or user_voice_id.",
      "example": "66dc3c1b7dc1f1c483cc5ab8"
    },
    "user_voice_id": {
      "type": "string",
      "description": "Custom cloned voice ID (MongoDB ObjectId). Used only when tts_id is absent.",
      "example": "66f1234567890abcdef12345"
    },
    "country": {
      "type": "string",
      "description": "Locale language part (e.g. 'en', 'zh', 'pt', 'ja'). Optional when using user_voice_id, defaults to 'en'. Combine with `region` to form locale like 'en-US'/'zh-CN'. Values should come from POST /api/v1/anchor/language_list (top-level `value`).",
      "example": "en",
      "default": "en"
    },
    "region": {
      "type": "string",
      "description": "Locale region part (e.g. 'US', 'CN', 'BR', 'JP'). Optional when using user_voice_id, defaults to 'US'. Combine with `country` to form locale like 'en-US'/'zh-CN'. Values should come from POST /api/v1/anchor/language_list (child `value`).",
      "example": "US",
      "default": "US"
    },
    "speechRate": {
      "type": "number",
      "description": "Speech speed multiplier (0.5-2.0)",
      "example": 1,
      "minimum": 0.5,
      "maximum": 2,
      "default": 1
    },
    "type": {
      "type": "string",
      "description": "Captcha type (required when captcha verification is needed)",
      "enum": [
        "turnstile",
        "aliyun_captcha"
      ],
      "example": "turnstile"
    },
    "turnstile_token": {
      "type": "string",
      "description": "Captcha token (required when type is 'turnstile')",
      "example": "0x4AAAAAAxxxxxxxxxxxxxxxxxx"
    },
    "captchaVerifyParam": {
      "type": "string",
      "description": "Captcha verification parameter (required when type is 'aliyun_captcha')",
      "example": "xxxxxxxxxxxx"
    }
  },
  "example": {
    "msg": "Welcome. Let's start your AI journey.",
    "tts_id": "66dc3c1b7dc1f1c483cc5ab8"
  }
}

Response Fields

code: integer
data: string
Generated audio preview URL
trace_id: string
Trace ID of this HTTP request. Include it when contacting support about this request. It is generated per request and is not a task identifier; use the returned task `_id` to query results.

Request Example

curl -X POST "https://headswap.app/api/v1/video/send_tts" \
  -H "Authorization: Bearer YOUR_API_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
  "msg": "Welcome. Let'\''s start your AI journey.",
  "tts_id": "66dc3c1b7dc1f1c483cc5ab8",
  "speechRate": 1
}'

Related Endpoints

Responses

200

Text-to-speech generation successful

400

Invalid request parameters

401

Unauthorized - Invalid or missing bearer token

500

Generated audio duration probe failed; msg includes the audio source failure reason.

Text to Speech API Documentation