Developer Docs/Voice Clone API

Voice Clone API

Build Voice Clone integrations with HeadSwap. Review authentication, request parameters, task creation, status, and result endpoints.

Overview

Convert text to natural-sounding speech using a library of built-in voices, or clone a custom voice from a short audio sample for consistent narration.

Primary Endpoint

POST/api/v1/userVoice/training

Create an asynchronous custom-voice training task from the submitted voice samples. - **Concurrency:** A user may submit only one training request at a time; concurrent submissions are rejected by the service guard. - **Charging:** The current voice_clone_model_costs entry for the selected model determines credits per clone; prices may be configured dynamically. The charge occurs at submission; a failed creation attempt rolls back it according to the documented failure flow. Read GET /api/v1/feature-pricing for current configured prices; a quote does not reserve credits. - **Status:** After creation, poll `GET /api/v1/userVoice/trainingRecord` and inspect `current_status` for progress. ### Request - Send a valid bearer token. The server evaluates the operation in the authenticated caller's access context. - Send an `application/json` body. Required fields: `name`, `voice_urls`. - API-token callers may include `webhook_url` and `webhook_token` for best-effort terminal-state notifications; ordinary JWT/cookie calls ignore these fields. ### Behavior - This is an asynchronous operation: a successful submission creates a task and returns before processing finishes. - Persist the returned task identifier and use the corresponding detail or list operation to observe progress. - Treat the detail endpoint as the source of truth even when webhook delivery is enabled. ### Response - A `200` response confirms task acceptance; it does not by itself mean media generation has completed. - Retain the returned identifier and wait for a documented terminal status before using output URLs. - JSON object responses, including error responses, normally carry a top-level `trace_id` string for this request; include it when contacting support. It is not a task identifier. - Do not infer undocumented fields or statuses; clients should tolerate additional response properties. ### Errors - `400` — Bad Request - Invalid training data. - `401` — Unauthorized - Invalid or missing JWT token. - `403` — Forbidden - Voice clone limit exceeded. ### Related Operations - `DELETE /api/v1/userVoice/{_id}` — Delete voice training record. - `GET /api/v1/userVoice/{_id}` — Get voice training record detail. - `PUT /api/v1/userVoice/{_id}` — Update voice training record name. Authentication: set header Authorization: Bearer <token> (supports user JWT or sk_ API token).

Request Parameters

NameTypeRequiredDescription
namestringYesName of the voice model
voice_urlsarray<string>YesArray of audio file URLs for training; minItems: 1
genderenum: female | maleNoVoice gender; default: "female"
denoisebooleanNoWhether to apply denoising
enhance_voice_similaritybooleanNoWhether to enhance voice similarity
modelenum: a2e | cartesia | minimax | elevenlabsNoVoice model to use; default: "a2e"
languagestringNoLanguage code from client (e.g. en-US, zh-CN). If omitted, server will infer from request headers.
webhook_urlstringNoHTTPS URL to receive task.completed / task.failed notifications. Best-effort delivery, single attempt, no retries; clients should treat the detail API as the source of truth.; maxLength: 2048
webhook_tokenstringNoOptional plaintext token returned in the X-A2e-Webhook-Token header so receivers can verify the request originated from a2e.; maxLength: 256
Request schema and conditional rules
{
  "allOf": [
    {
      "$ref": "#/components/schemas/UserVoiceTrainingRequest"
    },
    {
      "$ref": "#/components/schemas/WebhookInput"
    }
  ]
}

Response Fields

code: integer
data: object
data._id: string
Record id (MongoDB ObjectId)
data.name: string
Voice model name
data.voice_urls: array<string>
Audio sample URLs used for training
data.gender: enum: female | male
Voice gender
data.model: enum: a2e | cartesia | minimax | elevenlabs
Voice model provider
data.lang: string
Internal language code used by model
data.language: string
Original language code from client
data.current_status: enum: sent | pendding | processing | completed | failed
Training status
data.speaker_id: string
Speaker id returned by vendor
data.coins: number
Coins cost for this training (if any)
data.denoise: boolean
Whether denoising is enabled
data.enhance_voice_similarity: boolean
Whether enhance voice similarity is enabled
data.ttsRate: number
TTS rate (coins per 10 seconds)
data.hasRefundCoin: boolean
Whether coins have been refunded for failed training
data.migration_required: boolean
Whether this is a legacy Qwen VC voice that should be re-cloned before the vendor retires it
data.user_id: string
Owner user id
data.createdAt: string
Created time
data.updatedAt: string
Updated time
trace_id: string
Trace ID of this HTTP request. Include it when contacting support about this request. It is generated per request and is not a task identifier; use the returned task `_id` to query results.

Request Example

curl -X POST "https://headswap.app/api/v1/userVoice/training" \
  -H "Authorization: Bearer YOUR_API_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
  "name": "My Custom Voice",
  "voice_urls": [
    "https://example.com/audio1.wav",
    "https://example.com/audio2.wav"
  ]
}'

Related Endpoints

Responses

200

Voice training record created successfully

400

Bad Request - Invalid training data

401

Unauthorized - Invalid or missing bearer token

403

Forbidden - Voice clone limit exceeded

Voice Clone API Documentation