Overview
Convert text to natural-sounding speech using a library of built-in voices, or clone a custom voice from a short audio sample for consistent narration.
Primary Endpoint
/api/v1/userVoice/trainingCreate an asynchronous custom-voice training task from the submitted voice samples. - **Concurrency:** A user may submit only one training request at a time; concurrent submissions are rejected by the service guard. - **Charging:** The current voice_clone_model_costs entry for the selected model determines credits per clone; prices may be configured dynamically. The charge occurs at submission; a failed creation attempt rolls back it according to the documented failure flow. Read GET /api/v1/feature-pricing for current configured prices; a quote does not reserve credits. - **Status:** After creation, poll `GET /api/v1/userVoice/trainingRecord` and inspect `current_status` for progress. ### Request - Send a valid bearer token. The server evaluates the operation in the authenticated caller's access context. - Send an `application/json` body. Required fields: `name`, `voice_urls`. - API-token callers may include `webhook_url` and `webhook_token` for best-effort terminal-state notifications; ordinary JWT/cookie calls ignore these fields. ### Behavior - This is an asynchronous operation: a successful submission creates a task and returns before processing finishes. - Persist the returned task identifier and use the corresponding detail or list operation to observe progress. - Treat the detail endpoint as the source of truth even when webhook delivery is enabled. ### Response - A `200` response confirms task acceptance; it does not by itself mean media generation has completed. - Retain the returned identifier and wait for a documented terminal status before using output URLs. - JSON object responses, including error responses, normally carry a top-level `trace_id` string for this request; include it when contacting support. It is not a task identifier. - Do not infer undocumented fields or statuses; clients should tolerate additional response properties. ### Errors - `400` — Bad Request - Invalid training data. - `401` — Unauthorized - Invalid or missing JWT token. - `403` — Forbidden - Voice clone limit exceeded. ### Related Operations - `DELETE /api/v1/userVoice/{_id}` — Delete voice training record. - `GET /api/v1/userVoice/{_id}` — Get voice training record detail. - `PUT /api/v1/userVoice/{_id}` — Update voice training record name. Authentication: set header Authorization: Bearer <token> (supports user JWT or sk_ API token).
Request Parameters
| Name | Type | Required | Description |
|---|---|---|---|
| name | string | Yes | Name of the voice model |
| voice_urls | array<string> | Yes | Array of audio file URLs for training; minItems: 1 |
| gender | enum: female | male | No | Voice gender; default: "female" |
| denoise | boolean | No | Whether to apply denoising |
| enhance_voice_similarity | boolean | No | Whether to enhance voice similarity |
| model | enum: a2e | cartesia | minimax | elevenlabs | No | Voice model to use; default: "a2e" |
| language | string | No | Language code from client (e.g. en-US, zh-CN). If omitted, server will infer from request headers. |
| webhook_url | string | No | HTTPS URL to receive task.completed / task.failed notifications. Best-effort delivery, single attempt, no retries; clients should treat the detail API as the source of truth.; maxLength: 2048 |
| webhook_token | string | No | Optional plaintext token returned in the X-A2e-Webhook-Token header so receivers can verify the request originated from a2e.; maxLength: 256 |
Request schema and conditional rules
{
"allOf": [
{
"$ref": "#/components/schemas/UserVoiceTrainingRequest"
},
{
"$ref": "#/components/schemas/WebhookInput"
}
]
}Response Fields
- code: integer
- data: object
- data._id: string
- Record id (MongoDB ObjectId)
- data.name: string
- Voice model name
- data.voice_urls: array<string>
- Audio sample URLs used for training
- data.gender: enum: female | male
- Voice gender
- data.model: enum: a2e | cartesia | minimax | elevenlabs
- Voice model provider
- data.lang: string
- Internal language code used by model
- data.language: string
- Original language code from client
- data.current_status: enum: sent | pendding | processing | completed | failed
- Training status
- data.speaker_id: string
- Speaker id returned by vendor
- data.coins: number
- Coins cost for this training (if any)
- data.denoise: boolean
- Whether denoising is enabled
- data.enhance_voice_similarity: boolean
- Whether enhance voice similarity is enabled
- data.ttsRate: number
- TTS rate (coins per 10 seconds)
- data.hasRefundCoin: boolean
- Whether coins have been refunded for failed training
- data.migration_required: boolean
- Whether this is a legacy Qwen VC voice that should be re-cloned before the vendor retires it
- data.user_id: string
- Owner user id
- data.createdAt: string
- Created time
- data.updatedAt: string
- Updated time
- trace_id: string
- Trace ID of this HTTP request. Include it when contacting support about this request. It is generated per request and is not a task identifier; use the returned task `_id` to query results.
Request Example
curl -X POST "https://headswap.app/api/v1/userVoice/training" \
-H "Authorization: Bearer YOUR_API_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"name": "My Custom Voice",
"voice_urls": [
"https://example.com/audio1.wav",
"https://example.com/audio2.wav"
]
}'Related Endpoints
/api/v1/userVoice/{_id}Get voice training record detail
/api/v1/userVoice/trainingStart voice training
/api/v1/userVoice/trainingRecordGet voice training records
/api/v1/userVoice/completedRecordGet completed voice training records
/api/v1/userVoice/{_id}Delete voice training record
/api/v1/userVoice/{_id}Update voice training record name
Responses
Voice training record created successfully
Bad Request - Invalid training data
Unauthorized - Invalid or missing bearer token
Forbidden - Voice clone limit exceeded