Overview
Convert a still image into an AI-animated video. Control motion intensity, duration, and aspect ratio to produce smooth, high-quality video clips from a single reference image.
Primary Endpoint
/api/v1/userImage2Video/startConvert an image to video using AI generation with custom prompts. ### Request - Send a valid bearer token. The server evaluates the operation in the authenticated caller's access context. - Send an `application/json` body when using the optional controls documented in the request schema. - API-token callers may include `webhook_url` and `webhook_token` for best-effort terminal-state notifications; ordinary JWT/cookie calls ignore these fields. ### Behavior - This is an asynchronous operation: a successful submission creates a task and returns before processing finishes. - Persist the returned task identifier and use the corresponding detail or list operation to observe progress. - Treat the detail endpoint as the source of truth even when webhook delivery is enabled. ### Response - A `200` response confirms task acceptance; it does not by itself mean media generation has completed. - Retain the returned identifier and wait for a documented terminal status before using output URLs. - JSON object responses, including error responses, normally carry a top-level `trace_id` string for this request; include it when contacting support. It is not a task identifier. - Do not infer undocumented fields or statuses; clients should tolerate additional response properties. ### Errors - `400` — Bad Request - Invalid parameters. - `401` — Unauthorized - Invalid or missing JWT token. ### Related Operations - `GET /api/v1/userImage2Video/allRecords` — List image to video tasks. - `GET /api/v1/userImage2Video/{_id}` — Get image to video task detail. - `DELETE /api/v1/userImage2Video/{_id}` — Delete an image to video task. Authentication: set header Authorization: Bearer <token> (supports user JWT or sk_ API token).
Request Parameters
| Name | Type | Required | Description |
|---|---|---|---|
| name | string | No | Name of the image to video task |
| image_url | string | No | URL of the source image. Required for image-to-video; omit for text-to-video and reference-to-video. |
| prompt | string | No | Generation prompt |
| negative_prompt | string | No | Negative prompt to avoid unwanted features |
| lora | string | No | Lora ID. If provided, prompt/negative_prompt can be omitted (prompt will be overwritten by lora preset on backend). |
| generation_mode | enum: image-to-video | text-to-video | reference-to-video | No | Generation mode. Text and reference modes require a2e-v2 or a2e-v2-flash and a nonempty prompt up to 2000 characters. Reference mode accepts video_time of 3, 4, 5, 10 or 15; text mode accepts 5, 10 or 15.; default: "image-to-video" |
| generation_params | object | No | Text/reference parameters. Omit for image-to-video. Reference mode requires at least one asset; text mode accepts none. |
| generation_params.aspect_ratio | enum: 1:1 | 2:3 | 3:2 | 3:4 | 4:3 | 9:16 | 16:9 | 21:9 | No | default: "16:9" |
| generation_params.reference_assets | array<object> | No | At most 9 images, 3 videos and 3 audio files. Video and audio clips must each last 2–15 seconds, with at most 15 seconds total per type.; maxItems: 12 |
| generation_params.reference_assets[].asset_id | string | Yes | Unique reference identifier, such as image_1, video_1 or audio_1. |
| generation_params.reference_assets[].type | enum: image | video | audio | Yes | - |
| generation_params.reference_assets[].url | string | Yes | - |
| generation_params.reference_assets[].duration_seconds | number | No | Required for audio. Video duration is measured by the server.; minimum: 2; maximum: 15 |
| model_type | enum: GENERAL | FLF2V | No | Model type for generation; default: "GENERAL" |
| model_version | enum: a2e | a2e-v2 | a2e-v2-flash | No | Algorithm model version. When omitted, Ultra users default to a2e-v2 and other roles default to a2e. a2e-v2 costs more per second; a2e-v2-flash uses a2e pricing. V2 models are not covered by plan waivers. |
| auto_add_audio | boolean | No | Whether to add audio. Defaults to false for a2e (paid ThinkSound post-processing) and true for a2e-v2/a2e-v2-flash (native audio with no ThinkSound surcharge). |
| end_image_url | string | No | End image URL (required for FLF2V model) |
| extend_prompt | boolean | No | Whether to extend the prompt automatically; default: true |
| number_of_images | integer | No | Number of videos to generate at once; minimum: 1; maximum: 8; default: 1 |
| video_time | integer | No | Requested video time in integer seconds; billing uses this value. Image-to-video and first-last-frame (FLF2V) support 3-20 seconds; video extension supports 5-20 seconds. Reference mode supports 3, 4, 5, 10 or 15 seconds; text mode supports 5, 10 or 15. V1 3/4-second videos use 49/65 frames at 16fps; V1 durations above 5 seconds retain 5-second segment rounding. V2 output duration rounds up to the next supported 17k+5 frame count at 24fps.; minimum: 3; maximum: 20; default: 5 |
| video_length | integer | No | (Deprecated) Video length in frames. Prefer video_time. Backend converts frames to seconds internally.; minimum: 1; maximum: 1000 |
| skip_face_enhance | boolean | No | Whether to skip face similarity enhancement. Defaults to false (enhancing face similarity).; default: false |
| mask_face | boolean | No | Web NSFW preflight result. Set true only after the user confirms masking a detected human face. |
| minor_suspected_skip | boolean | No | Accepted for backward compatibility only. This endpoint has no backend CSAM detection; moderation comes from the algorithm upstream and this flag is not forwarded to it, so setting it does not bypass an upstream 1004.; default: false |
| webhook_url | string | No | HTTPS URL to receive task.completed / task.failed notifications. Best-effort delivery, single attempt, no retries; clients should treat the detail API as the source of truth.; maxLength: 2048 |
| webhook_token | string | No | Optional plaintext token returned in the X-A2e-Webhook-Token header so receivers can verify the request originated from a2e.; maxLength: 256 |
Request schema and conditional rules
{
"allOf": [
{
"type": "object",
"properties": {
"name": {
"type": "string",
"description": "Name of the image to video task",
"example": "My Image Animation"
},
"image_url": {
"type": "string",
"description": "URL of the source image. Required for image-to-video; omit for text-to-video and reference-to-video.",
"example": "https://example.com/image.jpg"
},
"prompt": {
"type": "string",
"description": "Generation prompt",
"example": "Make the person in the image wave their hand"
},
"negative_prompt": {
"type": "string",
"description": "Negative prompt to avoid unwanted features",
"example": "blurry, distorted, static"
},
"lora": {
"type": "string",
"description": "Lora ID. If provided, prompt/negative_prompt can be omitted (prompt will be overwritten by lora preset on backend)."
},
"generation_mode": {
"type": "string",
"enum": [
"image-to-video",
"text-to-video",
"reference-to-video"
],
"default": "image-to-video",
"description": "Generation mode. Text and reference modes require a2e-v2 or a2e-v2-flash and a nonempty prompt up to 2000 characters. Reference mode accepts video_time of 3, 4, 5, 10 or 15; text mode accepts 5, 10 or 15."
},
"generation_params": {
"type": "object",
"description": "Text/reference parameters. Omit for image-to-video. Reference mode requires at least one asset; text mode accepts none.",
"properties": {
"aspect_ratio": {
"type": "string",
"enum": [
"1:1",
"2:3",
"3:2",
"3:4",
"4:3",
"9:16",
"16:9",
"21:9"
],
"default": "16:9"
},
"reference_assets": {
"type": "array",
"maxItems": 12,
"description": "At most 9 images, 3 videos and 3 audio files. Video and audio clips must each last 2–15 seconds, with at most 15 seconds total per type.",
"items": {
"type": "object",
"required": [
"asset_id",
"type",
"url"
],
"properties": {
"asset_id": {
"type": "string",
"description": "Unique reference identifier, such as image_1, video_1 or audio_1."
},
"type": {
"type": "string",
"enum": [
"image",
"video",
"audio"
]
},
"url": {
"type": "string",
"format": "uri"
},
"duration_seconds": {
"type": "number",
"minimum": 2,
"maximum": 15,
"description": "Required for audio. Video duration is measured by the server."
}
}
}
}
}
},
"model_type": {
"type": "string",
"enum": [
"GENERAL",
"FLF2V"
],
"default": "GENERAL",
"description": "Model type for generation"
},
"model_version": {
"type": "string",
"enum": [
"a2e",
"a2e-v2",
"a2e-v2-flash"
],
"description": "Algorithm model version. When omitted, Ultra users default to a2e-v2 and other roles default to a2e. a2e-v2 costs more per second; a2e-v2-flash uses a2e pricing. V2 models are not covered by plan waivers."
},
"auto_add_audio": {
"type": "boolean",
"description": "Whether to add audio. Defaults to false for a2e (paid ThinkSound post-processing) and true for a2e-v2/a2e-v2-flash (native audio with no ThinkSound surcharge)."
},
"end_image_url": {
"type": "string",
"description": "End image URL (required for FLF2V model)",
"example": "https://example.com/end_image.jpg"
},
"extend_prompt": {
"type": "boolean",
"default": true,
"description": "Whether to extend the prompt automatically"
},
"number_of_images": {
"type": "integer",
"minimum": 1,
"maximum": 8,
"default": 1,
"description": "Number of videos to generate at once",
"example": 1
},
"video_time": {
"type": "integer",
"minimum": 3,
"maximum": 20,
"default": 5,
"description": "Requested video time in integer seconds; billing uses this value. Image-to-video and first-last-frame (FLF2V) support 3-20 seconds; video extension supports 5-20 seconds. Reference mode supports 3, 4, 5, 10 or 15 seconds; text mode supports 5, 10 or 15. V1 3/4-second videos use 49/65 frames at 16fps; V1 durations above 5 seconds retain 5-second segment rounding. V2 output duration rounds up to the next supported 17k+5 frame count at 24fps.",
"example": 5
},
"video_length": {
"type": "integer",
"minimum": 1,
"maximum": 1000,
"description": "(Deprecated) Video length in frames. Prefer video_time. Backend converts frames to seconds internally.",
"example": 81
},
"skip_face_enhance": {
"type": "boolean",
"default": false,
"description": "Whether to skip face similarity enhancement. Defaults to false (enhancing face similarity)."
},
"mask_face": {
"type": "boolean",
"description": "Web NSFW preflight result. Set true only after the user confirms masking a detected human face."
},
"minor_suspected_skip": {
"type": "boolean",
"default": false,
"description": "Accepted for backward compatibility only. This endpoint has no backend CSAM detection; moderation comes from the algorithm upstream and this flag is not forwarded to it, so setting it does not bypass an upstream 1004."
}
},
"required": []
},
{
"$ref": "#/components/schemas/WebhookInput"
}
]
}Response Fields
- code: integer
- data: object
- data._id: string
- data.name: string
- data.image_url: string
- data.current_status: string
- data.result_url: string
- data.cover_url: string
- data.mask_face: boolean
- Whether clients should display the privacy-processed input image as the task cover
- data.model_type: string
- data.end_image_url: string
- data.extend_prompt: boolean
- data.lora: string
- data.video_time: number
- data.video_length: number
- data.coins: number
- data.remainingDays: number
- data.expirationDate: string
- data.isExpired: boolean
- data.expirationDays: number
- data.total_videos: number
- Present only when number_of_images > 1
- data.all_records: array<object>
- Present only when number_of_images > 1
- trace_id: string
- Trace ID of this HTTP request. Include it when contacting support about this request. It is generated per request and is not a task identifier; use the returned task `_id` to query results.
Request Example
curl -X POST "https://headswap.app/api/v1/userImage2Video/start" \
-H "Authorization: Bearer YOUR_API_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"name": "My Image Animation",
"image_url": "https://example.com/image.jpg",
"prompt": "Make the person in the image wave their hand",
"negative_prompt": "blurry, distorted, static",
"model_type": "GENERAL",
"extend_prompt": true,
"number_of_images": 1,
"video_time": 5,
"skip_face_enhance": false,
"mask_face": false
}'Related Endpoints
/api/v1/userImage2Video/{_id}Get image to video task detail
/api/v1/userImage2Video/startStart image to video conversion
/api/v1/userImage2Video/allRecordsList image to video tasks
/api/v1/userImage2Video/batchDetailBatch query task details
/api/v1/userImage2Video/avgProcessingTimeGet image-to-video processing estimates
/api/v1/userImage2Video/unlimitedQueueLevelGet image-to-video unlimited queue level
/api/v1/userImage2Video/{_id}Delete an image to video task
/api/v1/userImage2Video/prompt_extensionExtend prompt for image to video
/api/v1/userImage2Video/flf2v_prompt_extensionExtend prompt for FLF2V image to video
Responses
Image to video task started successfully
Bad Request - Invalid parameters
Unauthorized - Invalid or missing bearer token