Upload Video
Click or drag a file to this area to upload
Format: mp4/mov, Duration: 3 seconds to 5 minutes
Resolution recommended 720P or 1080P, maximum 4K
Upload Image
Click or drag a file to this area to upload
Format: jpg/png, Size: up to 20M
Reusable digital twin
Train a reusable speaking avatar from one clear recording, then use it in future AI videos without filming the same performance again.
Price: 100 credits are required for one "Studio Avatar" session.









Continue training the deep learning model based on the provided video material to further improve the clarity and similarity of the generated faces. If the video material has good audio-visual synchronization, the model after continued training can generate lip movements with higher synchronization. If the audio and video are not synchronized or the sound quality is poor, please do not "Studio Avatar".
Video mode uses a speaking performance and supports continued training for improved lip-sync quality. If you only have a clear portrait, photo mode provides a shorter setup path for creating a reusable avatar.
The current uploader accepts MP4 or MOV recordings from 3 seconds to 5 minutes. A clear 720p or 1080p recording is recommended.
Use one visible speaker, keep the complete face in frame, avoid obstructions and background noise, and make sure the audio matches the lip movements.
Choose photo mode when you have one clear portrait and want a faster setup. Choose video mode when you can provide a speaking performance for training.
Yes. Upload only recordings you own or are authorized to process, and obtain the subject’s consent for avatar creation and its intended use.