Motion you can feel. Sound that belongs.
Kling 4.0 focuses on stable action and camera movement, expressive performances and synchronized sound. Two-channel stereo and improved lip-sync bring dialogue and singing scenes to life.
From a product reveal to a continuous story, explore a new generation of AI video with richer motion, sound and creative control.
Kling 4.0 focuses on stable action and camera movement, expressive performances and synchronized sound. Two-channel stereo and improved lip-sync bring dialogue and singing scenes to life.
Combine images, videos and subjects to guide identity, action and composition. The model also supports editing subjects, backgrounds and visual styles in existing footage, helping develop new variations from a shared creative idea.
Use up to 10 keyframe images to shape important moments across a 3–30-second video. Guide a character’s entrance, a scene change or a product reveal while giving the story room to unfold.
Explore product reveals and new settings while using reference assets to guide the product’s appearance.
Develop variations around a reference video’s pacing and camera language for different creative concepts.
Explore longer scenes with consistent subjects, controlled key moments and more natural performances.
Explore dialogue, accents and readable on-screen text for audiences across languages.
The full model brings together text-to-video, image-to-video, multimodal references, video editing and flexible keyframes.
3–30 seconds · 720p / 1080p / 4K
Flash is designed for frequent creative exploration, with faster generation and a focus on value for trying different ideas.
3–20 seconds · 720p · 8-bit SDR
Kling 4.0 is Kling AI’s next-generation video model, focused on audiovisual quality, multimodal references and controlled storytelling.
Explore product campaigns, social clips, character stories and multilingual content. Use references to guide products and characters, keyframes to shape scene changes and synchronized sound to enrich the story.
The full model supports text-to-video, image-to-video and up to 15 combined reference items: up to 10 images, 5 videos with a combined duration of 30 seconds, and 7 subjects, including at most 3 video-based subjects. Voice references are also supported; the combined limit still applies.
The full model supports 3–30-second videos at 720p, 1080p or 4K, with up to 10 keyframe images to guide important moments.
It supports two-channel stereo, improved lip-sync and multilingual dialogue, bringing motion, speech and singing together more naturally.
Flash targets faster, frequent creation with 3–20-second videos at 720p in 8-bit SDR. The full model supports 3–30 seconds and up to 4K for content that needs more time and visual detail.
Up to 30 seconds in 4K, with multimodal references, keyframe control and stereo audio to bring every idea to life.