30 s
Time on screen
A generation runs for up to 30 seconds without a cut; Kling 3.0 stopped at 15. A finished clip can then be extended, to a combined length of up to 2 minutes.
Coming soon
Kling 4.0 is not open for generation here yet
The controls on this page preview the layout only and submit nothing. Until Kling 4.0 opens on this site, the video models below are ready to use.
Available today
Kling 4.0
Text and image to video
Model capabilities
The limits of the model fall into four groups. Each group answers a planning question that comes up before a prompt is written.
30 s
A generation runs for up to 30 seconds without a cut; Kling 3.0 stopped at 15. A finished clip can then be extended, to a combined length of up to 2 minutes.
10 keyframes
As many as 10 still frames can be set along the clip. Every one is a picture the video is required to reach, which leaves the model to work out only how the scene travels from one to the next.
15 references
Each generation has room for 15 references. Images may fill up to 10 of those places, videos up to 5 and subjects up to 7. They hold a face, a product or a visual style steady throughout, and they can direct changes to footage that already exists.
4K HDR
Output goes as high as 4K, and both the 1080p and the 4K setting can be delivered in 10-bit HDR. Sound arrives with the video as a two-channel stereo track, and lip movement follows speech more closely than in earlier versions.
The figures are grouped by the part of a project in which they matter.
Keyframes and most references are still images, which means the look of a shot can be settled before any video is rendered. Work through the steps in order.
Create or collect the frames the shot must contain: the opening view, the turning point and the last picture. An image model is the quickest way to produce matching stills in one style.
Create stills with Nano Banana ProArrange the stills as keyframes in the sequence the audience will see them. Kling 4.0 takes up to 10, though a simple move from A to B is fully described by two.
Add references for whatever must remain recognizable: a person, a product, a location or a look. The allowance of 15 goes further when each file has a distinct purpose, so leave out near-duplicates.
Use the prompt for what stills cannot show: how subjects move, where the camera goes, what is said and what is heard. The 8,000-token allowance leaves space to write this in the order it happens.
Pick the aspect ratio for the destination, such as vertical for phone feeds or 21:9 for a cinematic frame, and then the resolution. 10-bit HDR at 1080p or 4K suits footage that will be color graded.
Match the requirement of the shot to the model. Each row names the deciding factor.
| The shot needs | Model | Why |
|---|---|---|
| A single take longer than 15 seconds, with specific frames fixed along the way | Kling 4.0 | Generations of up to 30 seconds and up to 10 keyframes. |
| A short cinematic clip with sound, in Standard or Pro quality | Kling 3 | Clips of 3, 5, 10 or 15 seconds with native audio, from text or a reference image. |
| A start frame and an end frame, or changes to an existing video | Kling o3 | An omni model with first and last frame control and video editing. |
| A large set of mixed reference material | Seedance 2.5 | Takes as many as 50 multimodal reference materials and produces up to 30 seconds at up to 4K. |
Kling 4.0 has a lighter sibling. Flash gives priority to speed and cost per clip, while the full model gives priority to realism and control.
Each plan lists the stills, the references and the prompt separately, because the model receives them as separate inputs.
Vertical · 30 seconds
21:9 · 30 seconds
Widescreen · 15 seconds
Longer clips give small inconsistencies more time to show. Five checks catch most of them.
Keyframes made with different lighting, lenses or palettes produce visible jumps. Generate them as a set and compare them side by side.
Prepare stills in the aspect ratio of the final video. A square still cannot fill a 21:9 frame without cropping or invented edges.
A reference that disagrees with a keyframe, such as a different jacket or hairstyle, leaves the outcome to chance. Remove one of the two.
State what should be heard and from which side. With a two-channel track, position is part of the description.
Give the last second a stable picture. Editing, looping and extension all start from that frame.
Practical answers for planning work around the model.
Kuaishou is the company; Kling AI is the name of its video generation platform and of the model line. Kling 4.0 is the fourth generation of that line and comes in two editions, the full model and Kling 4.0 Flash.
No. The 30 seconds are generated natively in one pass, so there are no joins to hide. Kling 3.0 produced at most 15 seconds per generation. When a scene needs more time, the finished clip can be extended, up to a combined 2 minutes.
No. Kling 4.0 also works from text alone or from a single image. Keyframes are optional and become useful when particular pictures have to appear at particular points; up to 10 can be supplied.
The limit applies to the total across all types. Within it, a generation can include up to 10 images, up to 5 videos and up to 7 subjects, in any combination that does not exceed 15.
A 10-bit signal records 1,024 levels per color channel where 8-bit records 256. Skies, shadows and skin therefore show smoother transitions, and the footage tolerates stronger color correction. On Kling 4.0 this HDR mode is available for both 1080p and 4K output.
Yes. Audio is produced together with the video as a two-channel stereo track. Speech is not limited to one language: English, Chinese, Spanish, Japanese and Korean are among those supported, along with regional accents and dialects, and lip movement tracks the words more accurately than in earlier Kling versions.
Widescreen, vertical and square output are all covered, and a 21:9 ultra-wide frame joins them for compositions in the proportion of cinema screens.
Yes, at 8,000 tokens. In practice that limit is rarely the constraint; clarity is. Write events in the order they occur and keep one instruction per sentence.
Flash is built for pace and economy rather than maximum fidelity. It returns results sooner and at a lower cost per clip, which makes it the practical option for trying ideas and for publishing in volume. Finished pieces that depend on realism and control belong with the full model.
Kling 3 belongs to the same family and generates clips of 3 to 15 seconds with native audio in Standard or Pro quality. Seedance 2.5 matches the 30-second length and the 4K ceiling and takes up to 50 multimodal reference materials. Kling o3 covers first and last frame control and video editing.
Video models here are paid for with credits. For Kling 3, for example, the amount depends on duration, quality mode and whether audio is enabled. Plans and credit packs are listed on the pricing page.
Yes. A keyframe is an ordinary still image, so it can be produced in an image model before the video stage. Generating all stills in one session, with the same style instructions, helps to keep them consistent with one another.