The world's first unified multimodal video model. Input anything. Understand everything. Generate any vision. A brand-new creative engine for creators to unlock endless possibilities.
Powered by the Multi-modal Visual Language (MVL) framework
The world's first unified multimodal video model integrates diverse video tasks into a single architecture — reference-based generation, text-to-video, keyframe interpolation (start/end frame), video inpainting, transformation, stylization, and video extension. Execute an end-to-end creative pipeline, from ideation to modification, all in one place.
Deep semantic understanding lets everything — images, videos, elements, texts — become part of your input. The model goes beyond modality limitations, integrating different perspectives to return outputs with pixel-perfect precision. Turn tedious post-production into simple conversations with prompts like "remove bystanders" or "change daytime to dusk."
Enhanced understanding of image and video inputs, with support for building elements from multiple angles. Like a human director, Kling O1 remembers your characters, props, and scenes to maintain consistency, accuracy, and continuity regardless of camera movement or scene development. Powerful multi-subject fusion delivers industrial-grade consistency for every character across every shot.
Not limited to single tasks — combine different tasks in one prompt. "Add a subject while modifying the background" or "change the style while using elements." Incorporate multiple creative ideas at once, exploring infinite creative possibilities with compound variations in a single pass.
Every shot needs its own duration for better story pacing. Kling O1 supports generations anywhere between 3-10 seconds, giving you complete control over how your story unfolds. Whether it's a fast-paced, impactful scene or a sustained narrative arc, you decide the rhythm of the shots.
Kling O1 is the world's first unified multimodal video model, developed by Kuaishou Technology. Built on the Multi-modal Visual Language (MVL) framework, Kling O1 integrates text, video, image, and subject inputs into a single, all-encompassing engine. This groundbreaking approach definitively resolves the "consistency challenge" in AI video generation, providing a deeply integrated, one-stop solution for film, television, social media, advertising, and e-commerce.
As the pioneer of unified multimodal video models, Kling O1 transcends the boundaries of traditional single-task video generation models by fusing a comprehensive spectrum of capabilities — including reference-based video generation, text-to-video generation, start and end frame generation, video in-painting, video modification and transformation, style re-rendering, and shot extension — into one versatile engine. This eliminates the need for creators to toggle between disparate models and tools; the entire creative lifecycle, from inception to refinement, is now a seamless, single-stream workflow.
Leveraging deep semantic reasoning, Kling O1 interprets all user inputs — whether images, video clips, specific subjects, or text — as executable prompts. By removing modality constraints, Kling O1 achieves a holistic understanding of elements from multiple perspectives, generating output with pixel-perfect precision. With its user-friendly multimodal prompt input interface, Kling O1 transforms complex post-production editing into a simple, conversational experience.
Upload 1-7 reference images or elements. Combine characters, items, outfits, scenes, and more, then use text prompts to define their interactions and bring static elements to life with precision and consistency. Prompt structure: [Element description] + [Interactions] + [Environment] + [Visual directions].
Comprehensive video editing capabilities: add or remove content, change angles or composition, modify subjects and backgrounds, restyle videos, recolor elements, change weather and environment, and green screen keying. Supported styles include American cartoon, Japanese anime, cyberpunk, pixel art, ink wash, watercolor, and clay.
Upload a 3-10s video as reference to generate previous or next shots within the same context. Reference video actions or camera movements to create completely new scenes with consistent motion and cinematography — generate next/previous shots, reference camera movements, or reference character actions.
Specify start and end frames to control the entire video from beginning to end. Describe scene transitions, camera movement, or character actions to achieve precise narrative control and cinematic storytelling, with full control over scene transitions, camera movements, and character actions.
Generate videos from pure text descriptions. Use natural language to describe your vision, and Kling O1 brings it to life with deep semantic understanding and pixel-perfect precision — natural language descriptions transform into cinematic videos.
Add flames to elements, freeze environments, apply facial textures or red-eye effects. Reimagine and redraw subjects to achieve more engaging visual effects with simple text commands — fire effects, freeze frames, facial effects, and subject reimagination.
World's first unified multimodal video model integrating generation and editing into one engine. No more switching between tools — complete your entire creative workflow in one place.
Transform complex post-production into simple conversations. No manual masking or keyframing needed — just describe what you want with natural language prompts.
"Director-like memory" maintains character, prop, and scene consistency across all shots. Multi-subject fusion ensures every element stays true to its identity, no matter how the scene evolves.
Generate videos between 3-10 seconds with complete control over pacing. Whether it's a quick impact or a sustained narrative arc, you decide the rhythm of your story.
| Duration | 3–10 seconds |
|---|---|
| Quality | High Definition |
| User Control | Full Duration Control |
| Images/Elements | 1–7 references |
|---|---|
| Video Reference | 3–10 seconds |
| Text Prompts | Natural Language |
| Framework | MVL |
|---|---|
| Architecture | Multimodal Transformer |
| Precision | Pixel-level |
| Generation | Text-to-Video, Image-to-Video, Reference-based Generation, Keyframe Interpolation |
|---|---|
| Editing | Content Addition/Removal, Subject Modification, Background Replacement, Localized Editing |
| Stylization | Video Restyle, Color Grading, Weather/Environment Changes, Green Screen Keying |
| Reference | Multi-element Fusion (1-7 inputs), Video Reference, Camera Movement Reference, Action Reference |
| Control | Frames Control, Duration Control (3-10s), Angle/Composition Changes, Creative Effects |
Lock in characters and props for each project with the Element Library. Generate multiple scenes with exceptional consistency and continuity, maintaining strict character, costume, and prop continuity across every shot to effortlessly create coherent cinematic sequences.
Mitigate the high costs and logistical friction of traditional offline advertising shoots. Upload product, model, and background images with simple prompts to rapidly generate multiple high-impact product showcase ads, significantly cutting production costs.
Create a 24/7 virtual runway. Upload model and clothing images with simple prompts to produce high-quality video lookbooks at scale. Flawlessly render fabric textures and details, solving the hassles of scheduling models and outfit changes.
Forget about tracking and masking. Post-production becomes as simple as having a conversation — input natural language like "remove the bystanders in the background" or "make the sky blue," and the model automatically completes pixel-level intelligent repair and reconstruction.
Rapidly create and edit engaging social media videos. Transform existing content with style changes, add trending effects, or generate fresh content from scratch — perfect for creators who need quick turnaround with professional quality.
While Kling O1 excels at multi-subject fusion, extremely complex interactions with many characters performing intricate coordinated actions may require multiple iterations to achieve the desired result. Breaking down complex scenes into simpler components can improve outcomes.
Certain extremely fine details such as intricate text, small logos, or highly detailed textures may not always render with perfect accuracy. For critical applications requiring pixel-perfect detail, manual review and potential touch-ups may be necessary.
Current generation is limited to 3-10 seconds per output. For longer videos, you'll need to generate multiple segments and potentially stitch them together. The model's consistency features help maintain continuity across segments when using reference videos.
While character and object consistency is excellent, maintaining exact stylistic elements (lighting, color grading, artistic style) across multiple independently generated shots may require careful prompt engineering and potentially using reference videos or images to guide the aesthetic.
Experience the world's first unified multimodal video model. Input anything, understand everything, generate any vision. Powered by Kuaishou Technology and the MVL framework.
Try Kling O1 now