Alibaba's newest video generation model, available on SharkFoto as a pure text-to-video tool. Pick any length from 1 to 15 seconds, any of six framings, and a 480P, 720P or 1080P render - all from a single written prompt.
Set any whole number of seconds from 1 to 15. Earlier Wan models on SharkFoto only offered fixed 5, 10 and 15 second steps, so a 7-second cut meant generating 10 and trimming. Here you ask for 7 and pay for 7.
16:9, 9:16, 1:1, 4:3 and 3:4 cover landscape, vertical, square and classic-camera shapes. Adaptive hands the framing decision to the model instead of pinning it to a fixed ratio, which suits prompts that describe a shot rather than a delivery format.
Wan 3.0 adds a 480P option below the usual HD tiers. Because credits scale with resolution, a 480P pass is the cheapest way to test whether a prompt reads the way you intended before committing to a 1080P render.
The prompt field accepts up to 2,500 characters - 500 more than Wan 2.6. That is enough to spell out subject, wardrobe, lighting, lens, camera move and pacing in one go instead of compressing the whole idea into a sentence.
Wan 3.0 is the current generation of Alibaba's Wan video model family, released publicly in August 2026. On SharkFoto it runs as a text-to-video model: you write a prompt, choose the shape, resolution and length of the clip, and the model renders the whole thing in one pass. No reference image, no storyboard, no editing timeline.
The practical difference from earlier Wan releases is control granularity. Duration is a continuous 1-15 second range rather than a set of preset buttons, the aspect ratio list grows from two options to six, and a 480P tier sits underneath 720P and 1080P so that iteration costs a fraction of delivery. The prompt window is larger as well, which matters more than it sounds: a video prompt has to carry subject, motion and camera direction all at once.
Wan 3.0 is built for single-prompt generation, so it plans the shot itself and holds subject appearance, spatial layout and motion consistent across the clip. That makes it a good fit for short social cuts, product and marketing b-roll, and quick visual concepts - anywhere the idea exists as a description rather than as footage you already have.
Duration is a slider across the full 1-15 second range, and every whole second is a valid request. A 2-second loop for a feed, an 8-second product beat, a 15-second full cut - each is asked for directly rather than trimmed out of a longer render.
Five fixed ratios cover the formats delivery actually requires: 16:9 for landscape and YouTube, 9:16 for Reels, Shorts and TikTok, 1:1 for feed posts, 4:3 and 3:4 for the more photographic shapes. Adaptive is the sixth option, for when the prompt should decide the frame.
480P, 720P and 1080P are separate renders, not an upscale chain. Credits scale with the tier, so the same prompt can be explored cheaply at 480P and then re-rendered once at 1080P when the wording is settled.
A long prompt window lets you direct rather than describe: subject and wardrobe, the environment and time of day, the lighting, the lens, how the camera moves, and what changes between the first and last second of the clip.
Wan 3.0 is designed to hold character detail, spatial relationships and motion coherent for the whole duration rather than drifting as the clip runs on, so faces, props and background geometry stay recognisable from the first frame to the last.
The model handles multi-shot structure natively, so a prompt that describes a small sequence does not have to be split into separate generations and cut together by hand afterwards.
| Resolution | 480P, 720P or 1080P (default 720P) |
|---|---|
| Clip Length | 1-15 seconds, any whole second (default 5s) |
| Aspect Ratios | Adaptive, 16:9, 9:16, 1:1, 4:3, 3:4 (default 16:9) |
| Generation | Single pass - no stitching or upscale step |
| Prompt | Required, up to 2,500 characters |
|---|---|
| Reference Image | Not accepted |
| Reference Video | Not accepted |
| Reference Audio | Not accepted |
| Documents | Not accepted |
| Tool | AI Video - Text to Video |
|---|---|
| Developer | Alibaba |
| Modes | Text to video only |
| Credits | Scale with both resolution and clip length |
| Other Modes | Use Wan 2.6 for image-to-video and video-to-video |
Generate 9:16 clips sized for Reels, Shorts and TikTok without cropping a landscape render. Short durations keep the cost of testing several hooks low, and the same prompt can be re-run at 1:1 or 16:9 for the other placements.
Fill the gaps in an edit - an establishing shot, a texture close-up, an ambient transition - without a shoot. Ask for the exact number of seconds the timeline needs rather than trimming a longer clip down.
Test whether an idea reads on screen before anyone books a camera. Iterate the wording at 480P where a pass is cheap, then render the version that works at 1080P.
A few seconds of motion at the top of a page or post does more than a static hero image, and 4:3 or 3:4 suits editorial layouts that a 16:9 clip would not sit in cleanly.
One written concept, rendered into landscape, square and vertical versions, so a campaign ships in every placement without re-framing work in an editor.
Both Wan models are available here. This is what actually differs between them in the generator, not on paper.
| Feature | Wan 2.6 | Wan 3.0NEW |
|---|---|---|
| Clip length | Fixed 5s, 10s or 15s | Any whole second, 1-15s |
| Resolution | 720p, 1080p | 480P, 720P, 1080P |
| Aspect ratios | 16:9, 9:16 | Adaptive + 16:9, 9:16, 1:1, 4:3, 3:4 |
| Prompt limit | 2,000 characters | 2,500 characters |
| Draft tier | None below 720p | 480P |
| Generation modes | Text, image and video to video | Text to video only |
| Credit model | Preset duration steps | Charged per second |
Wan 3.0 on SharkFoto accepts a written prompt and nothing else. There is no way to supply a reference image, a reference video, an audio track or a document. Work that has to start from an existing asset belongs on Wan 2.6 or another image-to-video model.
A single generation tops out at 15 seconds. Longer pieces have to be assembled from multiple clips in an editor, and continuity between separate generations is not guaranteed.
Credits scale sharply with resolution: a 1080P clip costs four times the same-length 480P clip, and 720P costs twice as much. Iterating at 1080P gets expensive quickly - settle the prompt at 480P first.
The model generates a new clip each time. It cannot extend an existing video, restyle footage, or edit a previous result - every run starts again from the prompt.
Describe the shot, set the length to the second, pick your framing, and let Wan 3.0 render it. Start at 480P to find the prompt, finish at 1080P.
Try Wan 3.0 now