Wan 3.0 AI Video Generator

Create 2-30 second 1080p videos with first and last frames, multimodal references, and generated sound.

Output duration
2-30 sec
Maximum resolution
1080p
Reference media
Up to 20
Generated sound
Same pass

Choose the input that fits the shot

Start with a prompt, control the first and last frame, or combine existing media as references.

Start with the scene

Describe the subject, action, setting, camera and audio. Wan 3.0 builds the shot from text.

No media required

Made with Wan 3.0

Four portrait prompts built for social feeds, product stories, food detail and fashion motion.

Talking straight to camera

A direct-to-camera product clip with natural delivery and soft room light.

View prompt

A woman in her late twenties sits in a bright bedroom holding a mint-green serum bottle with a gold cap up beside her face, talking straight to camera about it, turning the bottle to show the label. Vertical framing, handheld phone camera, soft daylight through window blinds behind her. Natural skin texture, a few loose strands of hair, an unforced smile between sentences. She speaks in a warm conversational voice.

Hands, texture and packaging

An overhead product sequence that gives the hands, powder and packaging distinct actions.

View prompt

Two hands open a mint-green matcha latte box on a pale pink tabletop, lift the lid away, then dip fingers into a glass jar of matcha powder and let it fall back through in a fine stream. Vertical framing, overhead and slightly angled, soft diffused studio light. Powder clings to the fingertips, a little dusts the table, the jar glass is slightly smudged. Quiet room tone.

Steam, sizzle and close detail

A food close-up that coordinates steam, hand movement and kitchen ambience.

View prompt

A chef's hands finish a bowl of ramen on a small counter kitchen pass: broth steaming, chopsticks laying two slices of chashu, a ladle of chilli oil breaking the surface into red rings. Vertical framing, handheld 35mm, warm overhead tungsten against a cold window edge, shallow focus. Steam hazes the lens, a thumbprint smudges the bowl rim, the scallion scatter is uneven. Ambient kitchen room tone, the clink of ceramic on steel.

Movement, fabric and light

A vertical fashion shot with a single follow move and strong backlight.

View prompt

A woman in her late twenties walks toward camera down a narrow city street at golden hour, a long camel coat catching the wind, boots on wet cobblestone. Vertical framing, handheld 50mm follow, backlit with a low flare across the frame. Loose strands of hair cross her face, one boot toe is scuffed, the pavement reflects unevenly. Street ambience with distant traffic.

Prompting guide

Write it like a shot brief

Separate subject, action, setting, camera and audio instead of stacking vague style adjectives.

Subject
Name who or what must remain recognizable.
Action
Use concrete verbs and give each beat enough time.
Setting
Describe the environment, time, light and useful texture.
Camera
Choose one main move such as push-in, orbit, pan or locked-off.
Audio
Name dialogue, foreground events and ambient layers separately.

A prompt that stays directed

Medium shot of a cyclist stopping beneath a train overpass at dusk. Rain reflects warm practical lights on the asphalt. The camera makes one slow push-in from chest height. The cyclist looks toward camera and says, "The city is still awake." Audio: distant traffic, light rain, one bicycle freewheel click.

Price the final format

Wan 3.0 generation is 30% off for a limited time. These discounted prices use a 5-second generation.

Compare plans

480p

17 credits (normally 23)

Fast draft passes for testing motion and prompt structure.

720p

33 credits (normally 47)

A practical working tier for social and in-product video.

1080p

66 credits (normally 94)

The finishing tier for larger playback surfaces and final delivery.

The workspace shows the exact limited-time quote before submission. Longer clips multiply the per-second cost.

Where Wan 3.0 fits

Current controls available in Vuppo. Choose by input structure, duration and output needs.

CapabilityWan 3.0Seedance 2.5Kling 3.0Grok Imagine 1.5
InputsText, first and last frames, image, video and audio referencesText, frame roles, image, video and audio referencesText, first and last framesText, start image or image references
Output duration2-30 seconds4-30 seconds3-15 seconds1-15 seconds
Maximum resolution1080p1080p4K1080p; references up to 720p
Generated soundOptionalOptionalOptionalIncluded
Reference scale10 images, 5 videos, 5 audio30 images, 10 videos, 10 audioFirst and last framesUp to 7 images

Comparison reflects the controls currently exposed in Vuppo. Provider limits, availability and pricing can change.

Wan 3.0 FAQ

Fast dance motion generated with Wan 3.0

Create the next shot with Wan 3.0

Open the Vuppo video workspace with Wan 3.0 and the General workflow already selected.