All articles
Wan 3.0 in Gensta: How to Create 30-Second AI Video with Audio

Wan 3.0 in Gensta: How to Create 30-Second AI Video with Audio

Wan 3.0 is available in Gensta for text-to-video, first- or first-and-last-frame animation, and multimodal reference workflows. This guide covers controls, limits, pricing, and prompt patterns.

Wan 3.0 is available in Gensta

Wan 3.0 is available in Gensta's public generator. You can select the model and run text-to-video, image-to-video (first frame or first-and-last frame), and multimodal reference-to-video jobs.

In the current Gensta integration: 480p, 720p, and 1080p; durations of 5, 10, 15, 20, or 30 seconds; synchronized audio (can be turned off with no price change); and an optional deep-thinking mode for complex prompts (same price, slower). One request accepts up to 10 images, 5 videos, and 5 audios, 20 files total. Video references are subject to an input-video duration limit; together with output duration, follow the limit shown in the form (Alibaba's official API: input-video + output ≤ 30 seconds).

Alibaba Cloud presents Wan 3.0 as one model for text, frame control, and multimodal inputs. Documents, web pages, dedicated editing, and extension are not currently exposed as separate public Wan 3.0 modes in Gensta.

Try Wan 3.0 in Gensta

What Alibaba officially documents for Wan 3.0

Wan 3.0 can begin from three broad kinds of input. Text-to-video builds a scene from a director-style prompt. Image-to-video can treat one picture as the first frame, or use two pictures to constrain the first and last frames. Reference-based generation uses images, clips, and audio as ingredients for identity, products, locations, motion, performance, or sound instead of forcing any one asset to be the literal opening frame.

Without a video reference, the official API supports 2–30 seconds of output at 480P, 720P, or 1080P and 30 fps. With video references, combined input-video duration plus output duration must be no more than 30 seconds. Alibaba describes joint audio-visual generation that can include dialogue, background music, and sound effects. Audio is optional at API level, and the vendor documentation says that switching it off does not change Alibaba's per-second charge.

A reference request may contain up to 10 images, five video clips, and five audio clips. Reference videos may total no more than 15 seconds, reference audio has the same combined-duration ceiling, and input-video total plus output duration must not exceed 30 seconds. Alibaba's broader workflow also accepts a document or a public web link and documents editing and extension. Those vendor features must not be represented as Gensta features unless Gensta exposes and verifies them separately.

Modes and limits available in Gensta

As of 2026-08-28, Gensta exposes three Wan 3.0 modes: text-to-video, one- or two-frame image-to-video, and multimodal reference-to-video. The form offers 480p / 720p / 1080p, 5–30 second durations, audio, and deep thinking. One request accepts up to 10 images, 5 videos, and 5 audios (20 files total); current configuration caps combined input video and audio duration at 15 seconds each.

Controls can change, so review the final form and quote before generating.

Choosing text, frame control, or multimodal references

Use text-to-video when the concept, action, and camera matter more than an exact pre-existing subject. A useful brief names the aspect ratio, duration, setting, principal subject, sequence of events, camera behavior, and sound plan.

Use a first frame when the composition is already settled in a product photograph, illustration, or portrait for which you have the required rights and lawful basis; obtain consent when applicable law and the service rules require it. The prompt should spend less space restating visible details and more space on what begins to move, how the light changes, where the camera travels, and how the shot resolves.

Use first-and-last-frame control for a deliberate transition. The two pictures establish endpoints, not the path between them. Describe that path explicitly and avoid endpoints whose geometry has no plausible transformation.

Use reference mode when different assets have different jobs. Alibaba's API uses Image 1, Video 1, and Audio 1, while the current Gensta UI inserts @Image1, @Video1, and @Audio1. Publish exact Gensta syntax only after a day-zero provider request verifies what is passed successfully. Until then, assign roles in plain language by upload order and avoid conflicting instructions for the same feature.

Read more

Five prompt recipes for Wan 3.0

Below are five practical prompt recipes for Wan 3.0 modes. Treat them as starting templates, not a guarantee of a perfect first take: for complex scenes, try a short 5–10 second pass first, then extend.

1. Thirty-second vertical product story — text-to-video 30-second vertical 9:16 product film, grounded studio realism. 0–6s: a sealed matte-white box stands on a pale stone table; slow push-in, soft room tone. 6–13s: the lid opens and a dark-green travel bottle rises into a narrow beam of warm light; one clean mechanical click. 13–21s: three match cuts show the same bottle in a train compartment, on a mountain path and beside a laptop; preserve its exact shape, cap and label placement. 21–27s: return to the table as a hand picks up the bottle; natural grip and believable shadows. 27–30s: stable centered hero shot with two seconds of stillness. No on-screen text, no extra products, no changing logo, no watermark. Audio: restrained electronic pulse, room tone and precise object sounds, no voiceover.

2. A single-character night narrative — text-to-video 30-second cinematic 16:9 short, natural motion, coherent night lighting. 0–8s: a bicycle courier waits under a closed cinema marquee in light rain; wide shot slowly moves to medium. 8–17s: a paper ticket slides from beneath the door; the courier notices it, picks it up and hears a projector start inside. 17–25s: the marquee lights turn on one row at a time while the camera makes a slow half-circle; keep the courier's jacket, bicycle and face consistent. 25–30s: the cinema door opens into warm light, but the camera stays outside as the courier steps in. Audio: rain, distant traffic, paper movement, projector hum, no dialogue.

3. Animate a product photograph — first-frame image-to-video Use the uploaded image as the exact first frame. Keep the sneaker's proportions, materials, stitching, sole pattern and color unchanged. Over 10 seconds, the camera performs a slow 25-degree orbit while a narrow band of light travels from heel to toe. Fine dust lifts from the surface and settles naturally; the sneaker itself does not deform or rotate. End on a three-quarter close-up with the logo area fully visible. Neutral studio sound, a soft fabric movement and one subtle bass hit; no speech, no text, no duplicated shoe.

4. Connect two controlled frames — first-and-last-frame image-to-video Use the first uploaded image as the exact opening frame and the second uploaded image as the exact final frame. Create a physically plausible 15-second transition between them. 0–5s: the empty daylight café remains still as the camera begins a slow forward move. 5–11s: evening arrives through continuous changes in outside light; staff enter naturally and set the tables without popping into existence. 11–15s: guests settle into the positions shown in the second uploaded image and the camera reaches its final mark. Preserve architecture, table layout and lens perspective. Audio evolves from quiet daytime room tone to a warm evening crowd; no abrupt cuts.

5. Separate references for character, product, motion, and music — reference-to-video Create a 20-second 9:16 social video. Use the first reference image for the actor's face, hairstyle and jacket only; do not copy its background. Use the second reference image for the exact bottle shape, cap and label colors. Use the first reference video for the pace and direction of the handheld camera, not the actor's identity. Use the first audio reference for rhythm and instrumental mood; do not copy any speech. For a 20-second result, all video references together must be no longer than 10 seconds so input-video duration plus output stays within 30 seconds. 0–7s: the actor enters a bright corner shop and takes the bottle from a refrigerator. 7–15s: follow at shoulder height as the actor walks outside and opens it; maintain face and product identity. 15–20s: one sip, a brief natural reaction, then a stable product close-up. Preserve realistic hands, bottle geometry and ambient light. No subtitles, no extra labels, no duplicated objects.

Wan 3.0 vs Seedance 2.5: a capability comparison, not a verdict

In Gensta's current configuration, both model paths are designed for text-to-video, one- or two-frame image-to-video, multimodal references, output up to 30 seconds, and audio generation. For Wan 3.0, the 30-second maximum is unconditional only without input video; a video-reference request must satisfy “input-video total + output ≤ 30 seconds”. The most important difference today is product status: Seedance 2.5 is publicly selectable in Gensta, while Wan 3.0 is hidden and has not completed a public production smoke test.

Gensta integration settings side by side: • Public status: Wan 3.0 — preparing, not user-selectable; Seedance 2.5 — available. • Modes: Wan 3.0 — t2v, i2v, ref2v planned; Seedance 2.5 — t2v, i2v, ref2v public. • Output duration: Wan 3.0 — 5, 10, 15, 20, 30 seconds planned, and with video references input-video total plus output must stay within 30 seconds; Seedance 2.5 — 5, 8, 10, 15, 20, 30 seconds. • Resolution choices: Wan 3.0 — 480p, 720p, 1080p planned; Seedance 2.5 — 480p, 720p, 1080p, and 4K in the Gensta integration. • Audio: Wan 3.0 — planned toggle; Seedance 2.5 — public toggle. • Reference limits: Wan 3.0 — planned 10 images + 5 videos + 5 audio, 20 files total; Seedance 2.5 — 9 images + 3 videos + 3 audio, 15 files total. • Prompt planning: Wan 3.0 — planned toggle; no matching toggle in the public sd25_ws config. • Comparable Gensta quality data: none for either route.

This table compares Gensta integration settings as of August 27, 2026, not each vendor's full model limits. Wan 3.0 is not public in Gensta, and identical-input quality has not been tested.

Gensta has not yet measured which model better preserves a face, product geometry, dialogue rhythm, or physical motion from identical inputs. A defensible answer requires a controlled comparison that keeps first attempts, settings, prices, timings, and failures. Until that exists, use a model that is currently available and reassess Wan 3.0 after its public launch.

Read the AI video guide

Limits and practical tips

Longer clips and large reference sets increase control and the chance of conflicts. Watch fine text, exact logos, complex hands, overlapping dialogue, abrupt geometry changes, and incompatible visual references especially carefully.

A practical order: start with a short t2v or i2v pass at 720p / 5–10 seconds, then first-and-last frame, then references. For ref2v, track combined input-video duration plus output duration up front. If you need exact brand copy or legally binding text, add it in post rather than relying on generation alone.

How to start with Wan 3.0 in Gensta

Open the video generator, select Wan 3.0, and start with a short 5–10 second 720p clip to check composition and audio. For image-to-video, describe motion and change rather than what is already visible. For references, name the files in the prompt and watch combined input-video duration.

A practical baseline for image-to-video at 720p and 5 seconds is about 200 Gensta credits under the current tariff. Audio and deep thinking do not change that price on this model. Always trust the live quote before submitting.

Try Wan 3.0 in Gensta

Wan 3.0 FAQ

Is Wan 3.0 available in Gensta now?

Yes. Wan 3.0 is available in Gensta's public model selector. Text-to-video, image-to-video, and reference modes are open. Check the live form for current controls and pricing.

Which Wan 3.0 modes are available in Gensta?

As of 2026-08-28, Gensta exposes text-to-video, image-to-video (one or two frames), and multimodal references. Documents, web links, dedicated editing, and extension are not yet separate public Wan 3.0 modes in Gensta.

Can Wan 3.0 really generate 30-second videos?

Yes. In Gensta, Wan 3.0 offers up to 30 seconds in all public modes. Without video references, that is a normal output-duration limit. With video references, combined input-video duration plus output duration must stay within the form's limit (Alibaba's official API: ≤ 30 seconds).

Which output resolutions does Wan 3.0 support?

The official API and Gensta's planned route list 480p, 720p, and 1080p. Gensta should not promise 2K or 4K without a documented and tested route.

Can a Wan 3.0 request combine images, video, and audio?

Alibaba's reference workflow accepts combinations of those media types: up to 10 images, five videos, and five audio clips. Combined reference video is capped at 15 seconds, combined reference audio has the same ceiling, and input-video total plus output must stay within 30 seconds. Gensta's planned configuration mirrors the file counts, but its duration constraints must be verified or implemented before public launch.

Can I send Wan 3.0 a document or a web page?

Alibaba's vendor API supports a file or public web link. Gensta's current integration accepts image, video, and audio references only, so the vendor feature must not be advertised as a Gensta input.

How much will Wan 3.0 cost in Gensta?

Public pricing has not been verified. The current code calculates 200 credits for its default 720p, five-second image-to-video configuration, but that number is not final until the public quote, actual charge, and production configuration agree. After launch, rely on the price shown before generation.

Is Wan 3.0 better than Seedance 2.5?

Gensta does not yet have a controlled comparison, so there is no supported winner. Both routes target long video, audio, and multimodal references, but their availability, resolution choices, and integration limits differ. Output quality requires identical tasks after Wan 3.0 launches.

Sources and methodology

This guide uses Alibaba Cloud's official documentation and Gensta's product configuration as of 2026-08-28. A controlled Seedance 2.5 benchmark has not been published; prompts here are templates, not output guarantees. Review live modes, limits, and pricing before generating.

All articles