Imagera AI - AI content creation platform for generating images, cloning voices, creating avatars, and enhancing videos. Privacy Policy | Terms

Blog Post

Is Wan 3.0 Real? What Alibaba Has Actually Shipped in 2026

As of 29 July 2026 there is no Wan 3.0. Here is every shipped Wan version, what has open weights, real duration limits, and where to run them.

By Imagera Team14 min readJuly 29, 2026Updated: August 2, 2026
Share:
Cinematic wide shot of a darkened post-production suite with a single glowing timeline on a large reference monitor

TL;DR

  • As of 29 July 2026 there is no Wan 3.0 — no GitHub repo, no HuggingFace weights, no Alibaba Cloud model id starting with wan3.
  • The newest Wan you can call is wan2.7: 2 to 15 seconds, 720P and 1080P, API-only on Model Studio, Together AI, Replicate and WaveSpeed.
  • The newest Wan with downloadable weights is Wan 2.2, Apache 2.0, documented at 5 seconds and 480P/720P.
  • Pages ranking for "Wan 3.0" contradict each other on parameters (60B vs 27B MoE vs 14B), duration (60s vs 30s vs 10s) and resolution (4K vs 1080p) — they cannot all be right.
  • The "27B total / 14B active" figure attributed to Wan 3.0 is Wan 2.2's own published MoE spec, quoted verbatim in Alibaba's repository.
  • No documented model generates 30 seconds in one pass. Veo 3.1 caps at 8s native (148s via 20 extensions, 720p only); Sora 2 at 20s native (120s via six extensions).
  • A 30-second cut today is six five-second shots planned as a scene — that workflow already runs in Imagera's video studios today.

Try it yourself — no setup

Turn prompts and images into cinematic AI video.

As of 29 July 2026, Wan 3.0 does not exist. Alibaba publishes no model, repository or API id under that name. The newest Wan you can call today is wan2.7, API-only. The newest Wan with open weights is Wan 2.2.

Now the interesting part: why you searched for it.

The promise attached to the words "Wan 3.0" is the most valuable thing anyone in video could ship right now — thirty seconds, generated in a single pass, at 4K. Not eight seconds. Not six five-second clips stitched together with a face that drifts at every seam. One continuous take long enough to hold a scene: an ad that plays start to finish, a product demo with a beginning and an end, a story beat that lands without a cut.

That is a real step change in format economics, and it is worth wanting. It is also, today, a claim rather than a product. Below are the receipts.

Cinematic film slate resting on a lit editing desk beside a rack of hard drives, dust in the light beam

1.Is Wan 3.0 real, or is it a rumour?

Wan 3.0 is not real as of 29 July 2026. There is no Wan 3.0 repository, no Wan 3.0 weights, and no Wan 3.0 API id on Alibaba Cloud. Every page describing it is describing something Alibaba has not published.

Three checks, all first-party, all repeatable in about ninety seconds:

The code. The Wan-Video GitHub organization shows "Showing 5 of 5 repositories": Wan2.2, Wan2.1, Wan-Dancer, Wan-skills, and a fork of diffusers. The highest-numbered repo in the org is Wan2.2. There is no Wan2.5, no Wan2.6, no Wan2.7 — and no Wan3.

The weights. The Wan-AI HuggingFace organization lists 24 models. The newest numbered line published there is Wan2.2 — Wan2.2-Animate-14B, Wan2.2-S2V-14B, Wan2.2-T2V-A14B, Wan2.2-TI2V-5B. Searching HuggingFace for "wan3" returns only unrelated third-party uploads (Mohammedf3/Wan3, Danielbsittler/wan3, Himmdy/Wan3D, intchous/wan3). None are from Wan-AI.

The API. On Alibaba Cloud Model Studio's model list (page last updated 15 July 2026), no model id begins with wan3. The newest Wan id listed is wan2.7-image-pro.

Absence of a repo is not proof a model will never exist; Alibaba could announce Wan 3.0 tomorrow. But today, anyone publishing a parameter count, duration ceiling or price for Wan 3.0 has no first-party source behind it.

2.What is the newest Wan model you can actually use today?

The newest usable Wan is wan2.7, served through Alibaba Cloud Model Studio and third-party hosts. Its documented ceiling is 15 seconds at up to 1080P. The newest Wan you can download and run yourself is Wan 2.2, released under Apache 2.0.

That split — API-only versus downloadable — is where the confusion starts. Versions 2.5, 2.6 and 2.7 exist as API model ids only; they do not appear in the GitHub org or the HuggingFace org. Version 2.2 and below exist as both.

The whole lineup, as documented on 29 July 2026:

VersionOpen weights?LicenceDocumented durationWhere to run it
Wan 2.1Yes — Wan-Video/Wan2.1Not quoted on pages checked-plus and -turbo: "Fixed at 5 seconds and cannot be changed."Self-hosted, Model Studio, Replicate
Wan 2.2Yes — 4 models on HuggingFace"Licensed under the Apache 2.0 License."5s at 480P and 720P (T2V-A14B)Self-hosted, Model Studio, Replicate, WaveSpeed
Wan 2.5Not in GitHub or HF orgn/awan2.5-t2v-preview: "Valid values are 5 and 10."Model Studio, Replicate, WaveSpeed
Wan 2.6Not in GitHub or HF orgn/a2 to 15 seconds, default 5Model Studio, WaveSpeed (absent from Replicate)
Wan 2.7Not in GitHub or HF orgn/at2v and i2v: 2–15s. videoedit: 2–10sModel Studio, Together AI, Replicate, WaveSpeed
Wan-Dancer-14BYes — HuggingFaceapache-2.0"Minute-scale" music-to-dance (card claim)Self-hosted
Wan 3.0No repo, no weights, no API idNowhere

Two details worth pulling out. First, wan2.7-videoedit accepts 2 to 10 seconds — a narrower range than text-to-video, per the video editing API reference. Editing is harder than generating, and Alibaba's own limits say so.

Second, Wan 2.7 image-to-video supports three modes: first-frame-to-video, first-and-last-frame-to-video, and video continuation. Continuation is how you get past 15 seconds. It is not native single-pass duration, and that difference matters enormously.

Two monitors side by side showing an identical film frame with subtly different colour grades, studio lighting

You don't need Wan 3.0 to ship this week

The thing people are actually waiting for is a 30-second cut that holds — same face, same wardrobe, same light, all the way through. That does not require an unreleased model. It requires a shot plan and an engine that keeps your look consistent across takes.

Imagera's cinematic video studio runs that workflow now, priced in credits instead of per second of output, so an experiment that misses costs you a retry rather than a bill. New engines are added as they become available — and the accounts already producing are the ones ready the day a longer ceiling lands.

See how credits work · Open the video studio

3.Where can I run Wan 2.7 outside Alibaba Cloud?

Wan 2.7 is served by at least three third-party hosts. Together AI runs it as Wan-AI/wan2.7-t2v, with "video outputs ranging from 2 to 15 seconds" and "720P and 1080P generation".

Replicate's wan-video collection carries Wan 2.7 entries spanning text-to-video, image-to-video, reference-to-video, video editing and image generation, alongside 2.5, 2.2 and 2.1. Notably, no Wan 2.6 model appears in Replicate's collection at all, while WaveSpeed AI lists collections for 2.1, 2.2, 2.5, 2.6 and 2.7.

That inconsistency is a useful signal in itself: providers pick up Wan versions unevenly. If a page tells you a version is "available on all major providers", the provider pages are one click away.

4.What do the "Wan 3.0" pages actually claim?

They claim mutually exclusive things. Pages ranking for "Wan 3.0" variously describe a 60B dense model, a 27B mixture-of-experts model with 14B active, and a 14B dense model. Some claim 60-second output, some 30, some 10. Some say 4K, some say 1080p. They cannot all be right.

This is the clearest tell available to a reader with no inside knowledge. You do not have to evaluate anyone's credibility — only notice that the specs disagree with each other, and that none appear on a page Alibaba controls.

Circulating claim about Wan 3.0What the first-party record shows (29 July 2026)
"60B parameters, dense"No parameter count for any Wan 3.0 model exists on GitHub, HuggingFace or Model Studio. There is no Wan 3.0 model card.
"27B MoE, 14B active per step"This is the published spec of Wan 2.2, verbatim, from its own README.
"14B dense"Wan 2.2's expert models are "about 14B parameters" each — that figure belongs to 2.2's architecture.
"60 seconds in one pass"The longest documented duration for any Wan id is 15 seconds (wan2.7-t2v, wan2.7-i2v).
"30 seconds, single pass, 4K"Wan 2.7's documented tiers are 720P and 1080P, including 1920*1080. No 4K tier is documented for any Wan id.
"Open weights on release"Wan 2.5, 2.6 and 2.7 all shipped as API ids with no repo in the GitHub org and no model in the HF org.
A specific release dateAlibaba has published no announcement, changelog entry or dated post for Wan 3.0.

To be explicit: that table compares claims against documentation. It is not an accusation against any publisher. Speculating about unreleased models is legitimate when it is labelled as speculation. The issue is that these specs circulate unlabelled, in contradictory forms, on a verification query — exactly where a reader most needs the label.

4.1Why does "27B MoE, 14B active" sound familiar?

Because it is Wan 2.2's real, published architecture. The Wan2.2 README states it plainly: "Each expert model has about 14B parameters, resulting in a total of 27B parameters but only 14B active parameters per step."

That is a shipped 2025 model with downloadable weights, described in Alibaba's own repository. When the same numbers reappear attributed to an unreleased Wan 3.0, the most economical explanation is that a published spec got re-labelled upstream and copied forward. Stated neutrally: the figure is real, and it belongs to 2.2.

Two more Wan 2.2 facts from that README that rarely survive into "Wan 3.0" coverage:

  • The high-compression Wan2.2-VAE "achieves a T×H×W compression ratio of 4×16×16, increasing the overall compression rate to 64 while maintaining high-quality video reconstruction." Note the axis order — 4×16×16 in time-height-width, not 16×16×4.
  • TI2V-5B "supports both text-to-video and image-to-video generation at 720P resolution with 24fps and can also run on consumer-grade graphics cards like 4090", and "can generate a 5-second 720P video in under 9 minutes on a single consumer-grade GPU."

The most recent dated entry in that repo's Latest News is 13 November 2025 (Wan2.2-Animate-14B integrated into Diffusers), preceded by 19 September and 26 August 2025. The open-weight flagship line has been quiet; the org has not. Wan-Dancer-14B landed on 13 July 2026 under apache-2.0, "a hierarchical framework for minute-scale coherent music-to-dance generation".

Close-up of a dancer mid-motion under hard key light, motion blur trailing from an outstretched arm

5.Can any model generate 30 seconds in a single pass right now?

No model with public documentation generates 30 seconds natively. The verified ceilings are 15 seconds (Wan 2.7), 20 seconds (Sora 2), 10 seconds (Hailuo 2.3 at 768P) and 8 seconds (Veo 3.1). Longer outputs exist, but every one is produced by extension, not by a single pass.

ModelDocumented native maximumExtension mechanicsResolution notes
Wan 2.7 (API)2–15s; videoedit 2–10sVideo continuation mode in i2v720P and 1080P
Wan 2.2 (open weights)5sNone documented480P and 720P
Google Veo 3.14, 6 or 8s — must be 8 "when using extension, reference images or with 1080p and 4k resolutions"+7s per extension, up to 20 times, "up to 148 seconds" combinedExtension is "720p" only
OpenAI Sora 2 / 2-pro"16- and 20-second generations""up to 20 seconds" per extension, six times, "maximum total length of 120 seconds"1080p exports need sora-2-pro
MiniMax Hailuo 2.310s at 768P, 6s at 1080P; default 6sNot documentedValues depend on model and resolution
Lightricks LTX-2 / 2.3No maximum stated in secondsNot documentedOnly rule on either card: "Frame count must be divisible by 8 + 1"
Tencent HunyuanVideo--video-length 129 frames by defaultNot documented

Sources: Google's Veo docs, OpenAI's video generation guide, and the vendor cards above. Kling, Runway and Seedance are deliberately absent: their first-party docs did not return retrievable spec pages during this check, and second-hand blog figures are not evidence.

Read the Veo row again, because it carries the whole lesson. Veo 3.1 will hand you 148 seconds — in 21 chunks, and the moment you use extension the resolution is capped at 720p. Sora will hand you 120 seconds, in seven chunks. Longer output is real. Longer single-pass output is not what those numbers describe.

One more current-state note: Google's Gemini API video overview (updated 30 June 2026) now tells developers to "Use Gemini Omni Flash as your default model for video generation" rather than naming Veo.

6.Why is long single-pass video so hard to build?

Four documented reasons, all structural rather than a matter of effort: quadratic attention cost, long-context memory, error accumulation across frames, and the fidelity penalty of temporal compression. Every long-video system trades one against another.

Attention scales quadratically. Research on long-context video states it directly: "scaling diffusion transformers to generate long-context videos is fundamentally limited by the quadratic cost of self-attention" (arXiv:2508.21058). Double the clip, quadruple the attention bill.

It is a memory problem too. The same work frames it as long-context retrieval: "models must retain and retrieve salient events across a long range without collapsing or drifting." The jacket has to still be the same jacket at second 26.

Errors compound. Later work lists the failure modes as "error accumulation, attribute drift, and the limited availability of long video data" (arXiv:2606.22370). Every frame conditions the next, so small mistakes get inherited and amplified.

Compressing time costs fidelity. The obvious fix is squeezing more seconds into fewer tokens, and that has a documented ceiling: pushing "beyond 4x without increasing channel capacity poses significant challenges", and low-compression encoding of subsampled video "surpasses that of high-compression encoders applied to original videos" (arXiv:2501.05442). Wan 2.2's VAE sits at exactly 4× on the temporal axis — right at the documented edge.

That is why "30 seconds, single pass, 4K" is such an attractive claim and such a hard build. When it lands, it will land with a paper.

Long strip of film frames laid out on a light table, later frames slightly warping in colour and shape

7.How do you ship a 30-second scene this week, without Wan 3.0?

You sequence it. A 30-second cut is not one 30-second generation — it is six shots of five seconds, planned as a scene, generated against locked references, and cut together. That workflow exists now, and it is how every polished 30-second AI spot you have watched was actually made.

The shot plan that works, and it is deliberately boring:

  1. Establishing (0–5s) — wide, sets place and light. Becomes your look reference for everything after.
  2. Subject entry (5–10s) — person, product or car enters frame. Same key light direction as shot 1.
  3. Detail insert (10–15s) — macro. Texture, label, hands, badge. Cheapest to get right, highest perceived production value.
  4. Motion beat (15–20s) — the only shot with real camera movement. Push in or track.
  5. Turn (20–25s) — the reveal, the reaction, the before and after.
  6. Payoff (25–30s) — hold the hero frame for the end card.

Six shots means six chances to reroll one shot instead of rerolling thirty seconds. Directors have cut scenes this way for a century for that exact reason.

On Imagera:

Start with shot 3, the detail insert. Smallest possible commitment, one generation, and it tells you within two minutes whether your look is right.

Storyboard cards pinned in a row above a monitor, each card a different framing of the same product scene

8.What happens the day Wan 3.0 actually ships?

New video engines are added to Imagera as they become available, and early access goes out through the studio itself. No date has been promised by anyone, including Alibaba — which is exactly why the useful move is to have your shot workflow already working before it lands, not after.

What is knowable: the last three Wan releases (2.5, 2.6, 2.7) all shipped as API endpoints first, and hosts picked them up within days. The integration path is short. Being already inside the studio, with a look you have dialled in and shot plans you have tested, is the difference between shipping on day one and starting on day one.

The cost of waiting runs one direction only. Six weeks producing 30-second cuts from sequenced shots gets you six weeks of published work and a house style. Six weeks refreshing a rumour page gets you a refreshed rumour page.

Imagera is priced in credits, not per second of output: 500 credits on Pro at $19.99, 1,500 on Business, 6,500 on Ultra — all currently shown at roughly half off. Monthly plan credits reset with each cycle, while add-on credit packs never expire, so capacity you top up for a launch is still there next quarter. Detail on the pricing page.

Editor's hands on a control surface, timeline glowing on the monitor above, warm rim light from a window

9.The status line, as of 29 July 2026

  • Wan 3.0: no repo, no weights, no API id, no announcement. Does not exist.
  • Newest Wan you can call: wan2.7 — 2 to 15 seconds, 720P and 1080P, API-only.
  • Newest Wan you can download: Wan 2.2, Apache 2.0, 5-second generations at 480P and 720P.
  • Newest Wan release of any kind: Wan-Dancer-14B, 13 July 2026, apache-2.0.
  • Longest documented single-pass output anywhere: 20 seconds (Sora 2). Everything longer is extension.

If that changes tomorrow, it changes on github.com/Wan-Video, huggingface.co/Wan-AI and Model Studio's model list before anywhere else. Those three pages take ninety seconds to check, and they outrank any blog — including this one.

Until then: the scene you were going to make with Wan 3.0 is six shots long, and you can start it today. Open the video studio.

Frequently Asked Questions

Is Wan 3.0 out yet?
No. As of 29 July 2026 Alibaba has published no Wan 3.0 model, repository, API id or announcement. The Wan-Video GitHub organization shows 5 of 5 repositories topping out at Wan2.2, the Wan-AI HuggingFace organization's 24 models top out at Wan2.2, and Model Studio's model list has no id beginning with wan3.
What is the latest Wan model right now?
The latest callable Wan is wan2.7, available as text-to-video, image-to-video, reference-to-video, video editing and image generation ids. Alibaba Cloud documents 2 to 15 seconds for wan2.7-t2v and wan2.7-i2v, with 720P and 1080P tiers including 1920*1080. wan2.7-videoedit is narrower at 2 to 10 seconds.
Is Wan 2.7 open source?
Wan 2.7 has no repository in the Wan-Video GitHub organization and no model in the Wan-AI HuggingFace organization. The same is true of Wan 2.5 and Wan 2.6. The most recent Wan versions with published open weights are Wan 2.2 (Apache 2.0) and Wan 2.1.
Can Wan generate a 30-second video?
Not in a single pass. The longest duration documented for any Wan model id is 15 seconds, on wan2.7-t2v and wan2.7-i2v. Wan 2.7 image-to-video does offer a video continuation mode, which chains generations together — that is extension, not native 30-second output.
Which AI video model has the longest single clip in 2026?
Of the models with retrievable first-party documentation, Sora 2 and Sora 2 Pro have the longest native generation at 20 seconds. Wan 2.7 reaches 15 seconds, MiniMax Hailuo 2.3 reaches 10 seconds at 768P, and Veo 3.1 caps at 8 seconds. Longer results everywhere come from extension.
Why do the Wan 3.0 specs I keep seeing disagree with each other?
Because none of them come from a page Alibaba controls. Circulating figures include 60B dense, 27B MoE with 14B active, and 14B dense — and 60s, 30s and 10s durations. The 27B/14B figure is Wan 2.2's real published architecture, which suggests a shipped spec got re-labelled and copied forward.
Is Wan 2.2 still worth running in 2026?
Yes, if you want weights you control. Wan 2.2 is Apache 2.0, its TI2V-5B model runs on consumer cards such as a 4090, and Alibaba's README states it can generate a 5-second 720P video in under 9 minutes on a single consumer-grade GPU.

Imagera Team

AI Content & Editorial Team

The Imagera AI editorial team brings together AI researchers, product specialists, and content strategists covering practical AI creation workflows.

Areas of Expertise:

AI Image GenerationAI Voice RecreationAI Avatar CreationContent Marketing

Put this guide to work

Turn prompts and images into cinematic AI video.

Turn prompts and images into cinematic AI video.