Half the models called "open" in 2026 are a download. The other half are an API key, a rate limit and a bill. The confusion costs real money: teams budget for a self-hosted pipeline, staff it, and find out eight weeks in that the version they actually wanted was never published as weights.
This is the ledger, checked against first-party repositories, model cards and vendor API docs as of 29 July 2026. Where a page does not state something, this piece says so instead of filling the hole.
One framing note: the reason the question got loud is duration. Nobody shops for weights out of principle. They shop because a 30-second scene that holds is now the format that sells, and they need to know whether the thing that produces it is a download or a subscription.

1.Which AI video models can you actually download in 2026?
The genuinely downloadable inference weights verified here are the Wan 2.2 family (Apache 2.0), Wan 2.1, Wan-Dancer-14B (Apache 2.0), Lightricks LTX-2 and LTX-2.3, and Tencent's HunyuanVideo. Everything else in this article is an API endpoint.
The Wan-Video GitHub organization reads "Showing 5 of 5 repositories": Wan2.2, Wan2.1, Wan-Dancer, Wan-skills, and a fork of diffusers. That is the entire public code surface. The Wan-AI Hugging Face organization lists 24 models, and the newest numbered line there is Wan2.2 — Animate-14B, S2V-14B, T2V-A14B, TI2V-5B.
| Model | Where the weights live | Licence, as named on the page | Duration the card documents |
|---|---|---|---|
| Wan2.2-T2V-A14B | Hugging Face Wan-AI | Apache 2.0 (stated in repo README) | "5s videos at both 480P and 720P" |
| Wan2.2-TI2V-5B | Hugging Face Wan-AI | Apache 2.0 (stated in repo README) | 5-second 720P at 24fps |
| Wan2.2-S2V-14B / Animate-14B | Hugging Face Wan-AI | Apache 2.0 (stated in repo README) | Not stated on pages checked |
| Wan-Dancer-14B | Hugging Face Wan-AI | apache-2.0 on the card | Minute-scale music-to-dance, per the card |
| Wan 2.1 | GitHub Wan-Video/Wan2.1 | Not verified in this pass | Not verified in this pass |
| Lightricks LTX-2 | Hugging Face Lightricks | No licence string verified here | No maximum duration stated |
| Lightricks LTX-2.3 | Hugging Face Lightricks | No licence string verified here | No maximum duration stated |
| Tencent HunyuanVideo | Hugging Face tencent | No licence string verified here | Default --video-length 129, shown as 5s |
Two details deserve pulling out. The LTX-2 model card states no maximum duration in seconds at all; the only length constraint it documents is "Frame count must be divisible by 8 + 1", and its inference example of 121 frames at 24fps is an example, not a ceiling. The LTX-2.3 card repeats that rule and likewise states no maximum duration, no fps ceiling and no frame limit. If you have read a blog asserting a hard LTX second limit, that figure did not come from the card. Lightricks also now ships LTX-2.3 alongside an LTX-2.3-22b IC-LoRA family, so check which version you pull.
Then Wan-Dancer, dated July 13, 2026, licensed apache-2.0, titled "A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation." The organization is still shipping open weights in mid-2026. It is just not shipping them on the numbered mainline.

2.Which models are marketed as open but ship API-only?
Wan 2.5, 2.6 and 2.7 are callable as API model ids while being absent from the project's own repositories. Google's Veo 3.1, OpenAI's Sora 2, MiniMax Hailuo and Luma's Ray models are documented purely as hosted endpoints in their vendors' developer docs.
The Wan case is sharpest, because one brand sits on both sides of the line.
| Model id | Callable at | Documented duration | Weights in the org's repos? |
|---|---|---|---|
wan2.7-t2v | Model Studio, Together AI, Replicate, WaveSpeed | "An integer from 2 to 15. Default: 5." | Not present |
wan2.7-i2v | Alibaba Cloud Model Studio | "an integer from 2 to 15" | Not present |
wan2.7-videoedit | Model Studio, Replicate | "an integer in the range [2, 10]" | Not present |
wan2.6-t2v | Model Studio, WaveSpeed | 2 to 15 seconds, default 5 | Not present |
wan2.5-t2v-preview | Alibaba Cloud Model Studio | "Valid values are 5 and 10." | Not present |
wan2.2-t2v-plus | Alibaba Cloud Model Studio | "Fixed at 5 seconds and cannot be changed." | Yes — Apache 2.0 weights |
wan2.1-t2v-plus / -turbo | Alibaba Cloud Model Studio | "Fixed at 5 seconds and cannot be changed." | Repo exists |
Read the top rows against the bottom rows and the shape is obvious. The hosted 2.7 line takes a 2–15 second duration parameter and outputs 720P or 1080P, including 1920×1080, per the Model Studio text-to-video reference (dated Jul 01, 2026). The image-to-video reference adds three 2.7 modes: first-frame-to-video, first-and-last-frame-to-video, and video continuation. The downloadable 2.2 endpoints are "fixed at 5 seconds and cannot be changed."
Distribution is uneven, which is a useful tell. Replicate's wan-video collection carries 2.7, 2.5, 2.2 and 2.1 — 2.7 entries spanning text-to-video, image-to-video, reference-to-video, video editing and image generation — but no 2.6. WaveSpeed AI lists collections across 2.1 through 2.7, including the 2.6 Replicate skips. Together AI serves Wan-AI/wan2.7-t2v at 2–15 seconds, 720P and 1080P, listed at $0.10 per second of generated video — the number your own GPU bill has to beat.
What this article will not tell you is why the gap exists. No page checked states that Alibaba stopped open-sourcing, or that 2.5 through 2.7 are deliberately closed. Three facts are verifiable: repos top out at 2.2, the model org tops out at 2.2, and 2.5/2.6/2.7 exist as API ids. Anyone telling you the motive is guessing.
The build-versus-buy answer, before you read the hardware section
If you want the weights to fine-tune, self-hosting is the right call and this article will tell you exactly which checkpoints are real. If what you actually want is finished video this month, the GPU is not the interesting part of the problem — the shot plan is.
Imagera's cinematic video studio gives you the second path with no cluster to rent and no checkpoint to babysit, priced in credits so cost tracks what you make rather than what you provision.
3.Is there a Wan 3.0 you can download?
No. There is no Wan3 model under the Wan-AI Hugging Face organization, and no model id beginning wan3 on Alibaba Cloud Model Studio's list, where the newest Wan id is wan2.7-image-pro.
A Hugging Face search for "wan3" returns only unrelated third-party repositories — Mohammedf3/Wan3, Danielbsittler/wan3, Himmdy/Wan3D, intchous/wan3 — none from Wan-AI. The Model Studio model list, dated Jul 15, 2026, tops out at 2.7. If someone offers you Wan 3.0 weights today, they are offering something whose provenance you cannot check.
Worth saying plainly, because the anticipation is real. Single-pass 30-second generation would change format economics more than any quality bump of the last two years, and new video engines are added to Imagera as they become available, with early access going out through the studio. No date to promise, no form to fill in. The teams that move first will be the ones whose account, shot workflow and credit balance already work — the practical argument for building in the cinematic video studio now rather than on an announcement later.

4.What does the licence fine print actually say?
Only two licences were verifiable by name here: the Wan2.2 repository states "The models in this repository are licensed under the Apache 2.0 License," and the Wan-Dancer-14B card carries apache-2.0. Every other open-weight card below needs you to read its licence yourself.
Three tiers turn up, and conflating them is how legal review kills a project in month three:
- A named permissive licence. Apache 2.0, as on the Wan2.2 README and the Wan-Dancer-14B card. You know where you stand.
- A community or research licence. Downloadable, often with acceptable-use terms, revenue thresholds or field-of-use limits attached. Downloadable is not the same as commercially unencumbered.
- "Open" with no licence named on the page you are reading. This is the one that bites. It does not mean the model is unlicensed; it means the page you used to decide did not tell you.
For the LTX and HunyuanVideo cards, this pass verified technical constraints and not a licence string — a gap in the reporting, not a verdict on the models.
5.What hardware does a 14B or 27B video checkpoint actually need?
The honest answer is that the cards checked publish almost no VRAM figures. The one hardware statement Wan makes is that TI2V-5B "can also run on consumer-grade graphics cards like 4090." No comparable statement exists for the 14B or 27B configurations.
What the architecture does tell you: Wan2.2 uses a mixture-of-experts design where "Each expert model has about 14B parameters, resulting in a total of 27B parameters but only 14B active parameters per step," per the Wan2.2 repository. Active parameters govern compute per step; total parameters govern what has to be resident or swapped. A rig sized on "14B active" is not the same purchase as one sized on "27B total."
The other lever is the VAE. Wan2.2 ships a high-compression VAE achieving "a T×H×W compression ratio of 4×16×16, increasing the overall compression rate to 64 while maintaining high-quality video reconstruction." Note the axis order — 4 in time, 16×16 in space, not the reverse, which secondhand write-ups routinely get backwards. That compression is why the 5B variant fits on one consumer card.
| What you want to know | What the sources actually state |
|---|---|
| VRAM for T2V-A14B | Not stated on the pages checked |
| VRAM for the 27B MoE configuration | Not stated on the pages checked |
| Consumer-GPU support | TI2V-5B "can also run on consumer-grade graphics cards like 4090" |
| Time per clip | "a 5-second 720P video in under 9 minutes on a single consumer-grade GPU" |
| Output ceiling per generation | TI2V-5B: 720P, 24fps, 5 seconds |
| Latent compression | 4×16×16 (T×H×W), overall rate 64 |

6.How long does a 30-second sequence take on your own GPU?
Take the one published timing figure — under 9 minutes for a 5-second 720P clip on a single consumer GPU — and a 30-second sequence is six clips, so roughly under an hour of GPU time. That is the clean run, before a single retry.
Nobody gets the clean run. Real work means rejected takes, prompt revisions, seed sweeps and a client note on shot four. Multiply by three to five and the honest number for one 30-second deliverable on one consumer card is most of a working day, on hardware that can do nothing else meanwhile. That arithmetic is derived from the published per-clip figure rather than quoted from it — but it decides build-vs-buy, so run it before you buy the card, not after.
This is also where people go looking for style control and end up in training, a separate discipline with its own cost curve. If that is your direction, the cinematic LoRA guide for Wan 2.2 covers it properly. This piece stays on inference weights.
7.Why can't open weights just generate a 30-second take?
Because long video is four hard problems at once: attention cost that scales badly, long-range memory, error accumulation across frames, and a temporal compression ceiling in the tokenizer. None is a resourcing problem you can solve with a bigger GPU.
The literature is blunt about each. On cost: "scaling diffusion transformers to generate long-context videos is fundamentally limited by the quadratic cost of self-attention" (arXiv:2508.21058). The same work frames memory as retrieval — "models must retain and retrieve salient events across a long range without collapsing or drifting." A 2026 survey adds that "generating long videos remains challenging due to error accumulation, attribute drift, and the limited availability of long video data" (arXiv:2606.22370). And on tokenizers, "extending state-of-the-art video tokenizers to achieve a temporal compression ratio beyond 4× without increasing channel capacity poses significant challenges" (arXiv:2501.05442).
Which is why the closed APIs answer duration with stitching, not generation. Google's Veo docs allow only 4, 6 or 8 seconds natively, and the value "Must be '8' when using extension, reference images or with 1080p and 4k resolutions." Extension adds 7 seconds at a time, "up to 20 times," giving "a single video combining the user input video and the generated extended video for up to 148 seconds" — but the resolution row is annotated "'720p' only for extension," so length costs resolution. OpenAI's docs state that "Both sora-2 and sora-2-pro support 16- and 20-second generations," each extension adding up to 20 seconds, six maximum, "for a maximum total length of 120 seconds." MiniMax defaults to 6 seconds and caps Hailuo-2.3 and Hailuo-02 at 10 seconds at 768P, 6 at 1080P. Luma supports extension, but "Extend is currently supported only for generated videos."
Native duration and extended duration are not the same product. One is a continuous take; the other is a chain with a seam at every joint. Anyone selling you 148 seconds is selling you 21 joins.

8.Should you self-host or buy hosted inference?
Self-host when the model is the product, when you have a real fine-tuning agenda, or when data cannot leave your network. Buy hosted inference when the video is the product and the deadline is fixed. Most commercial work is the second case.
Self-hosting is a legitimate, well-supported path in 2026. Apache 2.0 weights, working diffusers integrations and a 5B variant that runs on a 4090 are not a toy. If your inference cost compounds across millions of runs, owning the stack is correct and the maths will say so.
If you bill for delivery, the ledger reads differently. You are not buying compute, you are buying scheduled output. And the capability most buyers actually want — the 2.7-class 2–15 second range with a 1080P tier — is not downloadable at all today, so self-hosting it is not on the table.

9.What can you ship this week instead of waiting?
A 30-second cut is not one generation, it is a sequenced set of shots — and that workflow runs today. Six shots at five seconds each, cut in order, is the same structure the closed APIs assemble through extension, without the joins hidden from you. Start with the smallest possible commitment: one shot.
- Shot 1 — establishing. Wide, static or slow push. Set place and light in the video studio.
- Shot 2 — subject. Medium, matching the light from shot 1.
- Shot 3 — detail. Close, one movement only. This is the shot that sells texture.
- Shot 4 — the turn. Whatever the scene is actually about.
- Shot 5 — hero. For commercial work, use product reels or the universal reel maker, both driven from stills you already own.
- Shot 6 — resolve. Return to the shot-1 framing so the cut reads as one scene.
Already sitting on long footage? Go the other direction: AI video to reels cuts existing runtime into shorts. Still deciding which engine profile fits? The model comparison surface lays the options side by side.
Everything is priced in credits. Pro is 500 credits at $19.99 and the best value; Business is 1,500 credits at $49.99; Ultra is 6,500 credits at $199.99 with the lowest cost per credit — all shown at roughly half off. Monthly plan credits reset with each cycle, while add-on credit packs never expire — full plan detail is on the pricing page.
10.Status line: confirmed vs unverified, as of 29 July 2026
| Claim | Status |
|---|---|
| Wan2.2 weights are Apache 2.0 | Confirmed — stated in the repo README |
Wan-Dancer-14B is apache-2.0, dated July 13, 2026 | Confirmed — Hugging Face model card |
| Wan-AI publishes 24 models, newest line 2.2 | Confirmed — org model list |
wan2.7-t2v accepts 2–15 seconds, 720P/1080P | Confirmed — Model Studio reference, Jul 01 2026 |
| No Wan3 model exists under Wan-AI | Confirmed — HF search and Model Studio list |
| Wan 2.5–2.7 weights are deliberately withheld | Unverified — no source states a motive |
| LTX-2 has a hard second-based duration cap | Unverified — the card states no maximum |
| Kling, Seedance and Runway duration ceilings | Unverified — first-party docs returned no specs this pass |
| WaveSpeed's displayed Wan 2.7 price figures | Unverified — the billing unit is not stated |
Those last four rows matter as much as the first five. A ledger that lists only what it confirmed is marketing. A ledger that lists what it could not confirm is usable.




