Imagera AI - AI content creation platform for generating images, cloning voices, creating avatars, and enhancing videos. Privacy Policy | Terms

Blog Post

Open Weights or API-Only? What You Can Download in 2026

A verified 2026 ledger of open source AI video models: which weights you can download, the licences named, and where "open" really means API-only.

By Imagera Team14 min readJuly 29, 2026Updated: August 2, 2026
Share:
A split scene contrasting a physical GPU server rack with a sleek cloud workstation, representing downloadable AI video model weights versus API-only access in 2026

TL;DR

  • Genuinely downloadable as of 29 July 2026: Wan 2.2 (Apache 2.0), Wan 2.1, Wan-Dancer-14B (apache-2.0), Lightricks LTX-2 and LTX-2.3, Tencent HunyuanVideo.
  • Marketed under the same brand but API-only: wan2.5-t2v-preview, wan2.6-t2v and the whole wan2.7 line — the GitHub org shows 5 of 5 repos topping out at Wan2.2, and the Wan-AI Hugging Face org's 24 models do the same.
  • No Wan3 exists: no Wan3 model under the Wan-AI org, and no model id beginning wan3 on Alibaba Cloud Model Studio, where the newest is wan2.7-image-pro.
  • Licences split three ways: named permissive (Apache 2.0 on Wan2.2 and Wan-Dancer-14B), community/research terms, and cards that name no licence at all. Read the file before a contract depends on it.
  • Hardware truth: the cards publish almost no VRAM figures. Wan's only hardware statement is that TI2V-5B runs on consumer cards like the 4090, doing a 5-second 720P clip in under 9 minutes.
  • LTX-2 and LTX-2.3 state no maximum duration in seconds — only 'Frame count must be divisible by 8 + 1'. Any hard second cap you have read did not come from the card.
  • The closed APIs solve duration by stitching: Veo extends 7s at a time up to 148s but only at 720p; Sora 2 extends to 120s max. Native take length is a different product from extended length.
  • Build-vs-buy: self-host when the model is the product; buy hosted output when the video is the product and the deadline is real. Imagera's video studios run that six-shot 30-second workflow today, priced in credits.

Try it yourself — no setup

Turn prompts and images into cinematic AI video.

Half the models called "open" in 2026 are a download. The other half are an API key, a rate limit and a bill. The confusion costs real money: teams budget for a self-hosted pipeline, staff it, and find out eight weeks in that the version they actually wanted was never published as weights.

This is the ledger, checked against first-party repositories, model cards and vendor API docs as of 29 July 2026. Where a page does not state something, this piece says so instead of filling the hole.

One framing note: the reason the question got loud is duration. Nobody shops for weights out of principle. They shop because a 30-second scene that holds is now the format that sells, and they need to know whether the thing that produces it is a download or a subscription.

Rack of GPU servers glowing in a dark data hall, cool blue light, shallow depth of field

1.Which AI video models can you actually download in 2026?

The genuinely downloadable inference weights verified here are the Wan 2.2 family (Apache 2.0), Wan 2.1, Wan-Dancer-14B (Apache 2.0), Lightricks LTX-2 and LTX-2.3, and Tencent's HunyuanVideo. Everything else in this article is an API endpoint.

The Wan-Video GitHub organization reads "Showing 5 of 5 repositories": Wan2.2, Wan2.1, Wan-Dancer, Wan-skills, and a fork of diffusers. That is the entire public code surface. The Wan-AI Hugging Face organization lists 24 models, and the newest numbered line there is Wan2.2 — Animate-14B, S2V-14B, T2V-A14B, TI2V-5B.

ModelWhere the weights liveLicence, as named on the pageDuration the card documents
Wan2.2-T2V-A14BHugging Face Wan-AIApache 2.0 (stated in repo README)"5s videos at both 480P and 720P"
Wan2.2-TI2V-5BHugging Face Wan-AIApache 2.0 (stated in repo README)5-second 720P at 24fps
Wan2.2-S2V-14B / Animate-14BHugging Face Wan-AIApache 2.0 (stated in repo README)Not stated on pages checked
Wan-Dancer-14BHugging Face Wan-AIapache-2.0 on the cardMinute-scale music-to-dance, per the card
Wan 2.1GitHub Wan-Video/Wan2.1Not verified in this passNot verified in this pass
Lightricks LTX-2Hugging Face LightricksNo licence string verified hereNo maximum duration stated
Lightricks LTX-2.3Hugging Face LightricksNo licence string verified hereNo maximum duration stated
Tencent HunyuanVideoHugging Face tencentNo licence string verified hereDefault --video-length 129, shown as 5s

Two details deserve pulling out. The LTX-2 model card states no maximum duration in seconds at all; the only length constraint it documents is "Frame count must be divisible by 8 + 1", and its inference example of 121 frames at 24fps is an example, not a ceiling. The LTX-2.3 card repeats that rule and likewise states no maximum duration, no fps ceiling and no frame limit. If you have read a blog asserting a hard LTX second limit, that figure did not come from the card. Lightricks also now ships LTX-2.3 alongside an LTX-2.3-22b IC-LoRA family, so check which version you pull.

Then Wan-Dancer, dated July 13, 2026, licensed apache-2.0, titled "A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation." The organization is still shipping open weights in mid-2026. It is just not shipping them on the numbered mainline.

Close-up of a laptop screen showing a file manager with very large model checkpoint files, warm desk lamp light

2.Which models are marketed as open but ship API-only?

Wan 2.5, 2.6 and 2.7 are callable as API model ids while being absent from the project's own repositories. Google's Veo 3.1, OpenAI's Sora 2, MiniMax Hailuo and Luma's Ray models are documented purely as hosted endpoints in their vendors' developer docs.

The Wan case is sharpest, because one brand sits on both sides of the line.

Model idCallable atDocumented durationWeights in the org's repos?
wan2.7-t2vModel Studio, Together AI, Replicate, WaveSpeed"An integer from 2 to 15. Default: 5."Not present
wan2.7-i2vAlibaba Cloud Model Studio"an integer from 2 to 15"Not present
wan2.7-videoeditModel Studio, Replicate"an integer in the range [2, 10]"Not present
wan2.6-t2vModel Studio, WaveSpeed2 to 15 seconds, default 5Not present
wan2.5-t2v-previewAlibaba Cloud Model Studio"Valid values are 5 and 10."Not present
wan2.2-t2v-plusAlibaba Cloud Model Studio"Fixed at 5 seconds and cannot be changed."Yes — Apache 2.0 weights
wan2.1-t2v-plus / -turboAlibaba Cloud Model Studio"Fixed at 5 seconds and cannot be changed."Repo exists

Read the top rows against the bottom rows and the shape is obvious. The hosted 2.7 line takes a 2–15 second duration parameter and outputs 720P or 1080P, including 1920×1080, per the Model Studio text-to-video reference (dated Jul 01, 2026). The image-to-video reference adds three 2.7 modes: first-frame-to-video, first-and-last-frame-to-video, and video continuation. The downloadable 2.2 endpoints are "fixed at 5 seconds and cannot be changed."

Distribution is uneven, which is a useful tell. Replicate's wan-video collection carries 2.7, 2.5, 2.2 and 2.1 — 2.7 entries spanning text-to-video, image-to-video, reference-to-video, video editing and image generation — but no 2.6. WaveSpeed AI lists collections across 2.1 through 2.7, including the 2.6 Replicate skips. Together AI serves Wan-AI/wan2.7-t2v at 2–15 seconds, 720P and 1080P, listed at $0.10 per second of generated video — the number your own GPU bill has to beat.

What this article will not tell you is why the gap exists. No page checked states that Alibaba stopped open-sourcing, or that 2.5 through 2.7 are deliberately closed. Three facts are verifiable: repos top out at 2.2, the model org tops out at 2.2, and 2.5/2.6/2.7 exist as API ids. Anyone telling you the motive is guessing.

The build-versus-buy answer, before you read the hardware section

If you want the weights to fine-tune, self-hosting is the right call and this article will tell you exactly which checkpoints are real. If what you actually want is finished video this month, the GPU is not the interesting part of the problem — the shot plan is.

Imagera's cinematic video studio gives you the second path with no cluster to rent and no checkpoint to babysit, priced in credits so cost tracks what you make rather than what you provision.

Open the video studio · See credit plans

3.Is there a Wan 3.0 you can download?

No. There is no Wan3 model under the Wan-AI Hugging Face organization, and no model id beginning wan3 on Alibaba Cloud Model Studio's list, where the newest Wan id is wan2.7-image-pro.

A Hugging Face search for "wan3" returns only unrelated third-party repositories — Mohammedf3/Wan3, Danielbsittler/wan3, Himmdy/Wan3D, intchous/wan3 — none from Wan-AI. The Model Studio model list, dated Jul 15, 2026, tops out at 2.7. If someone offers you Wan 3.0 weights today, they are offering something whose provenance you cannot check.

Worth saying plainly, because the anticipation is real. Single-pass 30-second generation would change format economics more than any quality bump of the last two years, and new video engines are added to Imagera as they become available, with early access going out through the studio. No date to promise, no form to fill in. The teams that move first will be the ones whose account, shot workflow and credit balance already work — the practical argument for building in the cinematic video studio now rather than on an announcement later.

Storyboard cards pinned to a corkboard showing six sequential film shots, cinematic side light

4.What does the licence fine print actually say?

Only two licences were verifiable by name here: the Wan2.2 repository states "The models in this repository are licensed under the Apache 2.0 License," and the Wan-Dancer-14B card carries apache-2.0. Every other open-weight card below needs you to read its licence yourself.

Three tiers turn up, and conflating them is how legal review kills a project in month three:

  1. A named permissive licence. Apache 2.0, as on the Wan2.2 README and the Wan-Dancer-14B card. You know where you stand.
  2. A community or research licence. Downloadable, often with acceptable-use terms, revenue thresholds or field-of-use limits attached. Downloadable is not the same as commercially unencumbered.
  3. "Open" with no licence named on the page you are reading. This is the one that bites. It does not mean the model is unlicensed; it means the page you used to decide did not tell you.

For the LTX and HunyuanVideo cards, this pass verified technical constraints and not a licence string — a gap in the reporting, not a verdict on the models.

5.What hardware does a 14B or 27B video checkpoint actually need?

The honest answer is that the cards checked publish almost no VRAM figures. The one hardware statement Wan makes is that TI2V-5B "can also run on consumer-grade graphics cards like 4090." No comparable statement exists for the 14B or 27B configurations.

What the architecture does tell you: Wan2.2 uses a mixture-of-experts design where "Each expert model has about 14B parameters, resulting in a total of 27B parameters but only 14B active parameters per step," per the Wan2.2 repository. Active parameters govern compute per step; total parameters govern what has to be resident or swapped. A rig sized on "14B active" is not the same purchase as one sized on "27B total."

The other lever is the VAE. Wan2.2 ships a high-compression VAE achieving "a T×H×W compression ratio of 4×16×16, increasing the overall compression rate to 64 while maintaining high-quality video reconstruction." Note the axis order — 4 in time, 16×16 in space, not the reverse, which secondhand write-ups routinely get backwards. That compression is why the 5B variant fits on one consumer card.

What you want to knowWhat the sources actually state
VRAM for T2V-A14BNot stated on the pages checked
VRAM for the 27B MoE configurationNot stated on the pages checked
Consumer-GPU supportTI2V-5B "can also run on consumer-grade graphics cards like 4090"
Time per clip"a 5-second 720P video in under 9 minutes on a single consumer-grade GPU"
Output ceiling per generationTI2V-5B: 720P, 24fps, 5 seconds
Latent compression4×16×16 (T×H×W), overall rate 64

Single high-end graphics card on a workbench under focused light, dust motes in the air

6.How long does a 30-second sequence take on your own GPU?

Take the one published timing figure — under 9 minutes for a 5-second 720P clip on a single consumer GPU — and a 30-second sequence is six clips, so roughly under an hour of GPU time. That is the clean run, before a single retry.

Nobody gets the clean run. Real work means rejected takes, prompt revisions, seed sweeps and a client note on shot four. Multiply by three to five and the honest number for one 30-second deliverable on one consumer card is most of a working day, on hardware that can do nothing else meanwhile. That arithmetic is derived from the published per-clip figure rather than quoted from it — but it decides build-vs-buy, so run it before you buy the card, not after.

This is also where people go looking for style control and end up in training, a separate discipline with its own cost curve. If that is your direction, the cinematic LoRA guide for Wan 2.2 covers it properly. This piece stays on inference weights.

7.Why can't open weights just generate a 30-second take?

Because long video is four hard problems at once: attention cost that scales badly, long-range memory, error accumulation across frames, and a temporal compression ceiling in the tokenizer. None is a resourcing problem you can solve with a bigger GPU.

The literature is blunt about each. On cost: "scaling diffusion transformers to generate long-context videos is fundamentally limited by the quadratic cost of self-attention" (arXiv:2508.21058). The same work frames memory as retrieval — "models must retain and retrieve salient events across a long range without collapsing or drifting." A 2026 survey adds that "generating long videos remains challenging due to error accumulation, attribute drift, and the limited availability of long video data" (arXiv:2606.22370). And on tokenizers, "extending state-of-the-art video tokenizers to achieve a temporal compression ratio beyond 4× without increasing channel capacity poses significant challenges" (arXiv:2501.05442).

Which is why the closed APIs answer duration with stitching, not generation. Google's Veo docs allow only 4, 6 or 8 seconds natively, and the value "Must be '8' when using extension, reference images or with 1080p and 4k resolutions." Extension adds 7 seconds at a time, "up to 20 times," giving "a single video combining the user input video and the generated extended video for up to 148 seconds" — but the resolution row is annotated "'720p' only for extension," so length costs resolution. OpenAI's docs state that "Both sora-2 and sora-2-pro support 16- and 20-second generations," each extension adding up to 20 seconds, six maximum, "for a maximum total length of 120 seconds." MiniMax defaults to 6 seconds and caps Hailuo-2.3 and Hailuo-02 at 10 seconds at 768P, 6 at 1080P. Luma supports extension, but "Extend is currently supported only for generated videos."

Native duration and extended duration are not the same product. One is a continuous take; the other is a chain with a seam at every joint. Anyone selling you 148 seconds is selling you 21 joins.

Film editor's timeline on a monitor showing multiple clips joined end to end, dim studio

8.Should you self-host or buy hosted inference?

Self-host when the model is the product, when you have a real fine-tuning agenda, or when data cannot leave your network. Buy hosted inference when the video is the product and the deadline is fixed. Most commercial work is the second case.

Self-hosting is a legitimate, well-supported path in 2026. Apache 2.0 weights, working diffusers integrations and a 5B variant that runs on a 4090 are not a toy. If your inference cost compounds across millions of runs, owning the stack is correct and the maths will say so.

If you bill for delivery, the ledger reads differently. You are not buying compute, you are buying scheduled output. And the capability most buyers actually want — the 2.7-class 2–15 second range with a 1080P tier — is not downloadable at all today, so self-hosting it is not on the table.

Two desks side by side, one with a home GPU tower and one with a clean laptop workflow, cinematic contrast lighting

9.What can you ship this week instead of waiting?

A 30-second cut is not one generation, it is a sequenced set of shots — and that workflow runs today. Six shots at five seconds each, cut in order, is the same structure the closed APIs assemble through extension, without the joins hidden from you. Start with the smallest possible commitment: one shot.

  1. Shot 1 — establishing. Wide, static or slow push. Set place and light in the video studio.
  2. Shot 2 — subject. Medium, matching the light from shot 1.
  3. Shot 3 — detail. Close, one movement only. This is the shot that sells texture.
  4. Shot 4 — the turn. Whatever the scene is actually about.
  5. Shot 5 — hero. For commercial work, use product reels or the universal reel maker, both driven from stills you already own.
  6. Shot 6 — resolve. Return to the shot-1 framing so the cut reads as one scene.

Already sitting on long footage? Go the other direction: AI video to reels cuts existing runtime into shorts. Still deciding which engine profile fits? The model comparison surface lays the options side by side.

Everything is priced in credits. Pro is 500 credits at $19.99 and the best value; Business is 1,500 credits at $49.99; Ultra is 6,500 credits at $199.99 with the lowest cost per credit — all shown at roughly half off. Monthly plan credits reset with each cycle, while add-on credit packs never expire — full plan detail is on the pricing page.

10.Status line: confirmed vs unverified, as of 29 July 2026

ClaimStatus
Wan2.2 weights are Apache 2.0Confirmed — stated in the repo README
Wan-Dancer-14B is apache-2.0, dated July 13, 2026Confirmed — Hugging Face model card
Wan-AI publishes 24 models, newest line 2.2Confirmed — org model list
wan2.7-t2v accepts 2–15 seconds, 720P/1080PConfirmed — Model Studio reference, Jul 01 2026
No Wan3 model exists under Wan-AIConfirmed — HF search and Model Studio list
Wan 2.5–2.7 weights are deliberately withheldUnverified — no source states a motive
LTX-2 has a hard second-based duration capUnverified — the card states no maximum
Kling, Seedance and Runway duration ceilingsUnverified — first-party docs returned no specs this pass
WaveSpeed's displayed Wan 2.7 price figuresUnverified — the billing unit is not stated

Those last four rows matter as much as the first five. A ledger that lists only what it confirmed is marketing. A ledger that lists what it could not confirm is usable.

Ledger book open on a desk beside a monitor showing a video timeline, moody low-key lighting

Frequently Asked Questions

Which AI video model has genuinely downloadable weights in 2026?
Wan 2.2 is the clearest case: its repository states the models are licensed under Apache 2.0, and the Wan-AI Hugging Face organization hosts T2V-A14B, TI2V-5B, S2V-14B and Animate-14B. Wan-Dancer-14B, Lightricks LTX-2, LTX-2.3 and Tencent's HunyuanVideo are also published as downloads.
Can I download Wan 2.5, 2.6 or 2.7 weights?
Not from the project's own repositories. As of 29 July 2026 the Wan-Video GitHub org shows five repositories topping out at Wan2.2, and the Wan-AI Hugging Face org's 24 models top out at Wan2.2, while wan2.5-t2v-preview, wan2.6-t2v and the wan2.7 line exist as API model ids on Alibaba Cloud Model Studio.
Does Wan 3.0 exist yet?
No Wan3 model appears under the Wan-AI Hugging Face organization, and no model id beginning wan3 appears on Model Studio's model list, where the newest Wan id is wan2.7-image-pro. A Hugging Face search for wan3 returns only unrelated third-party repositories. Anything sold as Wan 3.0 weights today has provenance you cannot verify.
What GPU do I need to run an open-weight video model?
The only hardware statement verified here is Wan's: TI2V-5B "can also run on consumer-grade graphics cards like 4090," producing a 5-second 720P video in under 9 minutes on a single consumer-grade GPU. The model cards checked publish no VRAM figures for the 14B or 27B configurations.
Is Apache 2.0 the normal licence for open video models?
It is not a safe default assumption. Apache 2.0 was verifiable by name on the Wan2.2 repository README and the Wan-Dancer-14B model card. Other open-weight cards checked in this pass surfaced no licence string at all, so read the licence file yourself before a commercial project depends on it.
Why do open weights cap out around five seconds?
Long video generation runs into quadratic self-attention cost, long-range memory and drift, error accumulation across frames, and documented difficulty pushing video tokenizers past a 4x temporal compression ratio without adding channel capacity. Hosted APIs answer duration with extension chains rather than longer native takes.
How is Veo's 148 seconds different from a native 148-second generation?
It is a stitch, not a take. Veo allows only 4, 6 or 8 seconds natively, then extends by 7 seconds at a time up to 20 times to reach a combined 148 seconds, and the docs annotate the resolution row "720p only for extension" — so a long output costs you resolution and carries a seam at every joint.
Should I self-host or use a hosted studio?
Self-host when the model is the product, the fine-tuning agenda is real, or data cannot leave your network. Use a hosted studio when the video is the product and the deadline is fixed — including for the 2.7-class capabilities, which are not downloadable at all today.

Imagera Team

AI Content & Editorial Team

The Imagera AI editorial team brings together AI researchers, product specialists, and content strategists covering practical AI creation workflows.

Areas of Expertise:

AI Image GenerationAI Voice RecreationAI Avatar CreationContent Marketing

Put this guide to work

Turn prompts and images into cinematic AI video.

Turn prompts and images into cinematic AI video.