The first time I ran a video model on my own machine I left it going overnight and woke up to eleven seconds of footage and a room that was noticeably warmer. That is the honest starting point for open source video: nobody bills you per clip, and the bill arrives as electricity, VRAM and patience instead.

I still run a mix. Most of what ships for my brands goes through a paid API because it is faster and I am on a schedule. But there are three jobs where open weights win outright, and I want to be specific about which ones rather than pretending local is either magic or useless.

AI Video Generator Skool Community Banner

What "open source" actually buys you

Three things, and only three.

No per-generation cost. Once the weights are on disk, iteration is free. This matters more than people expect. On a paid API I write a careful prompt because each attempt costs money. Locally I run twelve variations of the same shot and pick one. That changes how you work, not just what you spend.

No content filter deciding for you. Hosted models refuse things constantly, and often for reasons that have nothing to do with the clip you asked for. A perfectly ordinary fashion shot got refused on me twice last week on a hosted model. Local weights do not refuse.

Your footage never leaves the machine. If you are making client work under NDA, that is not a nice-to-have.

What it does not buy you is quality parity at the top end. The best hosted models in 2026 — Veo 3.1, Kling 3.0, Seedance 2.0 — are still ahead of anything you can run at home, particularly on prompt adherence and on hands. Anyone telling you otherwise is comparing a cherry-picked local output against a bad hosted one.

The three models worth your disk space

Wan 2.2 (Alibaba). This is where I would start, and it is not close. Apache 2.0 licence, which means you can use the output commercially with no strings — the other licences in this space all have clauses worth reading twice. It does text-to-video and image-to-video from the same family, and the motion holds together on fast action better than anything else open. The catch is speed: it is the slowest of the three.

HunyuanVideo (Tencent). Held the "best open model" crown through early 2025 and still produces the most cinematic-looking output of the three at full precision. Full precision is also the problem — it wants 60GB+ VRAM before quantisation, which puts it out of reach of a single consumer card unless you accept the INT8 version.

LTX-Video. The speed play. Roughly two to three times faster than the other two at comparable settings, which makes it the one I would actually use for iteration. The newer LTX-2 line also generates synced audio in the same pass, which no other open model does yet.

Flat vector illustration of three stacked bars representing different VRAM capacities

The VRAM question, answered honestly

This is where most guides lie to you, so here are the numbers with the caveats attached.

Model Full precision Realistic quantised Licence
Wan 2.2 (14B) 80GB+ 720p on 16GB with Q5_K_M GGUF; 6-8GB with aggressive offloading Apache 2.0
Wan 2.2 (5B) ~16GB Fits on 8GB outright Apache 2.0
HunyuanVideo 60GB+ 24GB with INT8, visible quality cost Custom, read it
LTX-Video 13B 32GB (720p) 12-24GB at 720p with fp8 Open weights

The honest read: 16GB is the real entry point and 24GB is where it stops being annoying. Below 8GB you are technically running a video model, in the sense that a bicycle is technically a vehicle. The 6-8GB offloading numbers you see quoted are real, but they mean swapping weights in and out of system RAM for every step, and a clip that takes four minutes on a 4090 takes most of an hour.

Everything runs through ComfyUI in practice. It has become the default interface for open video the way Automatic1111 was for images, and every model above ships with community workflow files for it. Budget an evening for the first setup and about ten minutes for every model after that.

Where I still pay instead

Three cases, consistently:

  • Anything on a deadline. A hosted render comes back in two to five minutes. Local is fifteen to forty on good hardware. When I am producing three clips a day per brand, that difference is the whole day.
  • Anything where a face has to stay consistent. Hosted models with proper reference-element systems are simply better at this right now. I wrote up how I handle character consistency in consistent AI video characters.
  • Anything a client sees first. The quality gap shows up exactly where it hurts — hands, text, faces at distance.

For the free-but-hosted middle ground — no GPU, no bill — the options are different and worth their own look: see best free AI video generators and unlimited AI video generator.

Flat vector illustration of an open padlock beside a licence document with a ribbon seal

The licence trap nobody reads

Open weights are not the same thing as an open licence, and the difference bites when you monetise. Wan 2.2 ships Apache 2.0 — commercial use, no revenue cap, no attribution requirement. Several other well-known "open" video models ship custom licences with clauses about company size, revenue thresholds or acceptable use.

If you are putting output on a monetised channel or into a client deliverable, read the licence file before you read the benchmarks. I have seen people build a whole workflow on weights they were not allowed to sell from.

A realistic starting setup

If you have a 16GB card: Wan 2.2 14B via a Q5_K_M GGUF in ComfyUI, 720p, five-second clips. That is the best quality-to-hassle ratio available and it will run.

If you have 8-12GB: Wan 2.2 5B. It is genuinely weaker, but it runs at usable speed and you will learn the workflow, which is the actual bottleneck for most people starting out.

If you have no GPU worth the name: do not buy one to find out whether you like this. Use a hosted free tier for a month first. Most people discover the hard part was never the render — it was writing prompts that work, and that skill transfers to whatever you run later.

The part that decides whether any of this works

Model choice is maybe twenty percent of the result. The prompt, the reference image and the shot design are the rest, and none of that changes whether the weights are on your SSD or someone else's. I have watched people get better output from a 5B model with a good reference than from a top hosted model with a lazy prompt, repeatedly.

Wan is the model most people land on after reading this, and it is worth knowing what running it actually costs before you commit a weekend to the setup: Wan AI video generator: when open weights beat paying per clip.

Looking for something else? Browse all 72 AI video guides in one list.

That is most of what we work on together. I run a community where I post the actual prompts, the reference setups and the failures from producing daily AI video across two brands — including the local workflows when they are worth the trouble and the honest verdict when they are not. If you want to see what is really working rather than the highlight reel, come and join us inside the community.

Latest Stories

This section doesn’t currently include any content. Add content to this section using the sidebar.