The first time I ran a video model on my own machine I left it going overnight and woke up to eleven seconds of footage and a room that was noticeably warmer. That is the honest starting point for open source video: nobody bills you per clip, and the bill arrives as electricity, VRAM and patience instead.
I still run a mix. Most of what ships for my brands goes through a paid API because it is faster and I am on a schedule. But there are three jobs where open weights win outright, and I want to be specific about which ones rather than pretending local is either magic or useless.
What "open source" actually buys you
Three things, and only three.
No per-generation cost. Once the weights are on disk, iteration is free. This matters more than people expect. On a paid API I write a careful prompt because each attempt costs money. Locally I run twelve variations of the same shot and pick one. That changes how you work, not just what you spend.
No content filter deciding for you. Hosted models refuse things constantly, and often for reasons that have nothing to do with the clip you asked for. A perfectly ordinary fashion shot got refused on me twice last week on a hosted model. Local weights do not refuse.
Your footage never leaves the machine. If you are making client work under NDA, that is not a nice-to-have.
What it does not buy you is quality parity at the top end. The best hosted models in 2026 — Veo 3.1, Kling 3.0, Seedance 2.0 — are still ahead of anything you can run at home, particularly on prompt adherence and on hands. Anyone telling you otherwise is comparing a cherry-picked local output against a bad hosted one.
The three models worth your disk space
Wan 2.2 (Alibaba). This is where I would start, and it is not close. Apache 2.0 licence, which means you can use the output commercially with no strings — the other licences in this space all have clauses worth reading twice. It does text-to-video and image-to-video from the same family, and the motion holds together on fast action better than anything else open. The catch is speed: it is the slowest of the three.
HunyuanVideo (Tencent). Held the "best open model" crown through early 2025 and still produces the most cinematic-looking output of the three at full precision. Full precision is also the problem — it wants 60GB+ VRAM before quantisation, which puts it out of reach of a single consumer card unless you accept the INT8 version.
LTX-Video. The speed play. Roughly two to three times faster than the other two at comparable settings, which makes it the one I would actually use for iteration. The newer LTX-2 line also generates synced audio in the same pass, which no other open model does yet.

The VRAM question, answered honestly
This is where most guides lie to you, so here are the numbers with the caveats attached.
| Model | Full precision | Realistic quantised | Licence |
|---|---|---|---|
| Wan 2.2 (14B) | 80GB+ | 720p on 16GB with Q5_K_M GGUF; 6-8GB with aggressive offloading | Apache 2.0 |
| Wan 2.2 (5B) | ~16GB | Fits on 8GB outright | Apache 2.0 |
| HunyuanVideo | 60GB+ | 24GB with INT8, visible quality cost | Custom, read it |
| LTX-Video 13B | 32GB (720p) | 12-24GB at 720p with fp8 | Open weights |
The honest read: 16GB is the real entry point and 24GB is where it stops being annoying. Below 8GB you are technically running a video model, in the sense that a bicycle is technically a vehicle. The 6-8GB offloading numbers you see quoted are real, but they mean swapping weights in and out of system RAM for every step, and a clip that takes four minutes on a 4090 takes most of an hour.
Everything runs through ComfyUI in practice. It has become the default interface for open video the way Automatic1111 was for images, and every model above ships with community workflow files for it. Budget an evening for the first setup and about ten minutes for every model after that.
Where I still pay instead
Three cases, consistently:
- Anything on a deadline. A hosted render comes back in two to five minutes. Local is fifteen to forty on good hardware. When I am producing three clips a day per brand, that difference is the whole day.
- Anything where a face has to stay consistent. Hosted models with proper reference-element systems are simply better at this right now. I wrote up how I handle character consistency in consistent AI video characters.
- Anything a client sees first. The quality gap shows up exactly where it hurts — hands, text, faces at distance.
For the free-but-hosted middle ground — no GPU, no bill — the options are different and worth their own look: see best free AI video generators and unlimited AI video generator.

The licence trap nobody reads
Open weights are not the same thing as an open licence, and the difference bites when you monetise. Wan 2.2 ships Apache 2.0 — commercial use, no revenue cap, no attribution requirement. Several other well-known "open" video models ship custom licences with clauses about company size, revenue thresholds or acceptable use.
If you are putting output on a monetised channel or into a client deliverable, read the licence file before you read the benchmarks. I have seen people build a whole workflow on weights they were not allowed to sell from.
A realistic starting setup
If you have a 16GB card: Wan 2.2 14B via a Q5_K_M GGUF in ComfyUI, 720p, five-second clips. That is the best quality-to-hassle ratio available and it will run.
If you have 8-12GB: Wan 2.2 5B. It is genuinely weaker, but it runs at usable speed and you will learn the workflow, which is the actual bottleneck for most people starting out.
If you have no GPU worth the name: do not buy one to find out whether you like this. Use a hosted free tier for a month first. Most people discover the hard part was never the render — it was writing prompts that work, and that skill transfers to whatever you run later.
The part that decides whether any of this works
Model choice is maybe twenty percent of the result. The prompt, the reference image and the shot design are the rest, and none of that changes whether the weights are on your SSD or someone else's. I have watched people get better output from a 5B model with a good reference than from a top hosted model with a lazy prompt, repeatedly.
Wan is the model most people land on after reading this, and it is worth knowing what running it actually costs before you commit a weekend to the setup: Wan AI video generator: when open weights beat paying per clip.
Looking for something else? Browse all 72 AI video guides in one list.
That is most of what we work on together. I run a community where I post the actual prompts, the reference setups and the failures from producing daily AI video across two brands — including the local workflows when they are worth the trouble and the honest verdict when they are not. If you want to see what is really working rather than the highlight reel, come and join us inside the community.


Share:
AI Video Generator for Instagram Reels: The Setup I Actually Use
Luma AI Pricing: What Dream Machine Really Costs Per Clip