Most advice about A/B testing AI video assumes you have the budget to test one variable at a time. I don't, and you probably don't either. What I run is eight ads sharing one 5 EUR/day ad set, and I let the algorithm decide which of them gets shown.
That structure is not a compromise. It is the only design that produces a readable answer at this budget, and it has changed what I think testing is for. Here is what it has actually taught me, with the numbers from my own account rather than a case study.
The structure: eight ads, one budget, no per-ad spend limits
Every day I build one campaign holding one ad set. The budget sits on the ad set — never on the individual ads — and inside it go four video ads and four image ads for the same offer. They compete for the same money.
The reason is arithmetic. Splitting 5 EUR into four separate 1.25 EUR ad sets gives you four datasets too small to distinguish from noise, and it forces Meta into four separate learning phases. One ad set with eight creatives inside it stays in one learning phase and spends the whole budget on whatever is working that day.
The cost of this is that you do not get a clean controlled experiment. You get a ranking. For a small account, a ranking is worth more.

What the numbers looked like
Reading my own account over a recent seven-day window: 58 purchases at 4.84 EUR cost per purchase, across 29 active ads. The offer is a 9 USD/month community membership.
Two things in that sentence matter more than they look.
First, cost per purchase is the metric, not ROAS. Meta records only the first payment of a subscription. My measured revenue per purchase comes back as 7.92 EUR — exactly one month. So a subscription campaign's ROAS mathematically cannot clear 1.0 no matter how good the ad is. Judging it on ROAS means killing every winner you have. The right frame is payback: at 4.84 EUR against a 9 USD monthly price, a signup pays for itself inside month one.
Second, 29 active ads producing 58 purchases means most ads produce almost nothing. That is the expected shape and it is not a failure. The job of the seven losers is to make the winner identifiable.
| Metric | My reading | What I do with it |
|---|---|---|
| Cost per purchase | 4.84 EUR over 7 days | The decision metric |
| CTR of converting ads | 1.37% to 5.55% | Context only, never a kill signal |
| ROAS | Capped near 1.0 by design | Ignored for subscriptions |
| Impressions before judging | 1,000 minimum | Hard floor |
The lesson that cost me the most: CTR is not a kill signal
The obvious way to prune eight ads is to kill the ones nobody clicks. I built that rule and it was wrong.
In the window above, my converting ads spanned CTR from 1.37% to 5.55% — a four-fold spread among ads that were all producing sales. The lowest-CTR ad in that group still bought purchases at an acceptable price. If I had set a CTR floor anywhere sensible-looking, I would have killed a profitable ad.
So the rule became: an ad with a sale is never paused on CTR. CTR only earns a vote when an ad has spent real money and produced nothing, and even then it is compared against the worst CTR among the ads that are converting — not against a number I made up.
A live example from the same pull: one ad had 2,565 impressions, 8.32 EUR spent and zero purchases. On instinct, that gets cut. But its CTR was 2.26% — above the lowest converting ad on the account. People were clicking at a rate that has produced sales elsewhere, so the problem was not the creative. It stayed live.
Minimum spend before you are allowed to have an opinion
The most expensive mistake I have made in this pipeline was judging a campaign at 1.10 EUR of spend and 13 impressions, less than three hours old. There was no signal there at all — I was reading noise and calling it a verdict.
The floors I run now:
- 1,000 impressions before an ad can be evaluated on anything.
- Roughly 2x your target cost per purchase in spend before performance is judged at all. For a 9 USD target, that is around 18 USD of spend on that ad.
- Never judge on a single day. One bad window is weather, not climate.
These sound conservative. They are cheaper than the alternative, which is rebuilding a winner you deleted.

What is actually worth varying
With eight slots, spend them on differences big enough to survive a small sample. Four subtle variations of the same idea will return four statistically identical results and teach you nothing.
- Format, not polish. Video against static image is the single biggest split, which is why I run four of each. They behave differently enough that the answer is readable within days.
- The opening line. The first spoken sentence changes the audience the algorithm finds. Pull it from a proven structure rather than improvising — more on that in AI video hooks.
- The angle, not the wording. "Save time" and "save money" are a real test. Two rewrites of "save time" are not.
- Setting and cast. Cheap to change in a generated clip, and it moves results more than most people expect.
What I no longer bother varying: music, minor colour grading, small caption restyling. At this budget those differences disappear into the noise floor.
The thing that beat every creative test I ran
Worth saying plainly, because it undercuts the whole premise of creative testing: on my organic distribution, changing when a clip posted moved results by roughly twentyfold on the same file. No creative variation I have ever tested came close to that. The detail is in the posting slot that beat my creative by 20x.
The lesson generalises: before you spend a week testing hooks, check that distribution, placement and timing are not leaving a larger gain on the table. Creative testing has the best return once the cheaper structural wins are already taken.
A workflow you can copy
- One campaign per day, one ad set, budget on the ad set.
- Eight creatives inside it: four video, four static, same offer, same landing page.
- Vary format and angle, not polish. Render each clip at 15 seconds, vertical, and generate the statics from the same source images so the comparison is fair.
- Do not touch anything until each ad clears 1,000 impressions and the ad set has spent about 2x your target cost per purchase.
- Rank on cost per purchase. Pause the ads with real spend and no sales whose CTR sits below your worst converting ad. Leave everything else alone.
- Take the winner's structure into tomorrow's eight as the control, and change one thing to try to beat it.
Step six is where compounding happens, and it is the step people skip. A test that does not become tomorrow's starting point was just an expense.
The short version
At a small budget, do not run clean A/B tests — run one ad set with eight creatives and read the ranking. Judge on cost per purchase, not ROAS, and especially not on a subscription where Meta only ever sees month one. Never pause an ad that has a sale. Require 1,000 impressions and about 2x your target cost per purchase before forming an opinion. Vary format and angle rather than polish. And check your posting times before you blame the creative.
Related reading: AI video ad generator, AI UGC ads, what an AI video actually costs and how many AI videos per day.
For an agency the calculation is different again, because the constraint is throughput rather than ideas. Four services package well: performance creative at volume, localisation of one master clip into several markets, product video for the long tail of a catalogue nobody will ever shoot, and always-on organic posting. The trap is pricing against the generation cost — nobody is buying credits from you, they are buying the judgement about which clip is good, and that is the one line that does not scale down. Build the finishing pipeline once, keep a human on every approval, and put the disclosure in writing. What to sell, what it costs to deliver and what belongs in the contract is in the agency guide.
I publish the real ad structures, the daily numbers and the pause rules inside the community. If you want the working pipeline rather than the write-up, join us here — or if you would rather we just make the clips, we do that too: custom AI generated video.
Looking for something else? Browse all 128 AI video guides in one list.

Share:
Adobe Firefly Video Pricing: What You Get for Your Credits
AI Video Generator for Affiliate Marketing: Formats That Work Without the Product