Skillset Blog
How to Automate Facebook Ad Creative Testing With AI
A weekly creative testing loop you can hand to an AI: what to pull, how to group results, the kill and scale thresholds, and what to brief next.

Most Facebook creative testing does not fail because the ideas are bad. It fails because nobody keeps the books. Ads pile up with names like "v3 final final", results get compared across different budgets and dates, and the winner is whichever ad the person remembers. Automating creative testing with AI is mostly about forcing the bookkeeping to happen the same way every week, then spending your judgment on the briefs.
Here is the loop that works, and the specific places an assistant should own the work.
Step 1: Name creatives so a machine can group them
Before automation, fix naming. Every creative needs an angle label in its name, because angle is the unit you are testing, not the file. A hook about needle anxiety, a hook about price per month and a hook about a doctor's routine are three different bets. Five variations of the needle hook are one bet with five executions.
A workable convention is angle, format, version, for example needle-anxiety_ugc-selfie_v2. Once names carry the angle, an assistant can group 40 ads into 6 angles without you explaining anything.
Step 2: Pull the report at a fixed cadence and window
Pick one day and one window and never change them mid test. Weekly, trailing 14 days, purchases and cost per purchase at the ad level, plus spend, is enough for most accounts under $2,000 a day. Longer windows hide fatigue. Shorter ones let one lucky day promote a loser.
This is where an assistant earns its keep. It pulls the same fields every time, so week to week comparisons are actually comparable, and it does not get bored and change the date range.
Step 3: Throw out the noise before ranking anything
Set a minimum spend floor per creative and exclude everything under it. A reasonable floor is 1x your target cost per purchase: if you need a $60 CPA, a creative with $22 spent has told you nothing and should not appear in the ranking.
This single rule removes most bad decisions in creative testing. Half of what people call a winner is a creative with two purchases on $40 of spend.
Step 4: Rank by angle, then by execution
Aggregate spend and purchases by angle, compute cost per purchase at the angle level, and rank those. Then, inside the top angles only, look at individual creatives. You are asking two different questions in order: which story works, and which execution of that story works best.
Accounts that skip the first question end up with nine variations of a mediocre hook, because the best single ad happened to sit inside a weak angle.
Step 5: Apply written kill and scale rules
Write the thresholds down once, then let them run. A starting set:
- Kill a creative at 2x target CPA with zero purchases.
- Hold 48 more hours between 1x and 2x target CPA with at least one purchase.
- Graduate a creative to the scaling set after 3 purchases at or under target CPA, not after one.
- Retire an angle when its cost per purchase has risen 30 percent or more for two consecutive weeks, even if it is still profitable. That is fatigue, and it is cheaper to replace it early.
The numbers matter less than the fact that they are fixed in a file. If you decide the threshold during the review, you will find a reason to keep your favorite ad. Packaged versions of this rotation and graduation logic live under paid ads skills.
Step 6: Brief the next round from what won, not from scratch
The output of a testing week is not a report, it is a production order. For each surviving angle, brief three variations that change one element: hook line, opening visual, or proof. For each dead angle, do not replace it with a near copy.
An assistant is good at this when you give it the raw material instead of asking it to invent. Feed it reviews, support tickets and comments, and it will produce angles grounded in customer language. The paid ads catalog covers both halves of this: turning customer language into angles, and turning a chosen angle into executions.
Step 7: Close the loop with a log
Keep one row per creative per week: angle, spend, purchases, CPA, decision, reason. After six weeks this log is worth more than any course, because it shows which angles your specific audience responds to and how fast each one fatigues. It also makes the assistant better, since you can hand it last quarter's log and ask what pattern the winners share.
What to automate and what to keep
Automate the pulling, grouping, filtering, ranking and the mechanical kill and scale calls. Automate the first draft of briefs. Keep for yourself the judgment on offer changes, the decision to test a new audience, anything touching compliance, and the final read on whether a creative fits the brand.
The reason to put this in a skill file rather than a prompt is consistency. A prompt gives you one good review. A skill gives you the same review every Monday, including the week you are traveling and the week you are hung over, which is when accounts actually get wrecked. If you are new to what that file looks like, start with what a SKILL.md file is, then adapt the thresholds above to your own numbers.
The first week
Do not build the whole system. Rename creatives with angle labels, set the spend floor, and write your four thresholds. Run the review once by hand so you know what the output should look like, then hand the same procedure to your assistant and compare. If its output differs from yours, the file is missing a rule, and that missing rule is the most valuable thing you will write all month.
Questions
How much daily spend do you need before creative testing data means anything?
Enough for each creative to reach roughly your target cost per purchase in spend inside the test window. At a $60 target CPA and five creatives, that is about $300 of test budget before ranking is meaningful.
Should AI make the kill and scale decisions automatically?
It can apply written thresholds reliably, which is most of the work. Keep offer changes, new audience tests and compliance calls with a person.
How often should creative be reviewed?
Weekly on a fixed day, with a trailing 14 day window. Daily reviews overreact to noise and monthly reviews miss fatigue.