Case study · EdTech · Meta · Performance creative
Nobody can pick the winning ad in advance, so we stopped trying to.
The short answer
An EdTech client had no paid social history to build on. Over seven months we launched 704 ads, around 27 a week, each with a written hypothesis and a postmortem, because volume without either is just noise. Monthly spend scaled from $65,000 to a peak of $637,000, monthly revenue went from $377,000 to $1.74M, and cost per enrollment fell 46% as the budget grew almost ten-fold.
4.6x
Monthly revenue
$377K to $1.74M
9.8x
Monthly spend
$65K to $637K peak
−46%
Cost per enrollment
while spend scaled
704
Ad tests
461 new concepts, 243 variants
CHART PLACEHOLDER
Designer: replace this box with the ad tests chart (bars split into new concepts and variants, plus the revenue and spend lines).
Alt text to use: “Ad tests launched per month split into new concepts and variants, with monthly revenue and spend, December 2025 to June 2026”
| Month | New concepts | Variants | Total tests | Spend | Revenue |
|---|---|---|---|---|---|
| Dec 2025 | 8 | 7 | 15 | $65K | $377K |
| Jan 2026 | 24 | 25 | 49 | $185K | $644K |
| Feb 2026 | 41 | 73 | 114 | $322K | $795K |
| Mar 2026 | 90 | 20 | 110 | $438K | $1.34M |
| Apr 2026 | 72 | 20 | 92 | $568K | $1.44M |
| May 2026 | 127 | 70 | 197 | $637K | $1.17M |
| Jun 2026 | 99 | 28 | 127 | $360K | $1.74M |
| Total | 461 | 243 | 704 |
Isn’t 704 ad tests just spray and pray?
It’s the fair objection, and volume on its own would have produced nothing.
Most accounts fail in one of two directions. The first is over-investing in a handful of ads. Weeks of concepting, rounds of internal review, four beautiful launches. It wastes the opportunity, because nobody can identify the winning concept in advance. Not the strategist, not the designer, not the client. Limiting yourself to a few attempts is an expensive way to be confidently wrong.
The second is shipping everything to see what sticks. That produces the occasional winner and no knowledge, because without a stated hypothesis you can’t tell whether an ad won on the angle, the format, the offer or the timing. So you can’t do it again on purpose.
The split matters here. Of the 704 tests, 461 were new concepts or angles and 243 were variants of something already working. Two thirds of the output was a genuinely different idea rather than a resize or a colour change.
Nobody knows in advance what will win. That’s the argument for volume, and it’s also why every test needs a reason attached.
Who this is about
An EdTech business where the conversion that matters is an enrollment. Paid social hadn’t been used at any real scale before, so there was no library of proven ads, no history of which messaging worked, and nothing to build on.
What made this work
The design team sat under the performance team
One of the key factors behind volume, data driven hypotheses and fast testing is having the design under the performance team. Not a brand team operating on completely different KPIs. Not a creative agency that cares more about how the work looks than how it performs.
You need them working desk by desk, one Slack message apart. From the moment a data insight lands to the brief being written should take hours, not days or weeks.
At 27 ads a week, a week between reading the data and acting on it costs you a full round of tests.
Two senior people held the account every week, no rotation and no juniors. A performance lead who reads the whole funnel from CTR down to CAC and knows how the algorithm spends. A designer who turns the idea into something that does its job and builds statics and motion at test speed. Both answering to the same numbers, which in a lot of accounts is the whole difference.
Every hypothesis started from data
No spaghetti at the wall. Every concept started from something observed. Reddit threads and other external discussion where people talk about the decision in their own words. The client’s customer survey. The questionnaire applicants fill in when they apply, which is a large pile of first party language about what people want and what is stopping them.
Each of those sources produced several hypotheses rather than one, and we tested most of them. That is where a lot of the volume comes from. It is not 704 random ideas, it is a much smaller set of real signals worked through properly.
Every test got a postmortem
Results were read back against the hypothesis, not just against CPA. That’s the step that turns 704 launches into an accumulating body of knowledge instead of 704 separate events, and it’s the one almost everyone skips.
Questions we get asked
How many ads should you test at once?
Launch new ads as soon as you have enough signal to tell a loser from a winner. Not statistical confidence, directional data. There is no fixed number, it depends on how much signal your spend and conversion volume can carry. Ten ad tests in a week on $5,000 of spend and ten conversions is overkill: you spread the signal so thin that no ad gets the chance to show what it can do.
Is creative volume or creative quality more important?
Picking a side is how accounts stall. Over-thinking a handful of ads wastes the opportunity, because no one knows in advance what will win. Shipping everything without a hypothesis and a postmortem produces the occasional winner and no knowledge. This account grew because it did both at once: high output where every launch was a documented bet.
Who comes up with the concepts, the account lead or the designer?
Both, when they sit in the same team. The account lead is in the account daily, sees what is working and what is not, and reads the external and internal data. In our case they turn all of it into a dashboard that is friendly enough for designers to use, so the designers are looking at the same information. That is what lets them come up with concepts rather than execute briefs. It depends on the kind of designers you work with.
Client name withheld at the client’s request. All figures are actual reported results for the periods shown.
