Research
How we run a benchmark, and what we refuse to publish
Our rules for measuring AI marketing claims — what counts as a result, what gets thrown out, and why most published AI benchmarks are noise.
This post has no numbers in it. It is the method every post in this category follows, published first so that the ones with numbers can be checked against something.
We publish benchmarks because the AI marketing category is full of claims that nobody can reproduce. If we are going to add to that pile, the method has to be visible.
The problem with most AI marketing benchmarks
Nearly every "we tested X and got Y" post has at least one of these problems.
No control. A tool is credited with a lift that was really a seasonal change, a different audience, or a bigger budget running at the same time.
One sample. A single campaign, a single post, a single generation. Marketing outcomes have enormous variance between attempts. One result tells you where the distribution might be centred, and nothing about how wide it is.
The metric was chosen after the result. Impressions if impressions moved, engagement if engagement moved, "sentiment" if neither did.
No failure reported. A post that reports only what worked is an advertisement with a chart in it.
The five rules we hold ourselves to
1. The metric is written down before the run
Whatever we are measuring is fixed, in writing, before anything is generated or spent. If we later find a more interesting number, that becomes a new study with its own pre-registered metric, not a rewrite of this one.
2. Variance is measured, not assumed away
Anything that can be run more than once is run more than once. Where the same input produces very different outputs, the spread is reported alongside the average — because for a practitioner deciding whether to rely on something, the spread is the more useful number.
3. A result that does not survive a second run is not a result
If we cannot get a finding to reappear on fresh data, it does not get published as a finding. It might get published as "we looked for this and could not find it," which is a legitimate outcome and one we intend to publish more of.
4. The confounds are listed, including the ones we could not remove
Every study we run has them. Ad platforms optimise mid-flight. Audiences overlap. Creative fatigue exists. We name the ones we know about instead of implying a cleanliness the setup did not have.
5. Negative and boring results get published
The strongest signal that a benchmark programme is honest is that some of it is disappointing. If everything we ever measured happened to support the product we sell, nobody should believe any of it, and they would be right not to.
What we consider a strong enough result to act on
A finding has to clear three bars before it changes what we recommend or what we build:
- It reappears. Same shape on data that did not exist when the first run happened.
- It survives a plausible alternative explanation. We try to break it before we publish it, and say what we tried.
- It changes a decision. A statistically visible effect too small to change what anyone would do is a curiosity, and we will say so rather than dress it up.
What we will not do
- Publish a number without the method that produced it
- Publish a customer's data, aggregated or otherwise, without permission
- Report earnings or income claims of any kind, ours or anyone else's
- Present an estimate, a projection or a model output as a measurement
Reading the studies in this category
Every post in research states the question, the metric fixed in advance, how
the data was collected, the result, and a section on what the result does not
prove. If a post here is missing any of those, it is a mistake on our side and
we would like to know about it.
FAQ
Do you publish results that make Cook.ai look bad?
Yes, and that is the point of writing this policy down before the results exist. A benchmark programme that only ever confirms the product is a marketing channel, not research.
Can I reproduce your studies?
For anything that runs on public data or public platforms, we publish enough method to try. For anything involving customer accounts we publish the method but not the underlying data.
Why does this category have so few posts?
Because a study that meets these rules takes weeks, and most ideas do not survive the second run. Posts appear here at the rate real results do.
Start cooking
Cook.ai writes the content, runs the ads, builds the funnels and works the CRM — from one workspace.
Keep reading
VSLs are dead. Content is the funnel.
The long sales video stopped working because the audience changed, not because the script got worse. Here is what replaced it.
Read →How to build a one-person content engine in a week
A working setup for a solo business — what to publish, how to capture the lead, and how to tell which piece produced the money.
Read →How we build at Cook.ai
A small team, a lot of AI agents, and one rule that decides everything: a change is not finished until it is proven running.
Read →