Research

How we run a benchmark, and what we refuse to publish

Our rules for measuring AI marketing claims — what counts as a result, what gets thrown out, and why most published AI benchmarks are noise.

The Cook.ai team4 min read

This post has no numbers in it. It is the method every post in this category follows, published first so that the ones with numbers can be checked against something.

We publish benchmarks because the AI marketing category is full of claims that nobody can reproduce. If we are going to add to that pile, the method has to be visible.

The problem with most AI marketing benchmarks

Nearly every "we tested X and got Y" post has at least one of these problems.

No control. A tool is credited with a lift that was really a seasonal change, a different audience, or a bigger budget running at the same time.

One sample. A single campaign, a single post, a single generation. Marketing outcomes have enormous variance between attempts. One result tells you where the distribution might be centred, and nothing about how wide it is.

The metric was chosen after the result. Impressions if impressions moved, engagement if engagement moved, "sentiment" if neither did.

No failure reported. A post that reports only what worked is an advertisement with a chart in it.

The five rules we hold ourselves to

1. The metric is written down before the run

Whatever we are measuring is fixed, in writing, before anything is generated or spent. If we later find a more interesting number, that becomes a new study with its own pre-registered metric, not a rewrite of this one.

2. Variance is measured, not assumed away

Anything that can be run more than once is run more than once. Where the same input produces very different outputs, the spread is reported alongside the average — because for a practitioner deciding whether to rely on something, the spread is the more useful number.

3. A result that does not survive a second run is not a result

If we cannot get a finding to reappear on fresh data, it does not get published as a finding. It might get published as "we looked for this and could not find it," which is a legitimate outcome and one we intend to publish more of.

4. The confounds are listed, including the ones we could not remove

Every study we run has them. Ad platforms optimise mid-flight. Audiences overlap. Creative fatigue exists. We name the ones we know about instead of implying a cleanliness the setup did not have.

5. Negative and boring results get published

The strongest signal that a benchmark programme is honest is that some of it is disappointing. If everything we ever measured happened to support the product we sell, nobody should believe any of it, and they would be right not to.

What we consider a strong enough result to act on

A finding has to clear three bars before it changes what we recommend or what we build:

  • It reappears. Same shape on data that did not exist when the first run happened.
  • It survives a plausible alternative explanation. We try to break it before we publish it, and say what we tried.
  • It changes a decision. A statistically visible effect too small to change what anyone would do is a curiosity, and we will say so rather than dress it up.

What we will not do

  • Publish a number without the method that produced it
  • Publish a customer's data, aggregated or otherwise, without permission
  • Report earnings or income claims of any kind, ours or anyone else's
  • Present an estimate, a projection or a model output as a measurement

Reading the studies in this category

Every post in research states the question, the metric fixed in advance, how the data was collected, the result, and a section on what the result does not prove. If a post here is missing any of those, it is a mistake on our side and we would like to know about it.

FAQ

Do you publish results that make Cook.ai look bad?

Yes, and that is the point of writing this policy down before the results exist. A benchmark programme that only ever confirms the product is a marketing channel, not research.

Can I reproduce your studies?

For anything that runs on public data or public platforms, we publish enough method to try. For anything involving customer accounts we publish the method but not the underlying data.

Why does this category have so few posts?

Because a study that meets these rules takes weeks, and most ideas do not survive the second run. Posts appear here at the rate real results do.

researchbenchmarksmethod

Start cooking

Cook.ai writes the content, runs the ads, builds the funnels and works the CRM — from one workspace.

Keep reading