Cogbird is now in private beta.

How to tell whether an SEO change actually worked

Traffic went up after you shipped. That is not evidence the change caused it — here is what would be.

MeasurementKevin Larsson6 min read

In short

  • A before-and-after comparison cannot separate your change from the season, an algorithm update or a competitor who published on Thursday, because all of them act on the same pages in the same window.
  • The counterfactual has to come from other pages: hold back a comparable set of your own, change the rest, and take the difference between the two groups.
  • A control group chosen after the results are in is not a control group. Freeze the pages, the measurement window and the metric before the change ships.
  • A measurement returns three answers, not two — it worked, it did not, and the data cannot tell you. Reporting the third as either of the others is how a playbook fills up with noise.
  • Some changes cannot be measured on some sites, because there are no comparable pages left to hold back. Saying so is a better answer than a number nobody should trust.

You rewrote forty title tags on a Tuesday. Three weeks later the clicks on those pages are up eleven per cent. The rewrite worked — that is the sentence everybody writes in the retrospective, and there is nothing in the numbers so far that supports it.

The problem is not that eleven per cent is a small number. It is that the same eleven per cent would have appeared if the pages had been left alone and something else had moved. Something else is almost always moving.

Four things changed while you were watching one

Between the day you shipped and the day you measured, at least four other things acted on the same pages:

  • The season. Demand for most queries has a shape over the year, and for some of them the shape is violent. A page about tax deadlines does not need help in March.
  • Google itself. Ranking systems are updated continuously, and the named core updates are only the ones announced. A result page can be reordered without anything on your site having changed.
  • Your competitors. Somebody above you can be demoted, and somebody below you can publish. Either one moves your position without you touching anything.
  • You, elsewhere. Most teams ship more than one thing in three weeks. A release that made the site faster, a new internal link block, a page that got deleted — all of it lands on the same measurement window.

Any of these can produce eleven per cent on its own. Two of them working against each other can hide a change that really did work. This is the part that stings: before-and-after measurement does not merely risk crediting the wrong thing, it also throws away real wins by reporting them as flat.

Why the before-and-after comparison cannot separate them

A before-and-after comparison has one number on each side. Whatever happened in the world between those two numbers is inside the difference, and there is no arithmetic that gets it back out. You can be more careful — take a longer baseline, compare against the same weeks last year, exclude branded queries — and each of those helps a little, and none of them answers the actual question, which is what these pages would have done if you had left them alone.

That question has a name. It is the counterfactual, and no amount of care applied to the treated pages will produce it, because the treated pages were treated.

The counterfactual has to come from other pages

In a normal experiment you would randomise: half the users see the change, half do not, and the difference between the halves is the effect. You cannot do that here. There is one version of a page and Google sees it. Split-testing users tells you nothing about what a crawler was shown.

What you can randomise is pages. Take the set of pages you were going to change, hold some of them back, change the rest, and let both sets sit in the same weather. The held-back pages absorb the season, the algorithm update and the competitor who published on Thursday, because all of those things happen to them too. Whatever is left over on the changed pages is the change.

This is a difference-in-differences design, and it is old and boring and it is the thing that turns an anecdote into a measurement.

Choosing the pages you hold back

A control group is only as good as its resemblance to the treated group. The pages you hold back have to be pages that would have moved the same way, and picking them by hand is where most attempts quietly fail. Three properties matter more than the rest:

  1. They have to be the same kind of page. A product page and a blog post respond to different weather. Comparing across template types measures the templates.
  2. They have to have moved together historically. If two sets of pages have tracked each other for the last six months, they are likely to keep doing so — and if they have not, no amount of similarity on paper will fix that.
  3. They have to be big enough to have a signal. A control group of four pages with nine impressions between them is noise wearing a lab coat.

The uncomfortable consequence is that some changes are not measurable on some sites. If a site has one page of a given type, or if the pages that resemble each other are all in the treated set, then there is no control group to be had and the honest answer is that this change cannot be evaluated. That is a better answer than a number nobody should trust.

Freeze the group before you ship, not after

A control group picked after the results are in is not a control group. Once you can see which pages went up, choosing comparison pages becomes a choice about what answer you get, and it does not require any bad faith to go wrong — you will pick the ones that look most similar, and looking similar after the fact means looking similar to the outcome.

So the group is fixed and written down before the change goes out, along with the window you intend to measure over and the metric you intend to read. That record is what makes the result a finding rather than an opinion formed later.

Three honest verdicts, not two

A measurement like this returns one of three answers, and the third is the one most reporting leaves out:

  • It worked. The treated pages moved relative to the control by more than the two groups normally drift apart.
  • It did not. They moved together, or the treated set moved the wrong way. This is useful. It is the only way a playbook ever stops containing things that do not work.
  • Not enough evidence. The difference is inside the range the two groups wander in anyway. Not a failure of the change — a statement about how much the data can support.

Reporting the third as either of the first two is how teams end up with a list of best practices assembled entirely from noise.

How long to wait

Long enough for the pages to be recrawled and for the result to settle, and long enough for the weekly cycle to average out. In practice that is rarely less than three weeks and often more, and it depends on how often the site is crawled and how much traffic the pages get — a page with a thousand impressions a day resolves far sooner than one with thirty.

The temptation is to look on day four. Looking is fine. Concluding is not: early movement on a page that has just been recrawled is mostly the recrawl.

What this costs, and why it usually does not happen

Nothing in the above is difficult. It is, however, a set of steps that has to be taken every single time, before the interesting part, by somebody who would rather be shipping the next change. Pick a comparable set, freeze it, record the window, wait, difference the two groups, and be willing to report that the thing you spent a week on did nothing.

Which is why it is worth automating rather than remembering. Cogbird freezes a control group of your own untouched pages when a change is reported as shipped, watches both sets in Search Console, and returns one of those three verdicts when the window closes — including the one that says the data cannot tell you. How the verdict is calculated is written up in the documentation, and the short version is that there is no proprietary magic in it: it is the difference between two groups of your own pages, which is the only comparison available on a site you cannot randomise.

Whether you automate it or run it on a spreadsheet, the discipline is the same one. Decide what the pages would have done without you, before you find out what they did with you.

Kevin Larsson · Founder

Building Cogbird (your SEO agent) in public.

SEO that runs itself.

Start your 3-day trial.

Request access