Hi, experienced media buyers! Hi to the inexperienced ones too!
Can you guess which screenshots show a useless A/A test and which show an honest A/B test? Where the winner is obvious, and where you need to keep sending traffic? The answer is at the end of the post.

tldr;
A/B tests are often run incorrectly because of cognitive biases, tool limitations, and the limitations of the method itself. The decisions you get from those tests are low-quality too. A/A tests are a simple tool that improves the quality of decisions based on A/B tests.
—
Recently I helped set up landing-page testing. This isn’t the first time I’ve seen A/B tests run at random: someone turns one on, doesn’t collect enough data, decides it’s enough, and turns it off. The decisions end up random too. How do you deal with that? Read on.
What is A/B testing?
You probably know about A/B tests, but I’ll explain anyway. A/B tests, also known as split tests, are used to compare variants on live traffic and see which one is better. In affiliate media buying, this usually means testing landing pages or offers (in Russian), sometimes ads (in Russian), but advertising is a whole thing of its own. In other industries, it can be anything with a target metric you can compare: paywalls, checkouts, icons, onboarding, you name it. A/B tests have a lot of limitations. Affiliate media buying is one of the few industries where they apply relatively well because the funnels are usually very straightforward.
The target metric in affiliate media buying is revenue, profit, or something derived from them. You can keep an eye on other metrics—for example, sometimes it makes sense to look at CTR (in Russian)—but the primary metric is almost always about money: raw revenue, profit, ROI, EPV—any of them will do.
Why do we need A/B tests, and what is hard about them?
Because affiliate media buying is largely about constant experiments (in Russian). Experiments need to be done right (in Russian), otherwise your decisions will be shit. A/B is one of the tools that keeps you from pissing away every hypothesis.
Running a proper A/B test is not that easy. There is a whole set of questions that are not as simple as they may look: what counts as a hypothesis, a variant, and a target metric; how to compare variants; how the conditions under which you buy traffic affect the test; whether there are other hidden factors.
In this post I’ll focus on one question: how do you know how much traffic to send before an experiment is valid?
There are tools like statistical significance calculators that supposedly solve this, but they don’t. A statistical significance calculator is a limited tool with plenty of shortcomings, and you need to understand how and when to use it. There are other tools too, such as Bayesian tests or CUPED (looks like a reason for another post). They aren’t a panacea either, and they require fairly complex experiment design and execution. There are multi-armed bandits; Zhelty writes about them here (in Russian). I used multi-armed bandits in one of my homegrown trackers, something in the UCB family, and overall they work. But if you give them to media buyers, you get another problem: buyers don’t wait. Buyers think the algorithm isn’t aggressive enough and start meddling manually. I’ve also tried classic ML stuff like SVM. It’s an interesting tool, but a whole thing of its own. Not everyone can wire this into their tracking solutions themselves. That doesn’t change the fact that it’s a perfectly valid solution.
Tell a media buyer that making a decision after 20–30 conversions is a bad idea, and they don’t believe you. They don’t believe you when you show them statistical significance calculators, sequential tests, or Bayesian tests. They just don’t. Look, it’s obvious: 500 clicks to each landing page, 17 conversions here and 11 there. The winner is clear. Why are you bothering me? Come on. Standard story in affiliate marketing, ecommerce, installs (in Russian), and inside products.
Why? Because randomness is a hell of a force, for one thing, and because people have no intuition for probability and come with the full set of cognitive biases, for another. I’ll probably write a separate post about that. For now, moving on. Meanwhile, don’t forget to read Kahneman, Taleb, and some Yudkowsky (from before he moved into AI alarmism).
What are A/A tests?
Fine, A/B is clear (right?)—a useful tool, and not a simple one, even though it looks simple.
Then what is an A/A test, and why do you need it?
An A/A test is when you put two absolutely, hopelessly identical variants into a test. Two identical landing pages, two identical offers. Give them different names so you can tell them apart in the stats. Then what? You wait for identical results. Then comes the surprise. Say you have 200 conversions in total and a 20% difference between identical variants. How?
That is how: randomness is powerful, the sample is small, and the traffic is heterogeneous.
Besides, with tests like these, you need to think about the offers. If another three offers are rotating behind the landing pages, you may already have an A/B/C/D/E/F test.
There will almost always be differences. A 3–5% difference in an A/A test across hundreds of conversions is normal. If the difference is larger, something is going wrong.
Why do we need A/A tests?
Fine, you stuck an A/A test in there. What do you get back?
Here is the situation:
– A/B tests are a complicated tool that looks simple
– Quite often we do not realize how powerful randomness is, and we are subject to cognitive biases
– Tests need to be done right, not the usual way
If you are in this situation, reach for an A/A test. Used properly, A/A tests give you a pretty good picture of what is happening with an experiment and when the experiment becomes valid.
Important: at the start, you need to rule out a situation where you watch the stats and pick the moment when the experiment has randomly converged. The difference can grow again. For example, an A/A test may show similar conversion rates on a small amount of traffic—dozens of conversions—but if you wait for a few dozen more, the results may diverge.
Tests like these are a practical way to get a reasonably reliable experiment, with no magic, sign-up, or SMS.
Pros:
– easy to implement technically—just add another A/A alongside your other tests
– practical and indisputable results: until your two A variants converge, there is obviously not enough data for variant B
– can catch tracking screwups or hidden factors that affect conversion
Cons:
– You need to control the experiment (and yourself) in the early stages. Set rules for the minimum number of conversions before making decisions.
– You need more traffic for the test.
What to do in practice
In practice, as usual, it depends. You always need to think through your own experiments. A starting point:
– A good test takes time and volume. You have to accept that.
– Before an A/B test, you need to understand the baseline. Run 2–3 A/A tests at the very beginning: 200 conversions per variant. Do not do fewer. Why 200? It is an empirical number. More is better. Wait for a difference no greater than 5%. In your case, replace 5% with whatever difference you consider acceptable in A/B tests. If the difference is larger in that range, investigate whether there are technical errors, whether there is hidden variability that turns an A/A test into A/B or A/B/C/D, and whether there are other causes.
– Run an A/A/B test with an equal split. You can use an unequal split, but that makes control harder:
– If variant B has no conversions at all while A/A already has 10–15, check the technical side and keep sending traffic.
– Around 50 conversions and a 2–3x difference between variants is a borderline case. There is clearly not enough data, but the variants differ sharply. It makes sense to recheck the technical side and other sources of contamination, then run another 10–15 conversions. If the difference holds, stop.
– Fewer than 100 conversions on each variant AND variant B is 50% worse—stop.
– You can stop a worse result earlier; validate a better one for a few more days.
– A new geo, a new product, a new offer, a new approach—that is a new test. In my experience, results from past tests are useful more often than not, but you never know in advance. Use common sense.
– A test needs to run under the same conditions: variants need to be tested simultaneously, on the same traffic stream, on the same technical platform, and so on. Adding variants one after another makes for a bad test. Too many conditions will change, and the experiment will be meaningless.
– A test lasts at least a week to account for daily and weekly seasonality
– Once the test is over, it is good practice to keep A/A variants running as a permanent health check. That way you can tell when something goes wrong technically or new variability appears.
– Social proof in advertising (in Russian) is part of testing too
A/A/B Test Checklist:
– A/A tests for the baseline: 2–3 A/A tests, 200 conversions per variant, difference < 5%
– A/A/B tests with an equal split:
– 10–15 conversions on A and no conversions on B—check the technical side and keep sending traffic
– 50 conversions, 2–3x difference—check the technical side, run another 10–15 conversions, and stop
– 100 conversions, difference greater than 50%—stop
– New geo, product, or offer—new test
– Losing variant—stop; winning variant—validate at x2–x3
– The test runs under the same conditions
– Preferably lasts a week
– Variants start simultaneously
– Keep an A/A test based on the winner
Why A/A tests fail to converge is something I’ll write about separately.
Important: this tool isn’t a 100% guarantee of a good test. It’s a quick, practical mechanism that clearly improves the quality of an A/B test, but it doesn’t replace good experiment design.
Now the answer to the question at the beginning: all of the tests are A/A. There is not a single A/B test there. You can figure out A/B tests yourself now.

Finishing the post with some distrust is fine. Don’t believe me. Try it yourself.