Multi-armed bandit vs A/B testing: which should you use?

3 min read · updated

Use an A/B test when you need a clean, unbiased answer to one question and can afford to send half your traffic to the worse option until you know. Use a multi-armed bandit when you care more about conversions during the test than about a precise estimate, because it shifts traffic toward what is winning as evidence arrives. For a system that runs continuously, combine them: a bandit to choose, and a fixed holdout to prove.

What is a multi-armed bandit?

The name comes from a gambler facing a row of slot machines, each with an unknown payout. Every pull is a choice between exploiting the machine that has paid best so far and exploring one that might be better. A bandit algorithm makes that choice automatically, pull after pull.

In conversion optimisation the machines are page variants and a pull is a visitor. The algorithm tracks how each variant has done and sends the next visitor to the one it currently believes is best, while reserving some traffic for the variants it knows least about.

Two families dominate. Thompson sampling draws a plausible conversion rate for each variant from its posterior distribution and picks the highest draw. Upper confidence bound methods add a bonus to each variant's observed rate that shrinks as it collects evidence, and pick the highest total. Both explore more when unsure and less as data accumulates.

What does an A/B test give you that a bandit does not?

An unbiased estimate. Because traffic is split evenly and fixed in advance, the difference between the two groups is a fair measure of the effect, with a confidence interval you can defend.

A bandit gives that up on purpose. Once it favours one arm, the arms are observed at different times and in different amounts, so a naive comparison of their rates is skewed. That is fine if the goal is to earn conversions while learning. It is a problem if somebody later asks how much the winning variant was really worth.

When should you use each?

One big decision, such as a new pricing pageA/B test
Short lived content, such as a sale bannerBandit
Many variants and limited trafficBandit
You need a defensible lift number for financeA/B test, or a bandit with a holdout
Different visitors want different thingsContextual bandit
The loop runs continuously without a personContextual bandit with a fixed holdout
Choosing between them

What is a contextual bandit?

A plain bandit looks for one winner for everybody. A contextual bandit looks for the best option given what it knows about this visitor: where they came from, what they have read, whether they have been here before. The answer for a returning visitor stuck on pricing can differ from the answer for a first time visitor on a product page.

The context has to be coarse enough that each cell gets data. Seven well chosen intent contexts learn quickly. A thousand micro segments never converge.

What goes wrong with bandits in practice?

Delayed conversions are the first trap. A visitor who sees a variant today and buys on Thursday is counted as a failure until Thursday, so a bandit that updates every few minutes under rates every variant and can lock in on whichever one happens to convert fastest rather than most. The fix is an outcome window: a conversion is credited to the last exposure inside a set period, once, and arms are compared only on visitors whose window has closed.

Novelty is the second. A changed banner often lifts clicks for a few days simply because it is new. A bandit that trusts the first week will crown it, and the effect fades. Discounting old evidence, and keeping the holdout running after a winner is promoted, catches this.

The third is a moving target. Traffic mix changes with campaigns, seasons and weekdays. If the bandit never forgets, a winner from a sale week keeps winning long after the sale. Giving recent behaviour more weight than old behaviour keeps the policy current without throwing away what it learned.

The last is too many arms. Every extra variant splits the same traffic further. Five strong options learn faster than twenty weak ones, and an arm nobody would ship if it won should not be in the test.

Why does a bandit still need a holdout?

Because the bandit only compares its options with each other. If every option is worse than doing nothing, the bandit will still happily crown the least bad one. A holdout is a slice of visitors who never see any change, assigned by a stable hash so they stay put. Lift is then measured against them, and it stays honest however the bandit moves traffic.

IntellQ ranks strategies per intent context with a contextual UCB score and holds back 20 percent of visitors as control by default. The console shows the raw rate after a change next to the lift against control, and refuses to call a winner until each arm has the minimum sample.

Questions

Is a bandit faster than an A/B test?

It earns more during the test, because less traffic goes to losing variants. It is not faster at producing a precise estimate of the difference, and for that question a fixed split is usually quicker to a confident answer.

Thompson sampling or UCB?

Both work well in practice. Thompson sampling is randomised and handles delayed feedback gracefully. UCB is deterministic, which makes decisions easier to audit and explain. IntellQ uses a UCB style score for that reason.

What is regret in bandit testing?

Regret is the conversions lost by showing an option other than the best one. A/B tests have high regret because half the traffic sees the loser for the whole test. Bandits are designed to keep regret low.

How big should a holdout group be?

Large enough to measure lift within the time you care about. Ten to twenty percent is common. Smaller holdouts cost fewer conversions but take longer to show a real difference.