A/B Testing a Casino Lobby Without Fooling Yourself

Retention · 2026-07-19 · 9 min read · By CROCO Games

Lobby experiments fail in predictable ways: too small, contaminated, or measured on the wrong number. A working method for testing placement, tiles and rows.

Casino lobby tests have an unusually bad track record, and not because the teams running them are careless. The environment is genuinely hostile to clean experiments: traffic is unevenly distributed across markets and devices, the outcome variable is heavy-tailed to an extreme degree, players move between segments during the test, and the thing you changed is visible to everyone at once. Most published lobby "wins" would not survive a second run.

This is a working method: what to test, how to size it, the four ways the result gets contaminated, and what to do when the answer is "no difference".

Test the things that are actually testable

Good lobby experiments share a property: one visible element changes and everything else stays fixed. That rules out most of what teams want to test first ("our new lobby") and points at a short list that works:

What does not work as an A/B test: whole-lobby redesigns (too many simultaneous changes to attribute anything), and anything where the control group can see the treatment.

Sizing: the number that kills most tests

Casino outcome data is heavy-tailed. A handful of players contribute a disproportionate share of turnover, which means monetary metrics have enormous variance and require sample sizes most operators cannot reach for a single-position test in a reasonable window.

The practical response is to choose an outcome metric with lower variance, and accept that it is a proxy:

Metric Variance Sample needed Use when
Tile click-through Low Thousands of impressions Testing thumbnails, badges, position
Plays started per exposed session Low-medium Tens of thousands of sessions Testing rows and placement
Rounds per exposed session Medium Tens of thousands Testing whether interest converts to play
Net revenue per exposed player Very high Usually unreachable Long-run programme evaluation only

The honest framing: use click-through and plays-started to decide merchandising questions, and reserve revenue as a guardrail metric you check for catastrophes rather than as the thing you optimise. If a test shows a 3% revenue lift on 5,000 players, you have measured noise.

Decide the metric, the minimum effect worth acting on, and the stopping date before the test starts, and write them down. Peeking at a running heavy-tailed metric and stopping when it looks good is how organisations accumulate a portfolio of imaginary wins.

The four contaminations

1. Novelty. Any visible change produces a short-lived attention bump. Regulars notice that something moved, click it, and revert. A test that runs three days measures novelty; run at least two full weekly cycles so weekday and weekend behaviour are both represented, and inspect whether the effect decays across the window.

2. Cannibalisation. Promoting title A pushes title B down. Measuring only A shows a clean win while the lobby gained nothing. Always define the outcome at the row or lobby level, not the title level: total plays started from the shelf, not plays of the promoted game.

3. Selection. If assignment is by market, device or session rather than by player, the groups differ systematically before the test starts. Randomise at the player level where the platform allows it, and always check pre-period balance on the outcome metric — if the groups differed before the change, the difference afterwards means nothing.

4. Leakage. Players use multiple devices, share accounts, and read affiliate sites and streams that push the same titles. Leakage biases results toward zero, which makes real effects harder to detect and — worse — makes teams conclude that merchandising does not matter when the test simply could not see it.

The confound this industry cannot ignore: exposure begets exposure

There is a structural reason lobby tests mislead more than tests in other industries. Position drives play, and play then produces the numbers that justify position — a self-sealing loop, and it operates during your experiment too. If your platform's shelves are dynamically ranked by recent turnover, a title bumped into a better slot climbs further because of the bump. You are not measuring player preference; you are measuring your own ranking algorithm.

Before running any placement test, find out whether the shelf you are testing on is manually curated or algorithmically ranked, and freeze the ranking for the duration if it is the latter. Combined with the finding that half of a top ten is unchanged six weeks later, this also tells you why placement tests need explicit scheduling: shelves do not naturally produce variation to observe, so you have to create it deliberately.

What research adds — and where it stops

Laboratory work is genuinely useful for generating hypotheses about why a lobby change works. Graydon's 2018 study showing that pre-play signals shift game selection independently of payback tells you thumbnails and badges are worth testing at all. But a lab effect on 33 participants choosing between four machines does not predict the size of the effect on your traffic — it predicts the direction worth investigating. Treat published psychology as a hypothesis generator, and your own holdout as the evidence.

When the answer is "no difference"

Most lobby tests return no detectable effect, and this is the outcome teams handle worst — usually by slicing the data until a subgroup shows something. Three better responses:

First, check whether the test could have detected the effect you cared about. If your minimum effect of interest was 2% and the design could only detect 10%, the result is "we did not measure", not "no difference".

Second, treat genuine null results as licence to decide on other grounds — cost, contractual terms, portfolio balance. Knowing that position 4 and position 7 perform the same for a title is a useful, actionable finding.

Third, keep the log. A file of tests run, effects measured and decisions taken is what stops the same experiment being re-run every eighteen months by a new team, and it is the only durable asset an experimentation programme produces.

Frequently asked questions

What can you A/B test in a casino lobby?

Single visible elements with everything else fixed: thumbnail swaps on a set position, a title's position within a shelf, row order, row composition, and badge presence. Whole-lobby redesigns cannot be attributed and are not A/B tests.

How long should a casino lobby test run?

At least two complete weekly cycles so weekday and weekend behaviour are both captured, with the stopping date fixed in advance. Shorter windows mostly measure the novelty bump that follows any visible change.

Why do lobby tests so often show no result?

Usually because the outcome metric was too heavy-tailed for the sample — revenue per player needs far more traffic than a single position test provides — or because leakage and cannibalisation biased the estimate toward zero. Use click-through or plays-started as the decision metric.

Should revenue be the success metric for a lobby experiment?

Rarely. Revenue per exposed player has variance so high that realistic sample sizes cannot detect plausible effects. Use it as a guardrail to catch harm, and decide on lower-variance engagement metrics.

Key takeaways

Partner with CROCO Games

Running a fair test needs a provider that makes the variables visible. CROCO supplies per-title specs and multiple lobby-art formats, so a thumbnail or placement test changes exactly one thing — and free-rounds support lets you run exposure tests on a challenger title without rebuilding your promo stack.

Integration is a single REST API, live in about 24 hours, already serving 600+ operators across 50+ markets; published portfolio benchmarks (13.78% Day-2, 26.89% Day-7) give you a prior to test against instead of a blank slate. Bring us a position you want to challenge and we will help design the test around it.

Design a placement test What integration involves →