How We Measure Day-2 Retention: Methodology Behind the Benchmark
Retention · 2026-08-22 · 8 min read · By CROCO Games
A benchmark you quote should be attackable. Cohort definitions, exclusion rules, per-deployment design, honest limits - and a recipe to reproduce our number on your own BI.
We quote one number more than any other: 13.78% Day-2 retention for our Hold & Win portfolio. It appears on our retention technology page, in sales decks, and in comparisons against three competing providers (12.83%, 11.56% and 7.98%). A number used that hard deserves to be attackable — which means publishing exactly how it is produced, what it does and does not claim, and how you can reproduce it on your own data. This article is that methodology, in full.
The honest reason for writing it down goes beyond marketing hygiene. Retention numbers in iGaming are quoted constantly and defined almost never; the same "Day-2" label can describe metrics that differ by a factor of two purely through definition choices. If you are comparing providers — including us — the definitions below are the questions to ask everyone.
The metric, defined without wiggle room
Day-2 retention answers: of players who tried a game for the first time, what share came back to it the next calendar day? Every phrase in that sentence hides a decision, so here are ours:
Cohort entry. A player enters a title's cohort on their first-ever real-money session with that title on that deployment (operator). Demo play does not create a cohort entry; a player who tried the game months ago and returns is not "new" again. This is per-title cohorting — the same player can be day-0 for Coin Spark and day-45 for Coin Train.
The return window. "Day 2" means the next calendar day in the deployment's reporting timezone — day 0 is the first-session date, day 2 in the industry's 1-indexed convention is the day after it. We do not use rolling 24–48h windows: calendar-day logic matches how operator BI systems bucket sessions, which keeps our numbers reconcilable against yours. A return is any real-money round on the same title that day; there is no minimum bet count or session length filter, because every threshold is an editorial decision that inflates comparisons somewhere.
What is excluded. Days with known tracking gaps void the affected cohorts entirely (a partially observed day biases retention downward for that cohort and upward for the next). Cohorts smaller than a minimum size never enter the aggregate — small-sample days produce spectacular percentages and no information. Bonus-driven forced sessions (e.g., free-round campaigns landing on day 2) are flagged: campaign-inflated cohorts are reported, but separately.
Day-7 retention (26.89% for the same portfolio) follows identical rules with a seven-day window: returned on any of days 2–8, not "was still present at day 7".
Per-deployment aggregation: the design decision that matters most
The single largest distortion in cross-provider retention comparisons is audience mix. A provider whose games run mostly on sweepstake-style casual sites will post different raw retention than one running on high-stake regulated lobbies — regardless of game quality. Averaging across deployments and calling it a benchmark measures distribution, not product.
So our benchmark is computed per deployment, then compared within the deployment: on each operator where our Hold & Win titles run alongside competing providers' slots, we compute D2 for each provider's new-player cohorts on that same operator, over the same period, under the same lobby conditions — then aggregate the within-deployment comparisons. The 13.78% vs 12.83% / 11.56% / 7.98% figures are that construction: same shelves, same players, same weeks. It is as close to a controlled comparison as production data allows.
This is also why the competing providers stay anonymous. The comparison uses data observed on shared deployments; naming the providers would turn an aggregate product benchmark into specific commercial claims about named third parties' performance on identifiable operators — data those providers have not consented to publish. Anonymized, the benchmark says what we need it to say: under identical conditions, the retention ranking is consistent. (Our lobby-scan market data — section anatomy, provider counts — names names precisely because it observes only what casinos publicly display; deployment retention is a different data class and gets a different disclosure standard.)
What the benchmark does not claim: that every CROCO title beats every competitor title everywhere (it is a portfolio-level result); that 13.78% is a universal constant (it moves with market mix and season); or that D2 alone equals value (it is the earliest reliable signal, not the whole curve — the bridge from D2 to lifetime value is its own topic: from D2 to LTV).
Why we anchor on Day 2 at all
Three practical reasons. It is the earliest stable signal — day-1 same-day return is noisy with session-splitting artifacts, while D30 takes a month to read; a launch can be steered on D2 within a week (the playbook covers how). It is provider-comparable — bonus policies and CRM differ per operator, but within one deployment those conditions apply to every provider's games equally, which is exactly the construction above. And it is causally close to the game — the decision to return tomorrow is dominated by the session the game itself delivered, before CRM and reactivation machinery kick in. What a game does mechanically to earn that return — event cadence, feature proximity, session shape — is the subject of our retention-science overview.
Reproduce it on your own data
The whole point of a published methodology is that a sceptical operator can run it in an afternoon against their own BI. The recipe:
- Pick a period of at least 8 weeks and one title (or one provider's portfolio).
- Build the cohort: first-ever real-money session per player per title within the period. Exclude players first seen in the 30 days before the period starts (left-censoring guard).
- For each cohort day, count players with ≥1 real-money round on the same title the next calendar day, in your reporting timezone.
- Drop cohort days with tracking incidents and cohorts below your minimum size; flag days affected by free-round campaigns on the measured titles.
- D2 = returned ÷ cohort, aggregated across days by cohort-size weighting.
- To compare providers, compute the same number for each provider's new-player cohorts in the same lobby over the same weeks — never across different sites or periods.
If your numbers for our titles land near ours, the benchmark did its job. If they do not, the divergence is itself diagnostic — usually placement (cohorts recruited from a buried lobby position skew toward determined players and higher retention; hero placement recruits casual traffic and dilutes it). Which is why retention should always be read next to placement-aware KPIs, not in isolation.
The questions to ask any provider quoting retention
Ours or anyone's — the checklist is the same, and it is short: How is cohort entry defined, and is it per-title? Calendar-day or rolling window, and which timezone? What are the exclusion rules — small cohorts, tracking gaps, bonus-driven sessions? Is the comparison within-deployment or across different operator mixes? What is the sample scale and period, and does the number move seasonally? A provider who answers all five in writing is quoting a measurement. A provider who cannot is quoting a wish.
Frequently asked questions
What is Day-2 retention in slots?
The share of a game's first-time real-money players who play the same title again the next calendar day. In our convention day 0 is the first-session date and "day 2" is the following day, computed in the deployment's reporting timezone, with no minimum-activity threshold on the return.
Why do retention numbers differ so much between providers?
Mostly definitions and audience mix rather than game quality: per-title vs per-brand cohorts, rolling vs calendar windows, filtered vs unfiltered returns, and — biggest of all — averaging across different operator audiences. A meaningful comparison fixes the deployment: same lobby, same period, each provider's new players measured identically.
Is 13.78% Day-2 retention good?
Within our per-deployment benchmark it is the highest of the four slot providers measured (against 12.83%, 11.56% and 7.98%) — but the honest answer is that any absolute number is context-bound. Market mix, lobby placement and season all move it; the durable claim is the within-deployment ranking, not the absolute value.
How can an operator verify a provider's retention claims?
Run the published methodology on your own BI: first-ever-session cohorts per title, next-calendar-day returns, small-cohort and campaign exclusions, within-lobby provider comparison over identical weeks. Any provider whose claims survive that reproduction has earned the shelf position; any provider who will not share their definitions has answered a different way.
Key takeaways
- Our benchmark's definitions, in one line: per-title first-ever real-money cohorts, next-calendar-day return in the deployment timezone, no activity thresholds, small and compromised cohorts excluded.
- The load-bearing design choice is per-deployment comparison — same lobby, same weeks, every provider measured identically — which removes audience mix, the biggest distortion in cross-provider numbers.
- Competitors stay anonymous because deployment data is a different disclosure class than publicly visible lobby scans; the benchmark's claim is the within-deployment ranking, not named scalps.
- D2 is the anchor because it is the earliest stable, provider-comparable, game-attributable signal — steerable within a launch week.
- The methodology is fully reproducible on operator BI in an afternoon; divergence from our numbers is usually placement, and is diagnostic in itself.
Partner with CROCO Games
Judge us by the number after you have audited how it is made. The 13.78% Day-2 / 26.89% Day-7 Hold & Win benchmark ships with this methodology, per-title retention data under NDA, and our team will walk your analysts through reproducing it on your own BI during evaluation.
Certified by GLI, BMM, eCOGRA and iTech Labs, live with 600+ operators across 50+ markets, one REST API and about 24 hours to go live — then measure us on your own cohorts.