Back to Blog Listing

When 61% Of Users Are Low-Volume, Your Cohort Model Is Wrong

When 61% Of Users Are Low-Volume, Your Cohort Model Is Wrong
Kacper Osiewalski Aug 1, 2026 4 min read

Written by: Kacper Osiewalski, Lead Backend Engineer, Digital Colliers

A Pew Research study of 11,989 Polymarket users found that 61% placed fewer than 100 trades in the six-week window, and 58% won or lost less than $100 over the same stretch. That's not a fringe cohort. That's the majority of the book. And yet most operator analytics stacks I've seen still model engagement as if the power user curve is the whole curve.

This matters more in 2025 than it did in 2019, because the regulatory floor has moved. Affordability, RG, and AML obligations don't care that your revenue comes from the top decile. They care about the population you serve.

The whale-shaped analytics problem

Walk into most iGaming BI setups and you'll see dashboards optimised for the top 5%. GGR by segment, VIP retention, deposit velocity for high rollers. The low-volume majority shows up as a single blurred bucket labelled something like casual or new.

That framing was fine when the product team's only job was to grow LTV. It isn't fine now. The UK Gambling Commission's Remote Customer Interaction guidance came into force on 31 August 2022 and was expanded in 2024, and the £150 net deposit threshold over a rolling 30 days is a hard trigger for affordability checks. If you're grouping every player under £500 monthly into one cohort, you literally cannot see the population that lives closest to that trigger line.

The Pew numbers on Polymarket are a useful mirror. Real markets skew heavy toward low-stakes, low-frequency users. If your model doesn't reflect that, your product, your comms, and your RG interventions are all being tuned for the wrong shape.

What a realistic cohort model looks like

A cohort model that actually reflects the population treats low-volume, moderate-volume, and high-volume players as three separate curves with three separate behavioural signatures. Not one curve with tails.

Concretely, the operators I see getting this right tend to split roughly like this:

  • Low-volume: infrequent sessions, small stake sizes, long gaps between deposits. Their risk signal isn't spend velocity, it's sudden change. A dormant player depositing three times in a week is the interesting event.
  • Moderate-volume: regular sessions, stake sizes clustered tightly, deposits that hover near the affordability threshold. This is where the £150 rolling-30 line lives, and it's the cohort where RCI interventions actually earn their keep.
  • High-volume: the traditional VIP curve. Well understood, well instrumented, well over-served relative to their share of the user base.

The point isn't the exact cut points. The point is that each cohort needs its own definition of normal, its own set of anomaly signals, and its own intervention playbook. A single model trained on aggregate data will underfit two of the three groups by construction.

Tie it to the obligations you already have

Once you accept that three curves exist, the compliance side gets easier, not harder. Your affordability triggers can be cohort-aware. Your RG interventions can be cohort-appropriate. A pop-up that makes sense for a moderate-volume player is patronising for a whale and useless for someone who deposits £20 twice a quarter.

Around 1 in 4 UK-licensed operators fails to achieve a satisfactory AML rating on first assessment. That's not usually because the compliance team is asleep. It's often because the underlying data model can't answer questions like "show me every player who crossed the affordability threshold last month and what we did about it," segmented by cohort, without someone running SQL by hand.

The cost of not fixing this

The downside is priced in already. UK penalties for the most serious AML breaches reach up to 15% of gross gaming yield. Kindred Group publicly reported £14M in compliance-team cost for 2023, and that's a well-run operator with the model roughly right. If you're running lean and your cohort model is wrong, you're paying the same fixed compliance cost and getting worse detection out of it.

The operators shipping this properly in 2026 aren't the ones with the biggest data teams. They're the ones who stopped pretending the whale curve was the population, rebuilt cohorts around actual behaviour, and wired the affordability and RG triggers to the cohort a player actually belongs to. Everyone else is going to keep failing first assessments and wondering why.

Related Posts