Written by: Wiktor Stefański, Head of People & Operations, Digital Colliers
A New Mexico court just told Meta to put $567M into a fund for youth mental-health harms. If you advise a platform client, that ruling changed your discovery surface overnight. It's not just user data anymore. It's the machinery that decided what users saw, what got taken down, and who inside the company knew what a feature was doing to teenagers.
Most firms I talk to aren't ready to produce that machinery on a court's timeline. The data exists. It just lives in ten different systems owned by ten different teams, and nobody at the client has ever assembled it in one place before. That's the gap that's about to get expensive.
What's actually in scope now
The old model of platform discovery was user records, DMs, ad targeting, and maybe a few Slack threads. The new model is broader. Opposing counsel is asking for the artefacts that show intent and knowledge, not just outputs.
In practice that means:
- Experiment logs. Every A/B test the growth team ran on under-18 accounts, with hypotheses, cohort definitions, and outcome metrics.
- Moderation queues and rubrics. What got flagged, what got actioned, what got left up, and the policy that governed each decision at that point in time.
- Internal risk memos and red-team reports. The memos the safety team wrote that never left the wiki.
- Model cards and evaluation runs for any recommender or ranking system touching the affected cohort.
- Executive comms referencing any of the above, including the ones on Signal that someone forgot to disable disappearing messages on.
Each of these lives in a different tool. Experiment logs are in an internal platform nobody outside data science has heard of. Moderation records are in a vendor tool with a retention policy that may or may not have deleted the relevant window. Risk memos are in Notion, or Confluence, or someone's Google Doc.
The counsel-side workflow that actually works
The firms getting ahead of this are running a preservation and mapping exercise before the complaint lands, not after. The pattern looks something like this.
First, map the client's actual data topology. Not the org chart. The systems. Who owns the experiment platform, who owns the moderation vendor contract, who has admin on the safety team's wiki. This is a two-week engagement with engineering, not a memo from the GC.
Second, freeze retention. The default retention on most experiment platforms is 90 to 180 days. If the conduct at issue is two years old, you're already reconstructing from backups. Get written preservation holds into the systems, not just the people.
Third, build a query layer counsel can actually use. Most of these datasets require SQL and a working knowledge of the client's internal schema. If your associates are billing hours to learn a client's experiment schema every time a request comes in, you're burning margin. And under ABA Formal Opinion 512, you can't bill for hours that AI or automation actually saved, so the economics only work if the query layer is built once and reused.
That last point matters more than firms realise. Only about three of a lawyer's eight hours get billed on average. The hours you spend fighting a client's data warehouse are the hours you can't bill anyone for.
The cost of showing up unprepared
The New Mexico number is $567M. That's the ceiling on one case. The floor is the regulatory stack that's arriving on top of it. GDPR already reaches €20M or 4% of global turnover. The EU AI Act adds up to €15M or 3% of turnover for high-risk violations, and the Article 50 transparency obligations kick in on 2 August 2026, which means recommender transparency questions are about to become discoverable in a much more structured way.
Add the plaintiff bar noticing that experiment logs are producible, and the exposure math shifts. A firm that can't produce a coherent experiment history for its platform client in 60 days isn't just losing the motion. It's losing the client, because in-house counsel now knows another firm can.
What the prepared firms are doing
The firms that will handle the next wave well share a few habits. They've got at least one engineer or engineering-fluent associate embedded in the platform practice. They treat the client's data topology as a living document, refreshed twice a year. They've pre-negotiated preservation clauses into their engagement letters. And they've stopped pretending discovery in platform cases is a document review problem. It's a data engineering problem with a legal wrapper on it.
If your team is still running experiment-log requests through the client's data science team by email, the next $567M ruling is going to feel a lot closer than New Mexico.

