Written by: Wiktor Stefański, Head of People & Operations, Digital Colliers
The ScaleX finding in the LinkedIn post is the part that should keep partners up at night. Reviewers missed roughly a third of dangerous agent commands across 40,000 runs. That's not a training problem you can drill out. It's the base rate of human attention under time pressure, and it means "we had a human in the loop" is only worth something if you can prove what the human actually saw and decided.
For legal, this matters more than for most sectors. You're advising clients whose risk posture may hinge on their being able to show a regulator, a court, or an opposing party that AI output was checked by a qualified person before it went out. Without a log, that claim is testimony. With a log, it's evidence.
The minimum fields, and why each one earns its place
If you're helping a client stand up review discipline, or building it inside your own firm, the log needs at least five fields per event. Not ten. Five you'll actually fill in.
- Prompt. The exact input sent to the model, including any system prompt and retrieved context. If you can't reproduce the input, you can't defend the output.
- Response. The raw model output before any human edit. Store the pre-edit version alongside the final. The delta is the reviewer's actual work product.
- Reviewer. A named human with a role, not a shared inbox. "Legal Ops" is not a reviewer.
- Decision. Approved, rejected, edited-then-approved, or escalated. Four states, not a free text box.
- Latency. Time between response generation and reviewer decision. This is the field most firms skip and the one a regulator will ask about first.
The latency field is the tell. If your average review latency is eleven seconds on a 900-word contract summary, no one reviewed anything. You've built a rubber stamp with a log attached. Better to know that internally than to discover it in disclosure.
Why memory and Slack threads won't survive scrutiny
The regulatory floor is rising fast enough that ad-hoc review will stop being defensible on its own terms. The SRA issued its AI guidance to solicitors back in November 2023, and it leans heavily on competence and supervision. The ABA's Formal Opinion 512 in 2024 went further, including the point that lawyers can't bill for hours AI actually saved. Both regimes assume you can show your working.
Then there's the EU AI Act. Article 50 transparency obligations kick in on 2 August 2026, and the high-risk obligations follow on 2 December 2027. Penalties for high-risk violations reach up to €15M or 3% of global turnover. That's the ceiling, not the floor, but firms shipping AI-assisted work into EU-facing matters will be inside the perimeter one way or another.
And the courtroom signal is already loud. Stanford's tracker recorded AI-fabricated citations in court filings jumping from 87 cases to over 1,300 in eleven months across 2024. Every one of those is a lawyer who thought they'd checked the output. Some of them probably did check, briefly, and got unlucky. Without a log, you can't tell the difference between an unlucky reviewer and a fictional one.
What to tell clients this quarter
The advice you can give a general counsel today, without waiting for the next regulatory clarification, is boringly concrete.
- Pick the two or three workflows where AI touches client-facing or regulator-facing output. Not all of them. The obvious ones.
- For each workflow, capture the five fields above into something queryable. A database, not a spreadsheet, not a chat log.
- Set a latency floor. If the average review takes less than the time it takes to actually read the output, flag it. That's a real metric, not a vanity one.
- Run a monthly sample. Ten random events, reviewed by someone who wasn't the original reviewer. This is the audit trail a regulator will find persuasive.
The cost of not having the log
The cost isn't a fine, at least not first. It's a discovery request or a regulator letter that you can't answer inside the window. It's a matter where you settle because you can't prove the process. It's the associate who genuinely did the review being unable to demonstrate it two years later when they've moved firms.
Human-in-the-loop is a real defence. It just isn't a self-attesting one. The log is what turns the claim into a defence you can actually mount.

