Alphaweb Get the paper

Sector note · Legal

Legal AI pays for the second read, not the first draft

The legal AI conversation is about drafting. The hours, and the leaked money, are in review — and review is the part a supervising lawyer can actually check.

Alphaweb — Legal AI pays for the second read, not the first draft
What this note argues5 lessons · 40 seconds
  1. Legal AI earns its money in review, not in drafting.

    Reading documents other people wrote is the high-volume work, and it is the part software does dependably.

  2. A review output can be checked in seconds; a draft cannot.

    A claim about clause 14.3 is verified by reading clause 14.3. Checking a drafted paragraph means doing the work again.

  3. Your playbook does not exist in writing, whatever the team says.

    A precedent bank is not a standard. An agent needs the position, the fallback, the walk-away and approved wording per clause.

  4. The saving is worth nothing until the freed hour has an owner.

    Billed hours mean nothing until resold, and the person who saves them rarely sells them, so the work reads as a cost.

  5. The hard signature is on the documents the agent says are fine.

    If nobody will accept no issues found without re-reading the contract, you have added a step and raised cost per document.

2,000+ (as of 21 September 2026)rulings finding reliance on AI-hallucinated materialAI Hallucination Cases Database, Damien Charlotin
5.4%of contract value separating the contract-literateWorld Commerce & Contracting, with CCMI and Icertis
nearly 240 hoursexpected annual saving per legal professionalThomson Reuters Institute, Future of Professionals 2025

The money is in reading, not writing

Generation failures reach filings and sanctions; review failures quietly sign away contract value.

Almost every legal AI demonstration is a generation demonstration. Something gets drafted — a clause, an advice note, a letter before action — and the audience is invited to be impressed by fluency. The work that actually fills the day of a three-year associate, or of a commercial counsel sitting on four thousand live agreements, is the opposite. It is reading documents other people wrote and deciding whether they are acceptable. Review, not drafting. It is unglamorous, it is high-volume, and it is the part where software of this kind is genuinely dependable.

The asymmetry is now countable. As of 21 September 2026, a public database maintained by Damien Charlotin had logged over 2,000 court decisions worldwide — nearly 1,400 of them in the United States — in which a judge found that a party had relied on AI-hallucinated material such as fabricated case citations. Every one of those is a generation failure that reached a filing. Review failures never produce a sanctions hearing. They produce an indemnity cap nobody noticed, a change-of-control clause that surfaces during due diligence three years later, a renewal notice missed by eleven days. WorldCC research, produced with the Commerce and Contract Management Institute and CLM vendor Icertis, found that organisations treating contracts as sources of financial intelligence outperform their peers by an average of 5.4% of contract value, and that seventy per cent of respondents acknowledged the disconnect between contracts and financial oversight. That gap is made of documents that were signed and then never read again.

So the counter-intuitive position is a plain one. Attention goes to the first draft. Value goes to the second read.

You do not need to trust the model; you need the model to point at the page.

Why this work fits the software

Review decomposes into unit-versus-standard checks, so every output points back at the text that justified it.

Review has a shape that suits an agent, and drafting does not. A review task decomposes cleanly: take a document, break it into units — clauses, obligations, documents in a disclosure set — compare each unit against a written standard, and emit a decision that points back to the text justifying it. The standard is external and fixed: a negotiation playbook, the issues for disclosure agreed in a Disclosure Review Document, a schedule of regulatory obligations. The agent is never asked what the right answer is in the abstract. It is asked whether this particular text meets a stated position, and to show which words made it decide.

That is what makes the output auditable. A drafted paragraph is hard to check, because checking it means doing the work again. A review output is a claim of the form: clause 14.3 caps liability at one hundred per cent of fees paid, your playbook requires one hundred and fifty per cent or twelve months' fees, escalate. A supervising lawyer verifies that in fifteen seconds by reading clause 14.3. You do not need to trust the model; you need the model to point at the page. Built this way, the system fails loudly instead of quietly, which is the only failure mode a regulated profession can live with.

There is a second reason, less comfortable to say aloud. A human first-pass reviewer's judgement on document four hundred of a data room is not their judgement on document four. Everyone in the profession knows this and nobody says it in front of a client. Consistency across volume is precisely what the software is good at, and it is the dimension on which the incumbent process is weakest.

ProcessWhy it goes firstWhat the agent does
Contract review against a playbookThe highest-volume review task, the standard already half-exists, and every output is checkable in seconds against a numbered clause the reviewer has open in front of them.Ingests the counterparty's paper, maps each clause to a playbook position, marks it at-standard, fallback or escalate, drafts the redline using the approved alternative wording, and writes the negotiation note recording what was conceded and against which position. Anything the playbook does not cover goes to a named lawyer rather than being guessed.
Obligation extractionNobody is currently funded to do it, so there is no incumbent process to displace and no billable hour to cannibalise. This is where the 5.4% of contract value separating contract-literate organisations from their peers lives.Reads the executed set — master agreement, order forms, amendments, side letters — and extracts each operative obligation with its trigger, deadline, owner and source clause. Normalises against a house taxonomy, writes into the obligation register or CLM, and flags conflicts between the master agreement and documents executed under it.
Disclosure reviewThe standard is written and agreed between the parties in the Disclosure Review Document, proportionality is court-supervised, and PD 57AD already contemplates software and analytical tools — so the regime is receptive rather than hostile. It goes third because a wrong privilege call is expensive.Applies the agreed issues-for-disclosure coding to each document, ranks by relevance, proposes privilege and confidentiality flags with the passage that triggered them, tracks families and near-duplicates, and produces a defensible log of what was assessed and on what basis to support the disclosure certificate a named person has to sign.

Where to start, and where not to

Take contract review first, obligation extraction second, disclosure third; never start with research memos.

Three processes clear the bar: high volume, a standard that can be written down, and an output a supervising lawyer can check faster than they could redo it. Take them in the order below. Contract review first because the volume and the standard already exist in some form. Obligation extraction second because nobody is currently funded to do it, so there is no incumbent process to displace. Disclosure review third, not because it is unsuited, but because privilege calls are expensive to get wrong and the matter-level controls take longer to build than anyone expects.

Do not start with legal research or advice memoranda, however well they demo. Those are generation tasks, the failure mode is a confident invention, and the checking cost swallows the saving.

What actually stops this

Four blockers: no written playbook, no owner for the saving, no one to sign a clean report, a messy estate.

First, the playbook does not exist in writing. Every legal team says it has one. What exists is a precedent bank, a clause library of uncertain vintage, and a great deal of knowledge in the heads of four people. A playbook an agent can work from states, for each clause type, the standard position, the acceptable fallback, the walk-away and the approved alternative wording. Producing that is legal work, it takes several weeks, and it forces someone senior to commit in writing to positions they have been enjoying the freedom to take case by case. A good deal of the resistance that presents itself as scepticism about the technology is really this.

Second, nobody is scored on the saving. Thomson Reuters puts the expected gain at nearly 240 hours per professional per year, up from 200 in 2024, worth around $19,000 each. In a firm billing by the hour, those hours mean nothing until they are sold again, and the person who saves them is not the person who sells them. In-house, the saving usually lands in an external counsel budget owned by someone with no stake in the project. Until the freed capacity has a named destination, the work reads as a cost.

Third, no one will accept a clean report. The difficult signature is not on the escalations — it is on the documents the agent says are fine. If no partner or general counsel will accept no issues found without reading the contract themselves, you have added a step and raised your cost per document. The SRA answers this in principle: you remain responsible and accountable for the outputs from AI you are using. It has to be answered in practice, by a named individual with a stated tolerance, before the build rather than after it.

Fourth, and the only technical one, the document estate is a mess. Executed agreements sit as scanned PDFs across three systems, as email attachments, and in a contract lifecycle system whose metadata is wrong for the older half of the portfolio. Amendments are filed away from the agreements they amend. Conflicts walls and client confidentiality mean you cannot simply index everything into one store; retrieval has to respect matter boundaries. Expect to spend real effort here, and expect no one to find it interesting enough to fund.

The first ninety days

Spend the first month on legal work, not engineering, so a failure surfaces in week two rather than month six.

The sequence below deliberately spends its first month on legal work rather than engineering, and refuses to widen scope until the audit trail exists. It is designed so that failure arrives early and cheaply: if the fallback positions cannot be agreed in fifteen days, that is a finding about the organisation, and a far less expensive one than discovering it in month six.

What it is worth

Measure cost per document and rework rate, and staff the escalation queue permanently or throughput falls.

Anyone quoting a single percentage reduction across legal work is selling something. The unit that matters is cost per document reviewed to an agreed standard, and it moves with variables specific to you: how many documents of that type arrive each month, how much they vary, how tightly the playbook is drawn, and — decisively — what happens to the hour you free.

The costs that go unmentioned are worth stating. Inference is the smallest line. The larger ones are maintaining the evaluation set, re-running it every time the model or the playbook changes, and paying a qualified lawyer to work the escalation queue. Budget for that person permanently. An agent that escalates one clause in five across five hundred contracts a month generates a substantial standing volume of senior review, and if you have not staffed it, throughput falls rather than rises and the pilot is judged a failure for the wrong reason.

Measure rework rate and cycle time, not headcount. Our honest reading: for contract review at genuine volume against a genuine playbook, the arithmetic usually works within two quarters. For disclosure it depends entirely on the matter and on what the parties agree in the Disclosure Review Document. For obligation extraction the return is not time saved at all — it is leakage recovered, and you cannot know that number until you have read the obligations you have been ignoring. Anyone who gives you the figure before the work is done is guessing, including us.

What a supervisor will ask

In England and Wales the SRA is deliberately outcomes-focused — it does not specify which technologies firms should adopt — so the questions are about accountability rather than tooling, and they are answered by documents you either have or do not. Expect a supervisor to ask: which named individual is accountable for this output, given that you remain responsible and accountable for the outputs from AI you are using; what written standard was the document checked against, and can you produce the version of that playbook in force on the date of the review; what happened to the client's confidential material — which provider processed it, under what contractual terms, and was it retained or used for training; and can the supervising solicitor genuinely evaluate the output, or is the review a rubber stamp, which goes to competence and supervision under the Code of Conduct rather than to the technology. On the litigation side, PD 57AD gives the court power to require the use of specified software or analytical tools, so expect to justify your approach in the Disclosure Review Document and to be asked about recall, not just precision — the defensible question is what you might have missed. One question catches most teams unprepared: when you changed the model or the playbook, did earlier reviews get revisited, and how would you know which ones were affected. For in-house teams, the auditor's version of the same question is completeness of the obligations register and whether a material commitment could be sitting in an unread amendment.

A ninety-day sequence that survives contact

  1. Days 1-15 — Write the playbook downTake the twelve clause types you negotiate most often and get a named partner or general counsel to state, in writing, the standard position, the acceptable fallback, the walk-away and the approved alternative wording for each. This is legal work, not engineering, and it is the step everyone skips. If nobody will sign the fallbacks, stop here — you have learned something important for the price of two weeks.
  2. Days 16-30 — Build the gold setAssemble sixty to eighty previously reviewed agreements of one type from one counterparty class, together with the redlines and negotiation notes actually produced at the time. This is your evaluation set and your only defence against wishful thinking. Do not let a vendor assemble it, and do not let it consist of the easy ones. Include the three deals that went wrong.
  3. Days 31-55 — Run it in shadowThe agent reviews every incoming agreement of that type; a lawyer reviews the same document in the ordinary way; neither sees the other's output until both are complete. Measure disagreement rather than accuracy, and read every disagreement personally. Most will turn out to be playbook ambiguity rather than model error, and each one sends you back to step one — which is the point.
  4. Days 56-75 — Move the agent in front of the lawyerThe agent takes the first pass and the lawyer reviews its marked-up output rather than the raw contract. Keep a defined escalation route: anything outside the playbook goes to a named person under a service level. That queue, not the model's score, is the number that tells you whether this works — if it grows faster than it is cleared, the playbook is too narrow.
  5. Days 76-90 — Fix the record before you scaleVersion the playbook, pin the model, log every decision with the clause text it rested on and the playbook version in force, then re-run the gold set to establish a regression baseline. Only now add a second contract type. Adding one before the audit trail exists doubles your exposure and halves your ability to explain any of it.

What we read

The documents behind this note. Each entry says what it is, what it found, and why it should change what you do — then the link to the original.

Solicitors Regulation Authoritysource 1 of 5

Risk Outlook report: The use of artificial intelligence in the legal market

What it is
The regulator's risk report on AI use across the England and Wales legal market
What it says
Solicitors remain responsible and accountable for the outputs from AI they use, and the SRA does not specify which technologies firms should adopt.
Why it matters
Accountability, not tooling, is the regulated question, so the supervisor asks which named person signed and against which written standard.
Read the original →
Thomson Reuters Institutesource 2 of 5

The Future of Professionals 2025: Mind the gap

What it is
Annual survey of professionals on expected AI time savings and their value
What it says
Legal professionals are expected to free up nearly 240 hours a year, up from 200 in 2024, worth around $19,000 each.
Why it matters
It sizes the prize, and shows why the saving stalls: nobody in the firm is scored on hours that are freed but never resold.
Read the original →
Damien Charlotinsource 3 of 5

AI Hallucination Cases Database

What it is
A live public tally of court decisions involving AI-fabricated material
What it says
As of 21 September 2026 it logged over 2,000 decisions worldwide, nearly 1,400 of them American, where a judge found reliance on hallucinated material.
Why it matters
Every logged case is a generation failure that reached a filing, which is the asymmetry between drafting risk and review risk.
Read the original →
World Commerce & Contracting (with the Commerce and Contract Management Institute and Icertis)source 4 of 5

Stop the Leakage: WorldCC Report Provides Blueprint for Recovering 5.4% of Contract Value

What it is
Industry research on contract value leakage and contract-to-finance visibility
What it says
Organisations treating contracts as financial intelligence outperform peers by 5.4% of contract value, and 70% of respondents acknowledged the disconnect.
Why it matters
It puts a number on obligations nobody reads after signature, which is the case for extraction over any billable-hour saving.
Read the original →
Ministry of Justice (Civil Procedure Rules)source 5 of 5

Practice Direction 57AD - Disclosure in the Business and Property Courts

What it is
The disclosure rules for the Business and Property Courts of England and Wales
What it says
The court may give directions requiring the use of specified software or analytical tools, including technology assisted review, in carrying out disclosure.
Why it matters
Disclosure review is a receptive regime, but the standard is agreed with the other side, so recall and defensibility get argued in public.
Read the original →

The full paper

The white paper sets out the playbook schema an agent can actually work from, how to build and maintain a gold set that survives a model change, the escalation-queue staffing arithmetic, and the decision-log format a supervising solicitor can sign against.