Alphaweb Get the paper

Sector note · Public sector

In government, the decision is the cheapest part of the case

Attention goes to whether software can decide who qualifies. The audited evidence says the queue is made of something else entirely: missing documents, manual re-keying and work done twice.

Alphaweb — In government, the decision is the cheapest part of the case
What this note argues5 lessons · 40 seconds
  1. Public backlogs are a shortage of clerks, not of judgement

    Missing documents, manual re-keying and shadow spreadsheets build the queue, and clerical work is what this software is genuinely good at.

  2. Preventing a second pass is worth more than speeding the first

    In Access to Work, a reconsideration costs 241 minutes against 115 for a new application, so the expensive work is the case done twice.

  3. Start where the software cannot cause an adverse outcome

    If the worst failure is wasted human effort, you earn the audit record and caseworker trust needed before going near a determination.

  4. The rules you need to encode have never been written down

    Published guidance covers the ordinary case; accepted substitutes and local discretion live in practice, and two teams often differ.

  5. Faster is not fairer, and the record will prove which it was

    A well-built agent applies an inconsistent rule at speed and documents it cleanly, which is a gift to a claimant's solicitor.

109 daysAverage Access to Work time, against a 25-day targetNational Audit Office, The Access to Work scheme
241 vs 115 minutesCase manager time: reconsideration vs new applicationNational Audit Office, The Access to Work scheme
68%Asylum initial refusals that went on to appealNational Audit Office, An analysis of the asylum system

The queue is not made of decisions

Access to Work averaged 109 days against a 25-day target, and the cost sits in rework, not verdicts.

Ask a director of operations where AI belongs in a benefits or licensing service and the answer usually arrives as a question about judgement: can the software decide whether this applicant qualifies? It is the wrong first question, and the evidence against it sits in the audit trail rather than in any vendor's deck.

Take the Access to Work scheme. The National Audit Office reported that by November 2025 the Department for Work and Pensions was taking an average of 109 days to process an application against a 25-day target, with 62,100 applications awaiting a decision at the end of March 2025 and 31,700 payment requests outstanding. Buried further in the same report is the figure that actually explains the queue: case managers spend 241 minutes processing a reconsideration, compared with 115 minutes for a new application. The department spends more than twice as long redoing a case as it does doing it the first time. And the reason cases come back is rarely that the caseworker reached the wrong conclusion. It is that the file was incomplete, the evidence was the wrong evidence, or the reasoning was never captured in a form the next person could pick up.

The NAO also records that Access to Work case managers "have to transfer information manually from one system to the other and that they maintain management information on the full customer journey on separate spreadsheets". That is the real shape of a public-sector backlog. It is not a shortage of judgement. It is a shortage of clerks — and clerical work is the one thing this generation of software is genuinely good at.

Making a process faster does not make it fairer — a well-built agent will apply an unfair rule at speed, and produce a much cleaner record of having done so.

Why this work suits software that carries a case

The published test is specifiable; with almost 35% attrition, quality cannot live in caseworkers' heads.

Public administration has a property most commercial processes lack: the rules are published. Eligibility for a scheme sits in statute, in regulations and in guidance the department is obliged to make public and to apply consistently. A caseworker's job is to apply a written test to a stack of documents and record why. That is a shape software can carry — not because the model is clever, but because the specification already exists, and so does the obligation to show working.

The second property is that human capacity is unstable in a way the rules are not. The NAO's analysis of the asylum system found "almost 35% staff attrition rate" among asylum caseworkers in 2024-25, and that it "takes on average seven months for new staff to become fully effective". The same report found "42% of sampled decisions in a rolling twelve months to May 2025 having significant or fail errors". A process that loses a third of its practitioners every year cannot hold quality in people's heads. Encoding the checkable parts — is this the right document, is this field consistent with that one, does the claim meet the published test on its face — is not a threat to caseworker judgement. It is what makes the remaining judgement survivable.

Government is already building this way where it is honest about scope. MHCLG's Extract turns historic planning maps and documents into standardised data, and the team draws the boundary plainly: "It supports officers – it does not replace them", and "Officers review and confirm its accuracy before it is used." That is the correct posture for anything that touches a citizen's entitlement, and it is the posture an agent carrying a case should hold from the first day rather than the day the auditor arrives.

ProcessWhy it goes firstWhat the agent does
Document completeness check at intakeNo adverse decision attaches to it, so the worst failure is wasted human effort rather than a wronged applicant. It also attacks the dominant cause of rework: files that reach a caseworker incomplete and come back as a reconsideration.Opens each new application as it arrives, reads every attachment, and classifies each one against the scheme's published evidence requirements. Catches the specific failures that generate second passes: a payslip outside the required date range, an unsigned employer declaration, a medical letter naming a condition the application does not, an expired identity document. Drafts one consolidated request for everything missing, rather than the three sequential requests a queue normally produces, and writes the completeness assessment into the case record so the caseworker opens a file that is either ready or explicitly not.
Eligibility pre-assessment against published criteriaThe test is written down and publicly available, which makes it specifiable and auditable in a way discretionary judgement is not. The agent never issues the decision; it prepares it, so the legal decision-maker is unchanged and the approval conversation is about assistance rather than delegation.Works through the published criteria clause by clause against the assembled file and produces a recommendation in which every element carries two citations: the rule relied on, and the document and page that evidences it. Anything it cannot evidence is marked not determined rather than assumed, and routed to the officer with the gap named. The output is a decision pack a caseworker reads in minutes and signs, amends or rejects — with every disagreement logged as training data for the specification, not for the model.
Backlog triage and third-party chaseIt touches no decision at all, which makes it the easiest thing in this list to get approved, and it acts on the cases already waiting rather than only on new demand. It is also the work departments are least likely to staff, because chasing has no decision attached and therefore no productivity metric.Re-reads dormant cases and sorts them by what is actually blocking each one, which is usually a single missing item rather than caseworker capacity. Chases the third party holding it — employer, GP practice, local authority, previous department — on a schedule, and records each contact. Surfaces the cases that have become decidable since anyone last looked, flags those approaching or past a statutory or published service standard, and identifies duplicates and cases resolved elsewhere that are still sitting open in the count.

The three that go first

Completeness checks, eligibility pre-assessment and backlog chase: assistive, and none issues a decision.

The sequencing rule is simple and slightly unfashionable: start where the software cannot cause an adverse outcome. The first process should be one where the worst failure is that a human does some work the agent should have saved them. That buys you the operational evidence, the caseworker trust and the audit record you will need before anyone lets you near a determination.

All three below assume the same architecture: the agent works a case rather than answering a prompt, it writes everything it does into the case record in language a tribunal could read, and a named officer remains the decision-maker in law.

What actually stops these projects

Unwritten practice, shadow spreadsheets and the procurement gap stop more projects than the technology does.

The first obstacle is that the rules are not where you think they are. Published guidance covers the ordinary case; live practice is full of local convention that nobody has written down — which documents a particular team will accept in place of the named one, what counts as a reasonable explanation for a gap, when to exercise discretion. You cannot specify what has never been recorded. Budget several weeks of an analyst sitting beside caseworkers before anyone writes code, and expect the policy owner to discover that two teams have been applying the same regulation differently for years. That discovery is valuable and politically awkward, and it will slow you down.

The second is the data, and specifically the shadow data. The NAO found Home Office staff still using "additional spreadsheets" alongside the Atlas case management system, and noted that "poor-quality data and workarounds have been a long-standing characteristic". An agent that reads only the system of record reads an incomplete file and will confidently act on it. Either bring the spreadsheets into scope or accept that the agent's view of the case is partial and design the handover accordingly.

The third is procurement and the pilot trap. The OECD's Digital Government Outlook found AI in use in at least one area of government in 35 of 36 OECD countries, but only 21 of 36 (58%) providing central support for procuring AI goods and services. Skills gaps, it reports, are the most common challenge hindering adoption, while it attributes "a proliferation of pilots with little potential to scale" to the difficulty of measuring AI's impact. The pattern is familiar — a successful twelve-week proof of concept, then nine months in commercial and information assurance, by which point the sponsoring director has moved and the pilot quietly lapses. Decide the route to a production contract before the first sprint, not after the demo goes well.

The fourth is technical and, unusually, the easiest. Most departmental case management systems have no supported write interface. Driving them through the user interface works and then breaks on the next release. The honest answer is to treat the legacy system as read-mostly: the agent assembles, checks and drafts, and a human commits the result through the screens they already use, until an integration route is actually funded. Anything else buys a fragility you will be maintaining for years.

The first ninety days

Ninety days cannot change a statutory service; it can prove the thing works without touching a citizen.

Ninety days is not enough to change a statutory service, and anyone promising that is selling something. It is enough to know whether the thing is real, to have run against live cases without touching a citizen, and to have produced the evidence a senior responsible owner needs to fund the next stage or stop.

What it is worth

Count touches, not cases, and price the avoided appeals: 68% of asylum refusals went on to appeal.

We will not give you a percentage. Any number quoted before someone has read your guidance and counted your touches is a number from a different department, and senior public servants have heard enough of those. What we will give you is the arithmetic to do yourself.

Count touches, not cases. For one queue, work out how many separate times a human opens a file before it closes, and what each of those touches costs in loaded staff time. Then ask which of those touches exist only because something was missing, mis-filed or re-keyed. In Access to Work, the ratio of 241 minutes for a reconsideration against 115 for a new application tells you most of what you need: the expensive work is the second pass. Preventing a second pass is worth more than accelerating a first one.

Then look past the unit costs to the two things that dominate. The first is avoided downstream demand: the NAO found that 68% of people refused at initial decision in the asylum system went on to appeal, and appeals are an order of magnitude more expensive than the decision they contest. A completeness check that stops an application being refused for want of a document it could have had is worth vastly more than the minutes it saves the caseworker. The second is the cost of the queue itself — to the applicant waiting 109 days, and to the department carrying the correspondence, the complaints and the ministerial cases that a queue generates.

And the honest caution. Making a process faster does not make it fairer. If the underlying test is applied inconsistently across teams, or the guidance is out of date, or the reasons given to applicants are not the reasons actually relied on, then a well-built agent will do all of that at speed and at scale, and produce a much cleaner record of having done so. That record is a gift to a claimant's solicitor. Fix what you find during the specification work, or accept that automation will surface it for you at a time of someone else's choosing.

What a supervisor will ask

Expect the first question to be about legal responsibility, not accuracy: who is the decision-maker in law, and can they show they applied their own mind to this case rather than ratifying an output? In the EU, this is now explicit — Annex III of the AI Act (Regulation (EU) 2024/1689) classes as high-risk any system used by or on behalf of public authorities "to evaluate the eligibility of natural persons for essential public assistance benefits and services, including healthcare services, as well as to grant, reduce, revoke, or reclaim such benefits and services", which brings logging, human oversight, data governance and registration obligations with it. In the UK, the constraints arrive from several directions at once rather than one statute: the data protection regime's limits on decisions based solely on automated processing, the Algorithmic Transparency Recording Standard for central government, and ordinary public law — the duty to give reasons that are the actual reasons, the duty to make sufficient enquiry before deciding, and the public sector equality duty under section 149 of the Equality Act 2010, which a supervisor will read as a question about whether your error rate differs across protected groups. Practically, be ready for four demands: reproduce this decision on the version of guidance in force on the date it was made; show the evidence relied on and where it came from; show every case the system declined to determine and what happened to it; and show the override log, because a human-in-the-loop control with a near-zero override rate will be read as a rubber stamp rather than as evidence that the software is good. Assume the National Audit Office, the relevant ombudsman and the first judicial review will all ask for the same artefact, and build it as an output of the process rather than as a reporting exercise afterwards.

A ninety-day sequence that survives contact

  1. Days 1-15 — Instrument one queue and count touchesPick a single process, not a department. Measure how many separate times a human opens each file before it closes, what proportion of cases come back for a second pass, and why. Identify where information is re-keyed between systems and where the shadow spreadsheets live. This produces the baseline that every later claim will be judged against, and it is the step most often skipped because it produces no software.
  2. Days 16-30 — Write down the rules that are not written downSit an analyst beside caseworkers and extract the working practice that published guidance does not cover: accepted document substitutes, local discretion, escalation habits. Turn it into a decision specification, expose the places where teams differ, and have the policy owner sign it. If they will not sign it, you have found the real blocker in week four rather than week thirty.
  3. Days 31-50 — Build and run entirely in shadowThe agent processes live cases in parallel with the existing team, and its output reaches no applicant and enters no system of record. Compare its completeness assessments and recommendations with what the caseworkers actually did, case by case. Expect the first week to be poor and the disagreements to be informative about the specification more often than about the software.
  4. Days 51-75 — Human-in-the-loop go-live on a narrow sliceRelease one process, for one case type, with a named officer committing every output and able to reject it in one action. Log every override with a reason. Keep the slice small enough that the team can absorb a bad week, and resist the pressure to widen it the moment the numbers look good — early numbers on easy cases always look good.
  5. Days 76-90 — Write the record, then decide what stopsProduce what the auditor and the SRO will want: what the agent did, what a human changed, the error profile against the baseline, and the cases it refused to determine. Publish it internally, and, where a transparency obligation applies, externally. Then make an explicit decision about what widens, what stays human, and what is abandoned. A project with nothing in the abandoned column has not been honest with itself.

What we read

The documents behind this note. Each entry says what it is, what it found, and why it should change what you do — then the link to the original.

National Audit Officesource 1 of 5

The Access to Work scheme

What it is
NAO value-for-money audit of the DWP scheme that funds workplace support.
What it says
Processing averaged 109 days against a 25-day target by November 2025, with 62,100 applications awaiting a decision and reconsiderations taking 241 minutes.
Why it matters
It is the audited basis for treating rework rather than judgement as the driver of the queue, and it records the re-keying and spreadsheet workarounds.
Read the original →
National Audit Officesource 2 of 5

An analysis of the asylum system

What it is
NAO analysis of asylum casework capacity, decision quality and appeal volumes.
What it says
Caseworker attrition was almost 35% in 2024-25, new staff take seven months to become fully effective, and 42% of sampled decisions had significant errors.
Why it matters
It shows capacity and quality moving while the rules stay fixed, and prices the downstream demand: 68% of initial refusals went on to appeal.
Read the original →
Ministry of Housing, Communities and Local Government (MHCLG Digital)source 3 of 5

Extract: unlocking England's planning data

What it is
MHCLG Digital post on Extract, its tool for turning planning documents into data.
What it says
Extract converts historic planning maps and documents into standardised data, supports officers rather than replacing them, and has them confirm its accuracy.
Why it matters
Government's own tooling draws the assist-not-replace boundary, which is the posture to hold from day one rather than from the day the auditor arrives.
Read the original →
OECDsource 4 of 5

Adopting and governing AI in government, Digital Government Outlook

What it is
OECD survey chapter on how 36 member countries adopt and govern AI in government.
What it says
AI is used in at least one area of government in 35 of 36 countries, but only 21 of 36 provide central support for procuring AI goods and services.
Why it matters
It sizes the gap between running a pilot and reaching a production contract, which is where public-sector projects most often lapse.
Read the original →
European Unionsource 5 of 5

Regulation (EU) 2024/1689 (AI Act), Annex III, point 5(a)

What it is
Official EUR-Lex text of the AI Act, including the Annex III high-risk list.
What it says
Point 5(a) classes as high-risk any system used by or for public authorities to evaluate eligibility for essential public assistance benefits and services.
Why it matters
Eligibility work in the EU carries logging, human oversight, data governance and registration duties, so classification comes before any design choice.
Read the original →

The full paper

The paper sets out the decision specification method we use to turn published guidance and unwritten caseworker practice into an auditable rule set, the shadow-running protocol for validating an agent against live cases without touching an applicant, and the evidence pack template that answers what an auditor, an ombudsman and a judicial review will each ask for.