Skip to main content
Back to Read
Legal AI27 September 20264 min read

Legal AI evaluation: a practical checklist and synthetic test

Evaluate a legal AI workflow with fictional documents, a separate answer key, failure cases and a score sheet for accuracy, permissions and review effort.

Sketched sheets of paper fanned out, one torn and one blurred and two alike, with a drawn magnifying glass and a separate folded answer-key sheet.
Editorial illustration.

A useful demonstration includes the awkward inputs: a missing year, two conflicting dates, an unreadable page and a duplicate document. It shows whether the workflow preserves uncertainty and sends unresolved work to a person.

Use wholly fictional material first. Give the reviewer an expected-result sheet written before the system runs. That makes errors visible without treating real client files as convenient demonstration data.

Define the result before the demonstration

  1. 01Write the task
  2. 02Create fictional sources
  3. 03Keep the answer key separate
  4. 04Run and preserve the output
  5. 05Review errors and effort
  6. 06Record the next decision

Specify one task. For example: create a document register and propose chronology entries with a document identifier, version, page and the date exactly as supplied. Exclude legal conclusions, credibility assessments and external actions from this exercise.

Our synthetic disclosure pack contains five invented sources and a manually authored answer key. It has not been run through a model, and it supplies no accuracy or time-saving result. It is a starting exercise for evaluating behaviour.

Include failures that matter to the workflow

Input or eventExpected behaviour in this proposed exercise
Two sources give different datesPreserve both attributed accounts and flag the conflict
An event has no yearRetain the date text and the missing-year warning
A page cannot be readKeep it in the inventory and require manual review
A document repeats anotherMark a duplicate candidate and retain both sources
A user lacks matter accessWithhold the restricted source and its content
The reviewer is unavailableKeep the output pending and use the agreed escalation route

The downloadable pack covers the document examples. Permission and reviewer-availability checks require a separately configured test environment; reading the pack cannot establish that those controls work.

Choose failure cases that fit the practice. The immigration evidence workflow covers missing translations and inconsistent dates; the medical chronology workflow shows why a qualified observation must not become a definite diagnosis.

Keep inputs and expected answers separate

Give the system the task and source material. Keep the answer key for the evaluator. If the system receives the expected answers in its prompt, the exercise no longer demonstrates how it handles unfamiliar input.

Record the product and version where available, relevant settings, instructions, source versions and the output. Repeat the run to see whether material differences occur. Repetition is useful evidence, but a few successful runs do not establish a failure rate for all future matters.

Retain failed attempts in the report. If somebody changes the prompt after seeing an error, record the change and distinguish the revised run from the original result.

Measure the whole task

Count unsupported facts, missing entries, wrong references, unresolved exceptions and reviewer corrections. Measure preparation, checking and correction separately from automated processing or waiting time.

For a comparison with manual work, use comparable material and a consistent definition of completion. Familiarity with the answer key can make a later run artificially easier. Record who reviewed which sample and any limitations of the comparison.

Do not combine every issue into one reassuring score. A permission failure or an invented critical fact deserves an explicit decision even if the other entries are correct.

Use a result sheet that retains failed runs

FieldWhat to record
Run and configurationDate, product, available version, settings and instruction version
Input and expected resultSource IDs, task and evaluator's separate answer key
Observed resultCorrect entries, omissions, unsupported facts and wrong references
Human effortPreparation, checking and correction time
DecisionAccept for the next stage, revise and retest, or stop—with a named owner

For an invented ten-entry exercise, “nine entries correct” is not enough information. If the tenth exposes a restricted file or invents a material fact, that failure requires its own decision. Do not turn a small convenience sample into a general accuracy percentage.

The selection scorecard answers the earlier question of which task is worth testing. This result sheet records what happened once the test was chosen.

Agree what would stop progression

Before the exercise, have the responsible people agree which failures block a live pilot and which may be corrected and retested. Examples include unauthorised access, silent loss of a source or an unapproved external action. The appropriate criteria depend on the task and its consequences.

The SRA's supervision guidance requires appropriate human scrutiny of AI-assisted work and responsibility retained by an authorised individual. Passing an exercise does not transfer that responsibility to the software. SRA: effective supervision.

A decision to proceed should name the permitted use, reviewer, limits and monitoring arrangements. Supplier and data approval remain separate requirements; the confidentiality review helps prepare those questions. The legal hub and first-workflow guide help connect the evidence to a practical implementation decision.

Focused first step

Clarity before complexity

Get unstuck

We start with a focused clarity chat so the report is based on your real bottlenecks, current situation, and commercial priorities.

Full digital presence audit
AI opportunity assessment
Custom growth roadmap
Report after your clarity chat
Get unstuck

The report is prepared after the chat if there is a sensible fit.