Legal AI evaluation: a practical checklist and synthetic test
Evaluate a legal AI workflow with fictional documents, a separate answer key, failure cases and a score sheet for accuracy, permissions and review effort.

A useful demonstration includes the awkward inputs: a missing year, two conflicting dates, an unreadable page and a duplicate document. It shows whether the workflow preserves uncertainty and sends unresolved work to a person.
Use wholly fictional material first. Give the reviewer an expected-result sheet written before the system runs. That makes errors visible without treating real client files as convenient demonstration data.
Define the result before the demonstration
- 01Write the task
- 02Create fictional sources
- 03Keep the answer key separate
- 04Run and preserve the output
- 05Review errors and effort
- 06Record the next decision
Specify one task. For example: create a document register and propose chronology entries with a document identifier, version, page and the date exactly as supplied. Exclude legal conclusions, credibility assessments and external actions from this exercise.
Our synthetic disclosure pack contains five invented sources and a manually authored answer key. It has not been run through a model, and it supplies no accuracy or time-saving result. It is a starting exercise for evaluating behaviour.
Include failures that matter to the workflow
| Input or event | Expected behaviour in this proposed exercise |
|---|---|
| Two sources give different dates | Preserve both attributed accounts and flag the conflict |
| An event has no year | Retain the date text and the missing-year warning |
| A page cannot be read | Keep it in the inventory and require manual review |
| A document repeats another | Mark a duplicate candidate and retain both sources |
| A user lacks matter access | Withhold the restricted source and its content |
| The reviewer is unavailable | Keep the output pending and use the agreed escalation route |
The downloadable pack covers the document examples. Permission and reviewer-availability checks require a separately configured test environment; reading the pack cannot establish that those controls work.
Choose failure cases that fit the practice. The immigration evidence workflow covers missing translations and inconsistent dates; the medical chronology workflow shows why a qualified observation must not become a definite diagnosis.
Keep inputs and expected answers separate
Give the system the task and source material. Keep the answer key for the evaluator. If the system receives the expected answers in its prompt, the exercise no longer demonstrates how it handles unfamiliar input.
Record the product and version where available, relevant settings, instructions, source versions and the output. Repeat the run to see whether material differences occur. Repetition is useful evidence, but a few successful runs do not establish a failure rate for all future matters.
Retain failed attempts in the report. If somebody changes the prompt after seeing an error, record the change and distinguish the revised run from the original result.
Measure the whole task
Count unsupported facts, missing entries, wrong references, unresolved exceptions and reviewer corrections. Measure preparation, checking and correction separately from automated processing or waiting time.
For a comparison with manual work, use comparable material and a consistent definition of completion. Familiarity with the answer key can make a later run artificially easier. Record who reviewed which sample and any limitations of the comparison.
Do not combine every issue into one reassuring score. A permission failure or an invented critical fact deserves an explicit decision even if the other entries are correct.
Use a result sheet that retains failed runs
| Field | What to record |
|---|---|
| Run and configuration | Date, product, available version, settings and instruction version |
| Input and expected result | Source IDs, task and evaluator's separate answer key |
| Observed result | Correct entries, omissions, unsupported facts and wrong references |
| Human effort | Preparation, checking and correction time |
| Decision | Accept for the next stage, revise and retest, or stop—with a named owner |
For an invented ten-entry exercise, “nine entries correct” is not enough information. If the tenth exposes a restricted file or invents a material fact, that failure requires its own decision. Do not turn a small convenience sample into a general accuracy percentage.
The selection scorecard answers the earlier question of which task is worth testing. This result sheet records what happened once the test was chosen.
Agree what would stop progression
Before the exercise, have the responsible people agree which failures block a live pilot and which may be corrected and retested. Examples include unauthorised access, silent loss of a source or an unapproved external action. The appropriate criteria depend on the task and its consequences.
The SRA's supervision guidance requires appropriate human scrutiny of AI-assisted work and responsibility retained by an authorised individual. Passing an exercise does not transfer that responsibility to the software. SRA: effective supervision.
A decision to proceed should name the permitted use, reviewer, limits and monitoring arrangements. Supplier and data approval remain separate requirements; the confidentiality review helps prepare those questions. The legal hub and first-workflow guide help connect the evidence to a practical implementation decision.