In the only public benchmark that pitted legal AI platforms against practising lawyers, one task ended in a dead heat. Chronology generation: Harvey 80.2%, lawyer baseline 80.2%. Every other task in the Vals Legal AI Report had a winner. This one had a tie, with one difference: the tool averaged under a minute; the lawyers had two weeks.

That tie is the most honest thing anyone has said about an AI case chronology. The machine is not better than you at reading 800 emails and putting them in order. It is exactly as good, and it does not get bored at email 412. What it cannot do is decide which emails matter; what it will do, if you let it, is invent a date the scan could not read. Hence the two disciplines below: a source link in every cell, and a lawyer who opens the cells.

Why chronologies suit AI: extraction, not judgement

A chronology is the purest extraction task in litigation: pull the facts out of the documents, date them, attribute them, sort them. No law, no authority, no argument. Every failure mode that ends up in a sanctions order (invented cases, misgrounded quotes, jurisdictional drift) is switched off by the shape of the task.

One lawyer, whose employer limits AI access, listed the wish on r/Lawyertalk: “take a set of medical records from across 7 years and lay out a chronological timeline so I can see exactly what doctor a party consulted, when, and what they complained about”. The judgement stays with you: which events are relevant, which contradiction matters, which treatment gap the defence will exploit. The table ensures you have seen everything before you decide.

The review-table method: one row per document, one column per question

Whatever the tool, the structure is the same. Each document is a row; each question you would ask a paralegal is a column; each cell holds the answer, quoted, with a reference to the page, Bates number or message ID.

Chronology extraction columns for a batch of documents
For each document in this batch, extract one row with these columns:
Date (ISO) | Time | Author | Recipients | Document type | One-line summary in neutral language | Mentions [key issue 1]? (Yes/No plus the quoted words) | Mentions [key issue 2]? (Yes/No plus quote) | Contains an admission, instruction or promise? (quote it) | Bates or file reference.
Quote; do not paraphrase in the "admission" column.
If a date is missing or inconsistent with the content, write DATE UNCERTAIN and say why. Never infer a date from a neighbouring document.
Sort all rows by date.

Three rules make the table trustworthy. Factual columns are verbatim. A missing value is marked, not filled: “DATE UNCERTAIN” and “NOT PRESENT” are permitted answers, invention is not. And every row carries a reference, because an entry you cannot trace to a page is an assertion, not evidence.

Harvey’s 800-email example: 4,000 data points in minutes

Harvey’s walkthrough shows the method at platform scale. Upload the emails to Vault, create a review table, add questions as columns, “type your questions just like you are asking a colleague the questions”: “Who is the sender?”, “Who is the recipient?”, and, in the example matter, “Does the email mention travel on Air Force One?”. Each column gets an answer type (verbatim, free response or classification); up to 50 columns are allowed. Run it, and 800 emails yield “4,000 data points at once in minutes”, each cell linked to its source email. Sort by date and verify through the links.

Everlaw’s Storybuilder builds chronologies inside its e-discovery platform, and at ILTACON in September 2026 Reveal, DISCO and Nuix all launched agentic chronology or timeline features.

The small-firm version: NotebookLM sources or a Claude Project

Most firms do not have a Vault. The same method runs on tools costing nothing or twenty dollars a month, with two trade-offs: source links are only as good as your prompt, and scale is limited.

Tool Source link per cell Scale Tier to use
Harvey Vault / Legora tabular review Automatic, every cell Thousands of documents, up to 50 columns Enterprise contract
Everlaw Storybuilder / AI Assistant Automatic, page and line Full e-discovery populations Enterprise
Gemini Notebook (formerly NotebookLM) Clickable citations, answers only from uploads 50 sources free, 300 on Google AI Pro Workspace or Enterprise account; no Google BAA cover
Claude Project / ChatGPT Project Only if the prompt demands a reference per row Long context, but accuracy degrades as input grows Team, Enterprise or API
Spreadsheet plus a model, in batches Manual: you paste the reference column Unlimited, one batch at a time Whatever tier holds the batch

Gemini Notebook suits a small-matter chronology because it answers only from the files you upload and cites them. Attorney at Work’s litigator workflow: upload everything, ask for a one-page case brief “nailed to a specific passage”, then build the timeline.

Dated timeline in Gemini Notebook (or any source-grounded tool)
Build a dated timeline of events with people, documents, and significant notes. Flag contradictions and missing links.
Rules: one row per event; Date | Event | People involved | Source document and page | Note. Cite a source for every row; if no uploaded source supports an event, leave it out.
Where sources give different dates or accounts of the same event, list both and mark CONTRADICTION.
Where an event is referred to but no document evidences it, mark MISSING LINK and name the document you would expect to exist.

The first sentence is Attorney at Work’s verbatim prompt; the rest makes it auditable. The NotebookLM guide covers account settings.

A Claude Project does the same for a few hundred documents, with one warning: Anthropic’s own documentation says that as token count grows “accuracy and recall degrade” (see the context-window explainer), so batch the documents and merge the tables. Then run a gap pass on the merged table: every period of more than seven days with no documents, every person who appears as a recipient but never as an author, every pair of rows describing the same event with different details, and every DATE UNCERTAIN row grouped by reason.

Medical-record chronologies: gaps, pre-existing conditions, duplicates

EvenUp’s guide puts manual medical-record review at “10 to 20 hours” per moderately complex case and describes three stages: OCR and extraction of dates, diagnoses, treatments and providers; chronology; then categorisation, duplicate flagging and narrative summaries. Adjusters look for treatment gaps and pre-existing conditions, exactly what a column can flag.

Medical-record chronology with gap analysis
From the attached medical records (OCR'd; PHI under our BAA-covered tool), build a chronology: Date | Provider | Encounter type | Complaint or diagnosis (quoted) | Treatment | Work restrictions | Billed amount | Page.
Then: (1) treatment gaps over 30 days, with the page before and after each; (2) every mention of a pre-existing condition, quoted; (3) inconsistencies between providers; (4) duplicate records; (5) medications with start and stop dates.
Total the billed amounts and show the arithmetic. Do not estimate general damages or settlement value.

Bank statements and financial discovery in divorce

Family lawyers run the same method on money. A divorce practitioner on r/Lawyertalk describes “feeding it a couple years’ worth of bank & credit card statements and telling it to look for specific payment types or recipients, then having it segregate & export that information into an Excel spreadsheet”.

Transactions of interest from bank statements to CSV
From the attached statements (anonymised: [ACCOUNT_1], [PERSON_A]), extract every transaction matching any of: payee or reference containing [keywords]; amounts over [X]; transfers to accounts not in <known_accounts>; cash withdrawals over [Y]; recurring payments to the same payee.
Output CSV: Date | Account | Payee or reference (verbatim) | Amount | Criterion matched | Statement page. Then a summary table by payee with totals and date range.
State how many transactions you processed and how many pages you could not read.

Two rules. Anonymise before upload with the placeholder-and-key method in the anonymisation guide, and keep anything about children out of general tools. And redo every total in Excel: models are weak at arithmetic. The family law guide has the workflow end to end.

OCR first, or you get hallucinated dates

A scanned PDF is a picture; the model reads whatever text layer sits on top, and if that layer is garbage it will confidently produce garbage with a date on it.

“I pre-process everything with basic OCR cleanup before it hits any AI tool otherwise you get hallucinated dates, mangled drug names, and unusable citations.” — /u/Specific_Citron_8546, r/legaltech

So: run OCR, spot-check a few pages of digits against the image, de-duplicate, and split email threads into individual messages, or a thread quoted five times becomes five rows with one date.

Verifying an AI case chronology cell by cell

The tie at 80.2% matters: roughly one entry in five, from the tool or from a lawyer working alone, needs correction. The tool’s errors are confident and evenly distributed; yours cluster where you were tired.

A summary has no cells to check. A table does. The routine: open every row marked as an admission, instruction or promise and read the quote in the original; open every document you will exhibit; resolve every DATE UNCERTAIN row by hand; sample the rest. Reed Smith’s protocol for generative review (validate on 50 to 100 documents, expand to 500 to 1,000, then sample the full population) transfers directly; the e-discovery guide sets it out. Log who checked what and when.

Presenting the timeline: exhibits and cross-references

A verified table becomes three things: an internal working chronology with a reference column, where deposition preparation starts (contradictions between testimony and dated documents are a filter on this table); a narrative statement of facts, drafted with a row citation after every sentence and rewritten in your voice without the adjectives the model adds; and an exhibit list, since every row already names its document and page.

Keep the reference column in whatever you file; opposing counsel now cite-check the other side’s work for a living. Building and verifying the table is an exercise participants run in class in AI Lab for Lawyers, on anonymised material, because the verification step is where the skill lives and where the tool demos stop.

Where to go next: the deposition prep guide takes the chronology into transcript analysis, the pillar on how lawyers use AI places it on the reliability ladder, and the prompt library has the full prompt set; the rest of the task guides are in the use cases hub.

Frequently asked questions

Can AI build a litigation chronology?

Yes, and it is one of the best-evidenced uses. In the Vals Legal AI Report of February 2025, Harvey scored 80.2% on chronology generation, exactly the lawyer baseline, and ran six to eighty times faster. The reliable method is a review table with one row per document and one column per question, every cell quoted and linked to its source, followed by a human check of every entry you will rely on.

How do I make a timeline from emails with AI?

Convert threads to individual messages, OCR any scans, and upload the set to a review-table tool (Harvey Vault, Everlaw, Legora) or a document-grounded one (Gemini Notebook, a Claude Project). Ask column-by-column questions: date, sender, recipients, one-line summary, whether it mentions each key issue, and any admission or promise, quoted. Sort by date, mark uncertain dates, and open every email you intend to exhibit.

Can ChatGPT create a medical chronology?

Technically yes; ethically only on a tier with a signed business associate agreement, which consumer ChatGPT plans reportedly do not offer. Medical records are protected health information. Use a PI-specific tool such as EvenUp or an enterprise deployment under a BAA, OCR the records first, and ask for date, provider, complaint, treatment and page reference per row, plus treatment gaps over 30 days. Manual review runs 10 to 20 hours per moderately complex case, so the saving is real.

Which tools generate chronologies with source links?

Harvey Vault review tables link every cell to its source document; Everlaw's Storybuilder and AI Assistant cite page and line; Gemini Notebook (formerly NotebookLM) answers only from uploaded sources with clickable citations. At ILTACON 2026 Reveal, DISCO and Nuix all launched agentic chronology features. A Claude Project or ChatGPT will link nothing unless your prompt demands a page or Bates reference in every row.

How accurate are AI chronologies?

As accurate as the extraction and the OCR beneath it. Vals measured Harvey at 80.2%, level with lawyers, which means roughly one entry in five needs correction from either source. Scanned documents without OCR cleanup produce hallucinated dates and mangled names. Treat the table as a draft index: every admission, every date you will rely on and every DATE UNCERTAIN row gets opened against the original before it goes into a brief.

Written by

Dr. Niklas Schmidt, Partner at Wolf Theiss

Partner at Wolf Theiss Attorneys-at-Law, where he heads the firm-wide tax team; lawyer, author, TEDx speaker and technologist. He has spent well over 1,000 hours testing practical AI applications for legal work, runs a toolkit of roughly 80 AI tools in daily practice, founded the WT Crypto Academy (1,000+ participating lawyers) and has given around 450 talks over 20 years. He teaches the live course AI Lab for Lawyers on Maven.