ai contract review

Is AI Contract Review Accurate? What the Benchmarks Actually Measure

Adira EditorialLegal AI desk13 min read

Is AI contract review accurate? Accurate at a narrower job than the headlines suggest. The study behind most of those headlines measured something more specific than "AI reviews contracts better than lawyers." In 2018, LawGeex ran a well-designed test showing its AI hitting 94% accuracy spotting issues in NDAs, against 85% for experienced human lawyers. That result is real, and also routinely stretched into a claim it never made. This guide (published by Adira, which builds AI-assisted contract review and CLM software, so we have a commercial interest here, but this explainer is written to hold up on its own) walks through what the LawGeex study measured, what it did not, where AI review is genuinely reliable today, where it still fails on Indian law, and a test you can run yourself instead of trusting anyone's marketing, including ours.

What the LawGeex study actually measured

The study, run with academic advisors Dr Roland Vogl of Stanford Law School's CodeX centre and Professor Gillian Hadfield of USC, pitted LawGeex's AI against 20 US-trained corporate lawyers with backgrounds at firms and companies including Alston & Bird, K&L Gates, Goldman Sachs and Cisco. Each side reviewed the same five previously unseen NDAs, containing 153 paragraphs of legal text, and had to spot issues against a fixed rubric of roughly 30 defined risk categories, things like a missing mutuality clause, an unusual indemnification obligation, or an unlimited liability position.

The AI scored 94% accuracy. The lawyers averaged 85%, ranging from a low of 67% to a high of 94%, meaning the single best human lawyer tied the machine. The AI finished in 26 seconds; the lawyers took between 51 and 156 minutes, averaging 92 minutes.

The accuracy figure itself was calculated as an F-measure, the harmonic mean of precision and recall against the answer key. Precision asks: of everything flagged, how much was a real issue? Recall asks: of the real issues actually present, how many did you catch? A tool can score well by being cautious (high precision, weak recall, it misses things but rarely cries wolf) or aggressive (high recall, weak precision, it catches everything but drowns you in false alarms). The F-measure blends both into one number, useful for a controlled study but it hides the trade-off you actually care about in practice.

What the study did not measure

This is the part that gets dropped when the 94-vs-85 number gets recycled into a sales pitch. The task was closed-vocabulary issue-spotting on a single, highly standardised document type, an NDA, against a rubric the researchers had already defined. That is a real skill, but it is not the same as:

  • Novel or bespoke drafting. The lawyers were not asked to draft anything, negotiate a position, or judge what a fair resolution looked like. Spotting a missing mutuality carve-out is not the same task as deciding your walk-away position on liability in a live deal.
  • Judgement calls under ambiguity. A defined rubric means "right" and "wrong" were settled in advance. Real contract review is full of judgement calls where competent lawyers reasonably disagree, materiality, whether a risk is worth fighting for given the counterparty and deal size.
  • High-stakes, heavily negotiated agreements. An NDA is low-value and high-volume by design; the study's own framing was about "the review and approval of low-value, high-volume, day-to-day business contracts." Nothing in the result tells you how AI performs on a contested M&A share purchase agreement.
  • Real working conditions. Professor Hadfield herself noted the test conditions were close to ideal for the lawyers, no interruptions, no fatigue from the tenth contract of the day, which if anything means the study may understate the AI's speed advantage in a real office, while telling you nothing about accuracy under those conditions.

So the honest read of the LawGeex result is narrower and more useful than "AI beats lawyers": on a closed, well-defined issue-spotting task over a standardised document, a good AI tool can match or exceed the average accuracy of a busy human doing the same narrow task, dramatically faster. That is a genuinely strong result. It is not evidence that AI can replace legal judgement, and it says nothing at all about AI drafting a contract from scratch, a separate question with its own accuracy problems, covered in Can AI Draft a Contract?.

Where AI review is genuinely accurate

Three things AI does reliably well, and they map closely to what the LawGeex rubric actually tested:

  1. Consistent issue-spotting against a fixed checklist. If you tell an AI tool exactly what to look for, a missing termination clause, an unlimited liability cap, an assignment clause without consent language, it will apply that checklist the same way on document one and document three hundred. A tired human reviewer will not.
  2. Never getting fatigued. The lawyers' scores in the LawGeex study ranged from 67% to 94% on the exact same task under the exact same conditions. That spread is mostly humans being humans, some more careful, some rushed, some distracted. AI does not have a bad Friday afternoon.
  3. Catching clauses a skim-reader misses. On a long, repetitive contract, AI is genuinely good at flagging the one paragraph in section 14 that quietly changes governing law, because it reads every paragraph with the same attention it gave the first one.

Where AI review still gets it wrong

Four failure modes matter, and none of them show up in a closed-rubric NDA study:

Context. AI reviewing a clause in isolation cannot always tell whether an unusual term is a problem given who the counterparty is, what leverage each side has, or what was already agreed in a side letter. A "red flag" clause in a vendor contract with a Fortune 500 buyer may be entirely normal industry practice; the same clause with an unknown counterparty may be genuinely risky.

Materiality. A rubric-driven tool tends to flag deviations, not weigh them. It will flag a 45-day payment term with the same visual urgency as an uncapped indemnity, unless you have specifically taught it to weight severity. Frequency of flagging is not the same as importance.

India-specific law. This is the sharpest gap, because most general-purpose AI models are trained overwhelmingly on US and UK contract text and case law. Two concrete examples:

  • Non-compete clauses. Under Section 27 of the Indian Contract Act, 1872: "Every agreement by which any one is restrained from exercising a lawful profession, trade or business of any kind, is to that extent void," subject to narrow exceptions such as the sale of goodwill of a business. (Section 27, Indian Contract Act, 1872) A US-trained model defaults to analysing a post-employment non-compete for "reasonableness," duration, geography, scope, the test in most US states. In India, a blanket post-employment restraint is presumptively void, and reasonableness barely enters the analysis. A flag reading "consider narrowing the geography" instead of "this is very likely unenforceable in India" has applied the wrong legal framework entirely.
  • IP assignment clauses. Section 19(5) of the Copyright Act, 1957 provides that "If the period of assignment is not stated, it shall be deemed to be five years from the date of assignment," and Section 19(6) presumes the territorial extent to be India, unless the assignment states otherwise. (Section 19, Copyright Act, 1957) A model trained on US practice, where a silent assignment is usually read as broad and perpetual, tends to read a silent Indian clause the same way, backwards from what the statute says.

Hallucinated flags. AI tools built on large language models can generate a flag that quotes or references a sentence that is not actually in your document, or misattributes a clause number. This is not a rare edge case; it is a known behaviour of the underlying technology, and it is the single most dangerous failure mode because it looks exactly as confident as a correct flag.

There is also a document-level issue text-only review cannot see at all: whether the contract is validly executed under Indian law. Section 35 of the Indian Stamp Act, 1899 says an instrument chargeable with stamp duty "shall [not] be admitted in evidence for any purpose... unless such instrument is duly stamped." (Section 35, Indian Stamp Act, 1899) A contract can sail through clause-by-clause AI review with zero flags and still be functionally unenforceable in court because it was never properly stamped, a fact that lives outside the text of the clauses entirely.

Signs your AI review is trustworthy, and signs it is not

NormalRed flagWhy it matters
Every flag quotes the exact sentence or clause number from your documentFlags reference "the confidentiality clause" with no quote or citationAn unverifiable flag cannot be checked against your actual document, and is the first sign of a hallucination
Non-compete or restraint-of-trade flags mention Section 27 voidability in IndiaNon-compete analysed only on duration and geographic reasonablenessThat is the US test; applying it in India means the flag is built on the wrong legal framework
The tool distinguishes "unusual" from "high risk," with a stated reasonEvery deviation from a generic standard is marked equally urgentOver-flagging buries the two or three issues that actually matter under noise you will start ignoring
Execution and stamping are checked as a separate item, not assumedThe tool only ever discusses clause text, never mentions stamping or registrationAn unstamped agreement can be inadmissible in evidence in India regardless of how well the clauses are drafted
Numeric and defined-term consistency is checked (e.g. "30 days" vs "one month" used for the same thing)These silent inconsistencies are missedThis is the easiest class of error for a machine to catch; missing it suggests shallow review
The tool states uncertainty on genuinely ambiguous pointsThe tool answers every question with identical, flat confidenceReal contracts have grey areas; uniform certainty looks more like templated output than real analysis

A bad AI flag versus a better one

Bad: "The Non-Compete clause may pose an enforceability risk depending on jurisdiction. Consider negotiating the duration and geographic scope to be more reasonable."

What is wrong: no clause citation, no quoted text, and it applies a US-style reasonableness test to a clause governed by Indian law, where reasonableness is largely beside the point.

Better: "Clause 14.2 (Non-Compete): 'Employee shall not, for 24 months following termination, engage in any competing business within India.' Under Section 27 of the Indian Contract Act, 1872, a post-employment restraint of this kind is void to the extent it restrains a lawful profession, trade or business, subject to narrow exceptions such as the sale of a business's goodwill. As drafted, this clause is very likely unenforceable if the employment relationship is governed by Indian law. Recommend removing it or replacing it with a narrower, better-supported non-solicitation obligation instead."

What changed and why: the rewrite cites the exact clause and quotes the operative text, applies the correct statute for the governing law at issue, states a clear conclusion instead of a hedge, and gives an actionable next step.

How to measure accuracy for yourself

Do not take a vendor's benchmark, including ours, as the answer for your contracts. Run this instead:

  1. Pull 15 to 20 of your own past contracts of one type, say vendor MSAs, that a lawyer already reviewed, and note every real issue that review found.
  2. Run the same contracts through the AI tool with the same instructions you would give a junior reviewer.
  3. For each contract, count true positives (real issues the AI caught), false positives (things flagged that were not actually issues under your playbook), and false negatives (real issues it missed).
  4. Calculate precision as true positives divided by (true positives plus false positives), and recall as true positives divided by (true positives plus false negatives).
  5. Repeat the test whenever you change the tool, the prompt, or your playbook, and keep a running log so a drop in either number is visible immediately, not discovered after a bad flag slips through.

You do not need paid software to start this. You can mark up a first batch of clauses and note where the AI's flags matched or missed reality directly in your document, for free, in Weave, before deciding whether a fuller AI review workflow is worth paying for. For the fuller, disciplined workflow this test feeds into, structuring the ask clause by clause and demanding a citation for every flag, see How to Review a Contract With AI.

US and global contrast

The LawGeex study itself was run on US lawyers reviewing NDAs under a US-influenced rubric, and most general-purpose AI tools are trained overwhelmingly on US and UK legal text, which is exactly why the India gap above exists. On non-competes specifically, most US states use a reasonableness balancing test, weighing duration, geography and scope, though a handful of states, California among them, also treat most post-employment non-competes as void, closer to the Indian position than to the majority US rule. An AI tool defaulting to "reasonableness" language on an Indian contract is not wrong because reasonableness never matters anywhere, it is wrong because it silently imported the majority US test into a jurisdiction that does not use it.

FAQ

Does the 94% LawGeex number mean AI is generally more accurate than a lawyer? No. It means AI matched or beat the average human score on one narrow, closed-rubric issue-spotting task over standardised NDAs. It says nothing about drafting, negotiation, or performance on high-stakes or unusual contracts.

Can I trust an AI review tool's output without checking it myself? Not for anything that matters. Treat AI review as fast, structured triage that tells you where to look closely, and verify any flag with legal or financial consequences against the actual clause text yourself.

What is the difference between precision and recall in contract review, in plain terms? Precision tells you how much of what got flagged was actually a real issue, low precision means false alarms. Recall tells you how much of the real issues present actually got flagged, low recall means things slip through unseen. A useful tool needs to be reasonable on both.

Does AI review understand Indian law by default? Not reliably. Most general-purpose models default to US-style legal reasoning unless you specifically prompt for Indian statutory positions, which is why a non-compete or an IP assignment clause can get flagged using the wrong legal test entirely.

What is a hallucinated flag, and how do I catch one? A hallucinated flag quotes or references text that is not actually in your document, or misattributes it to the wrong clause. Catch it by spot-checking flags against your document text, and by occasionally asking the tool to find an issue you know is not present; if it invents one, treat that session's other output with more scepticism.

Is AI review reliable for a high-stakes or heavily negotiated contract? Use it as a fast first pass even there, it is genuinely good at catching missed clauses and inconsistent terms, but do not rely on it for the judgement calls or Indian-law-specific analysis a contract at that level actually needs. That is a job for a lawyer who knows your deal.

This guide gets you to a working, checkable understanding of what AI contract review accuracy studies actually show and how to test a tool against your own contracts. It does not tell you whether a specific AI review of a specific contract you are relying on is safe to act on, that depends on the tool, the contract, and the stakes involved, and is not legal advice. Have a lawyer review anything before you sign or walk away from it on the strength of an AI flag alone.

Frequently asked questions

Does the 94% LawGeex number mean AI is generally more accurate than a lawyer?
No. It means AI matched or beat the average human score on one narrow, closed-rubric issue-spotting task over five standardised NDAs, judged against roughly 30 predefined risk categories. It says nothing about drafting, negotiation, judgement calls, or performance on high-stakes or unusual contracts, which the study never tested.
Can I trust an AI review tool's output without checking it myself?
Not for anything that matters. Treat AI review as fast, structured triage that tells you where to look closely, and verify any flag with legal or financial consequences against the actual clause text yourself, especially given the known risk of hallucinated flags that cite text not actually in your document.
What is the difference between precision and recall in contract review, in plain terms?
Precision tells you how much of what got flagged was actually a real issue; low precision means you are wading through false alarms. Recall tells you how much of the real issues present actually got flagged; low recall means things are slipping through unseen. The LawGeex accuracy figure was an F-measure, the harmonic mean of both, which is useful for a single study score but hides the trade-off you should care about in practice.
Does AI review understand Indian law by default?
Not reliably. Most general-purpose AI models are trained overwhelmingly on US and UK legal text, so they default to US-style legal reasoning unless you specifically prompt for Indian statutory positions. That is why a non-compete clause can get analysed for 'reasonableness' instead of flagged as presumptively void under Section 27 of the Indian Contract Act, 1872, or why a silent IP assignment clause can get read as perpetual and worldwide instead of the deemed five-year, India-only default under Sections 19(5) and 19(6) of the Copyright Act, 1957.
What is a hallucinated flag, and how do I catch one?
A hallucinated flag is one that quotes or references text that is not actually in your document, or misattributes an issue to the wrong clause, a known behaviour of large-language-model-based tools. Catch it by spot-checking flags against your actual document text, and by occasionally asking the tool to find an issue you know is not present; if it invents one, treat that session's other output with more scepticism.
Is AI contract review reliable for a high-stakes or heavily negotiated contract?
It is still useful as a fast first pass, it is genuinely good at catching missed clauses and inconsistent terms, but it should not be relied on for the judgement calls or Indian-law-specific analysis that a high-stakes contract actually needs. That remains a job for a lawyer who understands the deal. To measure accuracy on your own contracts rather than trust a vendor's benchmark, pull 15 to 20 past contracts of one type that a lawyer already reviewed, run them through the tool, and calculate precision and recall against the lawyer's known findings.
Was this useful?

See how Adira drafts in your voice and reads contracts from your side.

Explore the showroom

Working through a contract like this? Weave is Adira’s free tool to read, mark up, and connect any contract in your browser — no account needed.

Try Weave — free