contract risk scoring

Contract Risk Scoring: What It Is and What It Misses

Adira EditorialLegal AI desk13 min read

Contract risk scoring takes a clause, or a whole contract, and assigns it a risk level, usually red, amber, or green, based on how far it deviates from a playbook or a known list of risky patterns. A red indemnity clause usually means "uncapped and one-sided." A green liability cap usually means "matches our standard position." The one thing people get wrong is treating the colour as a verdict rather than a prompt. A score tells you where to look first. It does not tell you whether the clause is actually a problem for this deal, with this counterparty, under Indian law. (Adira, which publishes this guide, builds contract review and CLM software, including risk scoring. We wrote this to be honest about what scoring does and does not do, whether or not you ever use ours.)

Risk scoring got popular for a good reason: a legal team reviewing forty contracts a month cannot read every clause with equal attention, and a score is a fast way to decide where that attention goes. But the same speed that makes scoring useful also makes it easy to over-trust. Below is how the mechanism works, what it is genuinely good for, and three specific ways it misses, including a red-green mistake Indian contract law makes unusually easy to fall into.

How risk scoring actually works

A risk-scoring engine is almost always one of two things, or a blend of both.

Rule-based scoring compares clause language against a playbook: pre-approved positions your legal team has written down. If your playbook says "liability cap must not exceed 12 months' fees" and a contract's cap says 24 months, the engine flags a deviation and scores it red or amber. This is deterministic and explainable, and only as good as the playbook: written for SaaS deals and run on a construction contract, half the flags will be noise.

AI-based scoring uses a language model to read the clause in context and judge risk against general patterns, not just your playbook. It can catch things a rule engine misses, an unusual definition that quietly narrows a warranty, for instance, but can be inconsistent between two similar clauses, and does not know your negotiating position unless told. Most commercial tools, Adira's engine included, blend both: rules for positions you have set, AI to catch what is not in the playbook and explain the deviation in plain language.

Either way, the mechanism is comparison against a standard someone chose in advance. That single fact explains almost everything the rest of this page covers.

What a score is actually useful for

Used correctly, a risk score does three jobs well: triage across a batch (forty contracts, two hours, a score tells you which five to open first, and even an imperfect score beats a random order), prioritising review inside one contract (a 20-page MSA might have three genuinely negotiable clauses and seventeen fine as drafted), and reporting to leadership ("reviewed 340 contracts, flagged 28 high risk, closed 25 before signature" fits a board update; a pile of PDFs does not).

None of these three uses require the score to be right about any single clause. They require it to be directionally useful across a large batch, a much lower bar that both rule-based and AI scoring clear reasonably well.

What it misses: context

A score compares a clause to a standard position. It does not know why this deal is different from the standard case the playbook was written for.

A liability cap of "fees paid in the preceding 12 months" might score amber on a playbook calibrated for high-value enterprise deals, correctly, because that cap is thin relative to typical exposure there. The same cap, on a low-value pilot contract where your actual exposure is small, might be entirely fine, but the engine has no way to know the deal is a pilot unless the playbook has been split by deal size, and most have not been.

The problem runs the other way too: an indemnity with a normal-looking cap can be genuinely dangerous if this counterparty has thin capitalisation and no insurance to back the promise. That fact lives in your CRM, not the contract text, so no clause-level score will ever see it.

What it misses: materiality versus frequency

A score usually treats "this pattern appears" as the trigger, not "this pattern will actually cost you money if it happens." Conflating the two produces two errors: over-flagging things that are common but low-stakes, and under-weighting things that are rare but catastrophic.

Liquidated damages clauses show why materiality needs its own analysis, not just pattern-matching. Section 74 of the Indian Contract Act, 1872 says that when a contract fixes a sum payable on breach, the complaining party is entitled to "reasonable compensation not exceeding the amount so named," whether or not actual loss is proved. Read the section on India Code or Indian Kanoon. A rules engine that flags "any liquidated damages clause without a stated methodology" as uniformly red is applying the wrong test. The Supreme Court in Oil and Natural Gas Corporation Ltd v Saw Pipes Ltd (2003) held that where a genuine pre-estimate of loss is hard to calculate and the parties have agreed a figure in advance, courts will generally enforce it as reasonable compensation without separate proof of actual loss, stepping in only where the figure looks like a penalty rather than a genuine estimate. What matters is whether the number reads as a punishment or a pre-estimate, a judgment a keyword match cannot make.

The frequency side runs the other way: a scorer that has seen "unlimited liability for IP infringement" carved out of the cap in most training contracts may treat that carve-out as low risk because it is common, when for a business that licenses its own software, that exact carve-out is the single largest number in the contract. Common is not the same as immaterial.

What it misses: India-specific enforceability

This is the sharpest miss, and the one most likely to embarrass a legal team that trusts the colour without reading the clause. A rules engine pattern-matches against structure, duration, and scope, things like "does a time period exist," "is geography defined." It is not, by default, checking a specific fact pattern against Indian statute.

Take a post-employment non-compete: "The Employee shall not, for a period of 12 months following termination, work for a directly competing business within India." To an engine calibrated on structure, this looks reasonable: bounded duration (12 months, not unlimited), defined geography (India, not worldwide), narrower than a worse version reading "24 months, worldwide, any affiliate." On those signals, plenty of playbooks would score this green, or at most amber, purely because it looks moderate next to the red version.

It is void. Section 27 of the Indian Contract Act, 1872 says plainly:

"Every agreement by which any one is restrained from exercising a lawful profession, trade or business of any kind, is to that extent void."

Read the section on India Code or Indian Kanoon. Unlike the reasonableness test used in many other jurisdictions, Section 27 does not ask whether 12 months is fair or India-only narrow enough. It asks one binary question: does the restraint operate after the relationship ends? If yes, it is void, regardless of how moderate the numbers look. The Delhi High Court applied exactly this reasoning on 25 June 2025 in Varun Tyagi v Daffodil Software Private Limited (FAO 167/2025), quashing an injunction against a departing employee and holding that a restrictive covenant operating after termination cannot be enforced, because Section 27 voids it outright regardless of how it is framed. A scoring engine that measures moderation on a spectrum, rather than checking the binary "during or after," will get this exact clause backwards, scoring the moderate-looking 12-month, India-only version green when a court would strike it exactly as fast as the worldwide, 24-month version.

This is not a one-off. The same gap shows up wherever Indian law uses a bright-line rule instead of a reasonableness spectrum: unstamped instruments under Section 35 of the Indian Stamp Act, 1899 are inadmissible regardless of how "standard" the stamp clause looks, and an unremarkable-looking unilateral variation clause can still be struck down as unconscionable under Section 23, the doctrine the Supreme Court applied in Central Inland Water Transport Corporation Ltd v Brojo Nath Ganguly (1986). Structural moderation and legal validity are different axes, and an engine trained only on the first will occasionally get the second exactly wrong.

Red flags in the scoring output itself

The table below is not about clause red flags, that ground is covered in How to Check a Contract for Red Flags. This one is about spotting when a risk score itself should not be trusted at face value.

NormalRed flagWhy it matters
Score explains which playbook rule triggered itScore with no visible reason, just a colourYou cannot calibrate what you cannot inspect; an unexplained score is a black box, not a tool
A moderate-looking non-compete still scores redShort duration, defined territory, but it scores greenSection 27 is binary (during vs after employment), not a spectrum; moderation does not equal enforceability
Liquidated damages flagged only when the amount looks disproportionateEvery such clause flagged red for "lacking a stated formula"Section 74 and ONGC v Saw Pipes ask whether the sum is a genuine pre-estimate, not whether a formula is spelled out
Score changes when you edit the playbook's stated positionScore never changes no matter what position you setMeans the tool is using a generic industry standard, not your actual negotiating position
High-risk flags cluster on clauses with real financial exposure (indemnity, liability, IP)High-risk flags cluster on formatting or defined-term inconsistenciesA scorer optimising for flag count, not materiality, will bury the clauses that actually cost money
Score is silent on stamping or jurisdiction-specific validityAn unstamped agreement scores the same as a properly stamped oneThese are pass/fail facts under Indian law, and a scorer blind to them will miss real exposure

Bad score, better score: the same clause, read two ways

The clause: "The Employee shall not, for a period of 12 months following the termination of this Agreement for any reason, directly or indirectly work for, consult for, or provide services to any business in India that competes with the Company."

The bad score (structure-only): A rules engine checks duration (12 months, bounded), geography (India, defined), and scope (competing businesses, not "any business"). Against a playbook that treats 24-months-worldwide as the red benchmark, this looks moderate. It scores amber, maybe green. Nothing in that pass asked the one question that actually decides the outcome under Indian law.

The better score (rule plus enforceability check): The engine asks a binary question before it measures degree: does the restraint operate during the relationship or after it? "Following the termination" answers that immediately, so the clause is flagged red, "post-termination restraint, void under Section 27 regardless of duration or geography," because the timing alone is dispositive, not because the numbers look aggressive. What changed is which question the score asked first. A structural pass asks "how extreme is this." An enforceability-aware pass asks "does Indian law allow this category of clause here at all," and only then looks at degree for categories where degree actually matters, like liability caps.

If you suspect a clause falls into this trap, paste the contract into Weave, Adira's free browser-based contract tool, and check the restraint language against Section 27 directly rather than relying on a generic colour.

How to calibrate a risk score to your own positions

A risk score is only as good as the playbook behind it, and most teams inherit a generic one rather than building their own. Three steps make a real difference:

  1. Write down your actual fallback positions, not your opening ask. A playbook encoding "we always demand uncapped indemnity" will flag every negotiated cap as a deviation, even ones your team is happy to accept.
  2. Segment by deal type and value. One liability threshold applied to both a ₹2 lakh pilot and a ₹2 crore enterprise deal will misfire on one every time. Use tiered playbooks if your tool supports them; otherwise treat scores on outlier-sized deals with extra scepticism.
  3. Add binary legal checks ahead of structural scoring wherever Indian law is binary, not a spectrum. Post-employment restraints (Section 27), unstamped instruments (Section 35), and clauses contracting out of a statutory right are pass/fail questions that should run before, not instead of, structural scoring.

Related reading: How to Do a Contract Risk Assessment walks through building the playbook this scoring depends on, and Is AI Contract Review Accurate covers the broader accuracy question for AI-assisted review, of which scoring is one piece.

US and global contrast

Risk scoring as a category originated largely in US and UK legal-operations practice, where playbooks are often built around a reasonableness spectrum because the underlying law is itself a reasonableness test. A US non-compete is commonly assessed on exactly the axes an Indian scoring engine wrongly applies here: is the duration reasonable, is the geography reasonable, is the scope no broader than necessary. That fits US non-compete law in the states that enforce them. It is the same spectrum logic, imported into an Indian playbook unchanged, that produces the green score on a void clause described above. The tool is not broken; it is answering the question US law asks, on a fact pattern governed by Indian law, which asks a different question.

FAQ

Can I trust a risk score to tell me a clause is fine if it comes back green? No, treat green as "did not match a known bad pattern," not "confirmed safe." A green score means nothing in the playbook caught it, which is different from a lawyer having reviewed it and agreed it is fine.

Should a legal team build its own playbook or use a vendor's default one? Start with a default to move fast, then replace it clause by clause with your actual fallback positions as you see real deals. An uncalibrated playbook flags things you do not care about and misses things specific to your risk profile.

Does AI-based scoring solve the enforceability problem rule-based scoring misses? Only if it is specifically prompted to check binary legal rules like Section 27 ahead of structural comparison. A general-purpose model asked to "assess risk" often defaults to the same reasonableness-spectrum thinking a rules engine uses. That is a design choice inside the tool, not something that comes free with "AI."

What is the single biggest category risk scoring gets wrong in an Indian context? Post-employment restraints, non-competes and broadly drafted non-solicitation clauses, scored on structure and moderation instead of the binary "during or after employment" question Section 27 actually asks.

How often should we re-check the playbook a scoring tool is using? At least every time a major deal type changes, a new jurisdiction is added, or a case like Varun Tyagi shifts how settled a category is. A playbook is a snapshot of legal positions at the time it was written, not a permanently correct standard.


This page explains what contract risk scoring measures, what it is genuinely good for, and where the mechanism itself, not any one vendor's implementation, tends to miss. It does not tell you whether a specific score on a specific contract in front of you is right. A risk score is a prompt to open the clause and look, not a verdict on whether it will hold up. For a clause where the score and your own reading disagree, or where real money is on the line, get a lawyer to look before you sign. This is not legal advice.

Frequently asked questions

Can I trust a risk score to tell me a clause is fine if it comes back green?
No, treat green as 'did not match a known bad pattern,' not 'confirmed safe.' A green score means nothing in the playbook caught it, which is different from a lawyer having reviewed it and agreed it is fine.
Should a legal team build its own playbook or use a vendor's default one?
Start with a default to move fast, then replace it clause by clause with your actual fallback positions as you see real deals. An uncalibrated playbook flags things you do not care about and misses things specific to your risk profile.
Does AI-based scoring solve the enforceability problem rule-based scoring misses?
Only if it is specifically prompted to check binary legal rules like Section 27 of the Indian Contract Act ahead of structural comparison. A general-purpose model asked to 'assess risk' often defaults to the same reasonableness-spectrum thinking a rules engine uses. That is a design choice inside the tool, not something that comes free with AI.
What is the single biggest category risk scoring gets wrong in an Indian context?
Post-employment restraints, non-competes and broadly drafted non-solicitation clauses, scored on structure and moderation instead of the binary 'during or after employment' question Section 27 actually asks. A moderate-looking version can score green while a court would strike it exactly as fast as an aggressive one.
How often should we re-check the playbook a scoring tool is using?
At least every time a major deal type changes, a new jurisdiction is added, or a case like Varun Tyagi v Daffodil Software (2025) shifts how settled a category is. A playbook is a snapshot of legal positions at the time it was written, not a permanently correct standard.
Does a risk score replace a lawyer's review?
No. It replaces the decision of where to spend limited review time first. Whether a specific flagged (or unflagged) clause actually holds up under Indian law is a judgment call the score does not make, especially for binary enforceability questions like Section 27 or Section 35 of the Stamp Act.
Was this useful?

See how Adira drafts in your voice and reads contracts from your side.

Explore the showroom

Working through a contract like this? Weave is Adira’s free tool to read, mark up, and connect any contract in your browser — no account needed.

Try Weave — free