ai contract data privacy

Does AI Legal Software Train on Your Contracts? (What to Ask)

Adira EditorialLegal AI desk13 min read

Every AI legal tool asks you to upload something confidential: a contract, a term sheet, a clause you are negotiating. The question that should come first is simple: does this tool learn from my document, in a way that could shape what it tells other customers, or does it just read it, answer, and forget it? Most vendors answer this vaguely, in marketing language that sounds reassuring and commits to nothing. (Adira, which publishes this guide, makes AI-assisted contract review and CLM software, so we sit on one side of this question ourselves. Adira states that it does not train any AI model on customer contract data. Do not take that on our word here either, verify it on Adira's security page and get it in writing, the way this guide asks you to verify every vendor's claim.) Below: what "training" actually means, what Indian law says about it, a real case working through this exact question in Delhi High Court right now, and the precise questions to put to a vendor before you upload a contract.

Training on your data versus processing your data

These are different things, and vendors lean on the confusion.

Processing is what happens when you ask a tool to summarise a clause, flag a risk, or answer a question about a document you uploaded. The model reads your text, generates an answer, and the interaction ends. Done properly, your document never becomes part of the model, it is input for one session, nothing more.

Training (including the narrower case of fine-tuning) is different in kind, not degree. It means your document's content is used to adjust the model's internal parameters, its weights, so the model's future behaviour, for you and possibly every other user, has been shaped by what it saw in your contract. A model trained on enough uploaded contracts could, in principle, reproduce a distinctive clause or commercial term from one customer's document in an answer given to a different customer. That is the risk a "no-training" promise is meant to rule out.

A third mechanism gets conflated with training: retrieval-augmented generation (RAG), where a system stores your document, or a mathematical representation of it, an "embedding", and pulls relevant chunks into a prompt each time you ask a question. RAG does not change the model's weights, so it is not training in the technical sense, but a copy of your data sits somewhere after the session ends. If a vendor says "we don't train on your data," ask separately whether they retain it for retrieval, and for how long. Those are two different promises, and answering only one has not answered your question.

Two vendors, two promises: the app and the model underneath it

Almost no legal AI product built the underlying language model itself. Vendors like Adira build the workflow and interface, then call a foundation model from a provider such as OpenAI, Anthropic, or Google. Two separate parties could, in theory, train on your data, and you need a written answer from both.

The app vendor's own promise covers whether the company you signed up with uses your documents to train its own models or fine-tune a base model for its product. This is the promise most legal-AI marketing pages address, and the one worth pinning to the contract, not a webpage.

The model provider's promise covers whether the underlying model itself retains or trains on data sent to it through the API. Major providers generally treat API traffic differently from their free consumer products: OpenAI states that API data is not used to train its models by default, and offers eligible enterprise customers a Zero Data Retention option, a meaningfully different policy from a free consumer chatbot, where content may be retained unless you opt out. Policies change, so ask your vendor which model provider it uses and get the current arrangement in writing rather than assumed from reputation. A vendor that cannot tell you what its own upstream provider does with your data has not answered the question.

The Indian legal backdrop: thinner than most buyers assume

A contract usually carries two different kinds of content, and Indian law treats them very differently.

Personal data inside the document (names, emails, signatures, sometimes ID numbers) falls under India's data protection framework. Two layers apply. The Digital Personal Data Protection Act, 2023 (DPDP Act) is notified but on a staggered commencement through 14 May 2027, and defines a "Data Fiduciary" who decides the purpose of processing and a "Data Processor" (Section 2(k)), permitted to process "only under a valid contract" under Section 8(2). Consent under Section 6(1) must "signify an agreement to the processing... for the specified purpose and be limited to such personal data as is necessary for such specified purpose"; using a document's personal data for a different purpose, such as training a general model, is exactly the kind of purpose expansion this section is built to catch once the Act's operative provisions come into force. (Full detail: Data Protection Clauses in Indian Contracts Under the DPDP Act.)

Independently, and in force today, the Information Technology Act, 2000 governs "sensitive personal data or information." Section 43A makes a body corporate liable to pay compensation where it is "negligent in implementing and maintaining reasonable security practices and procedures" and thereby causes "wrongful loss or wrongful gain to any person." (Section 43A, IT Act, 2000) The rules under that section, in force since 2011, are more directly on point: Rule 5(5) states that "the information collected shall be used for the purpose for which it has been collected." (Rule 5, IT SPDI Rules 2011) A customer who uploads a contract for review has given that data for review, not model training, a different purpose under this rule's own words.

The commercial terms in the document (pricing, deal structure, counterparty identity, negotiated positions) are a different story. Unless they identify a specific individual, they are not "personal data," and India has no general statute protecting confidential business information the way the DPDP Act protects personal data. That protection comes entirely from whatever confidentiality and data-use clause you negotiated. If it is silent on AI training, no statutory backstop fills the gap. This is the single most important fact here: for the part of a contract that is not personal data, the written promise in your vendor agreement is doing all the legal work.

A live fight over exactly this question: ANI Media v OpenAI

You do not have to take this on faith. Indian courts are actively deciding, right now, whether training an AI model on someone else's content without a licence is lawful, and the answer is neither settled nor generous by default to the party doing the training.

Asian News International (ANI) sued OpenAI in the Delhi High Court, alleging that ChatGPT was trained on its copyrighted news content without a licence, and sought an injunction plus roughly Rs 2 crore in damages. On 24 July 2026, after 32 hearings, Justice Amit Bansal declined to grant ANI's interim injunction, holding, on the facts before the court, that the use fell within India's fair-dealing exception. Section 52(1)(a) of the Copyright Act, 1957 excuses "a fair dealing with any work, not being a computer programme, for the purposes of private or personal use, including research," among other listed purposes. (Ani Media Pvt. Ltd v Open AI Opco LLC, CS(COMM) 1028/2024, Delhi High Court; judgment; statute.)

Two things matter here, and both cut against assuming the law will protect you. First, this is an interim order refusing an injunction, not a final verdict, and ANI can still appeal to a Division Bench. Second, the case is about published news assessed under a fair-dealing exception built around research and reporting, it says nothing about whether training on someone's unpublished, confidential contract would be treated the same way, that fact pattern has not been tested. The lesson is not "AI training is legal in India." It is that the question is contested and actively litigated. Do not rely on an unsettled fair-dealing argument to protect a contract you never intended as training data. Get an explicit, written promise instead.

The exact questions to ask a vendor

Put these in writing, in an email or an RFP, not a sales call, so the answers are on record:

  1. "Do you, or any sub-processor, use our documents, in whole or part, to train, fine-tune, or improve any AI model, including for other customers?" Cover the model provider as well as the vendor, and generated output as well as uploaded input, training on outputs is the gap most no-training clauses miss.
  2. "How long is our raw document retained after upload, and are embeddings deleted when the source document is deleted or the contract ends?"
  3. "Which foundation-model providers do you use, and what does your contract with them say about training on data we send through your product?" A Data Fiduciary stays responsible for what its Data Processor does; a vendor that will not name its own downstream processors is asking you to trust a chain you cannot see.
  4. "Where, geographically, is our data processed and stored, and will we get notice before that changes?" See Data Residency for Legal Software.
  5. "Can you share your current SOC 2 Type II report or ISO 27001 certificate, not just state that you hold one?"
  6. "Will you sign a Data Processing Agreement naming us as Data Fiduciary and you as Data Processor under the DPDP Act?" See Data Protection Clauses Under the DPDP Act.

Red flags

NormalRed flagWhy it matters
"We do not train any model on customer data, including via sub-processors""We use industry-leading AI to continuously improve our product"The second describes a benefit, not a data-use commitment, and is compatible with training on your documents
No-training commitment sits in the signed MSA or DPACommitment appears only on a marketing or trust-centre webpageA webpage can change without notice; only the signed contract binds the vendor
Vendor names its model provider(s) and states their training policyVendor will not disclose which model or provider it usesA training promise cannot be verified for a provider you are not told about
Clause covers "customer data" and "outputs generated using customer data"Clause covers only "documents you upload," silent on outputsTraining on the tool's own outputs is a real, separate mechanism; silence leaves a gap
Retention period and deletion process stated with a number"Data is retained as necessary to provide the service"An undefined period is not a limit, it is a placeholder that can mean indefinitely
Vendor shares a current SOC 2 or ISO 27001 report on requestVendor states it is "SOC 2 compliant," offers no reportA claimed standard nobody can inspect cannot be checked, and reports lapse
Sub-processor changes require notice or a right to objectVendor can add or change sub-processors silentlyYour data could reach a new AI provider, and training policy, without you knowing

Bad clause versus a better one

Bad: "Customer grants Vendor a licence to use Customer Data to operate, maintain, and improve Vendor's products and services."

What is wrong: "improve" is broad enough to include training a model, it is not limited to the vendor itself, and it says nothing about outputs, retention, or deletion.

Better: "Vendor shall not, and shall ensure that its sub-processors and any underlying AI model providers do not, use Customer Data, or any outputs generated using Customer Data, to train, fine-tune, or otherwise improve any AI or machine learning model, whether for Customer, any other customer, or general product improvement, without Customer's prior written consent. Vendor shall retain Customer Data only for so long as needed to provide the Services, shall delete Customer Data and any derived embeddings within 30 days of contract termination or a deletion request, and shall provide Customer with a current list of sub-processors, including underlying model providers, updated with 30 days' notice before any change."

What changed: "improve" is replaced with an explicit training prohibition reaching sub-processors and outputs, a deletion timeline replaces silence, and sub-processor visibility is built in rather than assumed.

Before you send a vendor's draft back, mark up its data-use clause against this checklist for free in Weave, Adira's browser-based contract tool, so you negotiate from a redline, not a vague objection.

US and global contrast

The EU's GDPR and newer EU AI Act are more prescriptive than the DPDP Act or IT Rules here. GDPR's Article 28 writes mandatory processor-contract terms directly into the statute, and the AI Act imposes transparency obligations on general-purpose model providers about training data. India's framework leaves nearly all of this to the contract: the DPDP Act covers personal data only, once in force, and no Indian statute requires disclosure of training practices for confidential content that is not personal data. The US has no single federal law here either, the picture is vendor-by-vendor, which is why the OpenAI API-versus-consumer-product distinction above matters. The result is similar everywhere for a buyer of legal AI software: the safety of your contract data depends far more on your vendor contract than on any privacy statute, so read the clause rather than assume the statute already did that work.

FAQ

If a vendor says its AI "learns" from my documents, does that always mean training? Not necessarily. "Learns" can mean the tool remembers context within your own account, closer to RAG, or can genuinely mean model training. Ask the vendor to say, in writing, whether your data updates model weights, and separately, how long it is retained for retrieval.

Does the DPDP Act stop an AI vendor from training on my contract? Only for the personal data inside the document, and only once its provisions are fully in force. Commercial terms, pricing, deal structure, counterparty positions, are not "personal data" and get no protection from the DPDP Act. Your confidentiality clause protects that content.

Does the ANI v OpenAI case mean AI training on copyrighted or confidential content is legal in India? No. It is an interim order about published news assessed under a fair-dealing exception, and ANI can still appeal. It does not decide whether training on someone's unpublished, confidential contract would be treated the same way, that fact pattern is untested.

What is the single most important document to ask for, beyond the contract clause? A current SOC 2 Type II report or ISO 27001 certificate you can read, plus a written, dated confirmation of the vendor's and its model provider's training policy. A website can change without notice; a signed clause is what you can enforce.

Should I trust Adira's claim that it does not train on customer contracts? Treat it as this guide asks you to treat any vendor's claim: verify it. Read Adira's security page, ask for the commitment in the signed DPA, and confirm which model providers Adira uses before relying on it.

This guide gets you to a working understanding of what "training on your data" means, what Indian law says and does not say about it, and the questions and clause language to demand before you sign. It does not tell you whether a specific vendor's contract meets your organisation's actual risk tolerance, that depends on facts a lawyer needs to review, and is not legal advice. Have a lawyer review the data-use and confidentiality terms in any AI vendor contract before you send it confidential material.

Frequently asked questions

If a vendor says its AI "learns" from my documents, does that always mean training?
Not necessarily. "Learns" can mean the tool remembers context within your own account, which is closer to retrieval-augmented generation (RAG) than training, or it can genuinely mean the document was used to update the model's weights. Ask the vendor to state, in writing, whether your data is used to update model weights, and separately, how long it is retained for retrieval purposes even if it is not used for training.
Does the DPDP Act stop an AI vendor from training on my contract?
Only for the personal data inside the document, such as names or emails, and only once the Act's substantive provisions are fully in force (commencement is staggered through 14 May 2027). Commercial terms in a contract, pricing, deal structure, counterparty positions, are not "personal data" under the Act and get no protection from it at all. Your contract's own confidentiality and data-use clause is what protects that content, since no Indian statute covers it.
Does the ANI v OpenAI case mean AI training on copyrighted or confidential content is legal in India?
No. It is an interim order refusing an injunction, in a case about published news content assessed under a specific fair-dealing exception in Section 52(1)(a) of the Copyright Act, 1957, and ANI can still appeal to a Division Bench. It does not decide whether training a model on someone's unpublished, confidential commercial contract would be treated the same way. That specific fact pattern has not been tested in an Indian court, so it should not be relied on as a general licence to assume AI training on private documents is safe.
What is the single most important document to ask a vendor for, beyond the contract clause itself?
A current SOC 2 Type II report or ISO 27001 certificate you can actually read, plus a written, dated confirmation of both the vendor's own training and retention policy and its underlying model provider's policy. A vendor's website or marketing page can change without notice; only a signed clause and a current, inspectable certification report are things you can actually check and enforce.
Should I trust Adira's claim that it does not train on customer contracts?
Treat it exactly as this guide asks you to treat any vendor's claim: verify it rather than accept it. Read Adira's security page, ask for the no-training commitment to be stated in the signed Data Processing Agreement rather than relying on a webpage, and confirm which underlying model providers Adira uses and what those providers' own training policies are before you rely on any of it.
Was this useful?

See how Adira drafts in your voice and reads contracts from your side.

Explore the showroom

Working through a contract like this? Weave is Adira’s free tool to read, mark up, and connect any contract in your browser — no account needed.

Try Weave — free