Summary. How to negotiate for data and models without buying rights the seller does not have.
Before the first call
Answer three questions internally. Negotiations that skip them go badly.
What are we actually buying? Data, a model, a hosted service, or some combination. Name the layers: the corpus, the model artifact, and the outputs. Each is a separate grant with separate terms.
What will we do with it? Not "use it for AI." Specifically: train a model from scratch? Fine-tune an existing one? Run inference only? Retrieve at query time? Embed the model in a distributed product? Distribute the model itself? Each of these is a different permission, and licensors price them differently.
What happens if we lose it? If the license terminates in three years and we must stop using models trained on the data, what does that cost? If the answer is "the business," then model survival on termination is not a secondary term — it is the deal.
PART ONE: DILIGENCE
Step 1 — Ask for documents, not assurances
The most important shift in this negotiation is to stop accepting representations and start requiring deliverables.
For a data license, request:
- A source manifest: every source in the corpus, with collection method and date range.
- Archived copies of the governing terms as they existed at collection. Not a description — the stored pages.
- The contracts under which any contributed or purchased data was obtained.
- Documentation of any machine-readable restrictions encountered and how they were handled.
- A statement of what personal information is present and the lawful basis for it.
- The filtering, deduplication, and exclusion processes applied.
- Whether copyright management information was preserved.
- Any claims, demands, or takedown requests received relating to the corpus.
For a model license, request:
- The actual license text, not the marketing page.
- The acceptable use policy and any field restrictions.
- Scale thresholds and what triggers a separate license.
- Whether fine-tuning is permitted and who owns the result.
- Whether outputs may be used to train other models.
- Flow-down obligations to your customers and distributors.
- The indemnity, with all conditions and exclusions.
- The provider's rights over your inputs and fine-tuning data.
The tell. A licensor who can produce these quickly has a real program. One who cannot produce archived terms — the single most revealing request — is warranting something it cannot document, and the price should reflect that.
Step 2 — Test what the licensor can actually grant
Run the chain of title.
- Who created the material? If not the licensor, under what agreement?
- If scraped, what did the source terms permit? Terms of service prohibitions are enforceable as contract regardless of copyright.
- If contributed by users or customers, do the contributor terms permit sublicensing for training?
- If purchased, does the upstream license permit onward licensing? Many prohibit redistribution outright.
- If generated internally, was any third-party material used in the generation?
Expect gaps. Most corpora of any size contain components the licensor cannot fully document. The productive response is not to abandon the deal but to segregate: exclude the undocumented components, or price and indemnify them separately.
Step 3 — Identify the deal-breakers early
Some findings should change the transaction rather than the price:
- The corpus contains material the licensor cannot license at all (third-party subscription data, another customer's confidential information).
- Personal information with no lawful basis and no feasible de-identification.
- A model license with a field restriction covering your intended use.
- An indemnity whose conditions exclude your actual deployment (fine-tuning, filter configuration, version pinning).
- No ability to segregate the corpus, making termination unmanageable.
Surface these in week one. Discovering them at signature costs the relationship as well as the deal.
PART TWO: THE GRANT
Step 4 — Name the permitted acts separately
A grant to "use the Data for artificial intelligence purposes" allocates nothing. Enumerate:
- Training a model from scratch
- Fine-tuning an existing model
- Evaluation and benchmarking
- Retrieval at inference time (technically distinct and often licensed differently)
- Internal use of resulting models
- Embedding a model in a product distributed to customers
- Distributing the model itself
- Sublicensing to affiliates, contractors, or customers
Licensors price these differently and should. A licensee that pays for training rights and then distributes the model has exceeded the grant.
Step 5 — Define the data with precision
- Scope: which datasets, which versions, which date ranges.
- Updates: does the license cover future additions, and at what price?
- Format and delivery: including the provenance manifest as a deliverable, not a representation.
- Segregation obligation: the licensee must maintain records of which models used which corpus versions. This is what makes every other provision enforceable.
- Exclusions: components carved out, and the licensee's obligation not to use them.
Step 6 — Handle the model layer
For hosted models:
- Version commitments and deprecation notice
- Availability and performance commitments
- Whether the provider may use your inputs, prompts, and outputs — opt out expressly
- Whether your fine-tuning data may be used for service improvement
- Data residency and subprocessor terms
- What happens on termination to fine-tuned artifacts
For weight releases:
- Confirm what the license actually permits; open weights are usually not open source
- Field and acceptable-use restrictions, with an interpretation in writing if your use is near a boundary
- Scale thresholds with objective triggers
- Restrictions on using outputs to train other models
- Attribution and notice obligations, and how they flow down
- Ownership of fine-tuned weights and adapters
- Survival of the license for models already deployed
Step 7 — Draft the output provisions honestly
Do not represent that outputs are copyrightable. Do:
- Assign to the customer whatever rights the provider has in outputs.
- Disclaim any representation as to copyrightability, since United States law requires human authorship.
- Confirm commercial use without field restriction.
- Acknowledge non-exclusivity — other customers may receive similar outputs. Say it rather than letting the customer discover it.
- Address residual licenses: does the provider retain any right to use customer outputs?
- Allocate output infringement risk explicitly. This is the indemnity discussion.
PART THREE: RISK ALLOCATION
Step 8 — Negotiate the conditions, not the grant
Indemnity headlines are similar across providers. The conditions differ enormously and determine value.
For each condition, ask: does our actual deployment satisfy it?
| Condition | Question to ask |
|---|---|
| Current model version required | Can we upgrade on the provider's schedule? Our stability requirements may conflict |
| Safety and filtering features enabled | Are we disabling anything for latency or accuracy? |
| No fine-tuning | We are fine-tuning. Does that void it? |
| No infringing input | What counts? Does a customer's uploaded document qualify? |
| Prompt attack exclusion | How is this defined, and who decides? |
| Notice period | Can we meet it operationally? Who is the named recipient? |
| Control of defense | Can the provider settle in a way that binds us? |
| Cap | Fees paid over what period? Against what realistic exposure? |
| Survival | Does it cover claims filed after termination for term-period use? |
A worked example of why this matters. A provider offers an output indemnity that sounds comprehensive. Conditions: current version, filters enabled, no fine-tuning, cap at twelve months of fees. The customer intends to fine-tune, pins versions for six months at a time, and pays $180,000 annually. The indemnity's realistic value is zero — every condition fails — and the customer should either change its deployment, negotiate the conditions, or price the risk as retained.
Step 9 — Negotiate the representations you can enforce
Ask for representations that are specific and verifiable:
- The corpus was collected by the methods described in the manifest
- The licensor holds rights sufficient to grant the enumerated permitted acts
- No component was obtained in violation of a written agreement known to the licensor
- Personal information is limited to what the manifest describes
- No claims or demands have been received relating to the corpus except as disclosed
Do not ask for representations no licensor can give, such as a warranty that the corpus infringes nothing anywhere. A representation the licensor cannot support will be negotiated to a knowledge qualifier, and a knowledge-qualified warranty on an undocumented corpus is worth very little. Better to narrow the scope to what can be documented and price the remainder.
Step 10 — Resolve termination before signature
This is the provision that most often gets deferred and should not be.
Decide, expressly:
- Do models trained during the term survive termination? (Usually the licensee's core requirement.)
- Must the corpus be deleted? Within what period? With certification?
- Do derived artifacts — embeddings, indices, fine-tuned weights — survive or die with the corpus?
- Is there a wind-down period long enough to retrain, if survival is not granted?
- Does termination for cause differ from termination for convenience?
- Does the indemnity survive for claims arising from term-period use?
- What records must the licensee retain to demonstrate compliance?
The commercial resolution that usually works: perpetual rights to models trained during the term, corpus deletion with certification, segregation records maintained for the survival period, and different treatment for termination for material breach.
The provision that makes it enforceable is the segregation obligation. Without a model-to-corpus mapping, neither party can prove which models are covered, and the survival right becomes a dispute rather than a term.
A negotiation, played out
Wrenfield Health operates ambulatory surgery centers. It wants to deploy a clinical documentation assistant that drafts operative notes from dictation. It has selected a vendor, Palisade Clinical AI, and the contract has arrived.
Wrenfield's deputy general counsel, Chukwuemeka Bartholomew, works through it over six weeks.
Week one: what are we actually buying?
Palisade's agreement is titled "Software as a Service Agreement" and treats everything as one product. Chukwuemeka separates the layers on a single page:
- A hosted model Wrenfield will not possess
- Wrenfield's clinical data flowing to Palisade as inputs
- Outputs — draft operative notes entering the medical record
- Fine-tuning, which Palisade's proposal describes as a "custom model tuned on your documentation"
Four layers, and the draft addresses two.
Week two: the diligence questions
He sends fourteen questions. The four that matter:
"What data was the base model trained on?" Palisade's answer references "publicly available medical literature and licensed clinical datasets." Chukwuemeka asks for the manifest. Palisade produces one for the licensed datasets and cannot produce archived terms for the "publicly available" component.
"May Palisade use our data to improve the service?" The draft says yes, by default, with an opt-out available "upon request." Wrenfield's data is protected health information. This is not a preference; it is a compliance requirement, and the default is backwards.
"Who owns the fine-tuned model?" The draft is silent. Palisade's business team says "you do." Palisade's counsel says "the tuned weights are part of the Service." These are different answers, and the second is the one in the document.
"What voids the indemnity?" The indemnity covers third-party IP claims arising from outputs. Conditions: current version, filters enabled, no customer fine-tuning. Palisade is proposing to fine-tune. The indemnity would be void from day one on the very deployment being sold.
Weeks three and four: restructuring
On training data. Wrenfield cannot verify the base corpus and Palisade cannot document part of it. Chukwuemeka does not walk away — the tool's value is real — but he changes the risk allocation: Palisade represents that it holds rights sufficient for the licensed components, discloses the undocumented component, and provides an IP indemnity for the model itself that is not subject to the fine-tuning exclusion, uncapped as to third-party IP claims arising from the base model. Palisade agrees, because the risk is one it created and can price.
On customer data. Rewritten to prohibit any use of Wrenfield data for training, evaluation, or service improvement, except within Wrenfield's own tenant for Wrenfield's own tuned model. Deletion on termination with certification. Subprocessor list and change notice. This is not a negotiation Wrenfield can lose and remain compliant, and Chukwuemeka says so plainly.
On the fine-tuned model. Wrenfield owns the tuned adapter; Palisade owns the base. On termination, Palisade delivers the adapter and certifies deletion of Wrenfield data. Whether the adapter is usable without the base is a practical limitation Chukwuemeka documents rather than pretends away — he negotiates a wind-down of twelve months instead.
On the indemnity. The fine-tuning exclusion is removed for tuning performed through Palisade's own service on Wrenfield's own data. It remains for tuning Wrenfield performs independently. This is a reasonable line and Palisade proposes it.
On outputs. Assignment of whatever rights exist; no representation of copyrightability; commercial use unrestricted; acknowledgment that other customers may receive similar outputs; and — the term Chukwuemeka adds that nobody proposed — a requirement that outputs be labeled as machine-generated within the record system, because clinical documentation attribution is a regulatory and liability issue independent of any license.
Weeks five and six: the two remaining fights
Version pinning versus indemnity conditions. Palisade's indemnity requires the current model version. Wrenfield's clinical validation process requires pinning a version for at least ninety days after validation. These are incompatible as drafted. Resolution: the indemnity covers any version released within the preceding twelve months, and Palisade gives ninety days' notice before deprecating a version.
Liability for clinical error. Not a licensing question, and the most important one. Palisade disclaims all liability for clinical outcomes. Wrenfield accepts this — the physician signs the note and remains responsible — but adds: a requirement that the interface present outputs as drafts requiring review, an obligation on Palisade to notify Wrenfield of any known systematic error mode, and confirmation from Wrenfield's insurer that the professional liability policy responds to documentation assisted in this way.
Signed in week seven.
What Chukwuemeka would tell the next person
The four layers page was the whole exercise. Every subsequent issue came from a layer the draft did not address.
The indemnity was worthless as drafted, and looked comprehensive. Reading the conditions against the actual deployment took twenty minutes and changed the economics of the deal.
The customer data default was backwards and would have been signed. Vendor forms default to broad service-improvement rights, and the opt-out is available to anyone who asks. Most customers do not ask.
The most valuable term was not in the contract at all. The insurer confirmation cost nothing and answered the question the board actually cared about.
A term sheet
| Term | Position |
|---|---|
| Scope of data | Named datasets and versions; manifest attached as a schedule |
| Permitted acts | Training / fine-tuning / evaluation / retrieval — each named |
| Field of use | [Defined]; no restriction on the licensee's product category |
| Distribution | Models embedded in Licensee's product: permitted; distributing the model itself: [permitted / not] |
| Term | [n] years |
| Model survival | Models trained during the term survive perpetually |
| Corpus deletion | Within 90 days of termination, with certification |
| Segregation | Licensee maintains model-to-corpus mapping for the survival period |
| Provenance | Manifest delivered at closing as a deliverable |
| Personal information | [Removed before delivery / lawful basis documented]; re-identification prohibited |
| Representations | Collection method; sufficiency of rights; no undisclosed claims |
| Licensor indemnity | IP and authority claims; cap [2×] fees; survives termination for term-period use |
| Licensee indemnity | Licensee's outputs and downstream use |
| Audit | Licensee compliance with segregation and deletion, on notice |
| Assignment | Permitted to an affiliate or in connection with a sale of the business |
Positions by seat
Data licensor. Grant narrowly and name the acts. Separate internal use from distribution. Require deletion and segregation. Warrant only what the manifest documents. Cap indemnity at a multiple of fees. Reserve the right to require exclusion of components if an upstream problem emerges, with a fee credit.
Data licensee. Require the manifest as a deliverable. Get perpetual model survival. Get an indemnity that survives termination. Get notice of any third-party claim relating to the corpus. Get the right to exclude components without losing the whole license. Negotiate assignment rights, because this asset must travel in a sale.
Model licensor. Treat weights as confidential information. Make acceptable use clear and flow it down. Decide about fine-tuning and say so. Set scale thresholds objectively. Write indemnity conditions you can actually verify — a condition you cannot audit is a condition you will waive in practice.
Model licensee. Read the license. Maintain a model inventory. Confirm what voids the indemnity before architecting the deployment. Reconcile version pinning against indemnity conditions. Get survival for deployed models. Confirm the provider's rights over your inputs and fine-tuning data, and opt out where possible.
When to walk away
- The licensor cannot produce archived source terms and the corpus is material to the business.
- A model license contains a field restriction covering your use and the licensor will not confirm an interpretation in writing.
- The indemnity's conditions exclude your actual deployment and the licensor will not change them or reduce the price.
- The licensor will not grant model survival and cannot offer a wind-down long enough to retrain.
- The corpus contains personal information with no lawful basis and no feasible de-identification.
- The corpus cannot be segregated, making every downstream obligation unenforceable.
Mistakes that recur
Licensing "the AI technology" as one thing. Three layers, three grants.
Accepting representations instead of documents. Provenance cannot be reconstructed later.
Deferring the termination question. It is the deal, not a boilerplate section.
Valuing an indemnity by its headline. The conditions determine whether it exists.
Assuming an open-weight model is open source. Field restrictions and scale thresholds are real and enforceable.
Fine-tuning without checking the indemnity conditions. This voids output coverage under most provider terms.
Ignoring the provider's rights over your inputs. Many hosted services reserve service-improvement rights by default.
Granting customers a broader indemnity than you received. The gap is retained risk, and it compounds across the customer base.
Skipping the segregation obligation. Every other provision depends on it.
Building an intake process
Organizations that evaluate more than a handful of these deals need a process, not a series of bespoke negotiations.
A one-page intake, completed by the business before legal engages:
AI PROCUREMENT INTAKE
1. WHAT IS BEING ACQUIRED?
[ ] Data corpus [ ] Model weights [ ] Hosted model service
[ ] Fine-tuning service [ ] Application built on a model
2. WHAT WILL WE DO WITH IT?
[ ] Train from scratch [ ] Fine-tune [ ] Inference only
[ ] Retrieval at query time [ ] Evaluation only
[ ] Embed in a product we distribute
[ ] Distribute the model itself
3. WHAT DATA WILL WE SEND IT?
[ ] None [ ] Internal, non-sensitive [ ] Customer data
[ ] Personal information [ ] Regulated data (health, financial)
[ ] Third-party confidential information
4. WHO SEES THE OUTPUTS?
[ ] Internal only [ ] Customers [ ] The public
Do outputs enter a record of legal or regulatory significance?
5. WHAT HAPPENS IF THE VENDOR DISAPPEARS?
Migration effort: Business impact:
6. BUDGET AND TIMELINE
Route by risk. A tool used internally on non-sensitive data with outputs nobody relies on needs a short review. A model producing outputs that enter a regulated record needs the full exercise. A single review standard applied to both wastes time on one and under-protects the other.
Maintain a model and data inventory. For every model and corpus in use: the license, the restrictions, the scale thresholds, the flow-down obligations, the indemnity conditions, the corpora used, and the deployed versions. This is the artificial intelligence analogue of an open source bill of materials, and organizations discover they need it at the worst moment — in diligence, or when a license term is asserted against them.
Reconcile upstream and downstream. Whatever indemnity the organization grants its customers should be no broader than what it received, plus consciously priced retained risk. Keeping a single table of upstream and downstream positions makes the gap visible.
Review annually. Licenses change, models are deprecated, restrictions are added at renewal, and deployments drift from what was approved. An annual reconciliation of the inventory against actual deployments catches the drift before someone else does.
Selling: the other side of the table
Vendors negotiating these deals face a mirror-image problem, and the ones who handle it well close faster.
Publish a provenance summary before anyone asks. A one-page description of where the training data came from, what was licensed, and what was excluded answers the first three diligence questions in every deal. Vendors who cannot produce one lose weeks per transaction and lose some deals entirely.
Make the customer-data default correct. Reserving broad service-improvement rights by default and offering an opt-out on request produces a predictable outcome: sophisticated customers strike it, unsophisticated customers sign it, and the vendor ends up training on data it should not have. Set the default to no use, and charge for the version that permits it.
Write indemnity conditions you can actually verify. A condition that the customer keep filters enabled is meaningful only if the vendor can tell whether they were. Conditions the vendor cannot audit are conditions the vendor will waive under pressure, which means they were never terms — they were negotiating positions that will be given away.
Decide about fine-tuning and say so. The most common friction point in these negotiations is a vendor whose sales team sells fine-tuning and whose contract voids the indemnity on fine-tuning. Align them.
Segment the offering. Internal use, embedded distribution, and model redistribution are different products with different risk profiles. Pricing them identically means the low-risk customers subsidize the high-risk ones and the high-risk ones are underpriced.
Address termination proactively. Offering perpetual rights to models trained during the term, with corpus deletion, resolves the customer's largest concern at almost no cost to the vendor — the customer keeps a model, not the data. Vendors that resist this on principle spend weeks negotiating something they will concede anyway.
Flow down what you must. A vendor building on a third-party model whose license imposes field restrictions and attribution obligations must pass them through. A customer who discovers an undisclosed flow-down obligation after deployment has a real claim.
Keep the license readable. Engineering teams make deployment decisions and rarely read a forty-page agreement. A two-page summary of the permitted uses and prohibitions, distributed with the license, prevents more breaches than any enforcement mechanism.
Pricing and structure
Commercial terms in these deals are less standardized than in software licensing, and a few structures recur.
Data licenses
| Structure | When it fits | Watch for |
|---|---|---|
| One-time fee, perpetual training rights | Static corpus; licensee wants certainty | Licensor gives away future value |
| Annual fee with updates | Corpus grows and matters | What happens to models at termination |
| Per-record or per-gigabyte | Highly variable volume | Measurement disputes |
| Revenue share on the resulting product | Licensor believes in the product | Audit burden; definition of attributable revenue |
| Equity or warrant | Early-stage licensee | Dilution; misalignment if the corpus is one of many inputs |
Model licenses
| Structure | When it fits | Watch for |
|---|---|---|
| Per-token or per-call | Variable usage; hosted service | Cost predictability at scale |
| Seat-based | Human-in-the-loop applications | Definition of a seat |
| Annual platform fee plus usage | Enterprise deployments | Minimum commitments |
| One-time weight license with support | On-premises or air-gapped requirements | Update rights; support term |
| Free below a scale threshold | Open-weight releases | The threshold, and what triggers it |
Terms that shift value more than the price does:
- Perpetual model survival on termination. This is often worth more than a discount and costs the licensor comparatively little.
- Uncapped indemnity for base-model IP claims, which the vendor created and can price, versus capped indemnity for output claims, which depend on customer behavior.
- Rate protection on renewal, particularly for usage-based pricing where volume will grow.
- Version support commitments, which matter enormously for regulated deployments.
- Exclusivity or field restrictions, which licensors sometimes grant cheaply and which can be very valuable.
A note on minimum commitments. Usage-based pricing with an annual minimum is the norm, and customers routinely overcommit in year one. Negotiate rollover of unused commitment, or a true-up mechanism, rather than accepting a use-it-or-lose-it floor set from a forecast nobody trusts.
Frequently asked questions
Should we insist on an audit right over the vendor's training data? Ask for one; expect resistance. Vendors reasonably decline to open training corpora to customer inspection. Workable substitutes: a third-party attestation, a provenance summary updated annually, a representation with a specific carve-out from the liability cap, and a notice obligation if the vendor learns of a claim relating to the corpus.
The vendor says its form is non-negotiable. Is it? For small deployments on standard hosted services, frequently yes, and the correct response is to assess whether the standard terms are acceptable rather than to spend weeks failing to change them. The terms most often movable even on a standard form are the customer-data use provisions and the deletion commitments, because vendors have compliant variants ready for enterprise customers who ask.
Do we need a separate agreement for a proof of concept? Yes, and a short one. The recurring failure is a pilot run under a click-through agreement that grants the vendor broad rights over the data used in the pilot — which is often real production data. A two-page pilot agreement restricting data use, limiting the term, and stating that no production deployment is authorized solves it.
How do we handle a vendor that uses a third-party foundation model? Ask which one, and ask for the flow-down obligations. The vendor's own license may impose field restrictions, attribution requirements, and scale thresholds that apply to you. A vendor that will not identify its base model cannot tell you what obligations you are inheriting, which is itself the answer.
How long should this negotiation take? A well-scoped data license with a cooperative counterparty: four to eight weeks. Diligence is the long pole, and it is the part worth taking time over.
What if the licensor refuses to provide a provenance manifest? Treat it as a finding, not an inconvenience. Either the licensor does not have the records — which is the answer — or it does not want to share them, which invites the same question. Price accordingly or decline.
Can we get an indemnity from an open-weight model provider? Usually not. Permissive open-weight releases typically disclaim everything. That is a legitimate model; it simply means you are self-insuring, and the decision should be made consciously and documented.
Is a de-identification commitment enough for personal information? It helps and is rarely sufficient alone. Confirm the technique, prohibit re-identification contractually, and address what happens if a data subject exercises deletion rights against material already in a trained model.
Should we ask for source code or weights in escrow? For hosted models where the business depends on continuity, yes — but escrow of weights without the serving infrastructure and the right to use them is of limited value. Scope it to what you could actually deploy.
What if our customers ask us for an AI indemnity? Answer it upstream first. Grant downstream no more than you received upstream, plus whatever retained risk you have consciously priced. This is one of the more common places organizations quietly assume large exposure.
Regulated and sensitive deployments
Some deployments carry obligations that no license negotiation addresses, and counsel should raise them even though they sit outside the contract.
Health care. Protected health information flowing to a vendor requires a business associate arrangement with specified terms. Training on that data is generally not permitted without a separate basis, and de-identification standards are specific. Clinical decision support raises regulatory questions about whether the tool is a medical device, which turns on how much the clinician can independently review the basis for the output.
Financial services. Data covered by 15 U.S.C. § 6801 and 15 U.S.C. § 6802 carries sharing restrictions. Model risk management expectations apply to models used in credit, pricing, and fraud decisions, and they require documentation of development, validation, and monitoring that a vendor may not maintain. Adverse action notice requirements demand explanations the model may not readily produce.
Employment. Tools used in hiring, promotion, and termination decisions face disparate impact analysis, growing state and municipal audit requirements, and notice obligations to candidates. A vendor's representation that its tool is "bias tested" is not a substitute for the deployer's own validation, and the legal exposure sits with the employer.
Insurance. Rate and underwriting models are subject to filing requirements and unfair discrimination prohibitions that vary by state and line of business.
Children. Any deployment reaching users under thirteen implicates specific consent and data minimization requirements, and several jurisdictions have added design-code obligations reaching older minors.
Biometrics. Facial, voice, and gait data are separately regulated in several states, some with private rights of action and statutory damages. Training on biometric identifiers without written consent is among the more expensive mistakes available.
Government use. Procurement rules, transparency obligations, and in some jurisdictions prohibitions on specific applications. A vendor selling into government may inherit obligations its commercial contract does not mention.
The practical instruction: run the regulatory analysis in parallel with the license negotiation, not after it. The contract terms that matter — data use restrictions, explainability support, audit cooperation, documentation deliverables — are the ones a regulated deployer must ask for during the negotiation and cannot obtain afterward.
Exit planning
Every one of these relationships ends, and the exit is negotiated at signature or not at all.
Know what you would lose. For each deployment, write down: the model, the data flowing to it, the fine-tuned artifacts, the integrations, and what the business does if the service stops next Tuesday. Organizations that have not done this cannot evaluate an exit provision because they do not know what they need.
Negotiate the exit deliverables.
| Deliverable | Why |
|---|---|
| Fine-tuned weights or adapters | Portability, even if imperfect |
| Training and evaluation datasets you provided | You should not have to reconstruct your own data |
| Embeddings and indices derived from your data | Expensive to regenerate |
| Configuration and prompt assets | Often the real accumulated work |
| Logs sufficient for audit and regulatory obligations | Retention duties may outlive the contract |
| Certification of deletion of your data | Compliance requirement |
| A transition period with continued service | Migration takes longer than anyone plans |
Address the practical asymmetry. Fine-tuned weights are frequently useless without the base model, and a vendor that delivers an adapter has satisfied the letter of a portability clause while delivering nothing usable. Where continuity genuinely matters, the alternatives are a longer wind-down, a source or weight escrow scoped to something deployable, or an architecture that keeps your data outside the model — retrieval rather than fine-tuning — so that switching vendors means switching an interface rather than rebuilding an asset.
Plan the architecture with exit in mind. This is the highest-leverage decision and it is made by engineering, early, usually without legal input. A system that keeps proprietary content in a retrieval layer and treats the model as an interchangeable component has an exit. One that bakes proprietary content into fine-tuned weights on a proprietary base does not.
Calendar the renewal. Usage-based agreements renew quietly with price changes, and the leverage to negotiate exists ninety days before renewal and not after. Set the reminder at signature.
Related documents
- Licensing Data and Models for Artificial Intelligence: Training Rights, Outputs, and Indemnities
- AI Licensing Diligence Checklist: A Practical Checklist
- AI Licensing Toolkit: Data Provenance Records, License Terms, and Indemnity Clauses
- Conducting a Fair Use Analysis: A Practical Guide
- Cloud and SaaS Agreements: Service Levels, Data Rights, Security, and Exit
- Drafting Software License Agreements: Key Terms and Negotiation Points
- AI Vendor Procurement and Governance Checklist: A Practical Checklist
- Indemnification and Limitation of Liability: The Risk Allocation Engine of Every Contract
- IP Transactions and Agreements Toolkit
