Summary. AI governance fails predictably: a policy is written, a committee is formed, and nobody ever learns what the company is actually running. This toolkit is built the other way around — discovery and inventory first, documentation as the output of a process rather than its substitute. It covers building an inventory that includes what procurement never saw, classifying systems by risk in a way that maps to the obligations that follow, impact assessments that produce decisions, oversight that can actually override a model, testing on a schedule, notice where required, and AI incidents that are not security incidents.


What this toolkit is for, and who should use it

Two years ago the question in most companies was whether to allow AI tools. That question is closed: they are already in use, in more places than the organization believes, procured on corporate cards, embedded in software the company already bought, and running inside features that vendors shipped without asking. Governance now means finding what exists, deciding what is high-stakes, and putting proportionate controls around that subset.

This toolkit is for the general counsel, privacy officer, CISO, or operations leader who has been handed "AI governance" and needs a program rather than a position paper. It is deliberately risk-tiered: the same controls applied everywhere would be both unaffordable and ineffective.

Roadmap at a glance

  1. Discovery — find what is actually running, including shadow AI.
  2. Inventory — record what governance needs to know about each system.
  3. Risk classification — tiers that map to obligations.
  4. Impact assessment — for the systems that warrant one.
  5. Controls by tier — oversight, testing, documentation, and logging.
  6. Notice and transparency — to individuals, customers, and regulators.
  7. Vendor and model supply chain.
  8. Incident response for AI failures.
  9. Policy, training, and culture.
  10. Assurance — monitoring, audit, board reporting, and annual review.

Stage 1 — Discovery

Start by assuming the inventory you were handed is incomplete, because it is. Run discovery through several independent channels, because each one misses different things:

  • Procurement and expense data — search for AI vendors in the AP ledger and in corporate card statements, including sub-$500 monthly charges that never touch procurement.
  • SSO and identity logs — every application anyone has authenticated to.
  • Network egress and DNS logs — traffic to model API endpoints.
  • Browser extension inventories and IDE plugin lists.
  • Existing software — ask each incumbent vendor, in writing, what AI features they have enabled by default, what model providers sit behind them, and whether customer data is used for training.
  • Ask people. A short, non-punitive survey asking what tools they use to do their work surfaces more than any scan, provided the answer does not get anyone in trouble.
  • Internal development — models built or fine-tuned in-house, including notebooks and prototypes that quietly reached production.

Illustration. A 300-person company's official inventory lists two AI systems. Discovery finds nineteen: a support summarizer enabled by default in the helpdesk platform, a resume-ranking feature switched on in the applicant tracking system, four departmental subscriptions on personal cards, a marketing agency's generative image workflow, and an engineer's fine-tuned model running in production behind an internal endpoint. Two of the nineteen — the resume ranker and one credit-related scoring feature — carry nearly all of the legal exposure.

Resources

Stage 2 — Inventory

For each system, record: name and vendor; business owner and technical owner; purpose in one sentence; inputs, including whether personal, sensitive, health, biometric, or third-party confidential data is involved; outputs and who consumes them; whether the output is advisory or decisional; affected populations; the underlying model and version; deployment date; risk tier; assessment status; last test date; and the contract reference.

Two fields do most of the work. "Advisory or decisional" determines whether human oversight controls are needed. "Affected populations" determines whether the regulated-decision rules apply. Everything else is context.

Keep the inventory in a system that supports review dates and ownership, not in a spreadsheet nobody opens. Tie it to procurement so that new systems enter automatically, and to offboarding so that departed owners do not leave orphaned entries.

Stage 3 — Risk classification

Use three or four tiers, defined by consequence rather than by technology.

  • Tier 1 — Consequential. The system materially influences a decision about a person's employment, credit, housing, insurance, education, healthcare, or access to an essential service; or it operates without meaningful human review in a safety-relevant context. Full controls apply.
  • Tier 2 — Elevated. The system processes personal or confidential data at scale, produces public-facing content, or informs significant business decisions. Selected controls apply.
  • Tier 3 — Standard. Internal productivity use on non-sensitive data. Acceptable use policy and training only.
  • Prohibited. Uses the organization will not permit at all — for example, emotion inference in hiring, social scoring, and any use expressly prohibited by an applicable statute.

Write the tier definitions so that a non-lawyer can apply them, and require the business owner to classify at intake with legal review for anything self-classified as Tier 3 that touches people.

Resources

Stage 4 — Impact assessment

For Tier 1 and selected Tier 2 systems, run a written assessment covering: the purpose and the alternatives considered; the data used and its provenance and lawful basis; known limitations and failure modes; the populations affected and the foreseeable harms, including disparate impact; the accuracy and fairness testing performed and its results; the human oversight design; the transparency provided to affected individuals; the security and privacy controls; the monitoring plan; and the residual risk with a named person accepting it.

The assessment is only useful if it can produce the answer "no." Build in an approval step with authority to reject a deployment, and record the conditions attached to any approval.

Note that several regimes now require an assessment by name: state privacy statutes require data protection assessments for profiling with legally significant effects, the Colorado AI Act requires impact assessments for deployers of high-risk systems, and the EU AI Act requires conformity assessments and, for certain deployers, a fundamental rights impact assessment.

Resources

Stage 5 — Controls by tier

Human oversight. For consequential systems, a human must review before the decision takes effect, must have the information and the time to evaluate it, and must have actual authority to override. Measure the override rate: a rate near zero across thousands of decisions means the review is nominal, and a regulator will read it that way too. Train reviewers on the system's known failure modes rather than on how to use the interface.

Accuracy and evaluation. Test on your own data before deployment, with a documented pass threshold. Re-test on a schedule and after every material model change. Retain the results — an undocumented test did not happen.

Bias testing. For systems touching employment, credit, housing, insurance, or public accommodations, test outcomes across protected characteristics and record both the method and the result. Understand the difference between disparate treatment and disparate impact, and that a facially neutral model producing disparate outcomes can violate Title VII, the ADA, the ADEA, ECOA, or the FHA. Run this analysis under privilege, and remediate what it finds.

Grounding and hallucination controls. For any system producing factual assertions, require retrieval grounding with citations to source, and require human verification for anything that leaves the building. Prohibit unverified AI output in filings, financial statements, safety documentation, and customer commitments.

Logging. Log prompts, outputs, model version, and the human decision for consequential systems, consistent with the retention schedule and any litigation hold. Logs are what make an incident investigable and an audit possible; they are also discoverable, so set retention deliberately rather than by default.

Access control. Confirm a retrieval-augmented system cannot surface documents the requesting user could not otherwise access. This is the single most common security defect in enterprise AI deployments.

Data controls. Enforce what may be entered: no customer confidential information, no third-party confidential information restricted by contract, no personal data outside approved systems, no source code in unapproved tools.

Resources

Stage 6 — Notice and transparency

Tell people when AI is involved in a decision about them, where the law requires it and where candor is warranted regardless.

  • Employment: pre-use notice for automated employment decision tools, an annual independent bias audit with published results where required (New York City Local Law 144), notice of AI video interview analysis in states that require it, and an accommodation path for candidates who cannot complete an automated assessment.
  • Consequential decisions generally: notice before use, a statement of the principal reasons for an adverse decision, a correction opportunity, and an appeal to a human, as the Colorado AI Act and analogous statutes require.
  • Credit: adverse action notices under the FCRA and ECOA with specific and accurate reasons — "our model declined you" is not a reason.
  • Synthetic content: disclose AI-generated or materially altered content where a statute, a platform rule, or an advertising standard requires it, and label synthetic likenesses and voices.
  • Chatbots: disclose that the user is interacting with a machine where required.
  • Privacy notices: describe AI processing, profiling, and any automated decision-making, and honor the opt-out and access rights that follow.

Resources

Stage 7 — Vendor and model supply chain

Most organizations are deployers, not developers, which means most of the risk arrives through contracts. Require, at minimum: no training on customer data without opt-in; disclosure of the model providers behind the product; notice of material model changes with a right to test; documented error rates and known limitations; an IP indemnity for outputs; security representations covering prompt injection and tenant isolation; incident notification covering AI-specific failures; and export and deletion on exit.

Then verify the configuration matches the contract. The default settings of a product are what govern in practice, and they frequently differ from what the enterprise agreement permits.

Ask about training data provenance for any model whose outputs the company will publish or rely on commercially, and about the vendor's own upstream indemnities. The copyright status of training data is actively litigated, and the exposure flows downstream through the products built on it.

Resources

Stage 8 — AI incident response

An AI incident is often not a security incident, so the security playbook does not fire. Define AI incidents explicitly: materially wrong outputs at scale, harmful or discriminatory outputs, confidential data leaked through outputs, a model change that degrades a validated system, prompt injection resulting in unintended action, and any use that violates policy.

Build the response path: detection and reporting channel, triage and severity, containment (which for AI usually means disabling the feature or reverting the model version), assessment of affected decisions and whether they must be re-made, notification obligations to individuals, customers, regulators, and insurers, and a post-incident review that feeds back into testing.

The step people miss is remediation of past decisions. If a screening tool was misconfigured for four months, the question is not only how to fix it but what to do about the applicants it rejected.

Illustration. A vendor silently upgrades the model behind a claims-triage tool. Accuracy on the insurer's specific claim types drops, and denials rise for two weeks before anyone notices. The contractual right to notice of model changes, the scheduled re-testing, and the logging that made the two-week window identifiable are what turn a systemic failure into a contained one.

Resources

Stage 9 — Policy, training, and culture

Publish a short AI acceptable use policy: approved tools, prohibited data, prohibited uses, the disclosure obligation for AI-assisted work products, verification requirements, and how to request a new tool. Make the request path easy — a slow approval process is the main cause of shadow AI.

Train by role. Engineers need secure-development and data-handling guidance; HR needs the employment rules; marketing needs disclosure and clearance rules; customer-facing staff need the verification requirement; managers need to understand what meaningful oversight means. Fifteen concrete minutes beats a ninety-minute survey of AI ethics.

Set the tone honestly. The goal is not to slow the organization down but to keep it out of the small number of uses that can cause serious harm. A program perceived as obstruction gets routed around, and routed-around programs are worse than none because they create false assurance.

Stage 10 — Assurance

  • Monitor production systems for drift, error rates, and output anomalies, with thresholds that trigger review.
  • Audit a sample of Tier 1 decisions annually — pull the logs, re-derive the decision, and confirm the human review was real.
  • Reassess each system at least annually, and immediately on a material model change, a new use case, or an incident.
  • Report to the board on the inventory by tier, the systems in Tier 1, the assessments completed, incidents, and the regulatory horizon. Directors' oversight duties extend to mission-critical risks, and AI is becoming one in regulated industries.
  • Map the program to a recognized framework — the NIST AI Risk Management Framework or ISO/IEC 42001 — so that the documentation answers a regulator's or a customer's questionnaire without a new project each time.
  • Retire systems properly: delete data, revoke access, terminate the contract, and remove the inventory entry with a record of the deletion.

Resources


Stage 11 — Frequently asked governance questions

"Do we need a committee?" You need a decision-maker and a documented path, which a committee can provide and often obstructs. For most organizations under a thousand people, a named owner in legal or privacy, a technical counterpart, and a standing monthly review with the business owners is enough. Escalate Tier 1 approvals to a small cross-functional group with authority to say no.

"Can we just ban it?" No, and the attempt is counterproductive. Prohibition drives use onto personal accounts and personal devices, where the company has no contract, no logging, no data protections, and no visibility. Approve a small set of tools with enterprise terms, make access easy, and enforce the boundary on data rather than on tools.

"Who owns AI output we publish?" Copyright protects human authorship; purely machine-generated material is not registrable, and a work containing AI-generated elements is protected only as to the human contribution, which must be disclosed on registration. Keep records of the human contribution for anything the company will need to enforce. Patent inventorship requires a human inventor. See Copyright Infringement Claims Against Generative AI.

"Is a prompt confidential?" Only if a contract says so. A standard NDA often defines confidential information in terms of documents and disclosures and never contemplates prompts or model outputs; check the definition and amend it.

"Are AI outputs discoverable?" Yes, and so are prompts, logs, and evaluation results. Set retention deliberately, apply litigation holds to AI systems the way you do to email, and assume that the record of what the model said and what the human did with it will be read aloud someday.

"What about attorney work product and privilege?" Entering client confidential information into a tool with training rights or broad human review can jeopardize confidentiality obligations. Route legal use through tools with enterprise terms and no training, and treat AI-assisted legal work under the same supervision duties that apply to any delegated work — verify every citation and every quotation before it leaves the office.

"How do we handle an employee who used AI badly?" The same way you handle any policy violation, but check first whether the policy was clear, whether the person was trained, and whether an approved tool existed for the task. Most violations are process failures wearing a disciplinary costume.

Resources


Master resource index

Articles

Checklists

Related toolkits

External and primary sources

This toolkit is educational and not legal advice. AI regulation is developing quickly and differs by jurisdiction and sector. Consult qualified counsel before deploying an AI system that affects people's rights, opportunities, or safety.