The risk is measurable. Verizon's 2026 Data Breach Investigations Report put third-party involvement in breaches at 48 percent, a figure counting all third parties, not AI vendors. IBM's 2025 Cost of a Data Breach report adds that among breached organizations, about one in five said the breach involved shadow AI; heavy shadow-AI users reported costs roughly 670,000 dollars higher than light or non-users, an association, not a proven cause. An unvetted AI tool is a third party with unusually broad access to privileged material.
Six questions have earned their place on my list.
Training data provenance. Where did the model's training data come from, and can the vendor show it had rights to use it?
Prompt retention and reuse. Are your prompts and documents stored, for how long, and used to train anything?
Subprocessors and hosting region. Who else touches the data, how far down the subprocessor chain, in which jurisdiction, and can deletion be verified rather than just promised?
IP indemnity. If an output infringes someone's rights, who carries that?
Audit trail. Can you reconstruct who asked what, and what the system returned?
Exit and data return. When the contract ends, how does your data come back, and in what format?
A vendor with clean answers welcomes these questions. Hesitation on any of them is itself an answer, and it arrives before you have signed anything. Treat it as a signal, not a verdict; the answers still have to land in the contract.
Sources
Next: When an AI workflow fails: an incident protocol for legal teams →