When you buy an AI product, you are deciding who else gets to know your business.

The demo shows you a dashboard, maybe a chat window. What it does not show you is the pipe behind it. Once the contract is signed, that product may maintain ongoing access to customer records, financial history, operational data, and whatever intellectual property lives in the systems it touches. It may read from that data repeatedly, write back to source systems, and create new artifacts from it that nobody in the demo mentioned.

I have sat on both sides of this table, and the gap I keep seeing is between how these products are sold and what they actually are. They are sold as software. What you are entering is a data relationship. And once the system is embedded in your operations, your negotiating position drops substantially, which is why the questions below belong before signature rather than at renewal.

If you run marketing at an association, this arrives on your desk in a specific shape. The AMS vendor adds an “AI member insights” module. The email platform adds generated subject lines and send-time prediction. A chatbot vendor asks for a feed of the member directory and the knowledge base. Each one is a data relationship, and the questions below apply to all three, whether the contract is $8,000 or $800,000. Your board is probably already asking about ethical AI. These questions come after that one, because a board can approve a use that is ethical and still sign a contract that gives the artifacts away. I wrote this for the CIO and the CISO as much as for you, because at most associations you are the one who has to ask these questions first. It assumes you have already made the broader case for AI at your association and are now choosing vendors.

A fair objection: most of this applies to any data-intensive SaaS product. True. AI did not create most of these risks. It amplifies them. AI products consume more kinds of data, create more derived data, combine more systems, and increasingly take actions rather than displaying information.

Traditional SaaS diligence asks what the vendor stores. AI diligence also has to ask what the vendor learns, derives, generates, and can act on.

Diagram titled More Than Software: enterprise AI buying is a data relationship, showing what the vendor receives, what it creates, and what it can do

Not every product deserves the full treatment. Diligence should scale with four things: how sensitive the data is, how broad the access is, how autonomously the system acts, and what a wrong output costs. A meeting summarizer and an agent with write access to your ERP should not face identical review. With that scaling in mind, the buyer’s job comes down to three questions.

Can we trust the vendor? Can we trace the data? Can we control what the system creates and does?

What they receive, what they create, what they can do with both

Diagram titled The Three Layers: what the vendor receives, what it creates, and what it can do with your data

Before the three questions, one idea that most procurement templates do not yet cover, because it is the idea the rest of the article hangs on.

Think of the relationship in three layers. There is what the vendor receives from you: the records, documents, and fields you connect. There is what the vendor creates from that: churn scores, customer classifications, forecasts, embeddings, behavioral profiles, recommendations, summaries, fine-tuned models, feature-store entries. And there is what the vendor, and the system, can do with both.

Most contracts handle the first layer. The defined term “Customer Data” usually means what you uploaded. The second layer is where things get thin. None of those artifacts existed before you connected. All of them describe your customers and your operations. And in many agreements, none of them are covered by the term that grants you your rights.

Sometimes the most sensitive thing an AI vendor holds is not what you gave them. It is what they made from it.

These artifacts are not equivalent, legally or technically. A model fine-tuned across thousands of customers does not make any particular weight yours. A profile, score, or summary built about your customers is a different matter. The contract should name the artifact types that matter to you and set out the rights around each. Ownership is often the wrong frame for model-generated artifacts. Retention, reuse, portability, and deletion are usually the rights that matter.

Hold that three-layer picture. It comes back in every section.

Can we trust the vendor?

Diagram titled Can We Trust the Vendor: security evidence including SOC 2 scope and exceptions, pen test remediation, DPA and subprocessor list, incident history, and cyber insurance

Ask directly: is this vendor secure and compliant? Those are two questions. A vendor can satisfy an audit within scope while carrying material risk outside it, so keep both open.

Then ask for evidence. Every vendor will tell you they follow industry best practices and that security is a top priority. Those are sentences. The principle is that evidence matters more than certification, and a certification is one form of evidence. What you want in hand:

  • A current SOC 2 Type II report, or other independent assurance appropriate to the vendor, deployment model, and industry (ISO 27001, HITRUST, and FedRAMP serve different purposes and are not interchangeable)
  • The scope of that assessment, and the material exceptions or findings
  • A recent penetration test summary, attestation, and remediation status for material findings
  • Security incident history and incident response procedures
  • The data processing agreement and subprocessor list
  • Cyber insurance appropriate to the risk and contract value, understanding that this speaks to financial resilience, not to whether the system is secure

A badge on a website is marketing. Read the actual report, and confirm that the product, infrastructure, and subprocessors you are buying sit inside its scope. SOC 2 is not pass/fail, so the exceptions matter as much as the cover letter.

One more question: in the past three to five years, has the company experienced any material security incident, unauthorized access, ransomware event, or breach involving customer systems or customer data? Ask it that broadly on purpose, because “breach” has legal definitions that vary by jurisdiction and you are asking what happened, not what was reportable. A yes does not automatically disqualify them. What matters is whether they can tell you what happened and what changed afterward.

A mature vendor may not hand over a raw penetration test or internal architecture without an NDA or a controlled review, and that is reasonable. What they should be able to do is give you enough evidence to assess the risk. Most will, and most of these questions will get boring, reassuring answers. Boring answers are what you want.

Can we trace the data?

Diagram titled Can We Trace the Data: source systems flowing into the vendor platform and out to subprocessors, with the data lifecycle from ingest to delete

This is the first layer, what the vendor receives, followed all the way through.

What it can reach. The safest integration is usually the one with the least access it needs to deliver the outcome. Before anything connects, know which systems it touches, which fields and records it can read, whether it can write back, whether sensitive fields can be excluded, what its service accounts are permitted to do, and whether it collects prompts, outputs, telemetry, or usage data beyond the core dataset.

Who can reach data through it. AI products can make an existing permissions problem worse, because they aggregate information across systems that were permissioned separately. Ask about SSO, SCIM provisioning, role-based access, admin controls, and audit logs. Then ask the question that matters most:

Can the AI surface information to a user that the user could not access directly in the source system?

If the answer is yes, or the vendor is not sure, the product is a new access path and needs to be governed like one. This is the most common failure I see in retrieval-augmented generation deployments: the retrieval index is built from documents without carrying the permissions those documents had in the source system.

Where it goes next. The company you contract with may be one layer of several. Ask which third parties process or access your data: cloud providers, foundation model providers, vector databases, analytics and logging platforms, support systems, enrichment providers. Keep “process” and “access” separate; a cloud provider can process data without its employees having routine access. For each subprocessor, and especially for foundation model providers, ask what it retains, for how long, and for what purpose; whether reduced- or zero-retention configurations are available; whether it can train on your data; and whether the primary vendor holds an enterprise agreement that prevents reuse. Ask whether subprocessors can change without notice to you. If the vendor cannot draw the path your data takes, there is a fair chance they have not drawn it for themselves.

Where it lives. Residency answers where data is stored. It does not necessarily answer where it is processed or who can access it. A vendor can store your data in the EU while support staff elsewhere have access and a subprocessor processes portions of it outside the region. Ask storage, processing, backups, and personnel access as separate questions. Regulated industries care first. Everyone else cares the day a regulator or a large customer asks.

When it leaves. Ask what happens when a record is deleted, when a user leaves, when an integration is disconnected, and when the contract ends. Deletion should cover more than the primary database: logs, backups, replicas, caches, embeddings, model artifacts, and derived datasets all need an answer. Be realistic about backups. Mature systems do not surgically remove one customer’s records from immutable backups on demand, so the useful question is when deleted data ages out of backups, logs, and disaster-recovery systems, and what prevents it from returning to production. “Deleted” and “no longer visible in the UI” are not the same thing, and vendors do sometimes mean the second when they say the first.

Can we control what the system creates and does?

Diagram titled Can We Control What the System Creates and Does: governance and control points, derived artifacts, and real-world actions

This is the second and third layers: what the vendor creates, and what the vendor and the system can do with it. Two groups of questions: first your rights over the data and what is derived from it, then your operational control over how the model behaves.

Rights over the data and what is derived from it

The vendor’s rights to use your information. Your contract should say you retain ownership of your data. It also needs to define the licenses you grant the vendor, what they cover, and what survives termination. Then go one layer further and ask whether any of the following can be used to train or improve models: source data, prompts, outputs, user interactions, metadata, embeddings, feedback, derived datasets. A reasonable default is that your data should not improve a model available to other customers unless you explicitly agree. If the vendor’s business model depends on that not being the default, you want to know now.

For multi-tenant systems, ask how your data is isolated from other customers across application, storage, retrieval, vector, inference, and training workflows. Physical separation is not always necessary. What matters is whether your prompts, context, outputs, and retrieved data remain logically segregated at each step. “We are multi-tenant” is a category, not an answer.

Your rights to what it creates. This is the three-layer idea becoming a clause. For every artifact type that matters, you want to know how long it is retained, whether it can be exported in a usable form, whether the vendor can reuse it in aggregate or otherwise, and what terms govern it if “Customer Data” does not. If the vendor’s paper does not mention derived artifacts at all, that is a gap in their paper, and it is the one to negotiate.

Operational control over what the system does

Whether the model is fit for your decisions. Model behavior deserves the same scrutiny as the software around it. For predictive systems, ask whether the training data resembles your customers, how the model was validated and on whose data, whether the performance metric fits the business decision, how drift is detected, how often it is retrained, and what a false positive and a false negative each cost you. For generative systems, add what the vendor can tell you about training-data provenance and known limitations, how outputs are grounded in your source material, and how output quality is evaluated. Foundation-model vendors often cannot disclose their exact training corpus, and that alone is not a red flag. For a bespoke model built on data like yours, ask much more directly.

A model can perform well in aggregate and still perform poorly on your data. The performance number in the deck is the vendor’s number. The one you care about is how the model performs on your population, against the decisions you will use it for.

What happens when it is wrong. It will be wrong sometimes. Part of what you are buying is how the system handles that. Can users override a recommendation? Are predictions presented with uncertainty, and has that uncertainty been calibrated and validated? Can the vendor provide the factors or evidence behind a prediction, or, for generative output, ground it in cited source material? Is there an audit trail? Can high-impact actions require human approval? Be careful with “explainability” as a checkbox. A model-generated explanation can be a plausible story written after the fact. Ask for the evidence.

What happens when it is manipulated. If the model can ingest untrusted content and call tools, untrusted data can become instructions to a system capable of taking action. A retrieved document or an inbound email is an input the model will read and may act on. Ask how the vendor prevents untrusted input from changing system behavior, escalating privileges, leaking data, or triggering unauthorized actions.

The higher the consequence of the decision, the more all of these controls matter. A wrong merchandising recommendation costs a click. A wrong fraud flag, pricing decision, or safety alert can cost a good deal more.

How you leave. Before signing, you should be able to describe the exit: export your original data, identify which derived artifacts are portable and secure export rights for the ones that matter (a churn-score table may travel easily; embeddings or model-specific artifacts may not be useful outside the vendor’s stack), revoke integrations, remove user access, delete retained information, and migrate without unreasonable friction.

Lock-in with an ordinary SaaS tool is an inconvenience. Lock-in with a vendor that has accumulated years of derived intelligence about your customers is a different kind of dependency, because what you would be leaving behind is knowledge about your own business.

Who is accountable

IT cannot carry this alone, and neither can procurement. Whoever owns vendor strategy, data architecture, and security is responsible for it, usually some combination of the CIO, the CISO, general counsel, and the head of procurement, with the business sponsor and the owner of the affected data or process at the same table. That last seat matters. The technical reviewers can approve a product that is safe to run without knowing whether the business use is appropriate. When those roles evaluate separately, each assumes another one is covering the data question. Often nobody is. If the organization has not settled whether it is ready to adopt AI at all, that conversation comes before this one.

The three questions

Before approving an AI, ML, or data product, you should be able to answer:

  1. Can we trust the vendor? Is it demonstrably secure and compliant, with evidence you have read rather than a badge you have seen?
  2. Can we trace the data? Do we know what happens to it from ingestion through deletion, including backups, logs, and subprocessors?
  3. Can we control what the system creates and does? Can our data, or anything derived from it, be retained, reused, or used to benefit another customer, vendor, or model, and can the system act in ways we have not authorized?

If the vendor cannot answer clearly, cannot provide independent evidence, or substitutes policy language for proof, treat that as a material buying risk and price it in. That may mean tighter contractual protections, narrower access, a limited pilot, a lower price, or deciding not to proceed.

You will not get every answer. Some vendors are early and their documentation will be incomplete. The important thing is knowing which answers are missing, what risk that creates, and who accepted it. That is the difference between choosing a risk and discovering later that you inherited one.

If you want a second set of eyes on a contract or a vendor’s security package before you sign, that is part of what the AI readiness audit covers.

References

  • AICPA, Trust Services Criteria (the basis for SOC 2 reports)
  • ISO/IEC 27001:2022, Information security management systems
  • NIST AI Risk Management Framework (AI RMF 1.0), January 2023
  • NIST Cybersecurity Framework 2.0, February 2024
  • OWASP Top 10 for Large Language Model Applications (OWASP GenAI Security Project)
  • Cloud Security Alliance, AI Controls Matrix (AICM)

More on AI for associations

Leave a Reply