title: "AI Vendor Questions: What to Ask Before Buying"
slug: ai-vendor-questions
excerpt: "AI vendor questions worth asking before you sign: does it train on your inputs, is your data isolated, where does it live, and who at the vendor can see it."
category: b2b
tags: [ai-vendor-questions, ai-procurement, ai-data-security, enterprise-ai, ai-compliance]
cover:
author_name: Nico
published_at: 2026-02-05

AI vendor questions are the checks a buyer runs before putting company data into an AI product, and there are 5 that matter more than everything else on a procurement form combined.

We get asked these at ChatFuse and we answer them in writing, which is the only form of answer worth having. If a vendor will not put an answer in a contract, the answer is no matter what the salesperson says.

AI vendor questions Five questions. If you cannot answer all 5, it is a consumer app.
1
Does it train on my inputs? Not "we respect your privacy". A written commitment that prompts, files and memory are never used for training.
2
Is my data isolated from other customers? Shared infrastructure is fine. Shared context is not. Ask which one they mean.
3
Where does it physically live? Region matters for law, not just latency. Ask about the subprocessors too, because that is where data actually travels.
4
Is there an audit trail? Who did what, when, retrievable later. Without it you cannot answer a regulator or investigate an incident.
5
Who at the vendor can see my data? The honest answer is rarely nobody. Ask what standing access staff have and what breaks the glass.
Compiled by ChatFuse from the questions buyers actually ask us.

Most procurement forms ask 40 questions and miss at least 2 of these. Length is not rigour.

Why do these 5 questions matter more than the rest?

Because a wrong answer to any of these cannot be undone later. A vendor with a clumsy interface is an annoyance. A vendor that trained on your client list cannot untrain, and a vendor with no audit trail cannot tell you what happened after the fact.

Everything else on a procurement checklist is a preference. These 5 are the ones where you cannot fix the problem later by switching supplier, because the damage is already distributed.

What is the right answer to "do you train on my inputs"?

A written no, covering prompts, uploaded files and stored memory, with the same commitment flowing down to the model providers behind the product. The last part is the one people forget to ask.

Most AI products are built on OpenAI, Anthropic, Google or Meta models, so the vendor's own policy is only half the answer. What matters is what their contract with those providers says. ChatFuse commits to zero training on your data and passes that through to the providers underneath, which is why we can state it plainly rather than in the conditional.

What does data isolation actually mean?

It means your context cannot appear in somebody else's session, which is a different claim from your data being in a separate database. Ask specifically whether memory, files and conversation history are scoped per customer, and how that scoping is enforced.

Two things called isolation Only one of them protects you.
Infrastructure isolation Often what is meant
  • Separate database rows or schemas
  • Encrypted at rest and in transit
  • Says nothing about retrieval
Context isolation What you need
  • Your memory only reachable by you
  • No cross customer retrieval, ever
  • Enforced in code, not policy
Ask which one they are describing. The answers sound identical in a sales call.

Why does physical location matter?

Because the law that applies to your data is the law of the place it sits, and of the places it passes through. A vendor hosted in one region may still use subprocessors elsewhere, and that is the part that catches people out during a compliance review.

Ask for the subprocessor list in writing. A vendor who cannot produce one has either not thought about it or does not want to say, and both are answers.

The list also tends to be longer than buyers expect. A single AI product can involve a model provider, a hosting provider, a vector store, an analytics tool and an error tracker, each in a different place. ChatFuse keeps that list short deliberately, because every entry on it is another jurisdiction and another contract somebody has to read.

What should an audit trail contain?

Who accessed what, when, and from where, kept long enough to be useful after an incident is discovered rather than while it is happening. Most incidents surface weeks later, so a 7 day retention window is decoration.

The related question is whether you can get the trail yourself or have to ask the vendor for it. Self service matters, because during an actual investigation you will not want to wait on a support ticket.

Who at the vendor can see my data?

The honest answer from most vendors is that some engineers can, under some conditions, and the good ones will tell you exactly which. Be suspicious of an unqualified nobody, because someone has to be able to debug production.

What separates vendors is whether access is standing or exceptional. ChatFuse is built so employees have no standing access to customer data, and the technical detail behind that claim is in zero trust AI data security. The question to ask any vendor is what has to happen for a human to see your data, and whether that event is logged where you can see it.

Do these questions apply to free AI tools too?

They apply more. Free consumer tools are typically funded by something other than your subscription, and the terms usually differ from the paid tier in exactly the areas these questions cover. Staff using a personal free account for company work is the most common version of this problem, and it does not appear on any procurement form because nobody bought anything.

Frequently asked questions

What should I ask an AI vendor about training on my data?

Ask for a written commitment that your prompts, files and stored memory are never used to train any model, and that the commitment extends to the model providers behind their product. A verbal assurance or a marketing page is not the same as a contractual term.

Is my data safe in a consumer AI tool?

It depends entirely on the tier and the terms, and consumer tiers frequently differ from business ones in exactly the ways that matter. If you cannot answer all 5 questions above for a tool your staff use, you are running an unassessed risk rather than a safe one.

What is a subprocessor list and why does it matter?

A subprocessor list names every third party that touches your data on the vendor's behalf, including the model providers. Your data can be legally exposed to a jurisdiction the vendor never mentions, and the list is the only place that shows up.

How long should AI audit logs be retained?

Long enough to investigate something discovered months later, which in practice means a year rather than a week. Ask about retention and about whether you can query the logs yourself without going through support.

Does using multiple AI models make vendor assessment harder?

It can, unless the models sit behind one platform with one set of terms. ChatFuse routes across more than 100 models under a single agreement and a single data policy, so the assessment is one vendor rather than a dozen. Consolidating that is often the quickest way to shorten a security review.

Print the 5 questions and take them to your next vendor call. The useful signal is not the answers, it is how quickly the vendor can give them and whether they will write them down.

If you work in a regulated field, patient data in ChatGPT walks through what these questions look like when the answer carries a legal obligation.

Start free with ChatFuse, or read the commitments on the security page before you do.

Back to Blog

Written by Nico

Share

Comments

Loading comments…

Secure signup continues in a new tab.