Skip to content
AI & Automation4 min read

How to Evaluate an AI Vendor's Claims Before You Sign

Twelve questions that separate AI vendors with working systems from vendors with good demonstrations, plus the contract terms that matter most.

The short answer

Evaluate an AI vendor by asking for a pilot on your own data with a pre-agreed accuracy target, a written description of how quality is measured and monitored, and clear terms on data use, model changes, and exit. A vendor who cannot show an evaluation method or explain what happens when the model is wrong is selling a demonstration.

Key takeaways

  • Demonstrations are curated. Insist on a pilot against your own messy data before you commit.
  • Ask how accuracy is measured and on what test set. A vendor without an evaluation method has no way to know their own quality.
  • Get in writing whether your data trains anyone's models, where it is processed, and what is retained.
  • Model changes are the hidden risk. Ask what happens to your results when the underlying model is upgraded or deprecated.
  • Price the exit before you sign. If you cannot describe how you would leave, you have no negotiating position at renewal.

AI vendor evaluation is harder than ordinary software evaluation for one structural reason: the product is probabilistic, and probabilistic systems demonstrate beautifully. A curated demo on clean examples tells you almost nothing about how the system behaves on your Tuesday-afternoon reality.

The questions below are the ones that produce informative answers. What matters is less the specific response than whether the vendor has clearly thought about it before you asked.

Questions about accuracy

How do you measure accuracy, and on what test set?

The answer you want describes a fixed set of labelled examples, a defined scoring method, and results tracked over time. The answer that should worry you is a percentage with no methodology attached, or a shrug toward customer satisfaction. A vendor without an evaluation harness cannot tell you their quality, because they do not know it.

What does the system do when it is uncertain?

Well-built systems have a confidence mechanism and an escalation path. Poorly built ones return their best guess with the same authority regardless. Ask to see what an escalation looks like in the interface.

Show me your worst cases

This is the most informative question on the list. A vendor with a real system in production has a catalogue of failure modes and will describe them readily, because they have spent months working on them. A vendor without one will deflect. The willingness to answer matters more than the answer.

Questions about your data

  • Is our data used to train or fine-tune any model, yours or a third party's? Get this in the contract, not in an email.
  • Which subprocessors touch our data, and in what jurisdictions is it processed and stored?
  • What is retained after processing, for how long, and can we require deletion?
  • If we terminate, what happens to our data and any model artefacts derived from it?
  • Is there an audit log of what the system did, sufficient to reconstruct a specific decision months later?

The audit question is the one most often overlooked and most often needed. When a customer disputes an outcome, or a regulator asks, or an internal investigation starts, you need to be able to say what the system saw and what it produced on a specific date. Systems that cannot do this create a category of problem that has nothing to do with accuracy.

Questions about change

Underlying models change. They are upgraded, retired, and re-tuned by their providers on schedules the vendor does not control. This is the most underexamined risk in the category.

AskA good answer sounds likeA concerning answer sounds like
What happens when the underlying model is upgraded?We re-run our evaluation suite and hold the previous version until the new one passesWe get the improvements automatically
Can we pin a version?Yes, with a defined support window and notice before deprecationYou are always on the latest
How will we know if quality drifts?Scheduled scored runs, alerting on a threshold, visible to youCustomers tell us
Who owns prompts and configuration built for us?You do, and we will export themThat is our proprietary layer

The last row deserves attention. If the accumulated configuration, prompts, and rules that make the system work on your business are the vendor's property, then your switching cost grows every month you use it and you have no leverage at renewal.

Questions about cost

  • How is this priced: per seat, per run, per token, or flat? Model it at three times current volume.
  • What happens if volume spikes unexpectedly? Is there a ceiling, and who is liable above it?
  • Are there charges for reprocessing when something fails and has to run again?
  • What is included in support, and what is billed?

Usage-based pricing is fair in principle and occasionally alarming in practice, because a workflow that suddenly runs ten times as often produces a bill nobody approved. Ask for a hard ceiling with alerting, and treat reluctance as information.

Questions about the exit

Price the exit before you sign. Not because you expect to leave, but because a relationship you cannot leave is a relationship you cannot negotiate.

  1. Can we export our data in a documented, usable format, and is there a charge?
  2. Do we keep the prompts, rules, and configuration developed for us?
  3. What is the notice period, and what happens operationally during it?
  4. If you are acquired or shut down, what are our rights, and is there an escrow arrangement?
If you cannot describe how you would leave a vendor, you are not a customer with options. Your renewal conversation will reflect that.

Signals that should slow you down

  • Accuracy quoted as a single number with no test set described.
  • A refusal to run a paid pilot on your own data.
  • No named failure modes, or visible discomfort at the question.
  • Quality monitoring described as something the customer does.
  • Contract language that is silent on training use of your data.
  • Configuration built for you that is claimed as the vendor's property.
  • Pricing that has no ceiling and no alerting.

None of these are automatically disqualifying. Early companies genuinely do have thin evaluation practice, and a strong team with a weak harness can be a good bet. But each one should move the price, the pilot conditions, or the contract terms in your favour, and a vendor who bristles at the questions is showing you how the relationship will run.

The shortest version

If you do only three things: insist on a paid pilot against your own data with a pre-agreed target, get data use and model-change behaviour in writing, and make sure the configuration built for your business belongs to you. Those three cover most of the ways these purchases go wrong.

Frequently asked questions

What should I ask an AI vendor about accuracy?

Ask how accuracy is measured, on what labelled test set, and how results are tracked over time. Ask what the system does when it is uncertain, and ask them to describe their worst failure modes. A vendor with a production system answers all three readily; a vendor without an evaluation method does not know their own quality.

Should I let an AI vendor pilot on my own data?

Yes, and you should insist on it. A demonstration runs on curated examples and tells you little about your real inputs. Agree an accuracy target in writing before the pilot starts, along with a stated consequence if it is missed.

What contract terms matter most when buying AI software?

Whether your data trains any model, which subprocessors and jurisdictions are involved, retention and deletion rights, audit logging sufficient to reconstruct a past decision, ownership of prompts and configuration built for you, version pinning and notice on model changes, and a documented data export on exit.

What is the risk when a vendor's underlying model changes?

Quality can shift without warning, because model providers upgrade and deprecate on schedules the vendor does not control. Ask whether the vendor re-runs an evaluation suite before adopting a new version, whether you can pin a version with a defined support window, and how you would be told if quality drifted.

How should AI software be priced?

Any model can be fair, but model it at three times your current volume before signing and insist on a hard ceiling with alerting. Ask specifically whether failed runs that must be reprocessed are billed again, since that charge tends to appear only after something goes wrong.

More questions are answered on the frequently asked questions page.

Your situation

General advice only gets you so far.

Describe the business and the constraint you are hitting. You will get a written assessment back within two business days from the person who would run the work.