Why this module exists
You are buying in the lemon market described in Module 2. Ordinary procurement instincts β issue a requirements document, compare feature matrices, negotiate on price β perform badly there, because the feature matrix is exactly the artefact that is free to fabricate.
Worse, the standard process actively selects for the wrong vendor. A long requirements list rewards whoever is most willing to tick boxes, and the honest vendor who writes "partial" loses to the one who writes "yes."
This module rebuilds procurement around verification. The method is not more scepticism; it is a different sequence.
7.1 Fit, value and risk are three questions
Collapsing them into one number hides the trade-off you are actually making, which is usually "we accepted more risk because it scored well on fit."
Fit is about your estate and your processes. Does it connect to the systems in your Module 4 integration map, at the frequency the slowest hop allows? Does it work the way the process actually runs, including the exceptions? Fit is the dimension most likely to be assessed optimistically, because vendors demonstrate the happy path and your process is mostly exceptions.
Value is about a quantified baseline. Not "improves productivity" β the specific number, today, that this is supposed to move, and by how much. If you cannot state the baseline, you cannot evaluate value from any vendor, and the entire exercise collapses into preference. This is why Module 10's metrics work must start before procurement, not after.
Risk has three components that behave differently:
- Regulatory β the Module 6 analysis: where data goes, what decisions are made, what must be disclosed.
- Operational β what happens when it is wrong, unavailable, or changed underneath you.
- Concentration β how much of your operation depends on one supplier, and what your position is when they reprice.
Score them separately. A vendor that is excellent on fit and value but unacceptable on regulatory risk is not an 8 out of 10; it is a no. Some dimensions are gates, not weights, and deciding which in advance is half the discipline.
7.2 A pilot that decides something
Most AI pilots do not fail. They succeed ambiguously, which is worse, because ambiguity is resolved in favour of whoever sponsored the pilot.
The fix is procedural and takes an hour. Pre-register the pilot before anything runs:
| Element | Question it answers |
|---|---|
| Metric | What single number are we moving? |
| Baseline | What is that number today, measured over what period? |
| Sample | Which cases, how many, chosen how β including the hard ones |
| Duration | How long, and why is that long enough? |
| Decision rule | What result means buy, what result means walk away |
| Custodian | Who holds this document, and it is not the sponsor |
Three failure modes this prevents:
Cherry-picked samples. If the sample is chosen after the vendor has seen the data, you are measuring their selection ability. Include the awkward cases deliberately β the ones your team finds hard. That is where the frontier is.
The moving metric. Without a pre-agreed number, "it wrote good letters" becomes the finding. Good is not a metric.
The sunk-cost close. A decision rule written in advance is the only defence against the argument that you have already invested three months.
One further design note: run the baseline arm. Where you can, have a comparable set handled the current way over the same period. Without it you cannot separate the tool's effect from the effect of paying attention to a process for the first time in years β which is real, large, and not something you need a vendor for.
7.3 Total cost of ownership
Four cost categories. Most business cases include the first and discover the others.
| Category | What it includes | Why it surprises |
|---|---|---|
| Licence | Per seat, per workflow, platform fee | The only one usually quoted |
| Consumption | Inference volume, and its growth | Grows with adoption β success raises the bill |
| Build and integrate | The movement and integration layers from Module 4 | Usually the largest line, and rarely the vendor's problem |
| Own and operate | Monitoring, evaluation, the person who owns it | Permanent; there is no version of this with nobody on it |
Two questions that reveal more than a spreadsheet:
"What does this cost at ten times today's volume?" Consumption pricing means your costs scale with your success. A business case built on pilot volumes can invert at production volumes. Ask for the curve, not the point.
"What does it cost to keep this working for three years?" Model versions change, upstream APIs change, your processes change. Somebody maintains the evaluation set and re-tests. If no name is attached to that, the true answer is that quality will drift downward until someone notices.
Set against these, note the counterweight from Module 1: at a fixed quality bar, inference prices have been falling by roughly an order of magnitude per year. Do not sign a long agreement at today's consumption rates without a repricing mechanism β you would be locking in the most rapidly deflating input in your stack.
7.4 Lock-in and exit
Ask the exit questions at the start, when you have leverage, rather than at renewal, when you have none.
- Where do your prompts and configurations live, and can you export them in a usable form?
- Where do your evaluation sets live? These are the most valuable asset you will build β a curated set of cases with known-good answers, representing years of accumulated judgement about your domain. Losing them means starting quality measurement from zero with the next vendor.
- Where do your traces live, and can you take them?
- Is the model swappable, or is the platform bound to one provider? Given the pace of the market, being unable to change model is a real cost.
- What survives termination? Specifically: for how long can you read your own history?
7.5 The contract terms that matter
Beyond ordinary commercial terms, six clauses are specific to this class of system.
| Term | What to secure | Why |
|---|---|---|
| Data use and training rights | Explicit prohibition on training on your data, or an explicit, bounded permission | The default is not always what you assume, and it varies by tier |
| Model change and deprecation | Notice period; the right to test before a forced change | A model change is a silent product change. Your evaluation set is how you detect it |
| Performance representation | Something measurable, tied to your pre-registered metric | Converts marketing into a term |
| Sub-processors | List, notice of change, right to object | Required for the Module 6 analysis; also concentration risk |
| Audit and evidence | Access to logs, or delivery of traces to you | You cannot evidence oversight you cannot see |
| Output indemnity | Position on intellectual property and third-party claims | Increasingly negotiable; almost never offered unprompted |
For anyone in scope of Module 6, add the data processing agreement and, where protected health information is involved, the business associate agreement. "We are compliant" is a marketing sentence. The document is the control.
7.6 Reference calls that produce truth
A reference supplied by a vendor is a rung-two claim (Module 2) until you have selected it yourself. Four questions produce more signal than an hour of general conversation:
- "What did you have to build yourselves that you expected to be included?" This surfaces the true scope boundary better than any statement of work.
- "What surprised you in month four?" Month one is onboarding enthusiasm. Month four is when the real behaviour appears.
- "What is your override rate, and has it moved?" Applies Module 5's test to someone else's deployment.
- "What would you not buy again?" Almost everyone answers this honestly, because it is not a question about the vendor.
And the one request that separates serious vendors from the rest: ask for a customer who stopped using the product. The refusal is itself information; the willingness is a strong signal; and the call, if it happens, is usually the most useful hour of the entire evaluation.
7.7 The scorecard
Two rules, and the second is the one that does the work.
Rule one: gates before weights. Some criteria are pass/fail β data residency, a required certification, the availability of a business associate agreement. Apply them first and eliminate. Do not let a strong score elsewhere buy a way past a gate.
Rule two: set the weights before you score anyone. This is the single most effective de-biasing step available in procurement, and it costs nothing. Weights chosen after you have met the vendors will, reliably and unconsciously, be the weights that favour the vendor you liked.
A workable structure:
| Dimension | Weight | Scored on |
|---|---|---|
| Fit β process | 20% | Handles the exceptions, not just the happy path |
| Fit β integration | 15% | Works within the slowest hop of your Module 4 map |
| Value β evidenced | 25% | Result of your pre-registered pilot, against baseline |
| Risk β regulatory | Gate | Module 6 analysis passes, or it does not |
| Risk β operational | 15% | Failure modes, support, and what happens when it is wrong |
| Risk β concentration and exit | 10% | Portability of prompts, evaluation sets and traces |
| Total cost at 10Γ volume | 15% | Three-year, all four categories |
Score 1 to 5 with a written justification per cell. The justifications, not the total, are what you will defend in six months β and writing them is what exposes the cells where you have no evidence at all.
Exercise β Complete the scorecard
Time: 90 minutes across two sittings. Produces the artefact for this module.
Sitting one, before you look at any vendor.
- Write your gates. Which criteria are pass/fail, and why.
- Set your weights, and have someone else sign off that they were set first.
- Write the pre-registration for the pilot: metric, baseline, sample, duration, decision rule, custodian.
Sitting two, after the pilot.
- Score two real options β and make one of them the assembly option, even if nobody has proposed it.
- Write the justification in every cell. Cells you cannot justify are the questions for your next vendor call.
- Write the one-paragraph recommendation you would sign your name to, including the strongest argument against it.
Self-check
- Why should weights be set before vendors are scored, and what goes wrong when they are not?
- Which criteria in your organisation should be gates rather than weighted dimensions?
- What is the difference between a pilot and a demonstration, expressed as a single procedural requirement?
- Why is the evaluation set the asset to protect in a contract, and where does yours currently live?
- Your business case assumes pilot-level consumption. What question tells you whether it survives production?
Further reading
- George A. Akerlof, The Market for "Lemons", 1970 β the remedies section, read as procurement design.
- ISO/IEC 42001:2023, AI management systems β as a certification signal in vendor assessment.
- Module 6 of this course, for the regulatory gates that precede scoring.