Enterprise AI that does the work: your first governed workflow, live in about two weeks.Map your first AI teammate
← The Operating Layer

Module 07 of 12 Β· 1.5 hours

Evaluating Platforms and Vendors

A procurement method for a market where claims are free and verification is expensive.

Artefact: A weighted evaluation scorecard, completed for two real options

Share

Why this module exists

You are buying in the lemon market described in Module 2. Ordinary procurement instincts β€” issue a requirements document, compare feature matrices, negotiate on price β€” perform badly there, because the feature matrix is exactly the artefact that is free to fabricate.

Worse, the standard process actively selects for the wrong vendor. A long requirements list rewards whoever is most willing to tick boxes, and the honest vendor who writes "partial" loses to the one who writes "yes."

This module rebuilds procurement around verification. The method is not more scepticism; it is a different sequence.

7.1 Fit, value and risk are three questions

Collapsing them into one number hides the trade-off you are actually making, which is usually "we accepted more risk because it scored well on fit."

Fit is about your estate and your processes. Does it connect to the systems in your Module 4 integration map, at the frequency the slowest hop allows? Does it work the way the process actually runs, including the exceptions? Fit is the dimension most likely to be assessed optimistically, because vendors demonstrate the happy path and your process is mostly exceptions.

Value is about a quantified baseline. Not "improves productivity" β€” the specific number, today, that this is supposed to move, and by how much. If you cannot state the baseline, you cannot evaluate value from any vendor, and the entire exercise collapses into preference. This is why Module 10's metrics work must start before procurement, not after.

Risk has three components that behave differently:

  • Regulatory β€” the Module 6 analysis: where data goes, what decisions are made, what must be disclosed.
  • Operational β€” what happens when it is wrong, unavailable, or changed underneath you.
  • Concentration β€” how much of your operation depends on one supplier, and what your position is when they reprice.

Score them separately. A vendor that is excellent on fit and value but unacceptable on regulatory risk is not an 8 out of 10; it is a no. Some dimensions are gates, not weights, and deciding which in advance is half the discipline.

7.2 A pilot that decides something

Most AI pilots do not fail. They succeed ambiguously, which is worse, because ambiguity is resolved in favour of whoever sponsored the pilot.

The fix is procedural and takes an hour. Pre-register the pilot before anything runs:

ElementQuestion it answers
MetricWhat single number are we moving?
BaselineWhat is that number today, measured over what period?
SampleWhich cases, how many, chosen how β€” including the hard ones
DurationHow long, and why is that long enough?
Decision ruleWhat result means buy, what result means walk away
CustodianWho holds this document, and it is not the sponsor

Three failure modes this prevents:

Cherry-picked samples. If the sample is chosen after the vendor has seen the data, you are measuring their selection ability. Include the awkward cases deliberately β€” the ones your team finds hard. That is where the frontier is.

The moving metric. Without a pre-agreed number, "it wrote good letters" becomes the finding. Good is not a metric.

The sunk-cost close. A decision rule written in advance is the only defence against the argument that you have already invested three months.

One further design note: run the baseline arm. Where you can, have a comparable set handled the current way over the same period. Without it you cannot separate the tool's effect from the effect of paying attention to a process for the first time in years β€” which is real, large, and not something you need a vendor for.

7.3 Total cost of ownership

Four cost categories. Most business cases include the first and discover the others.

CategoryWhat it includesWhy it surprises
LicencePer seat, per workflow, platform feeThe only one usually quoted
ConsumptionInference volume, and its growthGrows with adoption β€” success raises the bill
Build and integrateThe movement and integration layers from Module 4Usually the largest line, and rarely the vendor's problem
Own and operateMonitoring, evaluation, the person who owns itPermanent; there is no version of this with nobody on it

Two questions that reveal more than a spreadsheet:

"What does this cost at ten times today's volume?" Consumption pricing means your costs scale with your success. A business case built on pilot volumes can invert at production volumes. Ask for the curve, not the point.

"What does it cost to keep this working for three years?" Model versions change, upstream APIs change, your processes change. Somebody maintains the evaluation set and re-tests. If no name is attached to that, the true answer is that quality will drift downward until someone notices.

Set against these, note the counterweight from Module 1: at a fixed quality bar, inference prices have been falling by roughly an order of magnitude per year. Do not sign a long agreement at today's consumption rates without a repricing mechanism β€” you would be locking in the most rapidly deflating input in your stack.

7.4 Lock-in and exit

Ask the exit questions at the start, when you have leverage, rather than at renewal, when you have none.

  • Where do your prompts and configurations live, and can you export them in a usable form?
  • Where do your evaluation sets live? These are the most valuable asset you will build β€” a curated set of cases with known-good answers, representing years of accumulated judgement about your domain. Losing them means starting quality measurement from zero with the next vendor.
  • Where do your traces live, and can you take them?
  • Is the model swappable, or is the platform bound to one provider? Given the pace of the market, being unable to change model is a real cost.
  • What survives termination? Specifically: for how long can you read your own history?

7.5 The contract terms that matter

Beyond ordinary commercial terms, six clauses are specific to this class of system.

TermWhat to secureWhy
Data use and training rightsExplicit prohibition on training on your data, or an explicit, bounded permissionThe default is not always what you assume, and it varies by tier
Model change and deprecationNotice period; the right to test before a forced changeA model change is a silent product change. Your evaluation set is how you detect it
Performance representationSomething measurable, tied to your pre-registered metricConverts marketing into a term
Sub-processorsList, notice of change, right to objectRequired for the Module 6 analysis; also concentration risk
Audit and evidenceAccess to logs, or delivery of traces to youYou cannot evidence oversight you cannot see
Output indemnityPosition on intellectual property and third-party claimsIncreasingly negotiable; almost never offered unprompted

For anyone in scope of Module 6, add the data processing agreement and, where protected health information is involved, the business associate agreement. "We are compliant" is a marketing sentence. The document is the control.

7.6 Reference calls that produce truth

A reference supplied by a vendor is a rung-two claim (Module 2) until you have selected it yourself. Four questions produce more signal than an hour of general conversation:

  1. "What did you have to build yourselves that you expected to be included?" This surfaces the true scope boundary better than any statement of work.
  2. "What surprised you in month four?" Month one is onboarding enthusiasm. Month four is when the real behaviour appears.
  3. "What is your override rate, and has it moved?" Applies Module 5's test to someone else's deployment.
  4. "What would you not buy again?" Almost everyone answers this honestly, because it is not a question about the vendor.

And the one request that separates serious vendors from the rest: ask for a customer who stopped using the product. The refusal is itself information; the willingness is a strong signal; and the call, if it happens, is usually the most useful hour of the entire evaluation.

7.7 The scorecard

Two rules, and the second is the one that does the work.

Rule one: gates before weights. Some criteria are pass/fail β€” data residency, a required certification, the availability of a business associate agreement. Apply them first and eliminate. Do not let a strong score elsewhere buy a way past a gate.

Rule two: set the weights before you score anyone. This is the single most effective de-biasing step available in procurement, and it costs nothing. Weights chosen after you have met the vendors will, reliably and unconsciously, be the weights that favour the vendor you liked.

Vendors passing through gates before being scored on weighted criteriaSTEP 1 β€” GATES (PASS / FAIL)residencycertificationBAA / DPAexit termsSTEP 2 β€” WEIGHTS, SET BEFORE ANY VENDOR IS SEENfit 35%value 25%risk 25%cost 15%A strong score never buys a way past a gate β€” that is what makes it a gate.Weights chosen after meeting vendors will favour the vendor you liked.
Gates eliminate; weights rank. Doing it in the other order is how a strong feature score buys a way past a regulatory requirement.

A workable structure:

DimensionWeightScored on
Fit β€” process20%Handles the exceptions, not just the happy path
Fit β€” integration15%Works within the slowest hop of your Module 4 map
Value β€” evidenced25%Result of your pre-registered pilot, against baseline
Risk β€” regulatoryGateModule 6 analysis passes, or it does not
Risk β€” operational15%Failure modes, support, and what happens when it is wrong
Risk β€” concentration and exit10%Portability of prompts, evaluation sets and traces
Total cost at 10Γ— volume15%Three-year, all four categories

Score 1 to 5 with a written justification per cell. The justifications, not the total, are what you will defend in six months β€” and writing them is what exposes the cells where you have no evidence at all.

Exercise β€” Complete the scorecard

Time: 90 minutes across two sittings. Produces the artefact for this module.

Sitting one, before you look at any vendor.

  1. Write your gates. Which criteria are pass/fail, and why.
  2. Set your weights, and have someone else sign off that they were set first.
  3. Write the pre-registration for the pilot: metric, baseline, sample, duration, decision rule, custodian.

Sitting two, after the pilot.

  1. Score two real options β€” and make one of them the assembly option, even if nobody has proposed it.
  2. Write the justification in every cell. Cells you cannot justify are the questions for your next vendor call.
  3. Write the one-paragraph recommendation you would sign your name to, including the strongest argument against it.

Self-check

  1. Why should weights be set before vendors are scored, and what goes wrong when they are not?
  2. Which criteria in your organisation should be gates rather than weighted dimensions?
  3. What is the difference between a pilot and a demonstration, expressed as a single procedural requirement?
  4. Why is the evaluation set the asset to protect in a contract, and where does yours currently live?
  5. Your business case assumes pilot-level consumption. What question tells you whether it survives production?

Further reading

  • George A. Akerlof, The Market for "Lemons", 1970 β€” the remedies section, read as procurement design.
  • ISO/IEC 42001:2023, AI management systems β€” as a certification signal in vendor assessment.
  • Module 6 of this course, for the regulatory gates that precede scoring.

Working through this on a real portfolio?Book a 30-minute call and we will label the steps together β€” including the ones that turn out not to need a model.