Enterprise AI that does the work: your first governed workflow, live in about two weeks.Map your first AI teammate
← The Operating Layer

Module 02 of 12 Β· 1.5 hours

Hype, Evidence and the Lemon Market

Why the market behaves the way it does, and how to read a claim without a technical team.

Artefact: An evidence ladder, applied to a live proposal on your desk

Share

Why this module exists

Executives are told the problem is hype and the solution is scepticism. That advice is useless, because it supplies no method. It only changes the mood in the room β€” and a sceptical executive with no method makes exactly the same decisions as a credulous one, just more slowly.

This module treats the market as an economic system with knowable properties. Three forces explain almost all of the behaviour you are seeing, and none of them is stupidity. Once you can name them, you can evaluate a claim in minutes instead of quarters, and you can explain your reasoning to a board without sounding like a pessimist.

2.1 What the numbers actually say

Four findings frame the field. Take them together, and take their limits seriously β€” a course that quotes failure statistics uncritically is committing the error it is warning about.

Gartner, June 2025. More than 40% of agentic AI projects will be cancelled by the end of 2027, driven by escalating costs, unclear business value and inadequate risk controls. Gartner also names the supply-side behaviour β€” agent washing, the relabelling of existing chatbots and robotic process automation as agents β€” and estimates that of the thousands of vendors selling agentic AI, roughly 130 are real. Limitation: this is a forecast, not a measurement, and it is drawn substantially from a poll of self-selected webinar attendees.

MIT Project NANDA, July 2025. Roughly 95% of enterprise generative AI pilots produced no measurable business return, against thirty to forty billion dollars of spending. Limitation: "no measurable return" partly measures how few organisations set up measurement in the first place. That is a real finding, but it is a finding about measurement discipline as much as about the technology.

McKinsey, State of AI. The large majority of firms now use AI somewhere, yet only around 39% report EBIT impact at enterprise level, and the strongest correlate of impact is not model choice or vendor β€” it is fundamental workflow redesign. Limitation: self-reported attribution of profit to a technology is among the least reliable numbers in management research. Treat the ranking of factors as more robust than the level.

FactSet, Q1 2026. Of the 498 S&P 500 companies holding earnings calls, 337 mentioned "AI" β€” 68%, the highest in a decade against a ten-year average of 103. And companies that mentioned it outperformed those that did not over the following period. Limitation: correlation with share price is not causation, and the composition of the mentioning group is skewed towards technology.

Read the four together and one reading survives all the caveats:

The industry is not failing at AI. It is failing at placement, measurement and follow-through β€” while the market pays for the word.

That last clause is not cynicism, it is the FactSet data. Once saying the word is rewarded, everything downstream is people responding correctly to a price signal.

2.2 Force one: the demo stopped costing what the system costs

For every enterprise technology before this one, building a convincing demonstration was roughly as expensive as building the thing. You could not demo an ERP without an ERP. You could not fake a warehouse management system in a conference room; the pallets either moved or they did not. The demo was therefore evidence, because producing it required having solved the problem.

A language model severs that link. An afternoon and a well-crafted prompt produce something nearly indistinguishable from a working system, because in both cases the output is a paragraph of confident text.

The demo used to be evidence. Now it is a costume.

Notice who this disarms, because it is counterintuitive. A CFO can read a profit-and-loss statement. An operations director can walk a floor and tell you within ten minutes whether something is real. A CTO can read an architecture diagram and find the lie in it. But nobody in that room, at any level of seniority, can look at a chat window producing fluent paragraphs and tell you whether it will hold up on Tuesday against four thousand real messages.

Seniority normally buys pattern recognition. Here it buys almost nothing β€” and the pretence that it does is how expensive decisions get made quickly.

The practical response is not to refuse demos. It is to insist that a demo is the bottom of the evidence ladder, and to state out loud what would move it up.

2.3 Force two: you are buying in a lemon market

In 1970, George Akerlof asked a question about used cars that turned out to be about a great deal more: what happens to a market when the seller knows the quality and the buyer does not?

His answer runs as a mechanism, and it is worth following step by step because your own instincts are in it:

  1. Buyers cannot distinguish good from bad, so they will only pay a price reflecting average quality.
  2. Sellers of genuinely good goods, worth more than the average, decline to sell at that price and exit.
  3. Average quality therefore falls, so the price buyers will pay falls further.
  4. More good sellers exit. The market unravels toward the lemons.

Akerlof shared the Nobel Prize for this in 2001. Now place Gartner's estimate beside it β€” roughly 130 real vendors among thousands claiming agentic capability. That is not a scandal, and it is not a moral failing peculiar to this industry. It is the predicted equilibrium of a market in which making a claim is free and verifying one takes a quarter.

You are the buyer in that market. And here is the uncomfortable part: your instinctive defences β€” pay the average, hedge, run a small pilot with a small budget β€” are the textbook-correct response to information asymmetry, and they are precisely the behaviour that produces underfunded pilots that never leave the sandbox. Rational individual caution aggregates into the 95% statistic.

Akerlof did not end his paper with outrage. He ended it with remedies, and there were three:

Akerlof's remedyYour procurement equivalent
GuaranteesA pre-registered success metric on your data, with a decision rule agreed before the pilot
CertificationIndependent standards and audits β€” ISO/IEC 42001, SOC 2, and the AI Act conformity work of Module 6
ReputationReference customers you select, including one that churned

You cannot fix an information-asymmetric market by complaining about vendors. You fix it by becoming cheap to verify β€” a theme Module 7 turns into a scorecard.

2.4 Force three: the career arithmetic is asymmetric

Keynes wrote the line in 1936, about investment committees, and it has not aged:

Worldly wisdom teaches that it is better for reputation to fail conventionally than to succeed unconventionally.

Count the payoffs honestly inside your own organisation. An AI project that fails is a project that failed; everyone's failed, the market is hard, next item. Not having one, in 2026, is treated as a personality trait. Nobody has been called into a board meeting to explain a pilot. Plenty of people have been called in to explain the absence of one.

This is not a feeling. Two surveys make it measurable:

  • Dataiku's survey of chief executives found that 74% believe they could lose their job within two years if they do not deliver measurable AI gains.
  • WRITER, with Workplace Intelligence, surveyed 2,400 employees and C-suite leaders in April 2026: 75% of executives said their own company's AI strategy is "more for show" rather than genuine internal guidance, naming public relations and investor relations as the reason it exists. Thirty-nine percent had no formal plan for turning any of it into revenue.

Three quarters, admitted, by the people who commissioned the strategy.

Stack the three forces and the behaviour stops looking like madness and starts looking like an equilibrium: claims that are free to make and expensive to verify, evaluated by people who structurally cannot verify them, inside a system where being conventionally wrong costs nothing.

The leadership consequence is the useful part, and it is small and local. You cannot change the market. You can change what is rewarded in the one room you control β€” which is why Module 9 treats incentives as a governance instrument rather than a soft topic.

2.5 The evidence ladder

Here is the method the module promised. Every claim β€” from a vendor, a consultant, or your own team β€” sits on one of five rungs. The rung determines the size of commitment it justifies.

RungWhat it isWhat it justifies
1. DemonstrationA scripted walkthrough on the vendor's dataA second meeting. Nothing else
2. Case studyA written account you cannot independently checkShortlisting
3. ReferenceA customer you selected and called yourselfA paid pilot
4. PilotYour data, your volumes, a metric and decision rule fixed in advanceA production decision
5. Production evidenceA deployed system with published error rates over timeScaling, and reuse of the pattern elsewhere

Two rules make this operational.

Five rungs of evidence, each justifying a larger commitment5 Production evidencebuys scaling the pattern4 Pilot, pre-registeredbuys a production decision3 Reference you calledbuys a paid pilot2 Case studybuys shortlisting1 Demonstrationbuys a second meetingWEAKEST AT THE BOTTOMName the rung out loud. It turns a sales conversation into a shared workplan.
What each rung of evidence actually buys. Most vendor conversations start at the bottom and are treated as though they started near the top.

Rule one: name the rung out loud. Not as an accusation β€” as bookkeeping. "This is a rung-one claim, which is fine at this stage; here is what would take it to three." It converts a sales conversation into a shared workplan, and vendors who cannot climb the ladder disqualify themselves without an argument.

Rule two: pre-registration is what separates rung four from rung one. A pilot whose success metric is chosen after the results are in is not a pilot. It is a demonstration with a longer runway, and it will be interpreted favourably by whoever sponsored it. Write down the metric, the baseline, the sample and the decision rule before anything runs, and have someone who is not the sponsor hold the document.

2.6 Reading a claim in the meeting

Four questions, in order. They take about two minutes and work without any technical background.

  1. "What is the denominator?" β€” 40% better than what, measured over how many cases? A percentage without a base is decoration.
  2. "Whose data was it on?" β€” Yours, theirs, or a public benchmark? Benchmark performance transfers to your documents far less reliably than anyone hopes.
  3. "What happened to the cases it got wrong?" β€” A vendor who cannot describe their failure modes has not looked. This question separates rungs one and two from three and above faster than any other.
  4. "Who else has stopped using it, and why?" β€” Ask for a reference that churned. The refusal is itself information.

Exercise β€” Put a live proposal on the ladder

Time: 45 minutes. Produces the artefact for this module.

Take the most recent AI proposal that reached your desk β€” vendor or internal, it does not matter.

  1. List its load-bearing claims. The three or four that, if false, would change your decision. Ignore the rest.
  2. Place each on the ladder and mark the rung.
  3. Find the weakest load-bearing claim β€” the one carrying the most weight from the lowest rung. This is your project's actual risk, and it is rarely the one on the risk register.
  4. Write the single question that would move it up one rung, and send it today.
  5. Draft the pre-registration for the pilot you would run: metric, baseline, sample, decision rule, and who holds the document.

Keep the ladder. It applies unchanged to the vendor scorecard in Module 7 and to the evidence dimension of the capstone rubric.

Self-check

  1. A vendor reports "a 40% improvement." Name the two questions that most efficiently establish what that means.
  2. Why does individually rational buyer caution aggregate into the 95% pilot-failure statistic?
  3. What single procedural step distinguishes a pilot from a demonstration, and why must it come first?
  4. Your team says a competitor "has already deployed this." What would you need to see to place that claim above rung two?
  5. Which of Akerlof's three remedies is most available to you as a buyer this quarter, and what would you do with it?

Further reading

  • George A. Akerlof, The Market for "Lemons": Quality Uncertainty and the Market Mechanism, Quarterly Journal of Economics, 1970.
  • John Maynard Keynes, The General Theory of Employment, Interest and Money, Book IV, Chapter 12, 1936.
  • Gartner, Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027, June 2025.
  • MIT Project NANDA, The GenAI Divide: State of AI in Business 2025, July 2025.
  • FactSet Insight, Highest Number of S&P 500 Earnings Calls Citing "AI" Over the Past 10 Years, Q1 2026.

Working through this on a real portfolio?Book a 30-minute call and we will label the steps together β€” including the ones that turn out not to need a model.