promptrails
Executive course
Selling to Machines
E-commerce, retail and growth in the agentic era β visibility, protocols, operations and economics
A self-paced course on selling when the shopper may be a model or an agent: how generative engines choose what to cite, how to make a catalogue machine-legible, how to win the off-site citations you do not own, how to measure demand you cannot see, and what to do about agentic checkout. Thirteen modules, each producing an artefact.
- Modules
- 13
- Study time
- About 22 hours
- Level
- Practitioner to executive β technical depth explained, not skipped
- Edition
- 24 August 2026
https://promptrails.ai/courses/selling-to-machines
Learning outcomes
- Explain how generative engines and shopping agents select, cite and recommend products
- Distinguish the four AI surfaces in commerce and locate where your demand actually sits
- Audit a storefront for machine-legibility and produce a prioritised remediation specification
- Specify a product feed and structured-data layer an engine and an agent can both consume
- Set and defend an agent-access policy, balancing visibility against crawl cost and control
- Win and correct the third-party citations that ground answers about you, and stay the right side of the law that governs them
- Measure AI-influenced demand when the referrer is missing, and report it credibly to a CFO
- Evaluate the agentic checkout protocols and decide what to implement, when, and what to defer
- Design trust, fraud, dispute and returns handling for agent-initiated orders
- Adapt merchandising, pricing and assortment for a buyer that compares exhaustively
- Plan paid media across AI surfaces, separating live placements from announced ones
- Assess where durable advantage sits in agentic commerce, and what commoditises
- Reshape the team and the agency relationship around the work that actually changed
- Differentiate evidenced findings from vendor claims in a market generating both at speed
- Produce a twelve-month plan with sequencing, owners, measurement and a rollback position
Contents
- 01What Actually ChangedThree shifts in how demand reaches a retailer β and the much longer list of things that did not move at all.Artefact: A one-page brief separating what has changed from what has not1.5h
- 02The Four AI SurfacesAI search, on-platform assistants, third-party agents and your own β four different games, four different owners, one budget.Artefact: A surface map showing where your demand actually sits and who owns each surface1.5h
- 03LLM Visibility I β How Engines Choose What to CiteThe retrieval pipeline behind a generative answer, what the evidence says actually moves citation, and how to measure your position before you spend anything.Artefact: A prompt set and a dated baseline of your citation position2h
- 04LLM Visibility II β Making the Catalogue LegibleFeeds, structured data, rendering and access policy β the four layers that decide whether a machine can read what you sell.Artefact: A machine-legibility audit and a prioritised remediation specification2.5h
- 05LLM Visibility III β Winning the Citations You Do Not OwnMost of the sources grounding answers about you sit on somebody else's domain. This is the module about them.Artefact: A citation source register and a 90-day off-site plan with named owners2h
- 06Measurement When the Referrer DisappearsCounting AI-influenced demand you cannot see in analytics, and reporting it in a form a finance director will accept.Artefact: An AI-influenced demand model, with its stated confidence, that your finance team accepts2h
- 07Agentic Checkout and the Protocol StackACP, UCP and the payment rails beneath them β what exists, what it asks of you, and how to decide what to implement.Artefact: A protocol decision with a dated implementation position and a review trigger2h
- 08Trust, Fraud, Disputes and ReturnsWhat breaks operationally when the buyer is software, and the policy that has to exist before the first agent order arrives.Artefact: An agent-order policy covering authorisation, evidence, disputes and returns1.5h
- 09Merchandising for a Machine BuyerWhat changes when the shopper compares exhaustively, never sees a banner, and reads your specification instead of your copy.Artefact: A revised attribute, pricing and assortment specification1.5h
- 10Paid Media on AI SurfacesWhich placements exist, which are being sold as though they do, and what happens to performance marketing when the click is not the point.Artefact: A paid plan that separates live placements from roadmap items1.5h
- 11Advantage, Moats and the Operating ModelWhat commoditises, what compounds, and how the team and the agency relationship have to change.Artefact: An advantage thesis and a team and agency design1.5h
- 12Appendix: The VocabularyEvery term this course uses, defined once β because in a market this noisy, a shared definition is a negotiating position.Artefact: A vocabulary you can hold a vendor and an agency to0.5h
- 13Capstone: The Twelve-Month PlanOne assessed deliverable, assembled from the artefacts you built along the way.Artefact: The completed plan, assessed against a published rubric2h
The running case
A single running case, Marlowe & Finch, threads through every module: a fictional European home and lifestyle retailer of β¬182m revenue, with direct e-commerce, marketplaces, wholesale, eleven stores and an agency on retainer β chosen because it exercises every profile in the audience at once.
Module 01
What Actually Changed
Three shifts in how demand reaches a retailer β and the much longer list of things that did not move at all.
Artefact: A one-page brief separating what has changed from what has not
Why this module exists
Two failure modes bracket this subject, and both are expensive.
The first is treating agentic commerce as a rebrand of search β a new channel to bolt onto the acquisition deck, optimised by the same team with the same playbook. That reading fails because the buyer on the other side is no longer necessarily a person, and much of the playbook assumes one.
The second is treating it as a revolution in which everything is new. That reading fails faster, because it discards the parts of retail that have not moved an inch β margin, availability, fulfilment cost, returns rate, brand β and those are still where the money is made or lost.
This module draws the line between the two. Everything afterwards depends on drawing it in the right place.
1.1 The three shifts
Shift one: discovery moved from a list to an answer
For twenty-five years, discovery meant a ranked list of links and a human deciding among them. The retailer's job was to appear high on the list and win the click.
Generative engines return an answer instead β a synthesis, with a handful of sources cited beneath it. The list is still there underneath, but a growing share of shoppers stop at the answer. This is the shift most widely reported and most widely over-interpreted: the list has not disappeared, but the number of impressions that convert into a visit has changed, and so has which sources get named.
The operational consequence is precise: you are no longer competing for a rank, you are competing to be one of the few sources a model quotes. That is a different mechanism, measured differently, and Modules 3 and 4 take it apart.
Shift two: the buyer may not be a person
An agent can now hold a shopping mandate β a budget, a preference set, a constraint list β and act on it: comparing, selecting, and in a growing number of cases completing checkout without the human returning to the page.
This has been building through a series of concrete releases rather than as an abstraction:
| When | What |
|---|---|
| 29 September 2025 | OpenAI and Stripe publish the Agentic Commerce Protocol (ACP) under Apache 2.0, and ChatGPT Instant Checkout goes live for US buyers |
| AprilβOctober 2025 | Visa Intelligent Commerce and Mastercard Agent Pay are announced with pilots; Visa follows with the Trusted Agent Protocol in October |
| 11 January 2026 | Google announces the Universal Commerce Protocol (UCP) at NRF, co-developed with Shopify, Etsy, Wayfair, Target and Walmart |
| 16 February 2026 | OpenAI relaunches in-chat buying as "Buy it in ChatGPT," extended to more merchants |
| 4 March 2026 | OpenAI withdraws in-chat checkout, keeps discovery, and routes buying back to merchants and third-party apps |
Read that last row carefully, because it is the most instructive fact in this course. In-chat purchase β the capability every 2025 strategy deck was built around β was live for roughly five months and reached, by most accounts, around a dozen Shopify merchants before it was pulled. The stated reasons were not model quality. They were sales tax, fraud prevention and synchronising real-time inventory across millions of listings: Modules 4, 6 and 7 of this course.
The lesson is not that agentic commerce is hype. ACP continues, UCP launched, and the payment rails kept shipping. The lesson is that the boring operational layer is the binding constraint, and the company with the most capable model in the world could not route around it.
Module 7 covers what to do about this. Here, note only the change in category: some share of your orders will arrive from a piece of software acting under instruction, and your systems currently assume a human.
Shift three: your storefront became an API whether you built one or not
Agents and generative engines read your site. Not the way a shopper does β they consume your feed, your structured data, your rendered HTML, and whatever your robots policy permits.
That makes machine-legibility a commercial property rather than a technical one. Adobe's analysis of US retail sites in 2026 scored individual product pages at 66% on machine-readability, the worst-performing page type, with category pages at 74% and homepages at 75%. The pages carrying the actual product information were the least legible to the systems increasingly deciding which products get mentioned.
1.2 AI-influenced versus agent-executed
These get conflated constantly, and they demand different work from different teams. Hold them apart from the start.
| AI-influenced | Agent-executed | |
|---|---|---|
| What happens | A person asks an engine, reads the answer, then buys β possibly much later, possibly elsewhere | Software selects and transacts under a mandate |
| Who lands on your site | A human | A program, or nobody at all |
| Your lever | Being cited, and being legible when they arrive | Being reachable, priced and available through a protocol |
| Fails as | Invisibility you cannot see in analytics | Orders you cannot authenticate, dispute or return cleanly |
| Owned by | Growth, content, SEO/AEO | Commerce platform, payments, risk, operations |
| Volume today | The large majority | Small but compounding |
The practical instruction: work on the first, prepare for the second. Today's revenue is overwhelmingly AI-influenced, and that is where a quarter of effort returns something measurable. Agent-executed volume is small β but the work it requires (clean feeds, reliable availability, a dispute policy) takes quarters to build, and is largely the same work that improves the first case anyway.
That overlap is the single most useful fact in this course. It means you are rarely choosing between the two.
1.3 What has not changed
A vendor deck will tell you everything is new. Here is the list that is not, and every item on it still decides whether you make money.
Unit economics. Contribution margin after fulfilment, returns and payment costs is unchanged. An agent-sourced order at negative contribution is worse than no order, and easier to acquire in volume.
Availability. An agent that recommends an out-of-stock item does it once. Feed freshness is now a discovery input, not just an operational hygiene metric β Module 4 makes this concrete.
Returns. Nothing about a machine buyer reduces returns; several things about exhaustive comparison and reduced browsing may raise them. Your returns rate is still the difference between a good quarter and a bad one.
Brand. Models are trained and grounded on what the world says about you. Third-party mentions, reviews and coverage feed the systems now deciding whether to name you β which makes brand a retrieval input as well as a demand one.
Price and assortment. Comparison got cheaper and more exhaustive. That does not change what a good price is; it changes how quickly a bad one is found.
1.4 Reading this market without being taken
The numbers in agentic commerce are unusually noisy, and you will be shown a great many of them. Three habits are enough to protect you.
Insist on a date. This field moves monthly. Adobe's own data shows AI-referred traffic to US retail converting 38% worse than other traffic in March 2025, 42% better in March 2026, and 54% better by May 2026. All three are real. A claim without a date is not a claim.
Ask who is selling. A large share of published GEO and AI-visibility statistics come from firms selling GEO and AI-visibility tools. That does not make them wrong; it makes them unaudited. This course marks such figures explicitly, and you should too.
Separate traffic from money, and watch the second derivative. Percentage growth on a small base is the easiest impressive number in the industry. Adobe reported AI-referred traffic to US retail growing 393% year over year in Q1 2026 β and 138% year over year in May 2026. Both are real, both are large, and the deceleration between them is the more useful fact. A deck quoting only the first, five months later, is telling you something about the deck.
The question your finance team will ask is what share of revenue this represents today, which is a different number again β and Module 6 is about producing it honestly.
Exercise β The one-page brief
Time: 45 minutes. Produces the artefact for this module.
Write one page, for your own leadership, in three parts.
- What changed for us. Take the three shifts and say, specifically, which of them touches your business and where. Not "discovery is changing" β which queries, which categories, which margin.
- What did not. Name the five economics of your business that are unaffected, and state that they remain the constraint. This paragraph is what stops the programme becoming a technology project.
- The base rate. Find your current AI-identifiable referral share of sessions and of revenue. If your analytics cannot tell you, write that down β it is Module 6's problem and it is a finding, not a gap.
The brief is finished when someone who disagrees with you could argue with it. Vagueness is what makes a document unarguable.
Self-check
- A colleague says "AI traffic converts better, so we should shift budget." Which two questions establish whether that instruction is sound for your business?
- Give an example from your own catalogue of work that improves both AI-influenced and agent-executed outcomes at once.
- Which of the five unchanged economics is most exposed if agent-driven comparison becomes routine in your category?
- A vendor cites a 300% growth figure. What do you ask before it enters a plan?
- What share of your revenue is AI-identifiable today, and how confident are you in that number?
Further reading
- Adobe, Generative AI-powered shopping and traffic to US retail sites, Adobe Digital Insights, 2025β2026 β the quarterly series, read with the dates attached.
- Agentic Commerce Protocol specification, OpenAI and Stripe, from September 2025.
- Google, New tech and tools for retailers to succeed in an agentic shopping era, 11 January 2026 β the UCP announcement.
Module 02
The Four AI Surfaces
AI search, on-platform assistants, third-party agents and your own β four different games, four different owners, one budget.
Artefact: A surface map showing where your demand actually sits and who owns each surface
Why this module exists
"AI commerce" is not one channel. It is four, and they differ on every dimension that matters: who controls the surface, whether you can pay to appear, whether you can measure it, and what it costs to be there.
Teams that treat it as one thing produce a single undifferentiated workstream, usually owned by whoever is nearest to SEO, and are then surprised that half the work has no effect on the surface where their customers actually are.
This module separates them, and gives you an instrument for finding out which ones matter to you specifically β because the answer varies enormously by category and by market.
2.1 The four surfaces
Surface one: AI search
Generative answers on general search and assistant surfaces β Google AI Mode and AI Overviews, ChatGPT, Perplexity, Copilot, Gemini.
- Control: none. You influence citation, you do not place content.
- Payment: emerging and uneven. Sponsored formats exist in some surfaces and not in others; see Module 10.
- Measurement: poor by default, improvable. Module 6.
- Your lever: be citable, and be legible when the click lands. Modules 3 and 4.
This is where most of the addressable change sits today for most retailers, because it touches the large majority of demand that is still human.
Surface two: on-platform assistants
The AI inside a marketplace you already sell on. On Amazon this was Rufus, merged with Alexa+ on 13 May 2026 and rebranded Alexa for Shopping β a single assistant spanning the Shopping app, the website and Echo devices. In Europe the equivalents matter at least as much: bol.com in the Netherlands and Belgium, Zalando and Otto in Germany, each with its own assistant and its own listing rules.
Names matter here in a way they do not elsewhere in this course: a marketplace manager cannot find documentation, brief an agency or locate an ad setting for "the assistant."
- Control: none over the assistant, full over your listing.
- Payment: yes, and increasingly formalised β Amazon's Sponsored Prompts became billable on 25 March 2026, priced inside existing Sponsored Products and Sponsored Brands auctions rather than as a separate bid.
- Measurement: limited. Amazon does not expose assistant-specific attribution in Seller Central; brand-registered sellers see partial signals in Brand Analytics. Look for shifts in impression share on long, conversational queries.
- Your lever: structured attributes, review substance, Q&A and A+ content β the assistant reads all of it.
Amazon has cited around $12bn in incremental annualised sales attributed to the assistant, on its Q4 2025 earnings call, alongside a user base in the hundreds of millions.
Surface three: third-party agents
Software acting for a shopper: an assistant completing a purchase, a comparison agent, a procurement bot.
- Control: none, and it may not render your page at all.
- Payment: not meaningfully, yet.
- Measurement: poor; often indistinguishable from bot traffic unless you deliberately identify it.
- Your lever: protocol support, feed quality, availability accuracy, and an access policy that lets the right agents in. Modules 4, 6 and 7.
Volume here is small today. The work is long. That combination is why it belongs on a roadmap rather than in a quarter.
Surface four: your own agent
The assistant you run β on your site, in your app, in your service channels.
- Control: total.
- Payment: it is a cost, not a placement.
- Measurement: excellent; it is your own telemetry.
- Your lever: everything, which is exactly why it is the surface most often built first and justified last.
Note the asymmetry: this is the only surface you fully control, and the only one where the constraint is your own execution rather than someone else's platform. It is also the one where a poor implementation does direct brand damage, because the shopper knows it is you.
2.2 The comparison, in one table
| Control | Can you pay | Measurable | Time to value | Usual owner | |
|---|---|---|---|---|---|
| AI search | None | Partly, unevenly | Poor β fair | Weeks to months | Growth / SEO β often renamed AEO |
| On-platform | Listing only | Yes | Poor | Weeks | Marketplace team |
| Third-party agents | None | Not yet | Poor | Quarters | Nobody, usually |
| Your own agent | Total | N/A | Excellent | Months | Product / CX |
The column that decides your next quarter is time to value. The column that decides your next year is usual owner β and specifically the row where the honest answer is "nobody."
2.3 Finding out where your demand actually is
Category variance here is enormous, and the trade press reports averages that may describe nobody. Considered purchases with long research phases behave differently from replenishment. Regulated categories behave differently again. B2B differs from B2C on nearly every dimension.
Four measurements, each cheap:
1. The referral floor. What share of sessions and revenue arrives from identifiable AI sources today? Treat it as a floor rather than a total β a large share of AI referrals arrive without a usable referrer and land in "direct." Module 6 fixes this properly; for now, note the floor and the date.
2. The prompt test. Take twenty questions a real customer would ask before buying in your category. Run them across the engines your market uses. Record whether you appear, in what position, and what is said. That is your Module 3 baseline, and it usually reorders people's assumptions within an hour.
3. The marketplace check. Pull impression and conversion data for long, conversational, question-shaped queries and compare the trend against short keyword queries. Divergence is the assistant.
4. The agent check. Look at your server logs for declared agent traffic. Most retailers have never looked and are surprised in both directions β some find far more than expected, some find their edge provider has been blocking it entirely.
2.4 The surface with no owner
In almost every organisation that runs this exercise, third-party agents come back owned by nobody β while a decision about them is already being enforced, usually by an edge provider's default that nobody has reviewed.
That default may well be right. It should still be a policy, taken deliberately, by a named person. Module 4 Β§4.5 gives you the framework and the current landscape.
Exercise β Map your surfaces
Time: 60 minutes. Produces the artefact for this module.
Build a four-row table for your own business.
- Surface β the four.
- Present size β sessions and revenue where you can measure it; write "unmeasured" where you cannot, and do not estimate.
- Trend β direction over the last two quarters, with the source.
- Named owner β a person, not a team. Where the answer is nobody, write nobody.
- Next action β one, this quarter, or explicitly not this quarter.
Then run the twenty-prompt test in step two of 2.3 before your next planning meeting. It takes an hour and it reorders the agenda more reliably than any deck.
Self-check
- Which of the four surfaces gives you total control, and why is that not the same as it being the priority?
- Your marketplace revenue share is 30% and nobody in the AI programme works on marketplaces. What have you probably misallocated?
- Why is "we see very little agent traffic" an ambiguous finding rather than a reassuring one?
- Which surface currently has no owner in your organisation, and who should hold it?
- What would the twenty-prompt test most likely reveal in your category that your analytics does not?
Further reading
- Adobe Digital Insights, quarterly AI traffic reports for US retail, 2025β2026.
- Amazon's shopping-assistant developments through 2025β2026, including the March 2026 move to billable conversational placements and the May 2026 folding of the assistant into search.
- Cloudflare, Content Independence Day and subsequent AI-crawler policy posts, 2025β2026.
- EU AI Act, Article 50 transparency obligations, applicable from 2 August 2026.
Module 03
LLM Visibility I β How Engines Choose What to Cite
The retrieval pipeline behind a generative answer, what the evidence says actually moves citation, and how to measure your position before you spend anything.
Artefact: A prompt set and a dated baseline of your citation position
Why this module exists
Being named in a generative answer is now a demand channel. Almost nobody measures it, and a large industry has appeared to sell opinions about it.
This module does two things. It explains the mechanism β what actually happens between a shopper's question and the three sources listed under the answer β because you cannot influence a process you cannot picture. Then it establishes measurement, because everything in Module 4 is remediation, and remediation without a baseline is redecoration.
3.1 What happens between the question and the answer
A generative answer is not a ranking. It is a pipeline, and each stage is a different opportunity to be included or excluded.
Stage 1 β Query fan-out. The engine rarely searches the shopper's words. It decomposes the question into several sub-queries. "What's a good rug for a hallway with a dog?" becomes searches about durable rug materials, washable rugs, hallway sizing, pet-friendly fibres. You are competing against a query you never see, and typically several of them.
Stage 2 β Retrieval. Each sub-query fetches candidate passages β from a search index, a licensed corpus, the engine's own crawl, and increasingly from a merchant feed. Note that this is passage retrieval: the unit is a chunk of a page, not a page. A brilliant page whose relevant fact is buried in an image or loaded by script is retrievable in principle and absent in practice.
Stage 3 β Grounding. The retrieved passages are placed in the model's context. This is where a hard limit bites: only a fraction of retrieved material survives into the context window. Being retrieved is necessary and not sufficient.
Stage 4 β Synthesis. The model writes an answer using what it was given, weighted by what appears consistent, specific and corroborated across sources. Contradiction between sources is resolved, usually in favour of the majority or the more authoritative-seeming, and the loser may vanish.
Stage 5 β Citation. The engine attributes claims to a subset of the grounding material. Citation is chosen at the level of the sentence it supports, which is why a page can be retrieved, grounded, used β and still not named.
3.2 Why rank and citation are different outcomes
The overlap between what ranks in classical search and what gets cited in a generative answer is lower than practitioners expect. Vendor figures circulate putting it dramatically low; they are unaudited and this course does not repeat them, because the pipeline above is sufficient to explain the effect without them β and because your own prompt set, run against your own rankings, settles the question for your category in an afternoon.
Three structural reasons, all actionable:
Answers cite third parties disproportionately. A synthesised answer about which product to buy is often grounded in reviews, comparisons, forums and editorial β not in the retailer's own page, which the model reasonably treats as a party with an interest. This is the finding that most often reorders a budget: a significant portion of your citation exposure is on properties you do not own.
Specificity beats optimisation. A page that states a fact β the pile height, the fibre, the tested washing temperature, the actual guarantee period β gives the model something to lift and attribute. A page optimised to rank for a phrase often states nothing liftable at all.
Consistency across sources decides ties. When your specification says one thing and three retailers of your product say another, the model resolves the contradiction, and you may not win. Your product data on other people's sites is part of your visibility surface.
3.3 What the evidence actually says
Almost everything published on this topic is a vendor blog. There is one substantial peer-reviewed anchor, and it is worth knowing precisely.
Aggarwal et al., "GEO: Generative Engine Optimization", KDD 2024. The authors built GEO-bench β roughly 10,000 queries across nine datasets β and tested content modifications against generative-engine visibility, measured by how much of the answer a source's content occupies and how prominently it is cited.
What they found, in the form that matters to you:
| Modification | Effect |
|---|---|
| Adding statistics | Among the strongest, especially for factual and opinion-shaped questions |
| Adding quotations | Strong, particularly in explanatory and human-subject domains |
| Citing sources | Strong and broadly applicable |
| Authoritative phrasing | Positive |
| Keyword stuffing | Negligible to negative |
The paper's headline is a visibility improvement of up to 40% β a best case for a particular method in a particular domain, not a range to expect.
Two implications carry into your work:
- The tactics that work are the ones that make content more useful to a reader. Statistics, quotations and citations are not tricks; they are the properties of a source worth quoting. This is a rare and welcome alignment.
- Effects are domain-dependent. What lifts visibility in one category does not in another. This means your own measurement matters more than any published best-practice list, including this one.
3.4 What a prompt set is, and how to build one that isn't useless
A prompt set is your measurement instrument: a fixed list of questions, run on a schedule, against the engines your market uses. Everything else in visibility work is judged against it.
Most prompt sets fail in the same way β they are lists of keywords with question marks appended, which measures nothing a shopper would ever do. A usable set has four properties.
It follows real intent, across the whole journey.
| Stage | Example shape | Why include it |
|---|---|---|
| Problem | "how do I stop a rug slipping on wooden floors" | Where the category is entered; brand-free |
| Category | "best washable rugs for hallways" | Where shortlists form |
| Comparison | "X versus Y for a busy hallway" | Where you are lost or won against a named rival |
| Brand | "is [your brand] any good" | Where reputation is synthesised, often from sources you do not own |
| Post-purchase | "how do I return a rug to [your brand]" | Where service content decides repeat rate β and the most neglected |
It is fixed. A set you edit every month measures nothing over time. Change it deliberately, version it, and note the date.
It is sized honestly. Thirty to fifty prompts, measured properly, beats five hundred measured once. You are establishing a trend, not a census.
It records the answer, not just presence. Being mentioned is not the outcome. What was said is the outcome. Record the claim, the sentiment, and β critically β which source the engine attributed it to.
3.5 The three numbers β define them once
The course refers to these constantly and the capstone is assessed against them, so fix the definitions here and use no others.
| Metric | Definition | Worked |
|---|---|---|
| Mention rate | Prompts where your brand appears in the answer text Γ· total prompts | 23 of 40 = 57.5% |
| Owned-citation share | Citations pointing at properties you control Γ· all citations in those answers | 34 of 71 = 48% |
| Error count | Distinct factual errors about you across the set | 1 |
3.6 What to record
For each prompt, on each engine, on each run:
| Field | Why it earns its place |
|---|---|
| Date and engine version, where visible | The field moves monthly; an undated measurement is not a measurement |
| Are you mentioned | The crude signal |
| Position in the answer | First recommendation and fourth are different outcomes |
| What was said about you | Where the actual damage or advantage lives |
| Which sources were cited | Tells you whose page to fix β frequently not yours |
| Are competitors mentioned | Share of voice, and who is beating you where |
| Is anything wrong | Price, availability, specification, policy |
That last row is the one that pays for the exercise in the first month. Wrong information about you, repeated by an engine, is both a conversion problem and a trust problem, and it is usually traceable to a specific fixable source.
Exercise β Build the set and take the baseline
Time: 2 hours across two sittings. Produces the artefact for this module.
Sitting one β build the instrument.
- Write 30β50 prompts across the five journey stages. Use your own customer-service transcripts and site search logs for the wording; do not invent shopper language.
- Choose your engines by market share in your markets, not by prominence in the trade press.
- Fix the set. Version it. Date it.
Sitting two β take the baseline.
- Run every prompt, record all seven fields.
- Compute three numbers: mention rate, share of citations pointing at properties you own, and count of factual errors about you.
- Write the one paragraph that matters: what would have to change for the mention rate to move, and whether that change is on your site or somebody else's.
Keep the file. Module 4 is the remediation, and its success is judged against these three numbers.
Self-check
- Why can a page rank first in classical search and never be cited in a generative answer? Name two of the five pipeline stages that explain it.
- Which content modifications does the peer-reviewed evidence support, and which does it specifically not?
- Your prompt set shows a 20% mention rate and 70% of citations pointing at sites you do not own. Where does the first quarter of work go?
- Why is "what was said" a more useful field than "were we mentioned"?
- What is the risk of using a vendor's published citation benchmark as your target?
Further reading
- Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan, Ameet Deshpande, GEO: Generative Engine Optimization, KDD 2024 (arXiv:2311.09735) β the one substantial peer-reviewed anchor in this area.
- Adobe Digital Insights on retail machine-readability, 2026.
- Your own prompt-set file. In this subject it outranks everything above.
Module 04
LLM Visibility II β Making the Catalogue Legible
Feeds, structured data, rendering and access policy β the four layers that decide whether a machine can read what you sell.
Artefact: A machine-legibility audit and a prioritised remediation specification
Why this module exists
Module 3 established where you stand. This one is the work you can do on your own estate β Module 5 is the work you cannot.
It is also the module where a leader most needs enough technical grip to specify and challenge, without being able to implement. So it is written at that level: precise about what must be true, explicit about how to verify it, and silent about how to code it.
The scale of the opportunity is unusual. Adobe's 2026 analysis of US retail scored individual product pages at 66% on machine-readability β the worst-performing page type, below category pages at 74% and homepages at 75%. The pages that hold the facts an engine needs are the least readable. Most competitors are in the same position, which is why this work still returns disproportionately.
4.1 The four layers
| Layer | What it is | Fails as | Usually owned by |
|---|---|---|---|
| Feed | Your structured product data, syndicated to platforms | Missing attributes, stale price and stock | Merchandising / data ops |
| Structured data | Schema markup in the page | Present but invalid; thin coverage | Front-end / SEO |
| Rendering | What exists in the HTML a machine receives | Facts injected by script and never seen | Engineering |
| Access | Whether machines may fetch at all | Blocked by an inherited default | Nobody, then infrastructure |
Work the layers in that order. A perfect page nobody may fetch is worth nothing, but an accessible page with no facts in it is worth almost as little β and the feed reaches surfaces the page never will.
4.2 Layer one: the feed
For agentic surfaces the feed is not an export. It is the interface. Google's Universal Commerce Protocol, announced in January 2026, works from merchant catalogue data; agent surfaces increasingly answer from feeds rather than crawling pages. If a fact is not in the feed, it does not exist on those surfaces.
Three properties decide whether a feed is fit for machines.
Completeness beyond the required set. Platform-required fields get you listed. The optional ones get you answered. A shopper asking for a washable hallway rug under β¬200 in a dark colour is asking about four attributes; if you publish two, you cannot be the answer.
The attributes that carry disproportionate weight are the ones that appear in real questions: material and composition, dimensions, care and washability, colour family as well as marketing colour name, compatibility or fit, certifications, warranty period, country of origin, and energy or safety ratings where they apply.
Freshness, measured in minutes. Price and availability accuracy have become discovery inputs. An agent that recommends an out-of-stock product does not do it twice, and platforms increasingly demote merchants whose data does not match their site. Ask your team one question: what is the maximum age of the price an external surface can currently see? The answer is frequently "up to 24 hours," and is frequently a surprise to the commercial team.
Consistency with everywhere else. Your feed, your page, your marketplace listings and your distributors' listings should agree. Where they disagree, a synthesising model resolves the contradiction, and you do not control which way.
4.3 Layer two: structured data
Schema.org markup is how a page states, in a machine-parsable form, what it is describing.
The distinction that matters: presence is not validity, and validity is not coverage. Nearly every retailer has some Product markup. Far fewer have markup that validates without errors across the whole catalogue, and fewer still mark up the things that answer questions.
Prioritised for commerce:
| Type | Carries | Priority |
|---|---|---|
Product with Offer | Price, currency, availability, condition, GTIN/MPN/SKU | Essential |
AggregateRating and Review | Social proof an engine can cite | High |
FAQPage on product and service pages | The literal question-answer pairs an engine wants | High |
BreadcrumbList | Category context | Medium |
Organization with sameAs | Entity identity across the web β who you are | High and neglected |
MerchantReturnPolicy, ShippingDetails | Returns and delivery, the two most-asked post-purchase questions | High and almost universally missing |
Three rules that separate a real implementation from a checkbox:
Mark up what a customer asks, not what a validator accepts. Returns policy and delivery terms are among the most-asked questions in commerce and among the least-marked-up entities on the web. Module 3's prompt set almost always shows post-purchase prompts answered from a wrong third-party source; this is the fix.
Identifiers are not optional. GTIN, MPN and brand are how a machine knows your product is the same product it saw elsewhere. Without them you are a different item on every surface, and cannot win a comparison you are not recognised as being in.
Validate at catalogue scale, continuously. Spot-checking three pages proves nothing about 14,000. Validation belongs in your release process, with an error budget, like any other test.
4.4 Layer three: rendering
The question is simple and rarely asked: what does a machine receive when it fetches your page?
Many storefronts assemble the commercially important facts β price, availability, specifications, reviews β in the browser, after load, by script. A human sees them. A crawler that does not execute scripts, or gives up before they finish, does not.
Engines vary in what they execute and how patiently. Some render; some do not; some render on a delay measured in days. Designing for the least capable consumer is the safe position, because you cannot tell which one decided not to cite you.
The test, which anyone can run in two minutes: fetch a product page without executing JavaScript β view-source in a browser, or curl the URL β and search the returned text for the price, the availability, the top specification and one review. Whatever is missing is invisible to a meaningful share of machines.
Two further points that are routinely missed:
- Blocked subresources. If your robots policy blocks the paths serving your structured data or critical assets, a crawler that would have rendered cannot.
- Interstitials. Age gates and region redirects that a machine cannot pass turn every page behind them into an empty document. Consent walls are a special case and are dealt with below.
4.5 Layer four: access policy
Somebody has already decided which machines may read your catalogue. In most organisations it was decided by an edge provider's default, applied in a configuration nobody reviewed.
The landscape has hardened in the last two years. Cloudflare moved to blocking AI crawlers by default, introduced pay-per-crawl, and has continued tightening β from September 2026 its defaults block "mixed-use" crawlers on pages carrying ads. Cryptographic bot identification through Web Bot Auth is becoming the mechanism for telling a real agent from something wearing its user-agent string.
The decision is not binary, and framing it as "block or allow" is what produces bad policy. Four questions produce a defensible one:
- Which machines create demand? The crawler behind a shopping surface your customers use is a distribution channel. Blocking it is a decision to be absent.
- Which consume without returning anything? Bulk training crawlers and scrapers cost bandwidth and give nothing back. Different answer.
- What does it cost? Aggressive crawling is a real infrastructure line, and a large share of it re-fetches pages that have not changed.
- What must not be taken? Pricing intelligence and full catalogue extraction are competitive exposure, distinct from being discoverable.
A defensible default for most retailers: allow the crawlers behind shopping and answer surfaces your customers use, throttle or charge bulk training crawlers, require identification for agents that transact, and review the list quarterly β because the list changes faster than your infrastructure roadmap.
4.6 The audit, with pass criteria
Scored per layer. A criterion without a threshold is an opinion.
Two honest notes on the thresholds below. Where a platform publishes a requirement β identifiers, price-and-availability accuracy β the threshold follows it. Everywhere else these are the author's judgement about what is achievable and worth paying for, not a derived industry standard. Argue with them; what matters is that your version is written down before you measure, not after.
Freshness in particular has a cost. Moving from a nightly rebuild to sub-hour deltas is real engineering β an incremental pipeline, change-data capture on price and stock, and an availability service that can take the read load. Price that before you commit to the number.
| # | Check | Passes when |
|---|---|---|
| 1 | Feed attribute coverage | β₯90% of SKUs carry every attribute in your buying-question set |
| 2 | Feed freshness | Price and stock visible externally are β€15 minutes old |
| 3 | Feed/site consistency | <1% price or availability mismatch on a sampled 500 SKUs |
| 4 | Structured data validity | Within your stated error budget β most retailers can hold β€0.5% of SKUs, and the budget matters more than the number |
| 5 | Structured data coverage | Product, Offer, Rating, Returns and Shipping present on β₯95% of product pages |
| 6 | Identifiers | GTIN or MPN plus brand on β₯98% of applicable SKUs |
| 7 | Server-rendered facts | Price, availability, top specification and one review present without JavaScript |
| 8 | Access policy | A written policy, a named owner, and a review date |
Score honestly, publish the score, and re-run it quarterly against the Module 3 baseline. The pairing is the point: this table is the input, the prompt set is the outcome.
Exercise β Audit and specify
Time: 2β3 hours. Produces the artefact for this module.
- Score the eight checks on your own estate. Where you cannot measure a check, that is a finding β write "unmeasured," and note who would need to measure it.
- Run the buying-question test from 4.2 on your ten best sellers. Count how many questions are answerable from structured attributes.
- Run the rendering test from 4.4 on three product pages. List what is missing.
- Find out who owns your access policy today. If the honest answer is your CDN's default, write that sentence down; it is the most persuasive line in the document.
- Write the remediation specification: each item with an owner, a pass criterion, an effort estimate and a sequence position. Order by effect per week of engineering, not by score.
- Book the re-measurement against your Module 3 prompt set, with a date, before you start.
Self-check
- Why is the feed, rather than the website, the interface that decides visibility on several surfaces?
- Your structured data "is implemented." What two questions establish whether that is worth anything?
- What does the no-JavaScript fetch test tell you that a site audit tool usually does not?
- Give the four questions that turn an agent-access decision into a policy.
- Which of the eight checks would you expect your organisation to fail today, and who owns the fix?
Further reading
- Adobe, AI traffic grows but retail sites lag in AI search visibility, 2026 β the machine-readability scoring.
- Google, Universal Commerce Protocol and Merchant Center product-data documentation, 2026.
- Schema.org:
Product,Offer,MerchantReturnPolicy,ShippingDetails,FAQPage. - Cloudflare, AI crawler control and Web Bot Auth posts, 2025β2026.
Module 05
LLM Visibility III β Winning the Citations You Do Not Own
Most of the sources grounding answers about you sit on somebody else's domain. This is the module about them.
Artefact: A citation source register and a 90-day off-site plan with named owners
Why this module exists
Module 3 ended with a finding that reorders budgets, and Module 4 did not act on it: most of the citations grounding answers about you point at properties you do not control. Marlowe & Finch's baseline was 44 of 61.
Nearly every visibility programme in the market spends its entire budget on the other 17. The site work in Module 4 is necessary, it is cheaper to do because you own the estate, and on its own it caps out well below where you need to be β because a model that will not treat a retailer as a neutral source about its own products has to ground the claim somewhere, and it will.
This is the module about somewhere. It is also the one with an ethics boundary in it, because the tactics available here range from ordinary corporate housekeeping to things that will get you fined.
5.1 Sort your sources by controllability, not by channel
Marketing sorts sources by channel β PR, reviews, affiliates, community. That taxonomy is useless here, because it groups things with completely different response times and completely different levers.
Sort by how much control you have over the fact, which is the only variable that determines what you can do this quarter.
| Ring | Examples | Your lever | Typical time to move |
|---|---|---|---|
| Owned | Your site, feed, help centre, brand-owned publications | Direct edit β Module 4 | Days |
| Claimed | Marketplace listings, distributor and stockist pages, Google Business, retailer partner pages, app stores | Contractual and account-level: you can log in or you can ring someone with a contract | Weeks |
| Earned | Review publications, editorial comparisons, trade press, category guides | Relationship, and being worth writing about | A quarter, minimum |
| Ambient | Forums, community threads, social discussion, aggregators, wikis | Participation under the platform's rules, and correction of factual errors | Slow, and partly outside your reach |
Two things fall out of this table immediately.
The claimed ring is the highest-return work almost nobody does. These are pages about your products, containing your data, on domains you do not own β and you have either a login or a commercial relationship for most of them. When your specification page says one thing and four stockists say another, the model resolves the contradiction on volume and consistency, and you frequently lose to the majority version of yourself. That is fixable with a spreadsheet and a fortnight of emails, and it fixes stage 4 of the pipeline in Module 3.
The ambient ring is where the most-cited single source usually lives. Community threads and forum discussion are heavily represented in generative answers because they read as disinterested and specific. You cannot control them. You can correct a factual error, and you can be present under your real name where the platform allows it.
5.2 Correcting a wrong fact about you
This is the highest-return single activity in this module and it is almost never assigned to anyone.
From your Module 3 record, you have the source of every wrong claim. Work it as an operational process, not as a campaign.
Rank by citation frequency, not by how annoying it is. A mildly wrong price on a source cited in nine answers outranks a scathing review cited in one. Your prompt-set data gives you this ranking directly; nothing else does.
Match the route to the ring. A claimed-ring error is an account fix or a contractual conversation. An earned-ring error is an email to an editor with the correct figure and a source. An ambient-ring error is a correction posted under your own name, disclosed, once β and then left alone.
Give them something liftable. A correction that says "this is out of date" does less than one that says "the current price is β¬189 and the tested wash temperature is 30Β°C, both confirmed on our specification page." You are not only correcting a human; you are putting a specific, corroborated, more recent statement into the retrieval pool.
Then re-measure. The correction is not the outcome. The outcome is the next prompt-set run showing the error gone. Corrections propagate at the speed of the engine's crawl and index refresh, which is weeks, not hours β so measure at the next scheduled run and not the following morning.
5.3 The entity layer β being someone before being cited
Before a model can attribute a claim to you, it has to be confident that the various things called by your name are one company.
This is the least glamorous work in the module and it is upstream of everything else. Three parts:
Consistent identity facts everywhere. Legal name, trading name, address, founding date, ownership, category. Disagreement across your own properties is common β an About page, a marketplace seller profile and a companies-register entry that give three different founding years is enough to make an engine hedge.
Organization markup with sameAs. From Module 4, and here is why it earns its place: sameAs is how you assert that this site, that social profile, that knowledge-base entry and that marketplace store are the same entity. Without it, they are four entities that share a word.
Public knowledge-base presence, where you legitimately qualify. Wikidata and, where notability standards are met, Wikipedia are disproportionately represented in the corpora these systems were trained on and continue to retrieve from. If you qualify, the entry should exist and should be factually correct.
5.4 Getting cited on purpose
Module 3's evidence is directly actionable here, and it is the rare case where the tactic and the honest thing are the same tactic.
The GEO study found that adding statistics, quotations and citations to sources were among the strongest interventions on generative visibility, and that keyword stuffing was not. Read that as a commissioning brief:
Publish original data. You have information nobody else has: what your customers ask, what they return and why, how a category's sizes actually run, what fails after two years. A retailer that publishes a defensible annual figure β a returns-reason breakdown, a durability finding, a category price index β becomes the thing an answer can cite, because it is the only source of the number.
This is the single highest-leverage move in the module. It is also slow, and it requires you to publish something a competitor can quote.
Make claims quotable. A liftable sentence carries a subject, a number, a unit and a scope: "In 8,400 returns across 2025, 31% of rug returns cited colour mismatch" is liftable. "Colour is a common reason for returns" is not, and no amount of optimisation makes it so.
Be the source that other sources use. The earned ring cites something. If the trade press, the review publications and the category guides in your market are all working from the same three data sources, being the fourth is a durable position β and it is a Module 11 compounding asset rather than a campaign.
Go where the answer is already grounded. Your prompt set names the publications and communities that ground answers in your category. That list is your PR target list, and it will not match the one your agency is working from, because agencies target reach and this targets retrieval.
5.5 The line you do not cross
Everything above is legitimate. Two things adjacent to it are not, and both are enforced.
Undisclosed endorsement. Paying, incentivising or employing someone to recommend you without disclosing the relationship is illegal in most of the markets this course's readers sell in. In the EU it is an unfair commercial practice under the UCPD, and the 2024 Empowering Consumers Directive tightened the surrounding rules further; in the UK it falls under the DMCC Act, in force from 2025, which brought fake and incentivised reviews explicitly into scope with direct enforcement powers and turnover-based penalties; in the US the FTC's 2024 rule on consumer reviews and testimonials bans buying, selling and insider reviews outright.
The agentic-era wrinkle is that this is now cheaper to do at scale and easier to detect at scale, and the enforcement in all three jurisdictions has been moving toward the platform-level detection that finds patterns rather than instances.
Content produced to manipulate retrieval rather than to inform. Mass-generated review-shaped pages, fabricated comparison sites, sock-puppet community accounts. Beyond being fraud in several of the above senses, it is the exact pattern the platforms are building classifiers against, and the downside is asymmetric: the upside is a temporary visibility gain, the downside is being classified as a spam entity by systems whose decisions you cannot appeal.
5.6 Measuring off-site work
The measurement is already built. Module 3's record includes which sources were cited; you have simply been reading it as context rather than as a target list.
Turn it into a citation source register, and maintain it like a supplier list:
| Column | Why |
|---|---|
| Source | The domain or community |
| Ring | Owned, claimed, earned, ambient |
| Times cited in the last run | Your ranking, and the only one that matters |
| What it says about us | Correct, out of date, wrong, negative-and-fair |
| Owner | A person |
| Action and status | Correct, cultivate, monitor, or explicitly nothing |
Two metrics come off it, and both belong in the Module 6 reporting:
- Owned-citation share β the Module 3 definition. Rising is good, and it will move slowly.
- Error-carrying citations β the count of citations in your set that repeat something false about you. This is the one to drive to zero first, because it is achievable, it is attributable to specific work, and every one of them is costing conversion today.
Exercise β The register and the 90 days
Time: 2 hours. Produces the artefact for this module.
- Extract every cited source from your latest Module 3 run into one sheet. Do not filter yet.
- Tag each by ring β owned, claimed, earned, ambient.
- Count citations per source and sort descending. This ranking is your plan; resist reordering it by preference.
- Mark what each says about you, using the four values in Β§5.6. Be honest about negative-and-fair.
- Assign an owner to every row in the top ten. Rows without an owner do not move, and "marketing" is not an owner.
- Write the 90-day plan in three lines: every factual error corrected, every claimed-ring listing reconciled against your own data, and one earned-ring relationship cultivated with something worth citing.
- Set the re-measurement date β the next scheduled prompt-set run, not sooner. Corrections take weeks to propagate.
Self-check
- Why does site-only visibility work have a ceiling, and roughly where does your own baseline put yours?
- Which ring is most under-worked in your organisation, and who has the login or the contract for it?
- A source cited in eight answers says something mildly wrong; a source cited once says something damaging. Which do you fix first, and why?
- What original data could you publish that nobody else in your category holds?
- Which of your current tactics would you be uncomfortable seeing disclosed next to the citation?
Further reading
- Aggarwal et al., GEO: Generative Engine Optimization, KDD 2024 β Β§5.4 is its commissioning brief, read with the Module 3 caveats attached.
- US Federal Trade Commission, Rule on the Use of Consumer Reviews and Testimonials, 2024.
- UK Digital Markets, Competition and Consumers Act 2024, the consumer-protection provisions in force from 2025.
- Wikimedia conflict-of-interest editing guidance β read before anyone in your organisation touches an article.
- Your own citation source register. Like the prompt set, it outranks everything above.
Module 06
Measurement When the Referrer Disappears
Counting AI-influenced demand you cannot see in analytics, and reporting it in a form a finance director will accept.
Artefact: An AI-influenced demand model, with its stated confidence, that your finance team accepts
Why this module exists
Two claims are made about AI-referred traffic, both by serious people, and they cannot both be casually true: that it is a rounding error, and that it is the fastest-growing channel in commerce.
They are reconcilable. AI-referred traffic is small in absolute terms for most retailers, is growing very fast, and is systematically undercounted β because a large share of it arrives with no usable referrer and lands in "direct."
This module builds a measurement you can defend. Not a perfect one; a defensible one, with its error bars stated. That distinction is what gets a number into a board pack instead of an argument.
6.1 Why the number is wrong
Four mechanisms, each independent, all pushing the same direction.
Referrers are stripped or absent. Answers rendered in an app, a native client or behind a redirect frequently arrive with no referrer. The visit is real; the attribution is gone. It becomes "direct."
The influence and the visit are separated in time. A shopper asks an engine on Tuesday, reads the answer, and searches your brand on Thursday. The engine created the demand; brand search takes the credit. This is the largest effect and it is entirely invisible to last-click.
Some answers end the journey. The shopper gets what they needed β your delivery time, your returns window β and never visits. This is unmeasurable from your side by construction. Its existence is why session counts understate influence.
Agents do not look like sessions. A machine fetching a feed, or transacting through a protocol, generates no visit at all. It may generate an order.
The consequence is not "analytics is broken." It is narrower and more useful: your analytics reports a floor. Say so, and the number becomes usable.
6.2 Build a floor and a ceiling
The instinct is to find the number. Resist it β a single figure in this area is false precision, and the first person to poke it wins the meeting.
The floor: what you can see. Sessions and revenue from identifiable AI referrers, cleaned up as far as the mechanisms above allow. Improve it with three cheap moves:
- Maintain an explicit AI-source list in your analytics and update it monthly; the default channel groupings lag badly.
- Where an engine passes a parameter, capture it. Where it does not, note the absence rather than guessing.
- Segment your "direct" traffic by landing page. Direct traffic that lands on deep product and comparison pages, rather than the homepage, has a different origin than genuine type-ins.
The ceiling: what plausibly exists. Three triangulation methods, each weak alone and useful together:
| Method | What it catches | Weakness |
|---|---|---|
| Post-purchase survey β "where did you first hear about us / how did you research this?" | Time-separated influence, and answers that ended elsewhere | Self-report bias; needs volume |
| Brand-search lift modelling β model brand-query volume against your visibility work and known drivers | Demand pushed downstream by generative discovery | Confounded by campaigns and seasonality; needs a clean period |
| Direct-traffic decomposition β the trend in deep-landing direct traffic | Referrer-stripped visits | Confounded by app and email traffic |
Report both bounds, with the method attached. "Between 2.6% and roughly 7% of revenue is AI-influenced; the lower bound is measured, the upper is triangulated from a post-purchase survey with n=1,400 and should be treated as indicative." That sentence survives scrutiny. A single number does not.
6.3 The measurements that survive
Some metrics degrade as attribution degrades. Some do not. Weight your reporting toward the second group.
Robust:
- Mention rate and owned-citation share β the Module 3 metrics, in the exact senses defined there. Measured directly, independent of your analytics, and the leading indicators for everything else. Report both; the pair tells you whether engines are talking about you and whether they are using your own pages to do it.
- Brand search volume β the demand that generative discovery pushes downstream shows up here. Watch the trend against category and against your own spend.
- Assisted-conversion shape β whether journeys are getting shorter and more decisive; AI-influenced shoppers arrive more informed.
- Feed-surface impressions β where a platform reports them, this is machine-side demand that never becomes a session.
Fragile:
- Last-click channel revenue for AI sources.
- Session counts from AI referrers.
- Anything that requires a referrer to be intact.
6.4 Naming the counterfactual
Every AI measurement conversation eventually reaches the honest question: would that revenue have arrived anyway?
Frequently, yes. A shopper who asks an engine and then buys from you might have found you through search. Claiming the full value of those orders is the fastest way to lose the argument permanently.
Three defensible positions, in ascending order of strength:
- Incremental visibility. Measure citation share before and after remediation, and report the change in mention rate as the outcome. Honest, and it is what you actually influenced.
- Held-out categories. Remediate one product category and not a comparable one. Compare demand trends. Not perfectly clean, and far better than nothing.
- Pre-registered claim. State before the work what you expect to move, by how much, by when, and what result would mean it failed. Then report it either way.
The third is what turns a growth team into one the finance function believes on the next request.
6.5 What to instrument now
Six things, none of which takes long, all of which are painful to backfill:
- A maintained AI-source list in your analytics, reviewed monthly.
- A post-purchase survey question on research method, running continuously so you have a trend rather than a snapshot.
- Deep-landing direct traffic as a standing segment.
- The Module 3 prompt set on a fixed schedule, same engines, same recorder, recording all three metrics separately.
- Agent and crawler request logging, separated from human traffic at the edge.
- Order-source tagging for protocol orders before you have any, so the first one is already countable.
Exercise β Build the model
Time: 90 minutes. Produces the artefact for this module.
- Compute your floor from analytics, and write down what you know it excludes.
- Pick one ceiling method you can start this month, and start it.
- Choose your leading indicators β mention rate and owned-citation share, reported separately, unless you have something better.
- Design one counterfactual you can actually hold: a held-out category, a held-out market, or a pre-registered claim.
- Write the board paragraph in the form used in the case above: floor, ceiling with method and n, leading indicator, counterfactual, and the claim you will be judged on with its failure threshold.
- Take it to your finance business partner before the board, and note which sentence they attack. That sentence is your next month's work.
Self-check
- Name the four mechanisms that make analytics undercount AI-influenced demand. Which is largest, and which is unmeasurable in principle?
- Why is a floor-and-ceiling more persuasive than a single figure?
- Which metrics in your current reporting are fragile to referrer loss, and what would you promote in their place?
- What is the strongest counterfactual you could actually run next quarter?
- What failure threshold would you be willing to put in writing today?
Further reading
- Adobe Digital Insights, quarterly AI traffic reporting for US retail, 2025β2026 β for method as much as for figures.
- Module 3 of this course; your prompt set is the measurement instrument this module depends on.
Module 07
Agentic Checkout and the Protocol Stack
ACP, UCP and the payment rails beneath them β what exists, what it asks of you, and how to decide what to implement.
Artefact: A protocol decision with a dated implementation position and a review trigger
Why this module exists
In eighteen months, agent-initiated checkout went from a demo to two competing open standards backed by the largest commerce and payment companies in the world. Retail leaders are now being asked to commit engineering to something whose addressable volume is, today, small.
That is a genuinely hard decision, and it is made badly in both directions: by teams that implement everything because it is in the trade press, and by teams that dismiss it because the volume is small this quarter.
This module gives you the landscape as it stands, the questions that decide it for your business, and a way to hold a position that survives the next announcement.
7.1 What a protocol standardises
Strip the branding and every agentic commerce protocol addresses the same four problems:
| Problem | What the protocol defines |
|---|---|
| Discovery | How an agent learns what you sell, at what price, in what stock |
| Cart and intent | How a selection is assembled and confirmed on behalf of a buyer |
| Payment | How value moves, and with what credential β usually scoped to this purchase |
| Fulfilment and after | How order status, delivery, returns and support flow back to the agent |
What none of them standardises: your margin, your inventory truth, your returns policy, or who is liable when it goes wrong. Those remain yours, and Module 8 is about the last one.
7.2 The landscape, as at August 2026
Agentic Commerce Protocol (ACP). Published by OpenAI and Stripe on 29 September 2025 under Apache 2.0 and maintained in the open. It defines how an agent and a merchant complete a purchase, with the merchant remaining merchant of record. PayPal joined as a payment provider in October 2025; Stripe shipped supporting tooling in December 2025; purchasing inside ChatGPT relaunched for US users on 16 February 2026.
Then, on 4 March 2026, OpenAI withdrew in-chat checkout β retaining product discovery, and routing the purchase itself back to merchants and third-party apps. The reasons given were sales tax, fraud prevention and real-time inventory synchronisation.
Universal Commerce Protocol (UCP). Announced by Google at NRF on 11 January 2026, co-developed with Shopify, Etsy, Wayfair, Target and Walmart, released as an open specification under Apache 2.0 and offered to a standards body, and endorsed by a long list including Adyen, American Express, Mastercard, Visa, Stripe, Best Buy, Macy's, The Home Depot and Zalando. Microsoft announced Copilot support in the same period, which matters: UCP is not confined to Google's own surfaces. It covers discovery through checkout and post-purchase, and adds a cross-retailer cart with Google Pay, PayPal and buy-now-pay-later options.
The payment layer. Distinct from both, and moving on its own track. Visa Intelligent Commerce and Mastercard Agent Pay shipped agent-scoped credentials during 2025; Visa followed with a Trusted Agent Protocol, and Mastercard has continued with agent-identity and intent-verification work through 2026. Google's Agent Payments Protocol (AP2), published in September 2025 with sixty-plus payment and commerce partners, sits underneath UCP and contributes the piece the others leave implicit: a cryptographically signed mandate recording what the human actually authorised. Remember that word β Module 8 turns it into your dispute defence.
The direction across all of them is consistent: make an agent-initiated transaction distinguishable at the network level, so it can be scored, authorised and disputed differently.
7.3 The requirements that are common to all of them
This is the practical core of the module, and the reason a wait-and-see position is not the same as doing nothing.
Whichever protocol matters in your market, an agent transacting with you needs the same six things β and every one of them also improves your existing business:
- A machine-readable catalogue with reliable price and stock. Module 4, unchanged.
- Real-time availability. Not nightly. An agent that buys an item you cannot ship creates a refund, a dispute and a delisting risk.
- A programmatic path to create an order with an idempotency guarantee, so a retried request does not become two orders.
- Order status a machine can read β accepted, shipped, delivered, refunded.
- A returns and cancellation path that is not a phone call. Agents cannot use your call centre.
- An identity and authorisation model for who may transact on a customer's behalf, and within what limits.
Numbers 1, 2 and 4 you very likely need anyway. Numbers 3, 5 and 6 are the genuinely new engineering, and they are the ones to scope now even if you implement nothing.
7.5 What not to do
Do not build to a single vendor's proprietary integration unless the volume already exists. Open standards under permissive licences are a materially better bet than bespoke connectors.
Do not treat endorsement as adoption. A logo on an announcement slide tells you a company signed a press release. Ask the two questions from Β§7.2 β live or endorsing, and can a shopper complete a purchase in my market today β and note that OpenAI itself was live and then was not.
Do not let the protocol decision block the legibility work. Modules 3 and 4 pay off against demand that exists today, at every level of protocol adoption.
Do not skip the identity model. Of the six requirements, "who may transact on this customer's behalf, within what limits" is the one with legal consequences, and it is the one most often deferred.
Exercise β Take a position
Time: 90 minutes. Produces the artefact for this module.
- Answer the five questions in writing, with evidence for the first two rather than impressions.
- Score the six common requirements β have it, partial, do not have it β and identify your weakest.
- Write the position: implement, prepare, or wait, with the date it was taken.
- Write the trigger that changes it, and name who watches for it.
- Name the review date, and put it in the same governance forum as everything else.
- Ask your platform vendor the specific question from step two, in writing, this week.
Self-check
- What do agentic checkout protocols standardise, and what do they explicitly leave with you?
- Which of the six common requirements would you most likely fail today, and who owns it?
- Why is "prepare" a decision rather than a delay, and what makes the difference?
- A vendor cites forty named partners. What two questions establish what that means?
- What is your written trigger, and who is watching for it?
Further reading
- Agentic Commerce Protocol specification and repository, OpenAI and Stripe, from September 2025.
- Google, New tech and tools for retailers to succeed in an agentic shopping era, 11 January 2026, and subsequent UCP documentation.
- Visa Intelligent Commerce and Trusted Agent Protocol; Mastercard Agent Pay and agent-identity materials, 2025β2026.
Module 08
Trust, Fraud, Disputes and Returns
What breaks operationally when the buyer is software, and the policy that has to exist before the first agent order arrives.
Artefact: An agent-order policy covering authorisation, evidence, disputes and returns
Why this module exists
Every agentic commerce discussion is about acquisition. The losses are on the other side.
An agent-initiated order raises questions your operation has never had to answer: who authorised this, what evidence do you hold, who is liable when the cardholder says they did not want it, and how does a machine return something. These are not future problems in the way protocol adoption is a future problem β the first agent order arrives whether or not you have decided, and the answers are then made up by whoever is on shift.
This module is short and unglamorous, and it is the one most likely to save real money.
8.1 What is actually different
Three changes, each with an operational consequence.
The buyer and the payer may be different entities. A person delegates; software acts. Your fraud stack was built on the assumption that unusual behaviour indicates a bad actor β and agent behaviour is, by construction, unusual: fast, precise, comparison-heavy, arriving from unfamiliar infrastructure. Your first agentic commerce problem is likely to be false declines, not fraud.
The mandate is invisible to you. The customer told the agent "buy a rug under β¬200 in a dark colour." You never see that instruction. If they later dispute the purchase because they expected something else, the evidence sits with the agent platform.
Authorisation is scoped. The payment networks have moved deliberately here: Visa and Mastercard have shipped agent-scoped credentials and identity mechanisms through 2025β2026 precisely so an agent-initiated transaction can be distinguished at the network level, and therefore scored, authorised and disputed differently from a card-present or ordinary card-not-present sale.
8.2 The liability position, stated honestly
This is where a course must be careful, because the honest answer is that it is unsettled.
What is reasonably clear as at August 2026:
- Under existing scheme rules, a cardholder is generally responsible for what an authorised agent did within its mandate.
- Networks have introduced ways to identify agent transactions, which is the precondition for differentiated rules.
- Google's AP2, and the mandate concept it standardises, exist precisely because the signed record of what a human authorised is the missing evidence in this class of dispute.
What is not settled: binding, published chargeback rules specific to agent-initiated disputes. In their absence, disputes flow through existing reason codes, and merchants absorb the ambiguity.
8.3 The evidence to capture
This is the highest-return part of the module. Capture at order time, retain per your policy, and make it retrievable by order ID in one step.
| Evidence | Why it matters |
|---|---|
| Agent identity, cryptographically verified where the surface supports it | Distinguishes a legitimate agent from something wearing its name |
| Protocol and version used | Determines which rules and terms applied |
| Scope of the mandate, where passed | The single most useful artefact in a dispute over intent |
| Exactly what was presented β price, variant, delivery promise, terms | Defends "not as described" |
| Authorisation trail, including any step-up or confirmation | Defends "I did not authorise this" |
| Timestamps at each step | Establishes sequence when accounts differ |
| Fulfilment and delivery confirmation | Unchanged from today, and still decisive |
The test of whether you have this: pick one order and try to assemble the pack in under five minutes. If it takes an analyst an afternoon, you do not have an evidence process, you have data in several systems.
8.4 Returns, when the buyer cannot phone you
Returns are where agentic commerce most often meets an operation that cannot serve it.
Three requirements:
A programmatic returns path. An endpoint that can initiate a return, obtain a label and report status. If your returns process begins with a web form and a human, agent-sourced orders will generate support contacts rather than returns β which is worse for both parties.
A policy that survives comparison. Agents compare terms exhaustively and cheaply. A returns window materially shorter than the category norm is now a discovery-time disadvantage, not merely a post-purchase one. This is a commercial decision that has quietly moved from operations into merchandising.
Machine-readable terms. From Module 4: if your returns policy is not marked up, engines answer questions about it from wherever they can find something β frequently a marketplace listing with different terms. Wrong returns information is a conversion problem and a dispute generator at once.
8.5 The fraud rebalance
The specific advice for your risk team:
Do not treat agent traffic as bot traffic. Your existing bot mitigation will decline exactly the transactions you want, and the decline will be invisible in your funnel because it happens at the edge.
Segment your rules. Verified agent, unverified agent, and human should be scored separately. A verified agent's unusual velocity is expected; an unverified one's is not.
Watch the new abuse shapes. Mandate abuse (an agent operating beyond what the customer intended), scoped-credential replay, and promotion or pricing-error exploitation at machine speed β an agent finds a mispriced SKU faster than any human ever did, and tells others.
Measure false declines specifically. Add a decline reason segmented by agent status. If you cannot see it, you cannot know whether your risk stack is quietly closing the channel.
Exercise β Write the policy
Time: 60 minutes. Produces the artefact for this module.
One page, six sections.
- Definition. What counts as an agent order in your systems, and how it is flagged.
- Authorisation. What you require: verified identity, scoped credential, step-up conditions and value limits.
- Evidence. Which of the seven fields you capture, where, and the retention period.
- Disputes. Who assembles the pack, in what time, and against which reason codes.
- Returns. The path an agent uses, and the service level.
- Risk. How agent traffic is segmented, and how false declines are reported.
Then run the five-minute evidence test on one existing order. Whatever you cannot assemble is your first engineering ticket.
Self-check
- Why is your first agentic commerce problem more likely to be false declines than fraud?
- Which single piece of evidence most often decides a dispute about intent, and who holds it?
- What is the honest current position on agent-dispute liability, and what does that imply for your evidence practice?
- Why has your returns window become a discovery-time issue rather than only a post-purchase one?
- Can your team assemble a dispute pack for one order in five minutes today?
Further reading
- Visa Intelligent Commerce and Trusted Agent Protocol; Mastercard Agent Pay and agent-identity and intent materials, 2025β2026.
- PSD2 and the EU Consumer Rights Directive as amended by the Modernisation Directive β in particular the withdrawal-function requirement applying from 19 June 2026.
- OWASP Top 10 for LLM Applications β LLM01, prompt injection.
- Your card scheme's current operating regulations β the version with an effective date, not a vendor's summary.
Module 09
Merchandising for a Machine Buyer
What changes when the shopper compares exhaustively, never sees a banner, and reads your specification instead of your copy.
Artefact: A revised attribute, pricing and assortment specification
Why this module exists
Merchandising is built on a set of assumptions about how a shopper behaves: that comparison is expensive, that attention is scarce, that presentation influences choice, that most people see a fraction of the assortment.
A machine buyer violates all four. It compares exhaustively at near-zero cost, has no attention to capture, is unmoved by presentation, and can consider your entire range and everyone else's.
That does not mean merchandising stops mattering. It means the levers move β from presentation to specification, from position to attribute, from persuasion to fit. This module works through which levers move and what to do about each.
9.1 Which levers move
| Lever | Effect | Why |
|---|---|---|
| Merchandised position and placement | Weakens sharply | The agent does not see your grid |
| Photography and creative | Weakens for selection, holds for conversion | Still decides the human's final yes |
| Copy and persuasion | Weakens | Persuasion is not a retrieval signal |
| Structured attributes | Strengthens sharply | This is what the machine actually matches on |
| Price transparency and clarity | Strengthens | Comparison is now free and complete |
| Availability accuracy | Strengthens sharply | Recommending an out-of-stock item is punished |
| Reviews, specifically their content | Strengthens | Grounding material, quoted directly |
| Policy terms β returns, delivery, warranty | Strengthens | Compared as attributes, not read as small print |
| Bundling and cross-sell | Mixed | Machine-legible bundles work; visual merchandising does not |
The pattern: anything a machine can parse and compare gains weight; anything that required a human eye loses it. That is not an argument for worse creative β the human still converts β but it is an argument about where the marginal hour of merchandising effort now goes.
9.2 Attributes are the new merchandising
Module 4 treated attributes as a legibility problem. Here they are a commercial one.
Attributes must answer constraints, not describe products. Most retail taxonomies were built for navigation and reporting β category, sub-category, colour, size. Machine buyers arrive with constraints: washable, fits a 60cm gap, safe for dogs, arrives before Friday, under β¬200, made in Europe. If your data does not carry the constraint, you cannot be matched to it, no matter how well your page reads.
Build the taxonomy from questions, not from your PIM. The method is unglamorous and works:
- Pull six months of customer-service transcripts and on-site search queries.
- Extract every constraint a customer expressed.
- Rank by frequency.
- Check which of the top thirty exist as structured attributes.
Retailers typically find that between a third and a half of their most common buying constraints are not fields anywhere β they live in prose, in a photo, or in a colleague's head.
Attribute values need controlled vocabularies. "Charcoal," "dark grey," "anthracite" and "graphite" are four values for one filterable concept. Humans cope; matching does not. A colour family field alongside the marketing name is one of the cheapest wins available.
Negative attributes matter more than they used to. Not machine washable. Not suitable for underfloor heating. Stating what a product is not prevents a mismatch recommendation, and mismatch recommendations become returns β which is Module 8's cost, created here.
9.3 Pricing under exhaustive comparison
Three effects, distinct and often conflated.
Comparison is complete, not merely cheaper. A shopper who once checked three retailers now gets an answer synthesised across many. Being fourth-cheapest used to be survivable through obscurity; that obscurity is thinner now.
Comparison is total-cost, not headline. Agents compare delivered cost with terms attached: price plus delivery, minus the value of a longer returns window and a better warranty. This is genuinely good news for retailers who compete on service and have been unable to get credit for it at the point of comparison β but only if those terms are machine-readable. An unmarked-up 60-day returns policy is worth nothing at comparison time.
Errors propagate at machine speed. A mispriced SKU used to be found by a few people over hours. It is now found immediately and shared. Price-error controls are a live operational requirement, not a hygiene item.
What this does not mean is a race to the bottom. It means the differentiators must be expressed in fields. The retailer with a better guarantee, faster delivery or a longer returns window now has a way to win a comparison it previously lost on headline price β provided it publishes those facts as data.
9.4 Assortment and range
Two effects pull in opposite directions, and which dominates is category-specific.
Toward the long tail. Machine discovery is unusually good at finding the specific product that matches an unusual constraint. Obscure SKUs that never merited a merchandised slot become findable β if their attributes exist. Long-tail visibility is one of the clearest wins available, and it is almost entirely a data-completeness problem.
Toward consolidation. Exhaustive comparison is brutal on near-duplicates. Three similar SKUs at similar prices previously coexisted because shoppers saw one; now they compete, and the two losers absorb catalogue cost, data-maintenance cost and review dilution.
The practical instruction: enrich the tail, prune the near-duplicates. Both are merchandising decisions with a data prerequisite, and both are visible in your own returns and search data before any AI surface tells you.
Exercise β Rebuild the attribute set
Time: 2 hours. Produces the artefact for this module.
- Extract constraints from six months of service transcripts and site-search logs. Rank by frequency.
- Check the top thirty against your actual structured attributes. Mark each: exists, prose only, absent.
- Design the additions, including at least one negative attribute and one controlled vocabulary.
- Write the supplier rule that stops the problem recurring β no new SKU without the mandatory set.
- Check your policy terms β returns, delivery, warranty β are published as data, not only as page copy.
- List your near-duplicates and propose a consolidation.
Self-check
- Which merchandising levers weaken when the buyer is a machine, and which strengthen?
- Why does a constraint expressed in prose fail where the same fact in a field succeeds?
- How does a longer returns window become a comparison-time advantage, and what is the prerequisite?
- Which pulls harder in your category β long-tail enrichment or near-duplicate consolidation?
- Name one negative attribute that would prevent mismatched recommendations in your range.
Further reading
- Module 4 of this course, for the technical expression of everything above.
- Google Merchant Center product data specification, for the attribute vocabulary the largest surface consumes.
- Your own service transcripts. In this module they outrank every external source.
Module 10
Paid Media on AI Surfaces
Which placements exist, which are being sold as though they do, and what happens to performance marketing when the click is not the point.
Artefact: A paid plan that separates live placements from roadmap items
Why this module exists
Paid media is where the gap between what is announced and what you can actually buy is widest.
Every quarter brings news of advertising formats on AI surfaces. Some are live and billable. Some are pilots in one market. Some are a slide. An agency that cannot tell the difference β or has an incentive not to β will build you a plan that spends against the third category.
This module gives you the distinction, and the harder strategic question underneath it: what performance marketing becomes when a growing share of demand is mediated by something that does not click.
10.1 The classification
Before anything is bought, classify. Four questions, all answerable in a week:
| Question | If no |
|---|---|
| Can we buy it today, in our market, in our category? | It is not a plan line |
| Is it billable with reporting, or a managed pilot? | It is a test with a real cost |
| Is there any targeting or measurement control? | Budget it as brand, not performance |
| Can we turn it off independently? | It is a bundled placement, price it as such |
The most concrete example of a genuine transition is Amazon's: conversational placements inside its shopping assistant moved to a billable CPC format in March 2026 β that is, from an experiment to a line you can plan, with the reporting and controls following behind rather than arriving with it. That sequence is typical, and it is the sequence to look for.
10.2 How paid and organic interact here
On classical search, paid and organic were separate lanes on one page. On generative surfaces the relationship is different in a way that matters for budget.
The answer is synthesised from organic-ish material. What a model says about your product is grounded largely in retrieved content β your feed, your pages, third-party sources. Paid placement can put you in front of an answer; it does not usually change what the answer says about you. If the synthesis says your returns policy is short and your sizing runs small, paid spend buys attention for that.
Which means visibility work is upstream of paid efficiency. This inverts the usual sequencing argument. On generative surfaces, Modules 3 and 4 are not an alternative to paid β they are a prerequisite for it, because they determine the content of the answer your paid impression sits beside.
Feeds are increasingly the shared substrate. The same product data drives organic answers, agent surfaces and paid formats. A feed problem is simultaneously an organic problem, an agentic problem and a media-efficiency problem β which is why Module 4 keeps reappearing.
10.3 What changes about performance marketing
Four adjustments, in decreasing order of how quickly you should make them.
Attribution windows should widen. A shopper who researches with an assistant and buys three days later breaks a short window. If your bidding optimises to a click that no longer exists in the journey, it optimises to a shrinking sample. Module 6's measurement work is the input here.
Landing experiences meet a better-informed visitor. Someone arriving from a generative answer has often already compared. Adobe's May 2026 US retail data puts AI-referred visitors at 54% more likely to convert, spending 53% longer on site and viewing 23% more pages than other traffic β a visitor who is further along, not merely a different colour of traffic. Landing pages designed to educate from zero waste that.
Note the reversal: the same series had AI-referred traffic converting worse than average in early 2025. The visitor did not change; the population did, as the surfaces moved from early adopters to the mainstream. Which is the Module 1 habit again β the number is only meaningful with its date.
Brand terms behave differently. Generative discovery pushes demand downstream into brand search. Rising brand-term volume may be an output of your visibility work rather than an independent channel β which changes both how you value it and how you argue about its incrementality.
Creative for the human, data for the machine. The split is clean. The machine matched on your attributes; the human converts on your creative. Teams that respond to "AI changes everything" by degrading creative are optimising the wrong half.
10.4 Retail media and marketplaces
For most retailers this is where AI-adjacent paid money is real today, and it is frequently owned by a team not in the conversation at all β the pattern Module 2 flagged.
Three practical points:
On-platform assistants are already influencing organic placement inside marketplaces. Your listing quality β attributes, reviews, Q&A, structured content β feeds the assistant. That is unpaid work with paid consequences, since a better-matched listing lowers the cost of the paid placement beside it.
Conversational ad formats are early and moving fast. Expect reporting and controls to lag availability. Budget accordingly: treat early spend as learning with a cap, not as performance with a target.
If you operate a retail media network, the same questions arrive from the other side. Your advertisers will ask what you can offer on your own AI surfaces, and the honest answer will likely be "nothing yet, and here is the roadmap." Saying so is better than the alternative, because they will find out.
10.5 Briefing an agency
Four requirements, which together make vapour visible without an argument:
- Every line classified β live, pilot or announced β with evidence for anything marked live.
- Reporting interface named for each live line.
- Learning budget stated separately from performance budget, with a cap and a decision date.
- A pre-registered claim for each performance line: the metric, the baseline, and the result that would mean stopping.
An agency that can produce all four is worth keeping. One that cannot is charging you to find out.
Exercise β Build the plan
Time: 90 minutes. Produces the artefact for this module.
- List every AI-adjacent paid opportunity currently proposed to you, internally or by an agency.
- Run the four classification questions on each. Mark live, pilot, announced.
- Separate the budget into performance and learning, with the learning portion capped and dated.
- Identify your feed dependency β which paid lines get cheaper if Module 4's work lands. Say so explicitly; it is usually the strongest argument for that work.
- Write the agency brief using the four requirements above.
- Check who owns marketplace AI placements in your organisation, and whether they are in this plan.
Self-check
- What four questions classify an AI advertising placement, and which one most often exposes a pilot?
- Why is visibility work upstream of paid efficiency on generative surfaces rather than an alternative to it?
- Your attribution window is seven days. What does the research-then-buy pattern do to it?
- Why might rising brand-search volume be an output of your visibility work rather than a channel result?
- Which of your current paid lines would get cheaper if your feed improved, and by what mechanism?
Further reading
- Amazon advertising documentation on conversational and assistant placements, 2026.
- Adobe Digital Insights on AI-referred visitor behaviour β conversion, engagement, pages per visit β through May 2026.
- Google's commerce and ads announcements accompanying UCP, January 2026 onward.
Module 11
Advantage, Moats and the Operating Model
What commoditises, what compounds, and how the team and the agency relationship have to change.
Artefact: An advantage thesis and a team and agency design
Why this module exists
Everything so far has been capability: be legible, be measurable, be transactable. Capability is table stakes arriving on a schedule β your competitors are reading the same trade press and buying from the same vendors.
This module asks the strategy question. When the tooling has diffused, which of this is still worth something, and what does the organisation that holds it look like?
11.1 What commoditises
Be honest about this list, because vendors will sell all of it as differentiation.
Visibility tooling. Prompt-set monitoring is becoming a commodity subscription. Everyone will have it within a year.
Feed and schema hygiene. Necessary, urgent, and imitable in a quarter. Being legible is not an advantage; being illegible is a disadvantage. The asymmetry matters β you must do this work, and it will not distinguish you once done.
Protocol support. When it becomes a platform toggle, it stops being a differentiator on the day it ships.
Knowing that any of this is happening. Awareness is not strategy. It is table stakes with a lead time measured in months.
11.2 What compounds
Five categories, and they have a common shape: each is accumulated through operation rather than purchased.
Proprietary constraint data. The Module 9 work β knowing which constraints your customers actually express, and having them as structured attributes across a catalogue β is genuinely hard to copy. It was extracted from your service transcripts and your search logs. A competitor can copy the idea in an afternoon and the data not at all.
Third-party presence. Module 3's finding was that most citations point at properties you do not own. Relationships with the review publications, communities and editorial sources that ground answers in your category are slow to build, and they compound.
Your prompt-set time series. A dated record of how your citation position moved as you changed things is a proprietary dataset about your own market that nobody else has. It is also the only evidence in this field that is genuinely yours.
Operational truth. Real-time availability, accurate delivery promises, a returns path that works. These are expensive, boring, and exactly what machine surfaces reward and punish. A retailer whose stock data is right wins comparisons that a retailer with better copy loses.
Terms you can afford that others cannot. A longer returns window or better warranty, now expressed as machine-readable attributes, becomes a comparison-time advantage. If your cost structure supports terms a competitor cannot match, agentic comparison surfaces that advantage rather than hiding it β a genuine reversal, since these advantages were previously invisible at the point of choice.
11.3 The advantage thesis
One page, four parts. The fourth is the one usually missing.
- The asset. The specific thing this work makes more valuable in your business. Not a capability you can buy.
- The mechanism. How it compounds. Where does the loop close?
- The moat. Why a well-funded competitor cannot have it within eighteen months. If the honest answer is that they could, say so β that is a finding, and it should move your investment.
- The falsification clause. What observation, within a year, would show the thesis is wrong. Name the observation and the date you will check.
11.4 The ownership problem
Run the artefacts from this course against your organisation chart and a pattern appears: the work spans four teams and belongs to none.
| Work | Where it usually sits |
|---|---|
| Prompt set, citation measurement | SEO β often renamed AEO, same person, no extra time |
| Feed and attributes | Merchandising or data ops |
| Rendering and structured data | Engineering, behind a roadmap |
| Access policy | Infrastructure, by default, unreviewed |
| Protocol readiness | Commerce platform team |
| Agent orders, disputes, returns | Payments, risk, customer operations |
| Marketplace assistant optimisation | Marketplace team, usually absent from the programme |
Three viable models, with honest trade-offs:
Owner plus contributors. One named owner with a mandate and a standing forum; the work stays in the functions. Cheapest, fastest to start, and fails when the owner has no authority over the contributing teams' priorities.
A small dedicated pod. Three or four people β data, engineering, commercial β owning the artefacts end to end. Moves fastest, and risks becoming a silo the rest of the business routes around.
Embedded with a standard. Each function owns its layer against a central specification and a shared scorecard. Scales best, slowest to start, and requires the scorecard to have teeth.
For most mid-size retailers the first is the right start and the third is the right destination. What does not work is the common default: adding it to the SEO manager's objectives without changing anyone else's.
11.5 The agency relationship
Agencies are in this course's audience, so this section is written to be read from both sides.
What retainers were for is shrinking. Keyword research, content volume and link acquisition are the activities most exposed to both automation and to the shift away from ranking.
What clients now need is different work. Prompt-set design and disciplined measurement. Feed and attribute specification. Third-party source relationships β the citation surface nobody owns. Classification of paid opportunities into live, pilot and announced. Independent verification of vendor claims.
The honest repositioning is from "we improve your rankings" to "we own your machine-visibility measurement and remediation, and we tell you which vendor claims are real." That is a defensible retainer, and it is harder to commoditise than content volume.
For retailers, the test of an agency is now simple: ask what they measure that you could not measure yourself in a week. A good answer is a maintained prompt-set programme with a time series and a remediation record. A weak answer is a dashboard reselling a tool subscription.
Exercise β Thesis and design
Time: 2 hours. Produces the artefact for this module.
Part A β the thesis. One page, four parts, including the falsification clause with a date.
Part B β the operating model.
- Map the seven work items in 10.4 to named people. Mark every one that has no owner.
- Choose a model β owner-plus-contributors, pod, or embedded-with-standard β and write the argument for it.
- Name the forum and its cadence.
- Write the one sentence you would need to say to the team that currently has no idea it is in scope. It is usually the marketplace team.
Part C β the agency. List what you currently buy and what you now need. If the two lists barely overlap, that is the conversation.
Self-check
- Which of your current agentic commerce investments will be commodity within a year?
- What in your business compounds through operation rather than purchase?
- Why is being legible not an advantage while being illegible is a disadvantage?
- Which of the seven work items in your organisation has no owner today?
- What would your agency have to measure for the retainer to be defensible?
Further reading
- Modules 3, 4 and 8 of this course β the sources of the assets this module claims are durable.
- Your own service transcripts and prompt-set time series, which are the two proprietary datasets in play.
Module 12
Appendix: The Vocabulary
Every term this course uses, defined once β because in a market this noisy, a shared definition is a negotiating position.
Artefact: A vocabulary you can hold a vendor and an agency to
Why this appendix exists
This field generates vocabulary faster than it generates evidence. AEO, GEO, LLMO and AIO are sold as four disciplines and are largely one. "Agentic" is applied to a scripted chatbot and to a system that transacts under a mandate.
The cost is not aesthetic. When you and a vendor use the same word for different things, you buy something other than what you agreed. This appendix fixes the meanings this course uses, marks where the market disagrees, and is the vocabulary the final exam draws on.
A. Discovery and visibility
| Term | Definition |
|---|---|
| Generative engine | A system that answers a question with synthesised text and a small number of cited sources, rather than a ranked list |
| AI Overview / AI Mode | Google's generative answer formats within Search; distinct surfaces with distinct behaviour |
| GEO / AEO / LLMO / AIO | Four names for substantially one practice: making content and data more likely to be retrieved, used and cited by generative engines. This course uses LLM visibility for the outcome and GEO for the practice, following the peer-reviewed usage |
| Citation | A source attributed under or within a generative answer. The unit of visibility that matters |
| Mention | Your brand named in the answer text, with or without a citation. Weaker than a citation, and not nothing |
| Share of voice (generative) | Your mention or citation rate across a fixed prompt set, relative to named competitors |
| Prompt set | A fixed, versioned list of buying questions run on a schedule against chosen engines. The measurement instrument of Modules 3, 5 and 6 |
| Mention rate | Prompts in your set where your brand appears in the answer text, over total prompts. Module 3 |
| Owned-citation share | Citations pointing at properties you control, over all citations in those answers. Not the same as mention rate, and the distinction is load-bearing |
| Citation source register | The ranked list of sources grounding answers about you, tagged by ring β owned, claimed, earned, ambient β with an owner per row. Module 5 |
| Entity layer | The identity facts and sameAs assertions that let an engine know the things bearing your name are one company |
| Query fan-out | The engine's decomposition of one question into several retrieval queries. You compete against queries you never see |
| Retrieval | Fetching candidate passages for the model to work from. Roughly what "ranking" used to buy you |
| Grounding | Placing retrieved material in the model's context so the answer is built from it. Being retrieved does not guarantee being grounded |
| Hallucination | Confident output that is not true. In commerce it most often appears as a wrong price, stock status or returns policy |
| llms.txt | A proposed plain-text convention describing a site for language models. Not a ratified standard; support is inconsistent. Cheap to publish, not a strategy |
B. Machine legibility
| Term | Definition |
|---|---|
| Machine-readability | Whether a machine can extract the commercially relevant facts from your pages and data without executing your interface |
| Product feed | Structured product data syndicated to platforms. On several surfaces it is not an export but the interface |
| Feed freshness | The maximum age of price and availability visible on external surfaces. Now a discovery input, not only an operational metric |
| Structured data / schema markup | Machine-parsable statements in a page about what it describes β Product, Offer, MerchantReturnPolicy and others |
| GTIN / MPN | Global trade item number and manufacturer part number. How a machine knows your product is the same product it saw elsewhere |
| Controlled vocabulary | A fixed value list for an attribute β a colour family alongside forty marketing colour names |
| Negative attribute | An explicit statement of what a product is not suitable for. Prevents mismatched recommendations, and therefore returns |
| Server-side rendering | Facts present in the HTML a machine receives, rather than assembled later by script. The difference between visible and invisible to many crawlers |
| Constraint | What a shopper actually specifies β washable, fits 60cm, safe for pets. The unit a machine matches on, and usually absent from taxonomies built for navigation |
C. Agents and protocols
| Term | Definition |
|---|---|
| Shopping agent | Software that discovers, compares and sometimes transacts on a person's behalf under a mandate |
| Mandate | The instruction and limits a person gives an agent β budget, preferences, constraints. Usually invisible to the merchant, and the decisive evidence in a dispute |
| AI-influenced purchase | A human buys after consulting a generative engine. The large majority of AI-affected revenue today |
| Agent-executed purchase | Software completes the transaction. Small volume, long lead time to support |
| ACP β Agentic Commerce Protocol | Open standard for agent-initiated checkout published by OpenAI and Stripe, September 2025, Apache 2.0 |
| UCP β Universal Commerce Protocol | Google's open standard announced January 2026, covering discovery through post-purchase across its surfaces, co-developed with major retailers |
| Agent-scoped credential | A payment credential issued for an agent and limited in scope, so agent transactions are distinguishable at the network level |
| Idempotency | The guarantee that a retried request does not create a second order. Unremarkable until a machine retries |
| Web Bot Auth | Cryptographic identification of a crawler or agent, so a legitimate one can be told from something using its name |
| Pay per crawl | Charging automated clients for access to content rather than only allowing or blocking. Cloudflare's term for the model, and now the general one |
| AP2 β Agent Payments Protocol | Google's payment-layer specification, September 2025, contributing the cryptographically signed mandate record |
| SCA β strong customer authentication | The PSD2 requirement that applies to European transactions regardless of who initiated them. No protocol removes it |
| EU AI Act Article 50 | The transparency obligation, applicable since 2 August 2026, that a system interacting with people must disclose that it is AI |
| Prompt injection | Text placed in content an agent reads, crafted to be treated as instruction rather than as data. First on the OWASP LLM risk list, and with no reliable general defence |
| Access policy | Your deliberate decision about which machines may fetch what. In most organisations, currently an inherited CDN default |
D. Measurement and commerce
| Term | Definition |
|---|---|
| Referrer loss | Visits arriving without a usable source, misfiled as direct. A principal cause of AI undercounting |
| Floor and ceiling | Reporting measured and triangulated bounds instead of a single false-precision figure |
| Leading indicator | A measurement that moves before revenue does. Here, mention rate and owned-citation share |
| Counterfactual | What would have happened anyway. Naming it is what makes an improvement claim survivable |
| Pre-registration | Fixing metric, baseline and decision rule before the work starts |
| Revenue per visit | Revenue divided by sessions. More robust than conversion rate when traffic mix is shifting |
| False decline | A legitimate transaction refused by risk controls. The likely first cost of agent traffic |
| Evidence pack | The record assembled to defend a dispute: identity, mandate, what was presented, authorisation, timestamps, fulfilment |
| Retail media | Advertising sold by a retailer against its own audience β increasingly including its AI surfaces |
| Conversational placement | Paid inventory inside an assistant experience. Verify it is billable with reporting before it enters a plan |
E. Words this course uses carefully
"Agentic." Applied to everything from a rules-based chatbot to a system transacting under a mandate. Ask which of the two is meant, every time.
"AI traffic." Reported as a channel; in reality a floor with four known leaks. Prefer AI-influenced demand with stated bounds.
"Optimised for AI." Almost always means schema was added. Ask which of the four legibility layers was worked on and what the pass criterion was.
"Partner." In this market, frequently means "appeared on an announcement slide." Ask whether they are live, in your market, in your category.
Self-check
- A vendor offers "AEO services." What do you ask them to distinguish it from GEO, LLMO and ordinary structured-data work?
- Explain the difference between retrieval and grounding, and why being retrieved is not enough.
- What is a mandate, who holds the record of it, and why does that matter in a dispute?
- Why is AI-influenced demand a more defensible reporting term than AI traffic?
- Give an example of a negative attribute in your own catalogue that would prevent a mismatched recommendation.
Module 13
Capstone: The Twelve-Month Plan
One assessed deliverable, assembled from the artefacts you built along the way.
Artefact: The completed plan, assessed against a published rubric
The brief
Produce a twelve-month agentic commerce plan for your own business β or, if you would rather not work with internal data, for Marlowe & Finch, whose case file is below.
The capstone is assembled, not written from scratch. Every module produced a component. If you completed the exercises you already hold the plan; this is the work of making it one document that a board can approve and a finance director can hold you to.
Target length: eight to twelve pages.
| Section | Source | Length |
|---|---|---|
| 1. What changed for us, and what did not | Module 1 brief | 1 page |
| 2. Where our demand actually sits | Module 2 surface map | 1 page |
| 3. Visibility baseline | Module 3 prompt set and measurement | 1β2 pages |
| 4. Legibility audit and remediation | Module 4 audit and specification | 2 pages |
| 5. Off-site citation plan | Module 5 source register and 90-day plan | 1 page |
| 6. How we will measure and report | Module 6 model, with bounds and counterfactual | 1 page |
| 7. Protocol position | Module 7 decision, dated, with trigger | 1 page |
| 8. Agent-order policy | Module 8 policy | 1 page |
| 9. Merchandising changes | Module 9 attribute specification | 1 page |
| 10. Paid plan | Module 10 classification | 1 page |
| 11. Advantage thesis and operating model | Module 11 | 1β2 pages |
The Marlowe & Finch case file
Marlowe & Finch is fictional. It is a composite built to exercise every profile this course addresses; any resemblance to a specific retailer is coincidental.
The business. European home and lifestyle retailer, headquartered in Amsterdam. Revenue β¬182m. Channel mix: 61% direct e-commerce, 24% marketplaces, 15% wholesale and eleven owned stores. Markets: Netherlands, Germany, UK, Belgium, with Germany the growth priority. Catalogue 14,200 SKUs across rugs, lighting, textiles and small furniture. Gross margin 47%; returns rate 18% overall and 26% in rugs.
The pressure. Two consecutive quarters of flat direct traffic while revenue held, meaning the site is converting better on fewer visits β nobody is sure why. The board has read that AI traffic to retail grew several hundred percent and wants a strategy. A competitor issued a press release about "launching on ChatGPT."
What is known.
- AI-identifiable referrals: 1.9% of sessions, 2.6% of revenue, revenue per visit materially above site average, six-month trend steep.
- Prompt-set baseline: mentioned in 14 of 40 buying prompts; of 61 citations, 44 point at third-party sites and 17 at owned properties β an owned-citation share of 28%; six factual errors about the business; the most-cited single source lists a discontinued price 12% below current; all three post-purchase prompts return a wrong returns policy.
- Legibility audit: 3 of 8 checks pass. 47% of SKUs lack material composition; 71% lack care instructions; feed rebuilds nightly; 2,300 products emit invalid Offer markup; no returns or shipping markup; reviews and specifications rendered client-side; the CDN challenges all declared agents including two shopping surfaces their customers use.
- Attribute audit: seven of the top twelve customer constraints are not fields.
- Operations: availability updates nightly; no programmatic returns path; order idempotency untested; payments works.
- Eleven agent orders have already arrived. Nine were declined, one of the two completed was disputed and lost for want of evidence.
- Agency retainer β¬18k per month, currently weighted to content volume.
- Engineering capacity for this programme: roughly one and a half developers.
Deliberately absent. Competitor visibility data, category-level AI query volume, and the marginal economics of agent-sourced orders. You will have to state assumptions. Marking which of them your plan is sensitive to is part of what is assessed.
Assessment rubric
Scored 1β4 per dimension. A pass requires 3 or better on every dimension β a plan that is excellent on visibility and silent on disputes is not a plan you can execute.
| Dimension | 1 β Absent | 2 β Asserted | 3 β Evidenced | 4 β Falsifiable |
|---|---|---|---|---|
| Evidence discipline | Vendor figures repeated | Sources named | Sources named and dated, vendor claims labelled | Own measurement, with method and confidence stated |
| Visibility | "Improve AI presence" | Tactics listed | Prompt-set baseline with the three numbers | Re-measurement booked, with a threshold that would mean failure |
| Legibility | Not addressed | Schema mentioned | Eight checks scored with owners | Remediation ordered by effect per engineering week |
| Measurement | Traffic only | Floor reported | Floor and ceiling with method and n | Counterfactual held, claim pre-registered |
| Commercial realism | No numbers | Costs listed | Effort against actual capacity | Sensitivity to the assumptions you marked |
| Operations and risk | Not addressed | Policy referenced | Agent-order policy with evidence fields | Five-minute evidence test performed and reported |
| Sequencing | A wish list | Ordered by value | Ordered by value and readiness | Degrades gracefully at half the capacity |
A worked example
Take the visibility dimension.
Band 2 (asserted). "We will improve our presence in AI search through content and schema improvements."
Band 3 (evidenced). "Baseline on a fixed 40-prompt set across three engines, taken 12 August: mentioned in 14 of 40; 17 of 61 citations point at owned properties; six factual errors about us. Remediation specified across eight checks with named owners."
Band 4 (falsifiable). As band 3, plus: "Re-measured on the same set, same engines, same recorder, on 30 November. Target: 24 of 40 mentions, owned-citation share above 45%, zero factual errors. Below 18 of 40 we stop the remediation programme, report why, and move the budget to third-party source relationships β because that would indicate the constraint is off-site, not on our estate. The largest single contributor to the baseline errors was a third-party price listing, so we expect that hypothesis to be live."
The third version is longer because it contains the two sentences a board actually needs: what would change the answer, and what happens then.
How to use the rubric before submitting
Score your own draft honestly and write the sentence that moves each 3 to a 4. Most drafts arrive at 3s. The distance is usually one sentence per section β the threshold, the counterfactual, the trigger β and adding those seven sentences is the highest-value hour in this course.
Self-assessment questions
- Which section would collapse first under questioning from someone who sells AI visibility tooling for a living?
- What single measurable claim are you prepared to be judged on in twelve months, and who holds you to it?
- What did you decide not to do, and what specific observation would reverse it?
- Which assumption is load-bearing and unverified, and what is the cheapest way to test it in thirty days?
- If your engineering capacity halved next quarter, which half of the plan survives β and did you write that down before it happened?
Where to go next
The artefacts are instruments with a cadence, not one-time deliverables:
| Artefact | Review |
|---|---|
| Prompt set and citation measurement | Monthly, same engines, same method |
| Legibility audit | Quarterly |
| Protocol position and trigger | Every quarterly business review |
| Agent-order policy | On the first ten orders, then annually |
| Attribute set | Continuously, from service transcripts |
| Advantage thesis | Annually, against its own falsification clause |
The course ends here. The monthly measurement does not.