The brief
Produce a twelve-month agentic commerce plan for your own business β or, if you would rather not work with internal data, for Marlowe & Finch, whose case file is below.
The capstone is assembled, not written from scratch. Every module produced a component. If you completed the exercises you already hold the plan; this is the work of making it one document that a board can approve and a finance director can hold you to.
Target length: eight to twelve pages.
| Section | Source | Length |
|---|---|---|
| 1. What changed for us, and what did not | Module 1 brief | 1 page |
| 2. Where our demand actually sits | Module 2 surface map | 1 page |
| 3. Visibility baseline | Module 3 prompt set and measurement | 1β2 pages |
| 4. Legibility audit and remediation | Module 4 audit and specification | 2 pages |
| 5. Off-site citation plan | Module 5 source register and 90-day plan | 1 page |
| 6. How we will measure and report | Module 6 model, with bounds and counterfactual | 1 page |
| 7. Protocol position | Module 7 decision, dated, with trigger | 1 page |
| 8. Agent-order policy | Module 8 policy | 1 page |
| 9. Merchandising changes | Module 9 attribute specification | 1 page |
| 10. Paid plan | Module 10 classification | 1 page |
| 11. Advantage thesis and operating model | Module 11 | 1β2 pages |
The Marlowe & Finch case file
Marlowe & Finch is fictional. It is a composite built to exercise every profile this course addresses; any resemblance to a specific retailer is coincidental.
The business. European home and lifestyle retailer, headquartered in Amsterdam. Revenue β¬182m. Channel mix: 61% direct e-commerce, 24% marketplaces, 15% wholesale and eleven owned stores. Markets: Netherlands, Germany, UK, Belgium, with Germany the growth priority. Catalogue 14,200 SKUs across rugs, lighting, textiles and small furniture. Gross margin 47%; returns rate 18% overall and 26% in rugs.
The pressure. Two consecutive quarters of flat direct traffic while revenue held, meaning the site is converting better on fewer visits β nobody is sure why. The board has read that AI traffic to retail grew several hundred percent and wants a strategy. A competitor issued a press release about "launching on ChatGPT."
What is known.
- AI-identifiable referrals: 1.9% of sessions, 2.6% of revenue, revenue per visit materially above site average, six-month trend steep.
- Prompt-set baseline: mentioned in 14 of 40 buying prompts; of 61 citations, 44 point at third-party sites and 17 at owned properties β an owned-citation share of 28%; six factual errors about the business; the most-cited single source lists a discontinued price 12% below current; all three post-purchase prompts return a wrong returns policy.
- Legibility audit: 3 of 8 checks pass. 47% of SKUs lack material composition; 71% lack care instructions; feed rebuilds nightly; 2,300 products emit invalid Offer markup; no returns or shipping markup; reviews and specifications rendered client-side; the CDN challenges all declared agents including two shopping surfaces their customers use.
- Attribute audit: seven of the top twelve customer constraints are not fields.
- Operations: availability updates nightly; no programmatic returns path; order idempotency untested; payments works.
- Eleven agent orders have already arrived. Nine were declined, one of the two completed was disputed and lost for want of evidence.
- Agency retainer β¬18k per month, currently weighted to content volume.
- Engineering capacity for this programme: roughly one and a half developers.
Deliberately absent. Competitor visibility data, category-level AI query volume, and the marginal economics of agent-sourced orders. You will have to state assumptions. Marking which of them your plan is sensitive to is part of what is assessed.
Assessment rubric
Scored 1β4 per dimension. A pass requires 3 or better on every dimension β a plan that is excellent on visibility and silent on disputes is not a plan you can execute.
| Dimension | 1 β Absent | 2 β Asserted | 3 β Evidenced | 4 β Falsifiable |
|---|---|---|---|---|
| Evidence discipline | Vendor figures repeated | Sources named | Sources named and dated, vendor claims labelled | Own measurement, with method and confidence stated |
| Visibility | "Improve AI presence" | Tactics listed | Prompt-set baseline with the three numbers | Re-measurement booked, with a threshold that would mean failure |
| Legibility | Not addressed | Schema mentioned | Eight checks scored with owners | Remediation ordered by effect per engineering week |
| Measurement | Traffic only | Floor reported | Floor and ceiling with method and n | Counterfactual held, claim pre-registered |
| Commercial realism | No numbers | Costs listed | Effort against actual capacity | Sensitivity to the assumptions you marked |
| Operations and risk | Not addressed | Policy referenced | Agent-order policy with evidence fields | Five-minute evidence test performed and reported |
| Sequencing | A wish list | Ordered by value | Ordered by value and readiness | Degrades gracefully at half the capacity |
A worked example
Take the visibility dimension.
Band 2 (asserted). "We will improve our presence in AI search through content and schema improvements."
Band 3 (evidenced). "Baseline on a fixed 40-prompt set across three engines, taken 12 August: mentioned in 14 of 40; 17 of 61 citations point at owned properties; six factual errors about us. Remediation specified across eight checks with named owners."
Band 4 (falsifiable). As band 3, plus: "Re-measured on the same set, same engines, same recorder, on 30 November. Target: 24 of 40 mentions, owned-citation share above 45%, zero factual errors. Below 18 of 40 we stop the remediation programme, report why, and move the budget to third-party source relationships β because that would indicate the constraint is off-site, not on our estate. The largest single contributor to the baseline errors was a third-party price listing, so we expect that hypothesis to be live."
The third version is longer because it contains the two sentences a board actually needs: what would change the answer, and what happens then.
How to use the rubric before submitting
Score your own draft honestly and write the sentence that moves each 3 to a 4. Most drafts arrive at 3s. The distance is usually one sentence per section β the threshold, the counterfactual, the trigger β and adding those seven sentences is the highest-value hour in this course.
Self-assessment questions
- Which section would collapse first under questioning from someone who sells AI visibility tooling for a living?
- What single measurable claim are you prepared to be judged on in twelve months, and who holds you to it?
- What did you decide not to do, and what specific observation would reverse it?
- Which assumption is load-bearing and unverified, and what is the cheapest way to test it in thirty days?
- If your engineering capacity halved next quarter, which half of the plan survives β and did you write that down before it happened?
Where to go next
The artefacts are instruments with a cadence, not one-time deliverables:
| Artefact | Review |
|---|---|
| Prompt set and citation measurement | Monthly, same engines, same method |
| Legibility audit | Quarterly |
| Protocol position and trigger | Every quarterly business review |
| Agent-order policy | On the first ten orders, then annually |
| Attribute set | Continuously, from service transcripts |
| Advantage thesis | Annually, against its own falsification clause |
The course ends here. The monthly measurement does not.