Enterprise AI that does the work: your first governed workflow, live in about two weeks.Map your first AI teammate
← Selling to Machines

Module 13 of 13 Β· 2 hours

Capstone: The Twelve-Month Plan

One assessed deliverable, assembled from the artefacts you built along the way.

Artefact: The completed plan, assessed against a published rubric

Share

The brief

Produce a twelve-month agentic commerce plan for your own business β€” or, if you would rather not work with internal data, for Marlowe & Finch, whose case file is below.

The capstone is assembled, not written from scratch. Every module produced a component. If you completed the exercises you already hold the plan; this is the work of making it one document that a board can approve and a finance director can hold you to.

Target length: eight to twelve pages.

SectionSourceLength
1. What changed for us, and what did notModule 1 brief1 page
2. Where our demand actually sitsModule 2 surface map1 page
3. Visibility baselineModule 3 prompt set and measurement1–2 pages
4. Legibility audit and remediationModule 4 audit and specification2 pages
5. Off-site citation planModule 5 source register and 90-day plan1 page
6. How we will measure and reportModule 6 model, with bounds and counterfactual1 page
7. Protocol positionModule 7 decision, dated, with trigger1 page
8. Agent-order policyModule 8 policy1 page
9. Merchandising changesModule 9 attribute specification1 page
10. Paid planModule 10 classification1 page
11. Advantage thesis and operating modelModule 111–2 pages

The Marlowe & Finch case file

Marlowe & Finch is fictional. It is a composite built to exercise every profile this course addresses; any resemblance to a specific retailer is coincidental.

The business. European home and lifestyle retailer, headquartered in Amsterdam. Revenue €182m. Channel mix: 61% direct e-commerce, 24% marketplaces, 15% wholesale and eleven owned stores. Markets: Netherlands, Germany, UK, Belgium, with Germany the growth priority. Catalogue 14,200 SKUs across rugs, lighting, textiles and small furniture. Gross margin 47%; returns rate 18% overall and 26% in rugs.

The pressure. Two consecutive quarters of flat direct traffic while revenue held, meaning the site is converting better on fewer visits β€” nobody is sure why. The board has read that AI traffic to retail grew several hundred percent and wants a strategy. A competitor issued a press release about "launching on ChatGPT."

What is known.

  • AI-identifiable referrals: 1.9% of sessions, 2.6% of revenue, revenue per visit materially above site average, six-month trend steep.
  • Prompt-set baseline: mentioned in 14 of 40 buying prompts; of 61 citations, 44 point at third-party sites and 17 at owned properties β€” an owned-citation share of 28%; six factual errors about the business; the most-cited single source lists a discontinued price 12% below current; all three post-purchase prompts return a wrong returns policy.
  • Legibility audit: 3 of 8 checks pass. 47% of SKUs lack material composition; 71% lack care instructions; feed rebuilds nightly; 2,300 products emit invalid Offer markup; no returns or shipping markup; reviews and specifications rendered client-side; the CDN challenges all declared agents including two shopping surfaces their customers use.
  • Attribute audit: seven of the top twelve customer constraints are not fields.
  • Operations: availability updates nightly; no programmatic returns path; order idempotency untested; payments works.
  • Eleven agent orders have already arrived. Nine were declined, one of the two completed was disputed and lost for want of evidence.
  • Agency retainer €18k per month, currently weighted to content volume.
  • Engineering capacity for this programme: roughly one and a half developers.

Deliberately absent. Competitor visibility data, category-level AI query volume, and the marginal economics of agent-sourced orders. You will have to state assumptions. Marking which of them your plan is sensitive to is part of what is assessed.

Assessment rubric

Scored 1–4 per dimension. A pass requires 3 or better on every dimension β€” a plan that is excellent on visibility and silent on disputes is not a plan you can execute.

Dimension1 β€” Absent2 β€” Asserted3 β€” Evidenced4 β€” Falsifiable
Evidence disciplineVendor figures repeatedSources namedSources named and dated, vendor claims labelledOwn measurement, with method and confidence stated
Visibility"Improve AI presence"Tactics listedPrompt-set baseline with the three numbersRe-measurement booked, with a threshold that would mean failure
LegibilityNot addressedSchema mentionedEight checks scored with ownersRemediation ordered by effect per engineering week
MeasurementTraffic onlyFloor reportedFloor and ceiling with method and nCounterfactual held, claim pre-registered
Commercial realismNo numbersCosts listedEffort against actual capacitySensitivity to the assumptions you marked
Operations and riskNot addressedPolicy referencedAgent-order policy with evidence fieldsFive-minute evidence test performed and reported
SequencingA wish listOrdered by valueOrdered by value and readinessDegrades gracefully at half the capacity

A worked example

Take the visibility dimension.

Band 2 (asserted). "We will improve our presence in AI search through content and schema improvements."

Band 3 (evidenced). "Baseline on a fixed 40-prompt set across three engines, taken 12 August: mentioned in 14 of 40; 17 of 61 citations point at owned properties; six factual errors about us. Remediation specified across eight checks with named owners."

Band 4 (falsifiable). As band 3, plus: "Re-measured on the same set, same engines, same recorder, on 30 November. Target: 24 of 40 mentions, owned-citation share above 45%, zero factual errors. Below 18 of 40 we stop the remediation programme, report why, and move the budget to third-party source relationships β€” because that would indicate the constraint is off-site, not on our estate. The largest single contributor to the baseline errors was a third-party price listing, so we expect that hypothesis to be live."

The third version is longer because it contains the two sentences a board actually needs: what would change the answer, and what happens then.

How to use the rubric before submitting

Score your own draft honestly and write the sentence that moves each 3 to a 4. Most drafts arrive at 3s. The distance is usually one sentence per section β€” the threshold, the counterfactual, the trigger β€” and adding those seven sentences is the highest-value hour in this course.

Self-assessment questions

  1. Which section would collapse first under questioning from someone who sells AI visibility tooling for a living?
  2. What single measurable claim are you prepared to be judged on in twelve months, and who holds you to it?
  3. What did you decide not to do, and what specific observation would reverse it?
  4. Which assumption is load-bearing and unverified, and what is the cheapest way to test it in thirty days?
  5. If your engineering capacity halved next quarter, which half of the plan survives β€” and did you write that down before it happened?

Where to go next

The artefacts are instruments with a cadence, not one-time deliverables:

ArtefactReview
Prompt set and citation measurementMonthly, same engines, same method
Legibility auditQuarterly
Protocol position and triggerEvery quarterly business review
Agent-order policyOn the first ten orders, then annually
Attribute setContinuously, from service transcripts
Advantage thesisAnnually, against its own falsification clause

The course ends here. The monthly measurement does not.

That is the whole course. The exam is 25 questions in 40 minutes; pass it and your certificate arrives by email.Start the exam