Why this module exists
Being named in a generative answer is now a demand channel. Almost nobody measures it, and a large industry has appeared to sell opinions about it.
This module does two things. It explains the mechanism — what actually happens between a shopper's question and the three sources listed under the answer — because you cannot influence a process you cannot picture. Then it establishes measurement, because everything in Module 4 is remediation, and remediation without a baseline is redecoration.
3.1 What happens between the question and the answer
A generative answer is not a ranking. It is a pipeline, and each stage is a different opportunity to be included or excluded.
Stage 1 — Query fan-out. The engine rarely searches the shopper's words. It decomposes the question into several sub-queries. "What's a good rug for a hallway with a dog?" becomes searches about durable rug materials, washable rugs, hallway sizing, pet-friendly fibres. You are competing against a query you never see, and typically several of them.
Stage 2 — Retrieval. Each sub-query fetches candidate passages — from a search index, a licensed corpus, the engine's own crawl, and increasingly from a merchant feed. Note that this is passage retrieval: the unit is a chunk of a page, not a page. A brilliant page whose relevant fact is buried in an image or loaded by script is retrievable in principle and absent in practice.
Stage 3 — Grounding. The retrieved passages are placed in the model's context. This is where a hard limit bites: only a fraction of retrieved material survives into the context window. Being retrieved is necessary and not sufficient.
Stage 4 — Synthesis. The model writes an answer using what it was given, weighted by what appears consistent, specific and corroborated across sources. Contradiction between sources is resolved, usually in favour of the majority or the more authoritative-seeming, and the loser may vanish.
Stage 5 — Citation. The engine attributes claims to a subset of the grounding material. Citation is chosen at the level of the sentence it supports, which is why a page can be retrieved, grounded, used — and still not named.
3.2 Why rank and citation are different outcomes
The overlap between what ranks in classical search and what gets cited in a generative answer is lower than practitioners expect. Vendor figures circulate putting it dramatically low; they are unaudited and this course does not repeat them, because the pipeline above is sufficient to explain the effect without them — and because your own prompt set, run against your own rankings, settles the question for your category in an afternoon.
Three structural reasons, all actionable:
Answers cite third parties disproportionately. A synthesised answer about which product to buy is often grounded in reviews, comparisons, forums and editorial — not in the retailer's own page, which the model reasonably treats as a party with an interest. This is the finding that most often reorders a budget: a significant portion of your citation exposure is on properties you do not own.
Specificity beats optimisation. A page that states a fact — the pile height, the fibre, the tested washing temperature, the actual guarantee period — gives the model something to lift and attribute. A page optimised to rank for a phrase often states nothing liftable at all.
Consistency across sources decides ties. When your specification says one thing and three retailers of your product say another, the model resolves the contradiction, and you may not win. Your product data on other people's sites is part of your visibility surface.
3.3 What the evidence actually says
Almost everything published on this topic is a vendor blog. There is one substantial peer-reviewed anchor, and it is worth knowing precisely.
Aggarwal et al., "GEO: Generative Engine Optimization", KDD 2024. The authors built GEO-bench — roughly 10,000 queries across nine datasets — and tested content modifications against generative-engine visibility, measured by how much of the answer a source's content occupies and how prominently it is cited.
What they found, in the form that matters to you:
| Modification | Effect |
|---|---|
| Adding statistics | Among the strongest, especially for factual and opinion-shaped questions |
| Adding quotations | Strong, particularly in explanatory and human-subject domains |
| Citing sources | Strong and broadly applicable |
| Authoritative phrasing | Positive |
| Keyword stuffing | Negligible to negative |
The paper's headline is a visibility improvement of up to 40% — a best case for a particular method in a particular domain, not a range to expect.
Two implications carry into your work:
- The tactics that work are the ones that make content more useful to a reader. Statistics, quotations and citations are not tricks; they are the properties of a source worth quoting. This is a rare and welcome alignment.
- Effects are domain-dependent. What lifts visibility in one category does not in another. This means your own measurement matters more than any published best-practice list, including this one.
3.4 What a prompt set is, and how to build one that isn't useless
A prompt set is your measurement instrument: a fixed list of questions, run on a schedule, against the engines your market uses. Everything else in visibility work is judged against it.
Most prompt sets fail in the same way — they are lists of keywords with question marks appended, which measures nothing a shopper would ever do. A usable set has four properties.
It follows real intent, across the whole journey.
| Stage | Example shape | Why include it |
|---|---|---|
| Problem | "how do I stop a rug slipping on wooden floors" | Where the category is entered; brand-free |
| Category | "best washable rugs for hallways" | Where shortlists form |
| Comparison | "X versus Y for a busy hallway" | Where you are lost or won against a named rival |
| Brand | "is [your brand] any good" | Where reputation is synthesised, often from sources you do not own |
| Post-purchase | "how do I return a rug to [your brand]" | Where service content decides repeat rate — and the most neglected |
It is fixed. A set you edit every month measures nothing over time. Change it deliberately, version it, and note the date.
It is sized honestly. Thirty to fifty prompts, measured properly, beats five hundred measured once. You are establishing a trend, not a census.
It records the answer, not just presence. Being mentioned is not the outcome. What was said is the outcome. Record the claim, the sentiment, and — critically — which source the engine attributed it to.
3.5 The three numbers — define them once
The course refers to these constantly and the capstone is assessed against them, so fix the definitions here and use no others.
| Metric | Definition | Worked |
|---|---|---|
| Mention rate | Prompts where your brand appears in the answer text ÷ total prompts | 23 of 40 = 57.5% |
| Owned-citation share | Citations pointing at properties you control ÷ all citations in those answers | 34 of 71 = 48% |
| Error count | Distinct factual errors about you across the set | 1 |
3.6 What to record
For each prompt, on each engine, on each run:
| Field | Why it earns its place |
|---|---|
| Date and engine version, where visible | The field moves monthly; an undated measurement is not a measurement |
| Are you mentioned | The crude signal |
| Position in the answer | First recommendation and fourth are different outcomes |
| What was said about you | Where the actual damage or advantage lives |
| Which sources were cited | Tells you whose page to fix — frequently not yours |
| Are competitors mentioned | Share of voice, and who is beating you where |
| Is anything wrong | Price, availability, specification, policy |
That last row is the one that pays for the exercise in the first month. Wrong information about you, repeated by an engine, is both a conversion problem and a trust problem, and it is usually traceable to a specific fixable source.
Exercise — Build the set and take the baseline
Time: 2 hours across two sittings. Produces the artefact for this module.
Sitting one — build the instrument.
- Write 30–50 prompts across the five journey stages. Use your own customer-service transcripts and site search logs for the wording; do not invent shopper language.
- Choose your engines by market share in your markets, not by prominence in the trade press.
- Fix the set. Version it. Date it.
Sitting two — take the baseline.
- Run every prompt, record all seven fields.
- Compute three numbers: mention rate, share of citations pointing at properties you own, and count of factual errors about you.
- Write the one paragraph that matters: what would have to change for the mention rate to move, and whether that change is on your site or somebody else's.
Keep the file. Module 4 is the remediation, and its success is judged against these three numbers.
Self-check
- Why can a page rank first in classical search and never be cited in a generative answer? Name two of the five pipeline stages that explain it.
- Which content modifications does the peer-reviewed evidence support, and which does it specifically not?
- Your prompt set shows a 20% mention rate and 70% of citations pointing at sites you do not own. Where does the first quarter of work go?
- Why is "what was said" a more useful field than "were we mentioned"?
- What is the risk of using a vendor's published citation benchmark as your target?
Further reading
- Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan, Ameet Deshpande, GEO: Generative Engine Optimization, KDD 2024 (arXiv:2311.09735) — the one substantial peer-reviewed anchor in this area.
- Adobe Digital Insights on retail machine-readability, 2026.
- Your own prompt-set file. In this subject it outranks everything above.