There is no shortage of GEO advice right now. Publish more content. Earn media coverage. Structure your pages so the models can read them. Become a source the engines cite. Almost all of it aims at one thing: getting into the answer.
Getting in is the easy part. Knowing whether it worked is the hard part, and that is mostly missing from the conversation.
Yes, you can measure GEO. But not with one ranking, one prompt, or one dashboard number. Generative engines change their answers, and their sources shift by platform, location, wording, and time. That makes GEO measurement less tidy than search reporting, not impossible.
Data Story is a New Zealand GEO and AEO agency with one rule: if we cannot show whether the work moved something useful, we do not sell it as a strategy. We measure AI visibility across the customer journey, then read it against the business outcome at the end. The result is not a claim that one citation caused one sale. It is evidence of whether your brand is getting more visible, more accurately represented, and more likely to influence qualified demand.
Does GEO move a business outcome?
Visibility is an intermediate measure. The commercial measure still belongs at the end of the customer journey: qualified leads, bookings, sales, revenue, length of stay, applications, or another outcome the organisation already values.
This reframes a common alarm. As buyers move into AI, website sessions can fall even while demand holds, because more of the discovery now happens inside the answer. Read on its own, a session dip looks like decline. Read against AI visibility and revenue, it looks like a channel shift.
Illustrative: revenue lifts with AI visibility as sessions fall (indexed to 100)
The lines are indexed to 100 and illustrative. The point is the shape: revenue lifts with AI visibility, not with the falling session count. That is the correlation worth watching, and the reason not to cut a channel because one upstream metric dipped.
That is the difference between "we appear in ChatGPT" and "AI is sending us people who buy." One is a vanity metric. The other is an outcome.
Capturing the AI-referred part of that picture is easier than it was. Google Analytics now includes an AI Assistant default channel for recognised AI-assistant referrers such as ChatGPT, Gemini, Copilot, Perplexity, and Claude. It does not capture the whole influence of AI search. Google AI Overviews and AI Mode remain within Organic Search, and visits without a usable referrer may appear elsewhere or be unattributed.
Within those limits, GA4 can show how identified AI-referred visitors behave:
- engaged sessions and landing pages
- progression to high-intent content
- enquiries, bookings, purchases, or other key events
- conversion rate and value where tracking supports it
- performance relative to other acquisition channels
For longer buying journeys, we also look beyond last-click reporting. CRM data, assisted conversions, sales-call notes, brand search, Search Console, and the organisation's primary performance measure can add evidence.
The interpretation matters. AI visibility and revenue moving together is a correlation until the design supports a stronger causal claim. We use AI measures as evidence alongside established commercial and customer measures. We do not turn a changing answer engine into false certainty.
For a fuller walk-through of connecting AI visibility to revenue, see how to measure your brand in AI search.
So what does GEO correlate with across the funnel?
GEO measurement is strongest read next to the measures your business already trusts. At each stage of the journey, the prompt-level signal has established metrics around it that tend to move together when the work is landing.
| Journey stage | GEO / prompt signal | Read alongside |
|---|---|---|
| Discover | Discovery-prompt performance and share of voice | Meltwater reach, brand search volume, AI-bot crawl hits |
| Consider | Consideration prompt-set performance | Bing Webmaster AI report, Google GenAI impressions by page, AI-referral sessions in GA4, website sessions, Google impressions |
| Convert | Conversion-intent prompt performance, by product | Conversion rate, sessions, revenue, AOV, direct and brand-search traffic |
| Experience | Experience and service prompt-set performance | NPS, CSAT |
No single row proves causation. Together they show whether AI visibility is moving with the outcomes that matter, stage by stage.
What we measure, and why
That is the payoff. Now the detail: the metrics themselves, and why we care about each.
GEO is not one number, and that is the thing to get right early. A brand mention is not a citation. A citation is not a visit. A visit is not a conversion. Bundle them into a single "visibility score" and you bury the one thing that needs fixing.
So we keep them apart and check four things, one at a time:
- Presence: is your brand showing up in the answer?
- Citation: is your own site the source behind it?
- Accuracy: is the LLM getting you right?
- Access: can the engines even reach your content?
We started at the end, with the outcome. These four are what feed it.
Are you present in the answer?
The starting point is a prompt library built from the questions customers ask at each stage of their journey, from discovering a category to considering options, converting, and the experience after.
Those prompts are run across relevant answer engines. Depending on the market and question set, that can include ChatGPT, Google AI Overviews and AI Mode, Gemini, Claude, Perplexity, Copilot, and Meta AI. We record:
- whether the brand is mentioned
- how often it appears across repeated runs
- which competitors appear in the same answers
- whether the brand is recommended, compared, or listed without context
- which claims and attributes the answer associates with it
- which sources the engine cites
The unit of analysis is the topic, not an isolated prompt. One wording can bias an answer. A topic set should cover different customer types, needs, and phrasings without loading the desired brand or conclusion into the question.
For a destination, for example, asking "how long should I stay?" can encourage a longer itinerary. A better measurement set tests several neutral planning situations and reads the pattern across them. The same rule applies to products, professional services, healthcare, and other categories.
Is the engine citing your website?
Mentions and citations do different jobs.
| What it shows | When it matters most | |
|---|---|---|
| Mention | Your brand is named in the answer | Discovery, when the customer is building a shortlist |
| Citation | Your page helped ground the answer | Planning and comparison, when detail and proof decide it |
A citation is the stronger signal. It shows your content is grounding what the engine says, and it identifies the pages and subjects where your site has authority.
Citation value depends on the customer journey. At the discovery stage, being named may matter more than having a product page cited. During planning and comparison, owned citations become more important because the customer needs detail, evidence, and a path forward.
We therefore measure citations by topic and journey stage rather than treating total citations as a universal KPI. The useful questions are:
- Are your pages cited for the subjects you need to own?
- Are the cited pages current, accurate, and suited to the customer's next decision?
- Are third-party sources defining you instead?
- Which owned pages are gaining or losing source-share over time?
- Do citations support discovery, consideration, or conversion?
This is also why publishing more content is not a GEO strategy. A page has to perform a defined job in the answer and in the customer journey. The four numbers a GEO report uses to describe this, and why they are easy to confuse, are worth knowing first: mention, share of voice, citation, and source share.
Is the LLM getting your business right?
Visibility has limited value when the answer contains the wrong offer, location, price, policy, product detail, or brand association.
Accuracy measurement compares what the engines say with a controlled source of truth. We capture the claims in the answer, open the cited pages, and classify problems such as:
- outdated or incorrect information
- broken, redirected, or removed pages
- confusion with another organisation or product
- unsupported claims repeated from third-party sources
- missing information that changes the recommendation
- correct information attached to the wrong context
The output is an action register, not a collection of screenshots. Each issue should have an owner, source URL, evidence, priority, and next action. Repeated errors can then be traced back to the content, technical, or entity problem most likely to be causing them.
Can AI search systems access your content?
A site can perform well in Google Search and still restrict another service that retrieves web content.
The technical check covers robots.txt, canonical tags, redirects, response codes, JavaScript rendering, CDN rules, firewalls, and bot-management settings. Different services use different crawlers and user-triggered fetchers, and some publish IP ranges for verification. Perplexity, for example, distinguishes between PerplexityBot for search indexing and Perplexity-User for user-requested page access, and documents firewall configuration for both.
Access is not a blanket instruction to allow every bot. It is a decision about which systems should reach which content, verified through server logs and current platform documentation. User-agent strings alone are not sufficient proof that a request is legitimate.
A technical GEO review should establish:
- whether priority content can be retrieved
- whether the correct canonical page is returned
- whether security controls block a search crawler unintentionally
- whether important information is hidden behind rendering or interaction
- whether cited URLs still resolve to the intended content
This is a small part of GEO when it passes and a hard limit when it fails.
How do you track a changing answer?
A single check is a snapshot. It cannot show whether GEO is working.
We divide the prompt library into two sets.
The benchmark set stays fixed. These prompts are carefully defined and rerun on a consistent schedule. They provide a comparable baseline for mention rate, source-share, competitor share of voice, accuracy, and other agreed measures.
The investigation set can change. These prompts respond to new products, customer questions, engine behaviour, competitors, and findings from the data. They help diagnose opportunities without corrupting the benchmark.
This solves a basic tension in AI search optimisation. The market changes too quickly for a rigid prompt library, but a measurement system that changes every month cannot support a trend. Fixed prompts protect comparability. Adaptive prompts keep the work useful.
The reporting should also retain the conditions of each run, including engine, date, market, topic, and prompt version. Without that context, a movement in the chart can be mistaken for a movement in the market.
What belongs on a GEO dashboard?
The dashboard should follow the customer journey and keep the underlying detail available.
At discovery, the useful measure may be share of topic mentions or inclusion in a consideration set. During planning, it may be source-share for authoritative guides. Near a decision, it may be citation of product, service, or operator pages. After arrival, the measures move to engagement, conversion, revenue, or another business result.
A practical monthly view can include:
- benchmark prompt coverage and run conditions
- mention rate by topic and engine
- competitor share of voice within the same prompt set
- owned and third-party citation share
- cited URLs gained, retained, and lost
- accuracy issues opened and resolved
- crawler and retrieval health
- identified AI-assistant sessions and outcomes
- the established business measure each GEO indicator supports
⚠️ Warning
The calculations should be defined and checked before they reach the reporting layer. AI can help people explore the data. It should not be asked to improvise the metric each time someone opens a dashboard.
What do you get from GEO measurement?
Data Story's GEO Setup establishes the measurement system before optimisation work is judged.
You receive:
- a prompt library grounded in customer questions and search data
- a fixed benchmark set and an adaptive investigation set
- a baseline across the agreed answer engines and markets
- topic, journey-stage, competitor, mention, and citation measures
- an accuracy and cited-page review
- a technical retrieval and crawler-access check
- analytics reporting for identifiable AI-referred traffic
- a monthly view of movement, evidence, and recommended action
This gives the marketing team a defensible answer when the board asks whether GEO is working. It also shows where to act: content, technical access, brand definition, third-party authority, or the customer experience after the visit.
What should you look for in a GEO or AEO agency in NZ?
A GEO agency in New Zealand should be able to explain its measurement instrument, not only its content tactics. Ask how it controls prompt bias, preserves a benchmark, separates mentions from citations, verifies accuracy, and connects AI-search indicators to a commercial measure.
We will tell you when a number is soft. Certainty sells better than that. It is also how GEO gets oversold.
Data Story combines GEO measurement with SEO, analytics, customer-journey reporting, and conversion work. That lets us follow the evidence past the answer and into the result.
See what AI is saying about your brand
GEO Setup baselines your AI search visibility across the engines your customers use, then connects it to the outcome that matters to your business.
Rather talk it through first? Book a call. Want to keep up as the ground shifts? Get our newsletter.
Related reading
Frequently asked questions

Written by
Hannah Stuart
Head of Strategy
Hannah leads strategy at Data Story, working with clients to turn marketing data into growth roadmaps. She brings deep experience in full-funnel digital strategy, attribution modelling, and performance marketing across New Zealand and Australian markets.
LinkedInRelated articles

CRM
How we do CRM for destinations, experiences and tourism businesses
22 July 2026
For a tourism business, the difference between a CRM that holds contacts and one that drives revenue sits in a few specific places: attribution, audiences, automation and repeat bookings. How we refine each.
Read article
SEO & GEO
We have built a custom solution to the biggest challenge in destination marketing right now - AI Search Visibility
22 July 2026
See what AI tools say about your destination or experience today, then improve it. Our custom solution makes your visibility in AI search measurable across the customer journey, then gives you a plan to improve it.
Read article
SEO & GEO
Your destination is being recommended in AI. Can you see it?
21 July 2026
The trip now gets planned inside a chat, and most destinations have no way to see how they show up there. A GEO playbook for tourism boards, regions and luxury destinations: what to measure, and where to start this week.
Read article