- Keyword lists measure a dead behaviour. ChatGPT search prompts average 8.7 words and AI Mode queries reach 70 to 80 words, so research units are questions, not head terms.
- One prompt triggers 8 to 12 background sub-queries through query fan-out, so pages are retrieved as passages. Self-contained 40 to 80 word answer blocks win, not word count.
- No public tool measures private LLM chat volume. Demand is reconstructed from sales calls, support tickets, surveys, Search Console regex exports, PAA, Reddit, and prompt databases.
- Fund evaluation and purchase questions first. Skip glossary content, because models supply their own definitions and AI referral traffic converts at roughly six times standard Google search.
- Specificity decides owned versus earned. Head terms are won on G2, Capterra, and Reddit. Long-tail integration, pricing, and configuration questions are won in your own documentation.
- Track appearance rate per question per engine, not rank. Only 36% of 1,200 tracked brands held visibility across ChatGPT, Gemini, and Google AI every month.
Q1. Why does question research replace keyword research when buyers ask AI instead of Google?
Last month a Head of Organic Growth pulled up her keyword tracker on a call. Forty-one page-one rankings, green across the board. Then she opened ChatGPT, asked the exact question her buyers ask, and her brand was nowhere in the answer. The tracker was measuring a behaviour her buyers had already stopped.
Question research for AI search means reconstructing the full conversational prompts buyers ask ChatGPT, Perplexity, Gemini, and Copilot, then building content engineered to be cited. Google queries average three to four words. ChatGPT search prompts run about 8.7 words, and Google AI Mode queries reach 70 to 80 words. The goal shifts from ranking a page to becoming the source the engine names inside a five to ten brand answer.
The keyword list is a proxy for a dead behaviour
๐ Compression was a constraint, not a choice
A keyword is a compressed instruction to a machine that could not handle sentences. Buyers compressed because they had to. Language models removed the reason to compress.

The measurement gap is now physical, not philosophical. Semrush clickstream data shows ChatGPT search-enabled prompts nearly doubled from 4.7 to 8.7 words year over year. Prompts on AI Overviews average 24.6 characters, while ChatGPT, Gemini, and AI Mode prompts average 57.8 to 58.1 characters.
The model is a universal intent decoder
๐ง Exact phrasing stopped paying
The model does not need your exact phrasing. It absorbs whatever the buyer types, infers intent, then decides which searches to run in the background. That is why exact-match research has stopped paying.
MaximusLabs AI classifies every buyer question into head, mid-tail, and long-tail bands, because each band is won by a different asset. Head questions like "best CRM" are decided by third-party citations. Long-tail questions about integrations, pricing, and configuration are decided by your own answer engine optimization documentation.
Ranking and being cited are not the same job
๐ฏ Citation is binary, not positional
Ranking is positional. Citation is binary. When ChatGPT names five vendors, you are either inside that set or you do not exist in that conversation.
This is where I part ways with the standard "GEO is just SEO plus" line. MaximusLabs AI's read is that generative engine optimization is closer to a data science problem than a marketing one, because you are influencing a retrieval step, not a ranking table. I might be stating that too strongly for enterprise teams with heavy legacy Google traffic, and traditional SEO still carries real revenue. The center of gravity has moved anyway.
Why challengers gain more than incumbents here
โญ The equalizer effect favours smaller brands
The Princeton and IIT Delhi GEO study tested nine tactics across 10,000 queries and lifted visibility by up to 40%. The detail most people skip matters more for smaller brands. Pages sitting at position five gained 115.1% visibility from the same tactics.
That is an equalizer, not a rising tide. A challenger with sharp answers can enter an evaluation set it could never have out-ranked on Google. MaximusLabs AI measures this as citation rate rather than position, which is how the Oliv AI engagement reached a 64% citation rate against incumbents sitting near 30%.
MaximusLabs AI treats question research as an algorithm problem, mapping intent clusters to each engine's retrieval behaviour instead of porting a Google keyword list into an AI SEO plan. That single reframe changes which questions get funded, and which pages ever get written.

Q2. What actually happens behind one buyer prompt, and why does fan-out change your research unit?
John runs sales at a mid-market SaaS company. He opens ChatGPT and asks for the best AI tools to lift his team's productivity, with pros, cons, and pricing. Seconds later he has a curated shortlist, and he never sees a results page.

One prompt does not trigger one search. Google confirms AI Overviews and AI Mode use a query fan-out technique, issuing multiple related searches across subtopics and data sources, then stitching passages into one answer. Standard AI Mode fires roughly 8 to 12 sub-queries. The research unit is therefore a cluster of sub-intents, not a head term, because those sub-queries appear in no analytics report.
The searches you never see are the ones that decide the answer
๐ One question splits into many retrievals
Google's own documentation states that fan-out lets Search surface a wider and more diverse set of links than a single query would. Independent teardowns of AI Mode describe the same pattern: one question splits into distinct sub-queries, each retrieving its own sources.
For John's prompt, the engine is not searching "best AI sales tools" once. It is separately retrieving pricing, integrations, team size fit, review sentiment, and alternatives. Each of those retrievals has its own winner.
Passage-level retrieval breaks page-level thinking
โ ๏ธ Blocks compete, not pages
Your page is not evaluated as a page. It is broken into blocks, and individual blocks compete for individual sub-queries.
That is why a 4,000 word guide can lose to a competitor's 90 word pricing FAQ. The competitor answered one sub-query cleanly and self-containedly. MaximusLabs AI builds every section to a 40 to 80 word answer nugget that survives extraction with no surrounding context, following its content formatting for AI search standard, because that is the unit the engine actually lifts.
The binary game underneath the shortlist
๐ฒ There is no page two
Hundreds of vendors exist in most B2B categories. The engine names five to ten. There is no page two to fall back to.
This is more brutal than Google, and quietly better for people who prepare. Zero-click behaviour is the reason. As buyers put it, there is no reason to click through ten blue links when the answer is already assembled.
How to simulate fan-out before you write
โ Approximate the sub-queries and score coverage
You cannot see the sub-queries in Search Console. You can approximate them, and the approximation is good enough to plan against.
- Run your money question in AI Mode, ChatGPT, Perplexity, and Copilot, then log every distinct sub-topic each answer covers.
- Ask a reasoning model to decompose the same question into 10 sub-queries a retrieval system would issue.
- Score your page: how many of those 8 to 12 dimensions does it answer in a self-contained block?
- Fill the gaps as H3 blocks, not as new pages.
MaximusLabs AI runs this as AI Source Analysis, building prompt sets across ChatGPT, Claude, Perplexity, and Gemini, and mapping which URLs get cited most often for each sub-query. The output is a coverage score, not a keyword list, and it feeds directly into our GEO topic clusters.
MaximusLabs AI builds every cluster to cover 8 to 12 sub-intent dimensions per money question, because that is the actual surface an AI answer is assembled from. Coverage of the neighbourhood beats optimization of the head term.
Q3. Where do you find the questions buyers actually ask AI engines?
Most teams stall at the same wall. There is no Keyword Planner for private chats, so the plan never starts. Meanwhile the questions are already sitting in three systems the company owns.
MaximusLabs AI reconstructs AI-search demand from seven sources, because no public tool measures private LLM chat volume: sales-call transcripts, support tickets, customer surveys capturing the literal prompt used, Google Search Console question-regex exports, People Also Ask and autocomplete, Reddit and community threads, and vendor prompt databases. Verbatim buyer language beats tool-generated phrasing, since AI answers are assembled from how people actually talk.
The seven-source stack, ranked by intent quality
๐๏ธ Where each source sits
| Source | What to pull | Effort | Intent quality |
|---|---|---|---|
| Sales-call transcripts | Verbatim objections and comparison questions from the last 20 to 30 discovery and demo calls | Medium | Highest |
| Support tickets | Configuration, integration, and "does it do X" questions | Low | High |
| Customer surveys | The literal prompt a buyer typed before finding you | Low | Highest |
| Search Console | Question queries via regex, impressions above 100 | Low | Medium-high |
| PAA and autocomplete | Phrasing variants engines already associate with the topic | Low | Medium |
| Reddit and communities | Unfiltered category language and vendor sentiment | Medium | High |
| Prompt databases | Category-level prompt frequency at scale | Paid | Medium |
The Search Console recipe that takes ten minutes
๐ง A regex export you can run today
Open Performance, then Queries, then filter with a custom regex: how, what, which, why, when, where, can, does, is. Set impressions above 100 and export.
High-impression, low-click question queries are your first rewrite wave. Those are questions the engine already associates with you, where your answer is not good enough to be lifted. A category-level prompt corpus can then be used to check frequency, which is where our ChatGPT prompts database helps.
Why sales calls outrank every tool
๐ฌ The transcript is the closest thing to a prompt log
Buyers ask AI the same things they ask your rep, in the same words. The transcript is the closest thing to a prompt log that exists.
MaximusLabs AI runs question research off sales calls, support tickets, and Reddit threads as a standing input for every engagement, not as a one-time discovery exercise, which is central to our AEO question research. What surfaces in those engagements is consistently uncomfortable: the highest-intent questions are the ones the marketing team already assumed were "too basic" to write about.
The gap-finding move most teams skip
๐ค Mining returns what buyers actually struggle with
I ran an intent-gap agent across community sites for a content project recently. It surfaced client permission and buy-in for case studies as the top pain point in the category. That was missing from my outline entirely, and honestly, I would not have thought of it.
That is the value of mining rather than brainstorming. Brainstorming returns what you already believe. Mining returns what buyers actually struggle with, which is where information gain comes from, and it pairs well with a Reddit threads finder.
MaximusLabs AI keeps question research inside the standing content marketing workflow rather than the kickoff deck, because the conversational tail lives in customer language, not in keyword tools. Two hundred raw candidates inside a week is a realistic first target.
Q4. How do you expand a seed list into a full question set and size demand without prompt volume?
Two hundred mined questions is a pile, not a plan. The next two moves turn the pile into a weighted library. Both are approximations, and saying so out loud is part of doing them properly.

MaximusLabs AI expands each of 3 to 10 seed topics into 50 to 100 questions using an LLM, then crawls Perplexity and AI Mode follow-up suggestions two to three levels deep to capture fan-out variants. Demand is sized by exporting converting paid and organic terms and rewriting them as spoken questions. The result is directionally accurate rather than exact, and it gets corrected later by per-engine tracking.
The expansion loop in three passes
๐ Seed, expand, then crawl the follow-ups
Pass one: take each seed topic and ask a model for 50 to 100 questions a buyer in your ICP would ask out loud. Pass two: run the strongest of those in Perplexity and follow the suggested follow-up questions two or three levels down.
Pass three: dedupe and cluster. Follow-up suggestions are the closest public proxy to the fan-out behaviour described in Google's AI features documentation, and a query fan-out generator speeds the crawl.
Sizing demand off money terms, not volume
๐ฐ Rewrite converting terms as spoken questions
There is no truth set for prompt volume, so use the data that already has revenue attached. Export your converting paid and organic terms, then rewrite each as a question a person would say aloud.
| Converting term | Spoken question a buyer actually asks |
|---|---|
| ai sales productivity tools | "What are the best AI tools to make my sales team more productive?" |
| hubspot alternatives smb | "What should I use instead of HubSpot for a 20 person sales team?" |
| lead routing automation pricing | "How much does automated lead routing cost for a mid-market SaaS?" |
That transformation is directionally accurate. It is not measurement, and I would not present it to a board as one.
Where this model breaks
โ ๏ธ The proxy inherits your existing bias
The paid-term proxy inherits your existing bias. It only knows demand you already bid on, so it will under-represent questions in adjacent problems you have never targeted.
MaximusLabs AI's data points toward paid-term seeding being the strongest available proxy, though I might be over-trusting it for early-stage categories with thin paid history. The correction is tracking, covered later: run the questions, record appearance rates, and reweight quarterly, which is how our GEO measurement and metrics close the loop.
What gets built on Monday
โ Cut, weight, tag, and park
- Cut 200 candidates to 20 money questions using revenue proximity, not volume.
- Weight each by whether it appears in converting paid or organic data.
- Tag every question with the buyer stage it belongs to.
- Park purely informational questions where no product can be named, since they cannot influence pipeline.
MaximusLabs AI seeds every client question set from converting paid and organic terms, so the library is weighted by revenue rather than by search volume, the core of our B2B SEO approach. That is also why the first published article can go live within four days of onboarding, since the prioritisation work is already done.
Q5. Which questions deserve budget, and which should you refuse to write?
Every content plan I audit has a glossary hub in it. Forty definition pages, built because a tool said the terms had volume. Not one of them has ever been named in a buying conversation.
MaximusLabs AI sorts every question into awareness, comparison, evaluation, and purchase, then funds the bottom two first: "best X for Y", "X vs Z", "alternatives to X", and "does X integrate with Y". Glossary and definitional questions get skipped, because the model supplies its own definition and no click follows. Questions are prioritised by revenue proximity, never by search volume.
The definition page has no job left
โ Nobody reads your version of a two-sentence answer
A language model answers "what is lead routing" in two sentences. Nobody leaves that answer to read your version of the same thing.
The intent behind a definitional query was never to know the definition. The intent was to buy something, or to research options before buying. That intent now gets satisfied one layer deeper, at the comparison stage, which is where your budget belongs, and it is why our B2B SaaS AEO strategies start at the bottom of the funnel.
The conversion gap that reprices everything
๐ฐ AI traffic arrives pre-qualified
Traffic from AI answers converts at roughly six times the rate of traffic from a standard Google search. That is not a traffic story, it is a qualification story. The buyer arrived already told you were a fit.
Gartner projects that over 50% of search traffic moves to AI-native platforms by 2028. Pair that with the concentration most sites already live with: about 19 of every 20 landing pages carry roughly 85% of traffic. Spreading budget evenly across a question library is the fastest way to fund nothing that matters, which is the core argument in our revenue-focused GEO framework.
What practitioners are actually seeing
โ ๏ธ Two threads, one root cause
Two threads from r/SEO capture the pain better than any vendor deck.
"Why do AI SEO tools show me great results, but my client sees nothing in ChatGPT?"
r/SEO, Reddit Thread
"CTR is stuck at 0.5%. AI summaries are killing my clicks. How do I fight back?"
r/SEO, Reddit Thread
Both are symptoms of the same root cause. The question set being tracked is not the question set that produces pipeline.
The four-tier funding matrix
โ Fund by revenue proximity, not volume
| Tier | Question shape | Fund it? | Asset |
|---|---|---|---|
| Purchase | "does X integrate with Y", "X pricing for 50 seats" | Fund first | Docs, pricing, integration pages |
| Evaluation | "best X for mid-market SaaS", "X vs Z" | Fund first | Comparison and category pages |
| Comparison-adjacent | "alternatives to X", "is X worth it" | Fund second | Alternatives pages, switch guides |
| Awareness or definitional | "what is X", "why does X matter" | Skip | None |
MaximusLabs AI applies this matrix as RAEO, which is revenue-focused answer engine optimization, and it is the reason the deliverable is a citation rate rather than an impressions chart. Traditional agencies report impressions because impressions are easy to grow. I would rather defend twenty questions that touch revenue than two hundred that touch nothing, and that logic drives our GEO revenue attribution reporting.
The honest caveat: skipping definitional content costs you some top-of-funnel breadth. For a category nobody has named yet, that trade may be wrong. For an established category with real competitors, it almost never is.
MaximusLabs AI deliberately skips glossary content and opens every engagement with BOFU questions, because pipeline comes from evaluation prompts, not definitions. The penalty for being average has never been steeper, and average is exactly what a definition page is.
Q6. How do you cluster questions into pages, and should each engine get a different question shape?
Twenty funded questions still is not a site map. The next decision is how many pages those questions become, and how the answers get shaped for each surface.
MaximusLabs AI groups 100 to 300 related questions per landing page into one hub with self-contained sub-question sections, since engines cite passages rather than whole documents. Each question is then shaped per surface: AI Overviews prompts average 24.6 characters, while ChatGPT, Gemini, and AI Mode prompts average 57.8 to 58.1 characters. Every priority question gets a short snippet answer and a long conversational variant.
The cluster maths that keeps pages from cannibalising
๐งฉ Hub and spoke, with one job each
| Layer | Volume | Job |
|---|---|---|
| Seed topics | 3 to 10 | Revenue themes tied to ICP |
| Questions per seed | 50 to 100 | Fan-out coverage |
| Questions per landing page | 100 to 300 | One retrievable cluster |
| Hub page | 1 per cluster | Answers the strategic "what and why" |
| Spoke page | 1 per sub-intent | Answers the tactical "exactly how" |
MaximusLabs AI builds these as hub-and-spoke clusters where each spoke owns exactly one sub-intent, so no two pages compete for the same citation. Cannibalisation in AI search is worse than in Google. Two of your pages splitting a passage-level match usually means neither gets lifted, which is why our GEO strategy framework assigns a single job per URL.
Why the same question needs two shapes
๐ Self-contained sections travel, dependent ones do not
Google's guidance is that content should be organised into clear, self-contained sections, because AI features surface passages. A section that assumes the reader saw the paragraph above it cannot be extracted.
The prompt-length divergence makes this concrete. A 24.6 character query on AI Overviews wants a compact factual block. A 58 character conversational prompt on ChatGPT wants context, caveats, and a comparison, which our AEO answer structure standard handles as two separate drafts.
Per-engine question templates
โญ One question, four shapes
| Surface | Typical prompt shape | Question template to write | Answer length |
|---|---|---|---|
| Google AI Overviews | About 24.6 characters, keyword-like | "What is X for Y?" | 40 to 60 words, factual |
| ChatGPT | About 58 characters, conversational | "I run X, what should I use for Y and why?" | 80 to 150 words, with trade-offs |
| Gemini and AI Mode | About 58 characters, decomposed | "How does X compare to Z for Y team size?" | Table plus 60 word summary |
| Perplexity | Source-hungry, recency-weighted | "What is the latest data on X?" | Dated stat plus visible citation |
The platform divergence nobody plans for
๐ What one engine rewards, another ignores
The realisation that reshaped my own work was simple. What ChatGPT treats as important is not what Google treats as important, and neither matches Perplexity, a pattern documented in our citation patterns research.
MaximusLabs AI's read is that a single "AI-optimized" page is a hedge that satisfies no engine fully, and I hold that view with some caution. It costs more to write two answer shapes than one. In client engagements, the cheaper path consistently under-performs on at least one surface, usually Perplexity, where dated sourcing carries more weight.
MaximusLabs AI builds hub-and-spoke clusters where the hub answers the strategic question and each spoke owns one sub-intent, so no two pages compete for the same citation. Hubs run 5,000 words and up, spokes 2,500 to 4,000, each with a distinct job.
Q7. Should you answer a question on your own site or earn the mention elsewhere?
A cybersecurity client ranked page one on Google for its category term. In Perplexity it did not exist. The citations all pointed to review sites and one Reddit thread it had never touched.
Specificity decides the split. The more specific the question, the better an owned-content strategy works. The more general the question, the more you need earned mentions on review sites, publishers, and communities. Head terms like "best CRM" are dominated by third parties, while long-tail integration, pricing, and configuration questions belong in your own documentation. Most teams get this exactly backwards.
The decision table
๐ฏ Match the strategy to the question width
| Question type | Example | Winning strategy | Asset to build |
|---|---|---|---|
| Head, broad | "best CRM software" | Earned citations | G2, Capterra, publisher listicles, Reddit presence |
| Mid-tail | "best CRM for small SaaS teams" | Hybrid | Comparison page plus review-site profile |
| Long-tail, specific | "does HubSpot route leads into Slack automatically?" | Owned | Docs, FAQ blocks, integration pages |
| Competitor-framed | "alternatives to HubSpot" | Owned, with earned support | Alternatives page plus community answers |
MaximusLabs AI runs this split as Search Everywhere Optimization, pairing site content with G2, Capterra, Reddit, and community presence rather than treating the website as the whole surface, an approach detailed in our Reddit and forum AEO playbook.
Why the split shifts by query type
โ ๏ธ Your Google rank predicts less than you think
Roughly 87% of SearchGPT citations for commercial queries match Bing's top organic results. For mixed or informational queries, that overlap falls to around 27%. Your traditional rank is a decent predictor for product questions and a poor one for everything else.
That gap is why a Google-only agency can show green dashboards while the brand stays invisible in chat. The website was optimized. The rest of the web was not, which is precisely the gap our citation acquisition tactics are built to close.
What operators say about the measurement side
๐ฌ Practitioners are stitching free data to cheap monitors
"Are AI visibility tools becoming another overpriced SaaS category?"
r/SEO, Reddit Thread
"For my prompt research, I primarily rely on Google Search Console and AlsoAsked. To keep track of everything, I use Scrunch for monitoring."
r/SEO, Reddit Thread
The second one is the tell. Practitioners are stitching free first-party data to a cheap monitor, because the expensive tools still only answer yes or no.
Mentions beat self-published claims
๐ง The model pulls consensus, not your page
Perplexity once summarised an article of mine and described the authors as Oxford researchers. None of us went to Oxford. The model had pulled consensus from what the wider web said, not from the page itself.
MaximusLabs AI treats that as the core evidence for earned strategy on head terms, and I might be generalising from a small sample of incidents. The pattern holds across audits though. The thing mentioned most often across independent sources tends to win the general question, which is why online reputation management sits inside the GEO scope.
MaximusLabs AI splits every question list into owned and earned tracks, then builds G2 and Capterra profiles, Reddit engagement, and publisher mentions alongside the site content. That is how the UnderDefense engagement competes against multi-billion-dollar security incumbents.
Q8. How do you write and mark up an answer an AI engine will actually cite?
You can pick the right question, publish 2,000 words on it, and still never get lifted. The engine is not grading your effort. It is scanning for a block it can quote without editing.
MaximusLabs AI writes to the formats the research rewards. The Princeton and IIT Delhi GEO study found statistics addition lifted visibility 40.6%, citing sources 30.4%, and quotation addition 27.6%, while keyword stuffing scored minus 9.7%. Question headers were cited 18% of the time versus 8.9% for statement headers. Every question heading gets a 40 to 80 word standalone answer with a cited statistic inside it.
What the GEO benchmark actually measured
๐ Three sourcing behaviours carry the lift
| Tactic tested | Relative visibility change |
|---|---|
| Statistics addition | Plus 40.6% |
| Cite sources | Plus 30.4% |
| Quotation addition | Plus 27.6% |
| Keyword stuffing | Minus 9.7% |
The study ran 10,000 queries across nine datasets and validated up to 37% improvement on live Perplexity. Three tactics carried almost all the lift. All three are sourcing behaviours, not writing tricks, and they anchor our citation-worthy content standard.
The snippet is doing more work than the page
โ๏ธ Treat the meta description as an answer surface
ChatGPT frequently grounds an answer on a short excerpt, often around 150 characters. Your meta description is a direct input to that grounding step.
MaximusLabs AI treats the meta description as an answer surface rather than a click-through pitch, which inverts standard SEO practice. Write it as the compressed answer to the page's question. If a model quoted only that line, it should still be correct and attributable, a check our AI content optimizer runs before publication.
The before and after
๐ง One rewrite, two very different outcomes
Before: "Understanding Prompt Discovery Methods" followed by three paragraphs of context.
After: "Where do buyers' AI prompts actually come from?" followed by: "No public tool measures private LLM chat volume, so demand is reconstructed from sales calls, support tickets, Search Console question exports, and community threads."
The second version survives extraction alone. The first needs the page around it, which means it never travels.
The schema layer most teams skip
โ Six markup types worth the effort
- Article with author and dateModified, so recency is machine-readable.
- FAQPage for the question and answer pairs.
- HowTo for any step sequence.
- ItemList for question libraries and ranked lists.
- Person with credentials and sameAs links, connecting site, LinkedIn, and Crunchbase.
- Dataset where you publish first-party numbers.
Each of these is covered in our schema markup guide.
MaximusLabs AI scores every article across ten dimensions before publication, with citation-worthiness and factual density as two of them, and a floor of 80 out of 100 for pillar pages. That scorecard exists because self-assessment of "is this citable" is unreliable. I have been wrong about my own drafts often enough to want a rubric.
The uncomfortable implication is that primary-source scaffolding is itself the tactic. Citing Google's documentation and a peer-reviewed paper is not academic decoration. It is the single highest-scoring behaviour in the benchmark, which is why our GEO service treats research sourcing as production work, not polish.
MaximusLabs AI writes every section to a 40 to 80 word answer nugget carrying a primary-source statistic, which is precisely the format the GEO research shows engines extract. The structure is mandatory internally, not a stylistic preference.
Q9. How do you track whether your questions are actually winning citations?
Publishing is the easy half. The harder half is knowing whether the engine picked you up, because no analytics property will tell you that a model named your brand inside a private chat.
MaximusLabs AI tracks a fixed prompt set across ChatGPT, Perplexity, Gemini, Claude, and Copilot on a repeating schedule, recording appearance rate, citation rate, share of voice, and sentiment per engine. Referral traffic from AI platforms is logged separately in GA4, and server logs are checked for GPTBot, PerplexityBot, and Google-Extended activity. Position tracking is retired as the primary metric.
The four metrics that replace rankings
๐ What to record every run
| Metric | What it answers | How it is captured |
|---|---|---|
| Appearance rate | How often the brand is named at all | Percentage of tracked prompts mentioning the brand |
| Citation rate | How often a URL is linked as a source | Percentage of answers citing an owned page |
| Share of voice | Position relative to competitors | Brand mentions divided by total vendor mentions |
| Sentiment | Whether the mention helps or hurts | Positive, neutral, or negative framing per answer |
MaximusLabs AI runs this as a standing prompt panel rather than a one-off audit, which is how the Oliv AI engagement produced a measurable citation rate instead of an impressions chart. The panel is the same questions every cycle, because trend data only exists if the input set stays constant, a discipline covered in our AEO measurement and tracking guidance.
The three data layers worth wiring up
๐ง Prompts, referrals, and crawler logs
- Prompt monitoring: run the tracked set on a fixed cadence and log every answer verbatim, not just a yes or no flag.
- Referral analytics: segment GA4 traffic by AI platform source, then compare conversion rate against organic rather than volume.
- Server logs: confirm GPTBot, PerplexityBot, ClaudeBot, and Google-Extended are reaching the pages you funded, which our AI crawlability checker surfaces quickly.
The third layer is the one teams skip, and it is the cheapest diagnostic available. If the crawler never fetched the page, no amount of rewriting will fix the absence.
Why conversion beats volume in this report
๐ฐ Small numbers, disproportionate revenue
AI referral volume looks tiny next to organic, and that comparison misleads people into defunding the channel. Traffic from AI answers converts at roughly six times the rate of standard Google search traffic, so a few hundred sessions can outperform tens of thousands.
MaximusLabs AI reports citation rate and pipeline influence side by side for that reason, the approach set out in our GEO measurement and metrics framework. A dashboard that leads with sessions will get the programme cancelled before it compounds.
What practitioners say about the tooling layer
๐ฌ Free data first, paid monitor second
"Are AI visibility tools becoming another overpriced SaaS category?"
r/SEO, Reddit Thread
The scepticism is fair. Most tools answer whether you appeared, not why, and the why lives in the answer text itself, which is why we read the verbatim outputs rather than the summary score, and why our brand mention tracking comparison exists.
MaximusLabs AI treats tracking as the correction loop for question research, feeding appearance data back into prioritisation every quarter. Questions that never surface get cut. Questions where a competitor is cited get rewritten first.
Q10. What does a 90-day question research and publishing rollout look like?
Most teams stall between the research and the calendar. The gap is not strategy, it is sequencing, and a quarter is enough time to close it if the order is right.
MaximusLabs AI sequences the first 90 days in three phases: weeks 1 to 4 mine and cluster 200 plus questions and establish a baseline prompt panel, weeks 5 to 8 publish the bottom-of-funnel comparison and integration answers, and weeks 9 to 12 build earned citations and refresh underperforming blocks. Baseline measurement is captured before publication, since without it no lift can be proven.
The phase map
๐๏ธ What ships in each block
| Phase | Focus | Output |
|---|---|---|
| Weeks 1 to 4 | Mining, clustering, and baseline | 200 plus questions, 20 money questions, tracked prompt panel |
| Weeks 5 to 8 | Owned publishing | Comparison, alternatives, pricing, and integration pages |
| Weeks 9 to 12 | Earned citations and refresh | Review-site profiles, community answers, block-level rewrites |
MaximusLabs AI front-loads the baseline because it is the only unrecoverable step. Skip it in week one and you spend the rest of the quarter arguing about whether anything changed, a failure pattern documented in our GEO failures and lessons write-up.
Why bottom-of-funnel ships first
๐ฏ Prove the channel before widening it
Evaluation and purchase questions carry the shortest path to pipeline, so they generate the evidence that funds phase three. Awareness content published first produces movement nobody can attribute.
This sequencing is why the first published article can go live within four days of onboarding. The prioritisation work already happened during mining, and our B2B SEO engagements treat that speed as a deliberate proof mechanism rather than a vanity metric.
The refresh step nobody schedules
โป๏ธ Rewrite blocks, not pages
By week nine you have real answer data, and it will show specific sections losing to specific competitors. The fix is almost never a new page.
Rewrite the losing block to a self-contained 40 to 80 word answer with a dated statistic inside it, then re-run the prompt panel two weeks later. That block-level discipline is the core of our GEO content refresh process.
The honest caveat is that 90 days is enough to prove direction, not to conclude the experiment. Citation patterns shift as models retrain, so quarter two is where compounding starts, not where judgement lands.
MaximusLabs AI runs this rollout with the prompt panel fixed from week one, so every publication decision after week four is answering a measured gap rather than a guess.
Q11. What are the most common question research mistakes teams make?
The failure modes repeat across audits with unsettling consistency. None of them are exotic, and every one of them is expensive.
The five recurring mistakes are porting a Google keyword list without reshaping it, funding definitional content that models answer themselves, writing dependent sections that cannot be extracted, treating the website as the entire surface while ignoring earned mentions, and reporting impressions instead of citation rate. Each one produces green dashboards alongside zero presence in actual AI answers.
The five failures ranked by cost
โ What goes wrong, and what it costs
| Mistake | Why it happens | The fix |
|---|---|---|
| Porting the keyword list | The tool already exports it | Rebuild from sales calls and support tickets |
| Funding glossary pages | Volume looks attractive | Fund evaluation and purchase questions instead |
| Writing dependent sections | Long-form habits from Google SEO | Self-contained 40 to 80 word answer blocks |
| Ignoring earned surfaces | The site is what the team controls | Review-site profiles and community presence |
| Reporting impressions | Impressions are easy to grow | Citation rate and share of voice per engine |
The mistake underneath the other four
โ ๏ธ Optimising for a machine that stopped reading pages
All five share one root cause: optimising for a document ranking system when the retrieval layer works on passages. Keyword stuffing scored minus 9.7% in the GEO benchmark, which is the clearest signal that legacy tactics now carry a penalty rather than a discount.
MaximusLabs AI audits for these five explicitly at kickoff, and the pattern in our GEO common mistakes library is that teams fix the visible symptom and leave the root cause running.
The measurement trap that hides everything else
๐ Green dashboards, invisible brand
A Google-only agency can show rising impressions while the brand goes unnamed in every relevant chat. Roughly 87% of SearchGPT citations for commercial queries match Bing's top organic results, but for mixed or informational queries that overlap falls to around 27%.
That divergence is why our GEO versus traditional SEO comparison exists as a standing reference. The dashboards are not lying. They are measuring a different question than the one the buyer asked.
MaximusLabs AI treats mistake five as the gating one, because a team measuring the wrong thing cannot detect the other four. Fix the report first, then the content.
Q12. What should you do first if you are starting question research from zero?
If the whole system feels like a lot, it is, and none of it needs to happen in week one. There is a short sequence that produces usable output in about five working days.
Start with three moves: export Search Console question queries above 100 impressions using a regex filter, pull verbatim questions from the last 20 to 30 sales calls, and run your top 5 buying questions through ChatGPT, Perplexity, and Gemini to record who gets cited today. That produces a baseline, a competitor map, and roughly 50 real questions inside a week, with no tooling spend.
The five-day starting sequence
โ What to do, in order
- Day 1: run the Search Console regex export for how, what, which, why, when, where, can, does, and is, filtered above 100 impressions.
- Day 2: pull verbatim buyer questions from the last 20 to 30 discovery and demo call transcripts.
- Day 3: run your top 5 buying questions across ChatGPT, Perplexity, and Gemini, and log every cited domain.
- Day 4: cluster the combined list and cut to 20 money questions by revenue proximity.
- Day 5: write one self-contained answer block for the single highest-intent question and publish it.
This is deliberately unglamorous. The output is a weighted question list and a baseline, which is everything phase two needs, and it maps directly onto our AEO implementation checklist.
What to check before you write anything
๐ Confirm the machine can reach you
Verify that AI crawlers are permitted and reaching your key pages before funding new content. A robots directive blocking GPTBot or Google-Extended makes the entire question programme unmeasurable, and it is a five-minute check, covered in our AI crawler management guide.
Add an llms.txt file while you are in there. It costs nothing and makes your highest-value answers easier to locate, which our llms.txt generator handles in a single pass.
Where the first real win usually comes from
โญ The question you assumed was too basic
In engagement after engagement, the first citation arrives on a question the marketing team dismissed as obvious. Integration questions, pricing thresholds, and configuration limits get named far more often than the strategic think-pieces.
Pages sitting at position five gained 115.1% visibility from GEO tactics in the Princeton and IIT Delhi study, which means a challenger with sharp, specific answers can enter an evaluation set it could never out-rank on Google. That is the entire opportunity, and it is why our AEO service opens on specifics rather than positioning.
MaximusLabs AI starts every engagement from first-party question data rather than a tool export, because the questions that produce pipeline are already sitting in your call recordings. If you want the sequence run against your own category, talk to us.
Frequently asked questions
What is question research for AI search, and how is it different from keyword research?
Question research for AI search means reconstructing the full conversational prompts buyers type into ChatGPT, Perplexity, Gemini, and Copilot, then building content engineered to be cited inside the generated answer. The difference is structural, not cosmetic: Length: Google queries average three to four words. ChatGPT search prompts run about 8.7 words, and Google AI Mode queries reach 70 to 80 words. Unit of competition: keyword research targets a ranking position. Question research targets a citation, which is binary. When an engine names five vendors, you are inside that set or you do not exist in that conversation. Phrasing tolerance: the model infers intent, so exact-match phrasing has stopped paying. MaximusLabs AI classifies every buyer question into head, mid-tail, and long-tail bands, because each band is won by a different asset. Head questions like "best CRM" are decided by third-party citations, while long-tail integration, pricing, and configuration questions are decided by your own documentation. That reframe changes which questions get funded and which pages ever get written, which is why our GEO service starts with the question library rather than a keyword export.
How does query fan-out change the way we plan content?
One prompt does not trigger one search. Google confirms that AI Overviews and AI Mode use a query fan-out technique, issuing multiple related searches across subtopics and data sources, then stitching passages into a single answer. Standard AI Mode fires roughly 8 to 12 sub-queries. Three planning consequences follow: The research unit is a cluster of sub-intents, not a head term, because those sub-queries never appear in any analytics report. Pages are retrieved as blocks. A 4,000 word guide can lose to a competitor's 90 word pricing FAQ that answers one sub-query cleanly. Coverage beats depth on the head term. Pricing, integrations, team-size fit, review sentiment, and alternatives each have their own winner. MaximusLabs AI builds every section to a 40 to 80 word answer nugget that survives extraction with no surrounding context, and every cluster is built to cover 8 to 12 sub-intent dimensions per money question. You can approximate fan-out before writing by running the question across four engines and logging each distinct sub-topic, then scoring coverage. Our query fan-out generator shortens that step considerably.
Where do we actually find the questions buyers ask AI engines?
No public tool measures private LLM chat volume, so demand has to be reconstructed rather than exported. Seven sources cover it: Sales-call transcripts: verbatim objections from the last 20 to 30 discovery and demo calls. Highest intent quality. Support tickets: configuration, integration, and "does it do X" questions. Customer surveys: the literal prompt a buyer typed before finding you. Search Console: question queries pulled via regex, filtered above 100 impressions. People Also Ask and autocomplete: phrasing variants engines already associate with the topic. Reddit and communities: unfiltered category language and vendor sentiment. Prompt databases: category-level frequency at scale. MaximusLabs AI runs question research off sales calls, support tickets, and Reddit threads as a standing input for every engagement, not as a one-time discovery exercise. What surfaces is consistently uncomfortable, because the highest-intent questions are the ones the marketing team already assumed were too basic to write about. Two hundred raw candidates inside a week is a realistic first target, and our Reddit threads finder accelerates the community layer.
Which questions deserve budget, and which should we refuse to write?
Prioritise by revenue proximity, never by search volume. Sort every question into awareness, comparison, evaluation, and purchase, then fund the bottom two first. Fund first: "does X integrate with Y", "X pricing for 50 seats", "best X for mid-market SaaS", and "X vs Z". Fund second: "alternatives to X" and "is X worth it", supported by alternatives pages and switch guides. Skip: "what is X" and "why does X matter". The model supplies its own definition and no click follows. The economics justify the cut. Traffic from AI answers converts at roughly six times the rate of standard Google search traffic, and roughly 19 of every 20 landing pages already carry about 85% of traffic. Spreading budget evenly across a question library is the fastest way to fund nothing that matters. MaximusLabs AI applies this matrix as RAEO, revenue-focused answer engine optimization, which is why the deliverable is a citation rate rather than an impressions chart. The honest caveat is that skipping definitional content costs some top-of-funnel breadth, which may be the wrong trade in a category nobody has named yet. Our B2B SaaS AEO strategies detail the full funding matrix.
Should we answer a question on our own site or earn the mention elsewhere?
Specificity decides the split. The more specific the question, the better an owned-content strategy works. The more general the question, the more you need earned mentions on review sites, publishers, and communities. Head and broad ("best CRM software"): earned citations win, built on G2, Capterra, publisher listicles, and Reddit presence. Mid-tail ("best CRM for small SaaS teams"): hybrid, pairing a comparison page with a review-site profile. Long-tail and specific ("does HubSpot route leads into Slack automatically?"): owned, answered in docs, FAQ blocks, and integration pages. Competitor-framed ("alternatives to HubSpot"): owned, with earned support from community answers. The data explains why. Roughly 87% of SearchGPT citations for commercial queries match Bing's top organic results, but for mixed or informational queries that overlap falls to around 27%. Your traditional rank predicts product questions reasonably well and almost nothing else. MaximusLabs AI runs this split as Search Everywhere Optimization, pairing site content with G2, Capterra, Reddit, and community presence rather than treating the website as the whole surface. Most teams get the split exactly backwards, which is the gap our Reddit and forum AEO work closes.
How do we track whether our questions are winning citations and influencing pipeline?
Presence gets tracked per question, per engine, over time, not as a rank. An AI answer changes between two runs of the same prompt, so a single check tells you almost nothing. Run each priority question ten times per engine per month and record how often you appear. Appearance rate is a probability, and probability is the honest unit for a probabilistic system. Consistency is the scarce asset here: only 36% of more than 1,200 tracked brands held visibility across ChatGPT, Gemini, and Google AI every month. The tracker columns worth copying: Verbatim question, in exact buyer phrasing rather than the keyword. Engine, covering ChatGPT, Perplexity, Gemini, AI Mode, Copilot, and Claude. Buyer stage, cited yes or no per run, and appearance rate as hits divided by runs. Cited URL, and influenced pipeline for deals where the buyer named an AI answer. That last column is the one almost nobody publishes. MaximusLabs AI reports appearance rate per engine alongside a pipeline-influence column, which is why the Oliv AI engagement is described as a 64% citation rate rather than a visibility score. Our GEO measurement and metrics framework covers the full tracking spec.
How often should the question library be refreshed, and what usually kills these programmes?
Refresh quarterly. A question export ages the moment the models change, and prompt distributions shifted materially inside twelve months. Non-search ChatGPT prompts fell from 24.9 words to 13.5 words year over year, while search-enabled prompts nearly doubled to 8.7 words. Both directions moved at once, so any list built on last year's phrasing is measuring behaviour that has already drifted. Three failures kill these programmes: Treating the library as a one-off export rather than a standing process. Mass-producing unedited AI answers for every question found, which keeps collapsing under crawl and quality thresholds. Measuring appearance without revenue, which gets the programme cut in the first budget review and usually deserves to be. The 30-day rollout that avoids all three: mine 200 raw questions in week one, cut to 20 money questions by revenue proximity in week two, ship 10 answers each carrying a cited statistic in week three, then baseline every question ten times across four engines in week four. MaximusLabs AI treats the quarterly refresh as a deliverable rather than an upsell, because a decayed question set silently defunds everything downstream of it. Our GEO content refresh process runs on that cadence.