- Query fan-out is the retrieval step where an AI engine splits one prompt into several sub-queries, runs them in parallel, then merges retrieved passages into a single cited answer.
- Standard Google AI Mode runs roughly 8 to 12 parallel sub-queries. ChatGPT averages 4.7 per web search. Hundreds of searches applies only to Deep Search, capped near 20 iterations.
- Fan-out is conditional, not universal. Google says AI features may use it, and patent claims gate it on input length, query frequency, result quality, token count, and server load.
- Building one thin page per sub-query dilutes authority and risks Google's scaled content abuse policy. One dense page per question cluster wins more branches than six near-duplicates.
- Deeper retrieval is not better. Fact-check accuracy fell about 42% as retrieval scaled from 2 to 150 tool calls, so unambiguous self-contained passages beat sprawling coverage.
- Classic rank no longer predicts citation. Reliance on top-10 organic results for AI Overview citations fell from roughly 76% to about 38%, so citation share on revenue prompts is the real metric.
Q1. What is query fan-out, and what happens between your prompt and the answer?
Query fan-out is the retrieval technique where an AI search system decomposes one prompt into several model-generated sub-queries, runs them concurrently across the web and other sources, scores the returned passages, then synthesizes one cited answer. Google's VP of Search Liz Reid describes it as "breaking down your question into subtopics and issuing a multitude of queries simultaneously on your behalf" in the AI Mode announcement. One sentence goes in. A branching query tree comes out.
๐ The five stages between prompt and answer
Google Search Central confirms that both AI Mode and AI Overviews may use the technique, and that it surfaces "a wider and more diverse set of helpful links" than a standard results page. Underneath that sentence sits a five-stage pipeline.
- Analysis. The system reads your prompt for intent, complexity, and the kind of answer needed.
- Decomposition. It writes its own sub-queries covering subtopics you never typed.
- Filtering. It keeps a sub-query only if it is closely related and meaningfully different from ones already chosen, per US 2025/0117381 A1.
- Parallel retrieval. It runs the surviving sub-queries at the same time, not in sequence.
- Synthesis. It merges the retrieved passages into one answer and attaches citations.

Stage three is the one almost nobody explains. Diversity filtering means two near-identical pages do not both get retrieved. The system wants coverage, not repetition, which is the principle behind GEO topic clusters.
๐ The sub-queries do not look like your keywords
Here is the detail that reframes everything. iPullRank's reverse-engineering of AI Mode found sub-queries averaging 70 to 80 words, roughly 17 to 26 times the complexity of a traditional three or four word search.
So the engine is not searching "query fan-out." It is searching something closer to a full paragraph with constraints, context, and a stated use case baked in. Your page is being matched against a question a human would never type, which is why question research now replaces keyword research.
๐ฏ What this changes for a page you already own
Think of fan-out as a subtopic multiplexer. It takes one short human prompt and expands it into a branching tree covering definitions, comparisons, pricing, and use cases. Your page does not compete for the prompt. It competes, passage by passage, for individual branches.
That means a page can win branches it was never built for, and lose branches it should have owned. The unit of visibility is no longer the ranking URL. It is the extractable block inside it, which is the core premise of Answer Engine Optimization.
MaximusLabs AI rebuilt its research process around this pipeline instead of around keyword lists, tracing each claim back to the patents and platform docs that describe the retrieval behaviour rather than to blogs summarising them. The old way was one page per keyword, checked with a rank tracker. What changed is that the engine writes its own queries now, so we score client pages on passage extractability instead of keyword presence.
Q2. How is query fan-out different from query expansion, rewriting, and RAG?
Query expansion widens one retrieval by adding synonyms, so one query goes in and one comes out. Query rewriting replaces your query with a better-formed version of it. Query fan-out generates several distinct sub-queries, searches each one concurrently, then merges the results into a single answer. RAG (retrieval-augmented generation) is the broader architecture that fan-out feeds. Google never uses the phrase "query fan-out" in its patents, which describe query variants and thematic search instead.
๐งฉ Four techniques, side by side
These terms get used interchangeably in SEO writing, and the confusion has practical cost. If you think fan-out is just expansion with a new name, you will keep optimizing one page for one phrase.
| Concept | What it does | Queries in / out | Where it runs |
| Query expansion | Adds synonyms and related terms to widen one retrieval | 1 in / 1 out | Classic IR and search engines |
| Query rewriting | Replaces a messy query with a cleaner equivalent | 1 in / 1 out | Search front ends, voice assistants |
| Query fan-out | Generates several distinct sub-queries and runs them in parallel | 1 in / many out | AI Mode, ChatGPT search, Perplexity, Claude Research |
| RAG | Retrieves documents, then generates an answer grounded in them | Architecture, not a query step | Every AI answer engine |
Fan-out sits inside RAG. It is the step that decides what gets retrieved before anything gets generated.
๐ What Google's patents actually say
โ ๏ธ A warning worth keeping. Several patent numbers circulating in SEO articles as "Google's fan-out patent" belong to other assignees entirely, including Microsoft and Citibank. Check the assignee before you cite one.
The documents that genuinely describe this behaviour use different vocabulary:
- US 11,663,201 B2, "Generating query variants using a trained generative model," granted 30 May 2023, describes a model producing variant queries from a single input.
- US 11,769,017 B1, "Generative summaries for search results," covers synthesising an answer across retrieved content.
- The Thematic Search patent, filed December 2024, describes a research system for deep, broad, and complex queries that closely parallels AI Mode behaviour, as covered in Search Engine Journal's patent analysis.
- US 2025/0117381 A1, with October 2023 priority, is the one that spells out candidate sub-query generation with relatedness and diversity thresholds.
"Query fan-out" is Google's product language, introduced in marketing and documentation. "Query variants" and "thematic search" are the engineering language. Both describe the same family of behaviour, and both sit underneath Generative Engine Optimization as a discipline.
โ Why the distinction earns you trust
I spent a week last quarter chasing a patent number three separate blogs cited as definitive. It was assigned to a different company. That afternoon taught me more about this category's sourcing standards than any conference talk.
Getting the patent record right is the cheapest credibility you can buy with a skeptical SEO reader. It also protects you downstream. If your content strategy rests on a misattributed patent, every recommendation built on top of it inherits the error.
Q3. How many sub-queries fire, and does every prompt even fan out?
Standard Google AI Mode decomposes a prompt into roughly 8 to 12 parallel sub-queries, with 3 to 8 observed in live testing. ChatGPT averages 4.7 sub-queries per web search across 400 commercial and informational questions. "Hundreds of searches" applies only to Deep Search, which caps near 20 retrieval iterations. Fan-out is also conditional. Google says AI features "may use" it, and patent claims gate it on input length, query frequency, result quality, token count, and current server load.
๐ The measured numbers, with capture method attached
MaximusLabs AI logs sub-query counts per client prompt with the capture method recorded beside each entry, because a number without a method is not evidence. Here is what is actually published.
| Platform | Mechanism | Measured count | How it was captured |
| Google AI Mode | Parallel sub-queries | 8 to 12 (3 to 8 live) | Practitioner observation, not a Google disclosure |
| Google Deep Search | Multi-step reasoning loop | Up to ~20 iterations | Google product description |
| ChatGPT search | Sub-queries per web search | 4.7 average | Study of 400 questions, gpt-5.4-nano |
| Gemini (API) | Forced-grounding fan-outs | 10.7 average, range 3 to 28 | Seer Interactive, 501 prompts, 21 Nov 2025 |
| Claude Research | Parallel subagents | 3 to 5, each with 3+ tool calls | Anthropic engineering documentation |
๐ท๏ธ Label every number before you repeat it
โ The claim that AI Mode fires hundreds of concurrent searches for every query is false. It conflates standard AI Mode with Deep Search, which is a different pipeline doing a different job.
We use four labels internally, and I would recommend them to anyone reporting these figures: DOCUMENTED when a platform states it, TESTED when someone observed it with a stated method, INFERRED when it comes from an adjacent surface like an API, and UNKNOWN when nobody has published anything. By that standard, no verified sub-query count exists for Google AI Mode itself. Every figure in circulation is TESTED or INFERRED, which is why AEO measurement needs a stated method beside every number.
โ ๏ธ Sometimes the fan-out never happens
This is the part most guides skip. The patent gates sub-query generation behind conditions including input length, query frequency, result quality, token count, and server load. Simple factual prompts often run a single search nearly identical to the prompt.
MaximusLabs AI's logs point the same way, though I might be reading them too strongly given the sample size. Short navigational prompts show shallow or absent fan-out. Long comparison prompts show the widest trees. That pattern, if it holds, says your effort belongs on complex commercial questions rather than on simple definitional ones.
MaximusLabs AI sizes content plans against observed fan-out per engine, not against one industry figure applied everywhere. Agencies that quote a single number end up funding sub-queries nobody searches, and the retainer pays for coverage that never gets retrieved.
Q4. Which AI engines run query fan-out, and does it work the same way on each?
Google named the technique, but every major AI search system expands a prompt before retrieving. ChatGPT averages 4.7 sub-queries per web search. Claude's Research mode runs 3 to 5 parallel subagents, each issuing three or more tool calls, starting broad and narrowing. Perplexity displays its searches in the interface, which makes it the cheapest place to observe fan-out behaviour. Fan-out is also multimodal: Circle to Search and Lens split one image into objects and fire dozens of sub-queries.
๐ Same idea, four different mechanics
โ Treating fan-out as a Google-only phenomenon means optimizing for one of four retrieval surfaces your buyers already use.
| Engine | Expansion mechanism | Observability | What it rewards |
| Google AI Mode | Parallel sub-queries, multi-pass | Hidden, inferred only | Subtopic coverage, passage clarity |
| ChatGPT search | Model-generated sub-queries | Partially visible in tool calls | Self-contained Q&A blocks, expertise signals |
| Perplexity | Visible sub-searches | Fully visible in the UI | Recency, source transparency, readability |
| Claude Research | Parallel subagents, broad to narrow | Visible in the research trace | Long-form depth, methodology, citations |
| Copilot | Bing-grounded expansion | Hidden | Entity clarity, structured data |
Each column maps to a different brief, which is why per-platform work exists for ChatGPT, Perplexity, and Claude.
๐๏ธ Visual fan-out is already shipping
Google engineers describe the updated Circle to Search and Lens breaking a single image into multiple objects, then running simultaneous searches on each one, in Google's own explainer on visual search. They use the phrases "a dozen searches" and "dozens of sub-queries" for what looks to the user like one tap.
Almost no text-focused guide covers this. If your category is visual, meaning hardware, physical products, interfaces, or anything a buyer photographs, your image alt text, captions, and surrounding copy are retrieval surfaces now. Not decoration. The same logic drives multimodal search optimization.
๐ฏ The intersection is the real brief
Here is where it gets operationally awkward. Perplexity wants recent, readable, footnoted prose. Claude wants depth and methodology. ChatGPT wants self-contained conversational answers. AI Overviews wants a short answer-first block with clean E-E-A-T signals.
A page cannot be four different pages. It has to satisfy all four at once, which in practice means answer-first blocks, visible dated sources, plain language, and genuine depth below the fold. MaximusLabs AI's read is that the standard advice gets this backwards: most guides tell you to pick a platform and optimize for it, when the retrieval requirements overlap enough that building for the intersection costs less than building four times.
๐ฐ Why per-engine difference is a pipeline problem
Krishna Kaanth found this while running GTM at WiseMonk, where 98% of revenue came from search. What ChatGPT treats as a trust signal is not what Google treats as one, and neither matches Perplexity. The same page can be cited on one engine and absent on another with no change to its content.
That matters because the engine your buyer uses decides whether you enter their consideration set at all. When a buyer asks for the best tool in a category, five to ten names come back. There is no page two.
MaximusLabs AI optimizes client pages to the intersection of ChatGPT, AI Overviews, Perplexity, and Claude requirements rather than to a generic AI-friendly template, and tracks citations per engine daily. On the Oliv AI engagement that produced a 64% AI citation rate in 6 months, against roughly 30% for billion-dollar incumbents.
Q5. What types of sub-queries does a fan-out generate?
Fan-out trees cluster into eight variant types: exact match, semantic variation, navigational, comparison, definitional, process or how-to, list or recommendation, and contextual use-case. Each type needs a different passage shape. A comparison sub-query wants a table row. A definitional sub-query wants a 30 to 60 word block. A pricing sub-query wants explicit figures. Pages written as unbroken prose answer one or two types and forfeit the other six.
๐งพ One seed prompt, eight branches
Take a single buyer prompt: "best AI visibility tracking tool for a Series B SaaS company." Google's query variant patent, US 11,663,201 B2, describes a trained model generating variants from exactly this kind of input. Here is what the eight branches look like, and what each one needs from your page.
| Variant type | Fan-out example | Passage format it needs |
| Exact match | best AI visibility tracking tool | A direct answer block naming options |
| Semantic variation | top AI search monitoring software | The same answer using the buyer's other vocabulary |
| Navigational | [vendor name] pricing page | A clean, crawlable product or pricing page |
| Comparison | [vendor A] vs [vendor B] for B2B SaaS | A table row with named attributes |
| Definitional | what is AI visibility tracking | A 30 to 60 word standalone definition |
| Process / how-to | how to track brand citations in ChatGPT | Numbered steps, one action per step |
| List / recommendation | AI visibility tools for Series B companies | A scannable list with selection criteria |
| Contextual use-case | tool for a 3 person growth team | A qualifier sentence naming the context |
Mapping each branch to a format is the groundwork behind AEO question research.
โ ๏ธ Prose-only pages lose six of eight
Most pages carry one definitional block and a lot of narrative. That covers two branches. The comparison branch has no table to retrieve. The process branch has no steps. The contextual branch never names a company size or team shape, so the engine has nothing to match against "3 person growth team."
Format is not decoration here. It is the retrieval surface. A comparison sub-query looks for comparative structure, and an engine will pull a competitor's table over your well-written paragraph, which is why content formatting for AI search now carries real weight.
โ Run the eight-row audit on one page
Pick your highest-value commercial page. Write the eight variant types down the left column of a sheet. Fill in the plausible fan-out query for each one. Then mark covered or missing based on whether a standalone, extractable block already answers it.
Most pages come back with three or four covered. The fix is rarely a new page. It is adding the missing formats to the page you already have: one table, one numbered sequence, one explicit definition, and one qualifier sentence naming the buyer's context. The AEO implementation checklist covers the same ground block by block.
๐ฐ Which branches deserve your time
Not all eight branches carry equal commercial weight. Definitional branches get answered by the engine itself and rarely send a buyer anywhere. Comparison, list, and contextual branches are where a purchase decision actually narrows.
My working rule, and I hold it loosely because the sample is still small, is to fully cover comparison, list, and contextual first on any bottom-of-funnel page. Definitional coverage can be a single sentence. The temptation runs the other way, because definitional content is easier to write and feels productive. It also tends to be the branch where you are least needed.
Q6. Why does query fan-out break the one-keyword, one-page model?
Fan-out decouples the query you optimized for from the queries actually executed. The engine rewrites the prompt into definitional, comparative, pricing, and use-case variants, then keeps a sub-query only if it clears a relatedness threshold and is meaningfully different from ones already selected, per US 2025/0117381 A1. It also checks its own work and issues more searches on gaps. Near-duplicate pages get filtered out. Cluster coverage, not keyword density, decides how many branches you appear in.
๐ The keyword map that worked until 2024
The old workflow was clean. Pull volume, pick a head term, build one page per term, and check position weekly. Every page had a row in a sheet, and every row had a number attached to it.
That model assumed one input matched one output. Fan-out removes the assumption. As Mark Williams-Cook put it, people type the same query wanting different things, and there are many ways to express one question. Once those get fanned out, you get a massive list. That gap is the practical difference between GEO and traditional SEO.

๐ Two mechanics most guides skip
Here is where it gets harder than "write more topics." Two mechanics inside the pipeline actively penalise the obvious response.
- Diversity filtering. The patent keeps a candidate sub-query only if it is closely related to the prompt and different from sub-queries already chosen. Redundancy gets dropped at the query level, and the same logic applies to the pages retrieved.
- The multi-pass loop. Google showed at I/O 2025 that the system checks its work and issues additional searches when it detects gaps. A single strong passage is not enough if the surrounding subtopics are thin.
Together these mean six near-identical pages perform worse than one dense page. The engine is not counting your URLs. It is looking for non-overlapping coverage.
โ ๏ธ Where the old scoring quietly fails
โ A keyword gap analysis will tell you a page is fully optimized when it covers two of eight sub-query intents. The sheet turns green. The citations never arrive.
I have watched this happen on pages ranking in the top three. Position was fine, keyword coverage was fine, and the page was absent from the answer because it had no comparative block, no pricing specifics, and no named use case. Nothing in the old scorecard looks for those, which is one of the most common GEO mistakes to avoid.
โ Swap the audit, not the content calendar
The replacement is a coverage audit rather than a gap analysis. Score a page on how many sub-query intents it answers in standalone blocks, each one readable with no surrounding context.
MaximusLabs AI replaced keyword gap analysis with fan-out coverage analysis across client engagements, scoring pages on extractable intent coverage instead of keyword presence. We also skip top-of-funnel definitional volume on purpose, because AI engines already answer those questions themselves, and the clicks they produce rarely reach a buying conversation.
๐ฐ Coverage is a pipeline decision, not a content one
The practical effect is budget reallocation. Keyword volume pushes you toward broad informational pages, which is where most retainers quietly spend their output. Intent coverage pushes you toward the bottom-of-funnel and middle-of-funnel pages your ideal customer profile actually reads before a demo.
That is the part a board cares about. Not how many pages shipped, but whether the pages that influence a purchase decision are present in the answers buyers see. The B2B SaaS buyer journey in AI search maps where those moments sit.
MaximusLabs AI builds client coverage maps against ideal customer profile questions, not keyword volume, which is why client pages get cited for comparison sub-queries they never had a keyword row for.
Q7. Should you build one pillar page or a page per sub-query?
Build one dense page per question cluster. Thin pages for each variant, such as separate URLs for "best 43-inch TV" and "compact smart TV," dilute authority and risk Google's scaled content abuse policy, which explicitly names fan-out queries as a manipulation pattern. The honest caveat: the Thematic Search patent suggests domains appearing across multiple distinct URLs for related sub-queries gain aggregate prominence. Consolidate by cluster, and expand only when a subtopic has genuinely independent demand.
๐ธ The volume pitch, and why it sells so well
The pitch arrives in a deck with a big number on it. Your prompt fans out into twelve sub-queries, so you need twelve pages, and here is a retainer that produces them.
It sells because it is countable. A marketing head can show the board forty new URLs. The problem shows up two quarters later, when the pages rank for nothing and get cited for less. That failure pattern repeats across GEO failures and lessons.
๐งฉ What a topic actually is
Ethan Smith, CEO of Graphite, draws the line plainly: "for me the definition of a topic is there's one page that should target this cluster of things questions or keywords... you have one page."
That definition does real work. "Best 43-inch TV" and "compact smart TV" are not two topics. They are two phrasings of one buying decision, and a single page answering both wins both branches. Split them, and each half-page competes with the other while neither carries enough depth to survive diversity filtering.
โ ๏ธ Google wrote the warning into the definition
This is the detail almost no optimization guide mentions. Google published its formal definition of query fan-out in May 2026 in its generative AI features documentation, and the same document warns against building pages for fan-out queries, naming it under the scaled content abuse policy.
The warning and the definition live in one place. If your strategy is mass sub-query page production, you are operating against the guidance in the document that taught you the term. The same tension shows up in programmatic SEO work, where scale and quality have to be balanced deliberately.
โ๏ธ The honest counter-evidence
I will not pretend this is settled. The Thematic Search patent describes aggregate domain prominence rising when a domain appears across multiple distinct URLs for related sub-queries. Read alone, that argues for more pages.
Both can be true. Multiple genuinely distinct pages help. Multiple near-duplicate pages get filtered. The deciding question is whether each URL would still make sense if the others did not exist.
โ The decision rule I use
Three checks, in order, before any new URL gets created.
- Independent demand. Does this subtopic get asked as a standalone question, not just as a phrasing variant?
- Distinct job. Would the page answer something the hub cannot answer without losing focus?
- No overlap. Can you state, in one sentence, what this page covers that no sibling page covers?
Two no answers means it belongs inside the existing page as a new section.
| Approach | Page count per cluster | What it optimizes | Main risk |
| Programmatic micro-pages | One per sub-query | Countable output | Scaled content abuse, diversity filtering |
| Single pillar only | One | Depth and consolidation | Misses subtopics with real independent demand |
| MaximusLabs AI hub-and-spoke | One hub, spokes only on independent demand | Non-overlapping coverage | Slower to show page-count growth |
MaximusLabs AI settles this with architecture rather than volume. Hubs own the full cluster, spokes exist only where a subtopic has independent demand, and every page is checked against its siblings for overlap before a word gets written.
Q8. Does ranking number one on Google still get you cited in a fan-out?
Less than it used to. After the Gemini 3 rollout on 27 January 2026, reliance on top-10 organic results for AI Overview citations fell from roughly 76% to about 38%, and roughly 31% of citations came from organic positions beyond 100. Fan-out evaluates passage extractability across subtopics, not SERP position. A page ranking fourth with eight clean answer blocks can out-cite a page ranking first with none.
๐ The number that breaks the board deck
Most monthly reports still open with average position and impressions. That chart can keep climbing while AI citations go the other direction, because the two are now measuring different systems.
Pew Research Center tracked 68,879 searches across 900 US adults and found click-through dropped from 15% to 8% when an AI summary appeared, with only 1% clicking a cited source. So even a won ranking delivers roughly half the clicks it used to on those queries, which is the mechanic behind the decline in search referral traffic.
โ ๏ธ Why rank and citation came apart
Rank is a whole-page judgment. Citation is a passage judgment. The engine is not asking which page is best overall. It is asking which block answers this specific sub-query most cleanly.
That is why a page sitting outside the top 100 can get cited. It happens to hold the clearest passage for one branch of the tree. Authority still helps you get retrieved, but it does not substitute for a block the engine can lift, which is the whole point of citation-worthy content.
โ What stops working
Three habits quietly stop earning their keep once fan-out is in play.
- Leading with average position. It no longer predicts whether buyers see you in an answer.
- Optimizing headings and meta tags alone. Fan-out retrieval does not read your title tag as the answer.
- Treating a number one ranking as coverage. One ranked page can still miss six of eight sub-query intents.
โ What to report instead
Swap the headline metric. Track citation rate on a fixed list of revenue prompts, segmented by engine, and reviewed monthly. Keep rankings as a secondary diagnostic, not the lead.
The measurement sheet is simple: the prompt, the engine, whether you were cited, which page got cited, and whether that page sits at the bottom or middle of your funnel. Five columns, reviewed against pipeline rather than impressions. GEO measurement and metrics covers how to keep that sheet honest over time.
๐ฐ Mentions are the means, pipeline is the outcome
Here is where I will push back on my own category. Citation rate is still a visibility metric. A high one that never touches a buying conversation is a prettier vanity number, not a business result.
MaximusLabs AI measures every engagement on pipeline, not on mentions, and tracks which cited pages sit in the buying path rather than the reading path. My read is that most AI-visibility reporting stops one column short of the thing the board asked for.
MaximusLabs AI reports citation rate per prompt rather than average position, then ties it to pipeline. On the Oliv AI engagement that produced 27 qualified leads, $47K+ pipeline, 26% close rate in 4 months, with 30-40% of inbound now AI-sourced.
Q9. Does a deeper fan-out produce a better answer?
No. Peer-reviewed work (arXiv:2605.06635) found fact-check accuracy fell about 42% on average as retrieval scaled from 2 to 150 tool calls. GPT-5.4 lost 62% accuracy. Claude Opus 4.6 lost 22%. Context dilution and conflicting sources compound re-ranking errors. A wider fan-out is therefore not a bigger opportunity surface. Precision-oriented retrieval rewards unambiguous, self-contained passages.
๐ The assumption everyone is running on
The prevailing logic is simple and intuitive. More searches mean more evidence, more evidence means a better answer, and a deeper fan-out means more chances for your page to get pulled in.
Every agentic search demo reinforces it. You watch the model fire twenty tool calls, and the thoroughness feels like quality. I believed this myself until the numbers came in, and it is one of the quieter AEO mistakes in circulation.
๐ What the accuracy curve actually shows
The study scaled retrieval from 2 tool calls to 150 and measured fact-check accuracy at each depth. Accuracy did not plateau. It fell, and the fall was not uniform across models.
| Model | Accuracy change from minimal to maximal depth |
| Average across models | About 42% drop |
| GPT-5.4 | 62% drop |
| Claude Opus 4.6 | 22% drop |
Two mechanisms drive it. Context dilution, where the genuinely relevant passage gets buried among weakly relevant ones. And conflicting sources, where contradictory retrieved claims force the model to pick, and it picks wrong more often as the pile grows.
โ ๏ธ Why this is uncomfortable for my own category
โ A lot of AEO advice implicitly assumes depth is your friend. Cover more subtopics, get retrieved in more branches, win more citations. The accuracy data suggests the engines have the opposite problem.
I want to hold this honestly. It is one recent study, and it measures fact-check accuracy rather than citation selection. The two are related but not identical. My hypothesis is that engines will trend toward shallower, higher-precision retrieval as a result, and I would want to retest this in six months before building a strategy on it. That caution shapes how we treat experimental GEO techniques.
โ What to change in the writing itself
If retrieval is getting more precision-oriented, then ambiguity is the real enemy. Three concrete habits follow from that.
- Repeat entity names verbatim. Write "Google AI Mode" every time, not "the platform" or "it." A passage extracted alone loses every antecedent.
- Kill pronoun-dependent sentences. If a block opens with "This means," the reader and the engine both need the previous paragraph. That block cannot be lifted.
- State the qualifier inside the claim. Not "accuracy drops significantly," but "accuracy drops about 42% between 2 and 150 tool calls."
Each of those habits is really entity optimization applied at the sentence level.
๐ฐ The practical read for a growth lead
This changes where effort goes. Chasing the widest possible subtopic coverage is a worse bet than making fewer passages genuinely unambiguous.
A page with six airtight, standalone blocks beats a page with twenty vague ones. That is a smaller content bill and a harder editorial standard, which is not the trade most retainers are structured to make. It is also the argument behind answer structure writing.
โฐ When depth still helps you
Deep Search and Claude Research are the exceptions. Those pipelines deliberately run long chains for complex research prompts, and buyers do use them for vendor evaluation.
So the answer splits by surface. Standard AI Mode and ChatGPT search reward precision. Deep research modes reward depth and methodology transparency, meaning visible sources, stated dates, and explained method. Writing for both means clean short blocks up top and genuine depth below, in that order, which is covered further in the Claude optimization guide.
Q10. How do you map your own fan-out tree now that background query logs are hidden?
You reconstruct it rather than read it. OpenAI, Anthropic, and Google now hide background sub-query logs, which killed crowdsourced scraping projects. The working method: run each seed prompt through Perplexity where searches are visible, capture the 5 to 8 follow-up questions it generates, cross-check against the AlsoAsked API, then map those exact sub-questions to H2 and H3 headings on your primary page. MaximusLabs AI runs this reconstruction weekly on client revenue prompts, with every observation dated and the capture method recorded. It is an approximation, and should be labelled as one.
๐ธ The tool that died when the logs closed
Mark Williams-Cook described building a Chrome extension at queryfan.com, designed to crowdsource the hidden background queries AI engines fire. Users would share their searches and get access to the database for free.
Then the infrastructure changed. In his words, "now we've lost that background query data... I don't see how that's possible". The project ended because the data stream it depended on was closed off at the source. Our own query fan-out generator exists because that stream never came back.
โ ๏ธ What that means for anyone buying a tracker
That failure is the clearest signal in this category. If a specialist practitioner could not keep access to background query logs, then any vendor claiming to show you real sub-query trees is showing you a reconstruction too.
Ask the question directly before you sign anything. Where does this number come from, and was it captured from a live logged-in session or an API? The same question belongs in any AEO tools comparison you run.
โ The five-step reconstruction workflow
Ethan Smith's method is the most reproducible one available: "put them in perplexity... start with your searcher cluster and then fan out using a perplexity to say tell me what the fan out is". Here is the full loop.
- List your seed prompts. Start with 10 to 20 prompts your ideal customer profile actually types before a demo, not keyword phrases.
- Run each through Perplexity. Its interface shows the searches it runs, which no other major engine does.
- Capture the follow-ups. Record the 5 to 8 related questions it surfaces per prompt, verbatim.
- Cross-check with AlsoAsked. Pull the question clusters for the same seed to catch branches Perplexity missed.
- Map to headings. Turn the surviving sub-questions into H2s and H3s on the page that already owns the cluster.
Step two is easier if you already know how Perplexity surfaces and ranks sources.
๐ Date every observation, always
MaximusLabs AI logs each client prompt observation with a date and the capture method beside it, because an undated sub-query list is useless three weeks later.
The reason is drift. When a model updates, the fan-out changes. Without dated entries, you cannot tell whether your citation moved because your content changed or because the engine did. A dashboard number gives you no way to separate those two, which is why AI search visibility tracking needs provenance attached.
โ ๏ธ The limit worth stating out loud
This is an approximation, and I would not sell it as anything else. Real fan-out trees are session-aware and personalized, shaped by account history and context you cannot replicate from the outside.
MaximusLabs AI's logs point one way on this, though the sample is still modest. Reconstructed trees overlap heavily with observed citations on commercial prompts, and much less on broad informational ones. So I trust the method most exactly where the money is.
MaximusLabs AI treats this reconstruction as standing weekly practice on client revenue prompts rather than a one-off audit, with capture method recorded per entry. A tracker gives a number with no provenance. A dated log tells you what actually moved.
Q11. How do you structure a page to win across multiple fan-out sub-queries?
Front-load answers and write for two retrieval systems at once. Pages with direct 30 to 60 word answer blocks in the first 30% of text show a 44.2% citation rate, because citation likelihood is front-loaded and decays through the body. Each block should carry verbatim entity names for lexical matching plus natural phrasing for dense embeddings, so Reciprocal Rank Fusion floats it into the candidate pool from both directions. MaximusLabs AI requires a 40 to 80 word standalone answer block under every H2 in client content.
๐ The ski-ramp: citations cluster at the top
Citation likelihood is not evenly spread down a page. It is front-loaded, with 44.2% of citation share concentrated in the first 30% of text, then decaying through the middle and bottom, per MaximusLabs AI's citation patterns research across ChatGPT, Perplexity, and Gemini.
So section order is a retrieval decision, not an editorial preference. Your most commercially important answer belongs in the first third, not saved for a strong finish.

๐ง Build passages that two systems can both find
AI retrieval runs two matching systems in parallel. BM25 does lexical matching on actual words. Dense vector search matches on semantic meaning. Reciprocal Rank Fusion (a method for merging two ranked lists, commonly with k set to 60) combines both pools.
A passage that only wins one side rarely surfaces. So write each block to satisfy both:
- โ Use the exact entity names, product names, and technical terms a buyer would type.
- โ Phrase the surrounding sentences naturally, the way a person would ask the question.
- โ Avoid pronoun chains and vague references that make the block meaningless alone.
That dual target is the practical core of citation optimization.
๐ Schema helps parsing, not persuasion
Schema markup makes attributes like price, plan, and policy machine-readable, which helps an engine match your page to a commercial sub-query. Google's generative AI guidance frames structured data as an aid to understanding, not a ranking or citation lever.
Treat it accordingly. Add Article, FAQPage, DefinedTerm, and Organization markup because they reduce ambiguity. Do not expect schema alone to earn a citation. The schema markup basics breakdown covers which types actually earn their place.
๐ Off-site surfaces are retrieval surfaces too
Fan-out branches do not only land on your domain. Comparison and recommendation branches pull review platforms, directories, and community threads, which is where a large share of buying-prompt citations actually originate.
That means the same passage discipline applies off-site. Your G2 category description, your directory listing, and your product page on a third-party roundup all need clean, extractable, factual blocks. MaximusLabs AI calls this Search Everywhere Optimization, and in client engagements it runs as two pillars: Owned AEO, where your own content becomes the answer, and Earned AEO, where reviews, communities, press, and listicles make engines trust you. The Reddit and forum AEO playbook handles the community half.
โ The Monday-morning build checklist
Run this on one page before touching anything else.
- Add a 40 to 80 word standalone answer directly under every H2.
- Move your highest-value commercial section into the first third of the page.
- Replace pronoun openers in every block with the literal entity name.
- Add one comparison table and one numbered process sequence.
- Check readability, because dense academic prose underperforms in retrieval.
- Fix the same three things on your two strongest third-party listings.
MaximusLabs AI made the answer block structural rather than stylistic. Every H2 in client content opens with a 40 to 80 word passage that must make sense extracted alone, and pages below Flesch 55 do not ship. Traditional agencies optimize headings and meta tags, which fan-out retrieval does not read as the answer.
Q12. Which prompts should you optimize first, and how do you prove it moved pipeline?
Start with commercial branches, not definitional ones. Fan-out on a buying prompt targets price, plans, refund policy, free tier, and versus-competitor specifics, so bottom-of-funnel comparison and pricing pages are the highest-yield surface. Measure citation share on a fixed list of tracked revenue prompts in the live interface, segmented by engine. More than 60 AEO tracking tools query vendor APIs, which do not reproduce the personalized, session-aware trees real users trigger. MaximusLabs AI measures every engagement on pipeline, not on mentions.
๐ฐ The report that keeps climbing while pipeline flattens
You know this meeting. The monthly deck shows rankings up, impressions up, a new content count, and then the board asks where the pipeline is. Nobody in the room can connect the two charts.
That gap is structural, not anecdotal. Agencies sold rankings and impressions on fixed-scope retainers, with strategy from one firm and execution from another, and content written for blue links. None of it reported against revenue, which is exactly what GEO revenue attribution sets out to fix.
๐ What changed, and why definitional content stopped paying
Buyers now ask ChatGPT, Perplexity, Gemini, Claude, and Google AI Overviews, and they get one answer naming five to ten brands. That list is the consideration set. There is no page two to recover on.
Definitional branches get answered by the engine itself, so winning them produces a mention and little else. Comparison, pricing, and use-case branches are where a purchase narrows. Semrush ran its own fan-out experiment across four updated articles and saw tracked-prompt citations move from 2 to roughly 5 in a month, volatile, with brand mentions dipping. That is the realistic shape of results, not a step change. The B2B SaaS AEO strategies guide sequences those branches in order of commercial weight.

โ ๏ธ Why most trackers cannot tell you the truth
Over 60 AEO tracking tools query vendor APIs instead of scraping live, logged-in interfaces. API responses do not replicate the personalized, session-aware fan-out trees real users trigger, and providers actively hide background sub-query logs.
So treat dashboard numbers as directional trend lines. โ Do not report them as what your buyers saw. โ Verify the prompts that matter most by hand, in the live interface, on the engines your ideal customer profile actually uses. The same caveat runs through every ChatGPT tracking tool on the market.
๐ The five-column measurement sheet
MaximusLabs AI tracks client visibility daily across every major engine, re-run and regression-tested, and reports it in five columns rather than a visibility score.
| Column | What it records |
| Prompt | The exact wording a buyer types |
| Engine | ChatGPT, AI Mode, Perplexity, Claude, or Copilot |
| Cited | Yes or no, verified in the live interface |
| Page cited | Which URL the engine pulled |
| Funnel stage | Bottom, middle, or top of funnel |
The fifth column does the work. A citation on a top-of-funnel definitional page is a mention. A citation on a comparison page is pipeline influence.
โญ Mentions are the means, not the finish line
MaximusLabs AI's read is that most AI-visibility reporting stops one column short of what the board asked for. Citation rate is still a visibility number. It earns its place only when you can show which cited pages sat in the buying path.
My own bias here is worth naming. I would rather report a lower citation rate concentrated on commercial prompts than a high one spread across definitional ones.
MaximusLabs AI runs AEO with SEO underneath, strategy and execution in one team, and reports pipeline influence from tracked prompts. Nidra Goods ranked first across Google, ChatGPT, and Perplexity for "best sleep mask" from one strategy. UnderDefense competes with multi-deca-billion-dollar incumbents on citation share rather than budget, the pattern documented across our answer engine optimization case studies.
Frequently asked questions
What is query fan-out in simple terms?
Query fan-out is the retrieval technique where an AI search system takes one prompt, writes several of its own sub-queries from it, runs those searches at the same time, scores the passages that come back, and then synthesizes a single cited answer. Google's own language for it is breaking your question into subtopics and issuing a multitude of queries simultaneously on your behalf. The practical effect is that the engine, not the searcher, decides what actually gets searched. Here is what happens in sequence: Analysis. The system reads intent, complexity, and the answer format needed. Decomposition. It generates sub-queries covering subtopics nobody typed. Filtering. It drops sub-queries that are redundant or weakly related. Parallel retrieval. Surviving sub-queries run concurrently. Synthesis. Retrieved passages merge into one answer with citations. That filtering step matters most and gets explained least. Because the engine favours diversity, two near-identical pages will not both be retrieved, so coverage of a whole question cluster beats repetition. This is the mechanic that makes Answer Engine Optimization different from keyword-led SEO: the unit of visibility is the extractable passage, not the ranking URL.
How many sub-queries does a query fan-out actually generate?
Standard Google AI Mode decomposes a prompt into roughly 8 to 12 parallel sub-queries, with 3 to 8 commonly seen in live testing. ChatGPT averages 4.7 sub-queries per web search across 400 commercial and informational questions. Claude's Research mode runs 3 to 5 parallel subagents, each issuing three or more tool calls. The widely repeated claim that AI Mode fires hundreds of concurrent searches is wrong. Hundreds applies only to Deep Search, a separate multi-step pipeline capped near 20 retrieval iterations. One honest caveat belongs on every one of these numbers. Google has never published a verified count observed inside AI Mode itself, so every figure in circulation is either a practitioner observation or an inference from an adjacent surface such as the Gemini API. MaximusLabs AI logs sub-query counts per client prompt with the capture method recorded beside each entry, because a number without a method is not evidence. We label each figure as documented, tested, inferred, or unknown before it reaches a client report. Why this matters commercially: sizing a content plan against an inflated count funds coverage nobody searches. Plan against observed fan-out per engine instead, using the approach in our AEO measurement framework .
Is query fan-out the same as query expansion or query rewriting?
No. These three techniques get used interchangeably in SEO writing, and the confusion leads to real strategy errors. Query expansion widens a single retrieval by adding synonyms or related terms. One query goes in, one comes out. Query rewriting replaces a messy query with a cleaner, better-formed equivalent. Again one in, one out. Query fan-out generates several distinct sub-queries, searches each concurrently, and merges the results into one answer. One in, many out. RAG , meaning retrieval-augmented generation, is the broader architecture that fan-out feeds. Fan-out is the step deciding what gets retrieved before anything is generated. A detail worth knowing: Google never writes the phrase query fan-out in its patents. The engineering documents describe query variants and thematic search. Several patent numbers circulating in blog posts as the fan-out patent are assigned to entirely different companies, so check the assignee before citing one. If you treat fan-out as expansion with a new name, you will keep optimizing one page for one phrase and lose the other branches. The structural difference between the old model and the new one is covered in our comparison of GEO and traditional SEO .
Does every prompt trigger a query fan-out?
No. Fan-out is conditional, which almost every guide on the topic misses. Google's own documentation says AI features may use the technique, not that they always do. The underlying patent gates sub-query generation behind measurable conditions including input length, query frequency, result quality, token count, and current server load. Candidate sub-queries then survive only if they clear a relatedness threshold and are meaningfully different from ones already selected. In practice that produces a pattern worth planning around: Short navigational or simple factual prompts often run a single search nearly identical to the prompt. Long comparison and evaluation prompts generate the widest trees. The same prompt can fan out differently depending on session context and personalization. MaximusLabs AI's prompt logs point the same way, though the sample is still modest and we hold the reading loosely. Commercial comparison prompts show deep fan-out. Definitional prompts frequently show none. The implication for budget is straightforward. Effort belongs on complex, high-value commercial questions rather than on simple definitional ones the engine answers by itself. That prioritization logic sits at the centre of our AEO strategy planning work.
Should I build a separate page for every fan-out sub-query?
No. Build one dense page per question cluster instead. Thin pages targeting each variant, such as separate URLs for best 43-inch TV and compact smart TV, dilute authority and compete with each other. Google published its formal definition of query fan-out in May 2026, and the same document warns against building pages for fan-out queries, naming the practice under the scaled content abuse policy. The definition and the warning live side by side. There is honest counter-evidence. The Thematic Search patent suggests domains appearing across multiple genuinely distinct URLs for related sub-queries gain aggregate prominence. Both things can be true: distinct pages help, near-duplicates get filtered out. Three checks before any new URL gets created: Independent demand. Is this subtopic asked as a standalone question, not just a rephrasing? Distinct job. Would it answer something the hub page cannot cover without losing focus? No overlap. Can you state in one sentence what this page covers that no sibling covers? MaximusLabs AI settles this with hub-and-spoke architecture rather than page volume, checking every page against its siblings for overlap before writing starts. The structure is explained in our guide to GEO topic clusters .
How do I find the fan-out queries for my own brand?
You reconstruct them rather than read them. OpenAI, Anthropic, and Google now hide background sub-query logs, which ended the crowdsourced scraping projects that once exposed them. The working five-step loop: List seed prompts. Use 10 to 20 prompts your ideal customer actually types before a demo, not keyword phrases. Run each through Perplexity. Its interface displays the searches it runs, which no other major engine does. Capture the follow-ups. Record the 5 to 8 related questions it surfaces, verbatim. Cross-check. Pull question clusters for the same seed from a question-research tool to catch missed branches. Map to headings. Turn surviving sub-questions into H2s and H3s on the page that already owns the cluster. Date every observation and record the capture method. When a model updates, the fan-out shifts, and undated entries make it impossible to tell whether your content moved the result or the engine did. State the limit plainly: real trees are session-aware and personalized, so this is an approximation. MaximusLabs AI runs this reconstruction weekly on client revenue prompts rather than as a one-off audit, and you can start the same process with our free query fan-out generator .
How do I measure whether fan-out optimization moved pipeline?
Track citation share on a fixed list of revenue prompts, verified in the live interface, segmented by engine, and reviewed monthly. Rankings and impressions no longer predict whether a buyer sees you in an answer. The scale of the disconnect is measurable. Reliance on top-10 organic results for AI Overview citations fell from roughly 76% to about 38% after the Gemini 3 rollout, with roughly 31% of citations coming from positions beyond 100. Click-through also dropped from 15% to 8% when an AI summary appeared. Use a five-column sheet rather than a visibility score: Prompt. The exact wording a buyer types. Engine. ChatGPT, AI Mode, Perplexity, Claude, or Copilot. Cited. Yes or no, verified by hand. Page cited. Which URL the engine pulled. Funnel stage. Bottom, middle, or top of funnel. The fifth column does the work. A citation on a definitional page is a mention. A citation on a comparison or pricing page is pipeline influence. MaximusLabs AI measures every engagement on pipeline rather than on mentions, tracking which cited pages sit in the buying path. On the Oliv AI engagement that produced 27 qualified leads, $47K+ pipeline, and a 26% close rate in 4 months.