- Multimodal search optimization structures images, video, audio, and text with metadata and schema so generative engines can retrieve, verify, and quote your media rather than merely rank it.
- AI engines rarely watch or listen. They read text representations: reviewed VTT captions, transcripts, OCR-readable on-screen text, alt text, filenames, EXIF, and ImageObject or VideoObject schema.
- The unit of retrieval is no longer the page. It is the roughly 150 character excerpt or the specific video timestamp an agent can lift and stand behind.
- Google gates video indexing hard: an indexed watch page, VideoObject markup, a stable thumbnail, a submitted sitemap, and fetchable files. Miss one gate and the rest stop mattering.
- Audio is the largest untapped surface. Episode pages with full HTML transcripts, named speakers, question-phrased timestamps, and PodcastEpisode schema turn invisible audio into quotable evidence.
- Sequence by commercial intent, not pillar order: BOFU images first, video captions and chapters second, audio transcripts and entity sameAs closure third, with CMS publish gates preventing decay.
Q1. What Is Multimodal Search Optimization, and Why Is It a GEO Problem Rather Than an Image SEO Problem?
A Head of Organic Growth I spoke with last quarter pulled up her image audit on a shared screen. Every alt tag was filled. Every file was compressed. Her brand still had zero presence in Gemini's visual answers. The audit was clean and the outcome was invisible, and that gap is the whole story.
Multimodal search optimization is the practice of structuring images, video, audio, and text with descriptive metadata and schema so search engines and generative AI can retrieve, understand, and cite them. It spans four pillars: image, video, audio, and cross-modal consistency. Traditional image SEO chased a thumbnail ranking. Multimodal GEO makes your media the evidence object the model quotes when it answers instead of linking.
๐ฏ The Four Pillars, Defined Plainly
Most teams treat this as alt-text housekeeping. That framing is why the work never shows up in pipeline.
The discipline breaks into four parts. Image (what the model can classify and describe), video (what it can index and timestamp), audio (what it can transcribe and quote), and cross-modal consistency (whether all three agree with your page copy). This is the terrain covered in our multimodal GEO knowledge base.
| Pillar | The retrievable unit | The failure mode |
| Image | Alt text, filename, ImageObject schema | Generic "product photo" alt text |
| Video | Transcript, chapter, thumbnail | Watch page not indexed |
| Audio | Episode transcript, speaker name | No on-domain text version |
| Cross-modal | Agreement across all surfaces | Narration contradicts page copy |
โ ๏ธ Why "Image SEO" Is the Wrong Mental Model

Image SEO was built for a results page with slots. You optimized to occupy a slot. Multimodal GEO is built for a synthesized answer with no slots at all.
Generative Engine Optimization (getting cited by AI engines, not just ranked by Google) treats the asset differently. The model is not choosing which of ten images to display. It is deciding whether your media contains a fact clean enough to repeat.
๐ก The Snippet Is the New Rank
Here is the reframe I keep coming back to. The atomic unit is no longer the page. It is the roughly 150-character excerpt, or the specific video timestamp, that the agent pulls.
If that fragment cannot stand alone, the brand gets skipped. Not ranked lower. Skipped.
Google's own video documentation makes the mechanic concrete: it defines Clip and SeekToAction structured data specifically so key moments inside a video become addressable. Google is telling you, in its own docs, that the moment is the unit.
๐ The Scale Nobody Budgets For
Google reports Lens now handles more than 20 billion visual searches every month, up from roughly 3 billion in 2021. That is not a side channel. That is a primary input surface most content plans still ignore entirely.
MaximusLabs AI's read is that the standard advice gets this backwards. The category tells you to optimize media so pages rank better. The inverse is closer to true now: pages exist to host media the model can lift. I might be reading the direction of travel too hard, but the client data keeps pointing the same way.
โ What Changes on Monday
Your deliverable changes shape. It stops being "publish the post" and becomes "publish the post plus its machine-readable media layer."
That means every asset ships with schema, a text representation, and a self-contained answer block attached to it. Nothing gets published half-legible.
MaximusLabs AI engineers every section to a 40-80 word answer nugget standard, because an asset that cannot survive extraction cannot be cited, and an uncited asset is invisible regardless of how good it looks.
Q2. How Do AI Engines Actually "See" Your Images, Video, and Audio?
There is a diagnostic I run before any strategy conversation. Open the client's product page, turn JavaScript off in the browser, and reload. Half the time the reviews vanish. So do the specs, the tabs, and sometimes the images themselves.

MaximusLabs AI's audits consistently find the same failure: AI engines rarely watch or listen, they read the text representation. That means transcripts, human-reviewed VTT captions, on-screen text lifted by OCR, alt text, filenames, EXIF, and schema. If no clean text key is pushed into the index, the model cannot retrieve your asset, cannot verify it, and cites the competitor who supplied one.
๐ The Text Key Principle
Think of every media asset as a locked box. The text representation is the key you hand the retrieval system.
No key, no access. It does not matter how good the footage is or how much the shoot cost.
This is why the transcribe-everything advice keeps surfacing. Transcribe every video, publish it as text on the page, and add chapters with timestamps so the crawler can index specific sections rather than one undifferentiated blob.
๐น On-Screen Text Is a Ranking Surface
Modern systems read the words burned into your video frames using OCR (optical character recognition, which converts pictures of text into machine-readable text). Your slide titles, your product labels, and your captions on screen all become retrievable strings.
Most video teams design typography for humans watching on a phone. Almost nobody designs it for a model parsing frames.
โ ๏ธ Auto-Captions Versus Reviewed VTT
Auto-generated captions are a draft, not an asset. They mangle product names, drop technical terms, and invent words that were never said.
Human-reviewed VTT caption files function as the semantic source of truth for the asset. When the transcript says something different from your page copy, retrieval confidence drops and the engine looks elsewhere.
| Signal | What the model does with it | Effort |
| Reviewed VTT | Treats as source of truth | Medium |
| Auto-captions | Parses, often wrong | None |
| On-screen text | Reads via OCR | Design-level |
| EXIF data | Reinforces provenance | Low |
| ImageObject schema | Confirms subject and context | Low |
๐งญ Provenance Is Part of the Read
Accurate EXIF data (the metadata your camera writes into the file) tells a classifier where an image came from. Original photography with intact provenance reads differently than a stock file that appears on 4,000 other domains.
I would not overstate this one. It is a supporting signal, not a switch, and I have not seen it flip an outcome on its own.
โ The Render-and-Extract Gate
Add one gate before publish. Render the page with JavaScript disabled, then confirm the media, transcript, and reviews are all present in the raw HTML. Our AI crawlability checker exists for exactly this test.
MaximusLabs AI runs this JavaScript-off render check on every client template, because asynchronously loaded reviews and media are invisible to the crawlers doing the summarising. What you see in the browser is not what the bot got.
Google's video documentation reinforces the point from the other direction: it requires that video content files be fetchable so the system can understand the contents and generate previews. Fetchable is the whole game.
MaximusLabs AI treats JavaScript minimisation and HTML-rendered critical content as a technical GEO implementation service line, not an afterthought, because extractability beats production value every single time.
Q3. Why Should a Founder Fund Multimodal Work Before the Next Blog Batch?
Every founder I meet has the same forecast model on the same spreadsheet tab. Keyword volume, times an assumed click-through rate, times a conversion rate. That model was built for a results page that no longer behaves the way it did.
Because a quarter of visual search is purchase-adjacent. Google reports Lens handles 20+ billion visual searches monthly, roughly 4x its 2021 volume, with 1 in 4 carrying commercial intent. AI Mode sessions run 93% zero-click. Text-only content competes for a shrinking click pool. Optimized media competes for the answer itself, where the surviving high-intent demand now concentrates.
๐ธ The Situation: A Forecast Built on a Curve That Moved
The CTR curve was the load-bearing assumption in organic forecasting for fifteen years. Position one earns roughly this much, position three roughly that much, and you plan headcount off it.
That curve has shifted underneath the model. Planning off the old one produces a number your board will eventually ask you to explain.
โ ๏ธ The Complication: What the Studies Actually Report
Independent measurement puts the organic click-through decline somewhere between 34.5% and 65% on queries where AI Overviews appear, depending on whose sample you use. Position-one CTR specifically fell from 1.41% to 0.64%, roughly a 54% loss. We track this pattern in detail in our work on AI search click-through rates.
Semrush's analysis of AI Mode sessions found a 93% zero-click rate, against 43% for standard search carrying AI Overviews and 34% without them.
| Surface | Zero-click rate | What it means for you |
| Standard search, no AIO | 34% | Old model roughly holds |
| Standard search with AIO | 43% | Clicks thin out |
| AI Mode | 93% | Citation is the outcome |
๐ The Counter-Evidence Nobody Publishes
Here is the number that complicates the panic narrative, and I think withholding it is a credibility mistake. When Semrush compared the same keywords before and after AI Overviews appeared, the zero-click rate actually dropped slightly, from 33.75% to 31.53%.
AI Overview prevalence itself was volatile through 2025, moving from 6.49% in January to a roughly 25% peak in July, then settling at 15.69% by November. This is not a smooth collapse. It is a redistribution, and the honest read is that impact is highly query-specific.
The forecasting fight has the same shape. Gartner projects search engine volume dropping 25%. SparkToro's clickstream data shows Google search grew around 21.6% in 2024. Both can hold if search is fragmenting rather than shrinking.
๐ฐ The Resolution: Reallocate, Do Not Defund
The move is not to cut organic. It is to stop spending the marginal hour on content that only competes for a click.
Segment your top 50 revenue keywords by AI Overview presence. Push production hours toward the visual-answer subsets and the BOFU pages where Lens commercial intent already lands.
MaximusLabs AI reports citation share by modality rather than sessions, because inside an AI answer there is no position two to fall back on. Our BOFU-first allocation exists for the same reason: TOFU is the first thing the engines absorbed, and our revenue-focused GEO framework is built around that constraint.
MaximusLabs AI skips top-of-funnel content deliberately and starts clients on bottom-of-funnel assets, because when the click pool shrinks, the surviving clicks are the ones closest to a purchase decision.
Q4. What Are the File-Level Specs for an AI-Ready Image, and Do Core Web Vitals Still Matter?
A technical SEO handed me a 47-page audit for a client last year. Forty of those pages were Core Web Vitals remediation. Zero pages addressed whether a single product image was extractable. The client paid for it, and it changed nothing.
An AI-ready image is original, 1200px or wider, served as WebP or AVIF under roughly 200KB, with a keyword-descriptive filename, explicit width and height attributes, async decoding, and lazy loading everywhere except the LCP element. Core Web Vitals are a rendering hygiene floor, not a citation lever: unset dimensions cause layout shift, but no CLS score has ever earned a citation.
๐ผ๏ธ The Format Layer
WebP and AVIF cut file size by roughly 25% to 50% compared with JPEG and PNG at equivalent quality. That is the single highest-return hour in this entire list.
The 1200px width floor matters for Google Discover eligibility. Below it, the asset simply is not considered.
| Spec | Target | Why |
| Width | 1200px minimum | Discover eligibility floor |
| Format | WebP or AVIF | 25-50% smaller than JPEG/PNG |
| Weight | Under ~200KB | Render speed, crawl budget |
| Filename | Descriptive keywords | Retrievable string |
| Dimensions | Explicit width/height | Prevents layout shift |
| Origin | Original, not stock | E-E-A-T signal |
โ ๏ธ Never Lazy-Load the Hero
This is the mistake I see most often. A team enables lazy loading site-wide, including on the Largest Contentful Paint element, and their LCP gets worse instead of better.
The rule is narrower than the plugin default. Preload the LCP image, set fetchpriority="high" on it, and lazy-load only what sits below the fold.
๐ The Thresholds, and Where to Fix Them
Interaction to Next Paint should sit at or under 200ms, and Cumulative Layout Shift at or under 0.1. Both are template-level problems, not URL-level problems, which is why our technical SEO and website audit work starts at the template.
Fix the blog template once and you fix a thousand posts. Chasing individual URLs in a crawler report is how teams burn a quarter.
โ The Honest Verdict on Core Web Vitals
Now the part most technical audits will not say out loud. In fifteen years of this work, I have never seen a Core Web Vitals improvement, by itself, drive a traffic increase.
Technical audits function as a security blanket. They produce a thick PDF, everyone feels diligent, and the actual constraint (whether the model can extract a fact from your asset) goes untouched.
That does not make the work worthless. Unset dimensions genuinely break layout, and a page that renders badly for a bot renders badly for retrieval. It makes it a floor, not a strategy.
โ Where the Remaining Hours Belong
Fix rendering so bots can read the asset. Then stop.
Every hour after that belongs to extractability: alt text quality, schema coverage, transcripts, and surrounding context. That is where citations actually come from.
MaximusLabs AI compresses technical remediation into a week-one sprint, then redirects the budget to extractability, and our first optimised asset typically ships by day four rather than after a two-month audit cycle.
Q5. How Do You Make Images Discoverable in Google Lens and AI Answers?
Write alt text as Subject plus Action plus Context plus Detail, never "image of product". Surround each image with 150 words of topical text and full ImageObject schema, then submit an image sitemap and test the query directly in Google Lens. Original photography beats stock, since AI visual classifiers weigh sharpness, lighting, and origin, and accurate EXIF reinforces provenance.
๐๏ธ The Four-Part Alt Text Formula, Shown Not Told
Alt text quality is easiest to teach by contrast. Here are two rewrites from real audit work.
Before: alt="dashboard screenshot"
After: alt="Sales rep reviewing an Oliv AI call-scoring dashboard, filtered to Q3 enterprise deals, showing a 64% citation-rate panel"
Before: alt="black jacket"
After: alt="Waterproof shell jacket with a magnetic storm-flap closure, recycled nylon fabric, and a high funnel neck, photographed on a model in rain"
The second version of each carries a subject, an action or state, a context, and a specific detail. That is four retrievable facts instead of zero. The same discipline runs through our GEO content optimization guide.
๐ธ Originality Is a Trust Signal, Not a Design Preference
Original imagery functions as an explicit Experience signal inside E-E-A-T (Experience, Expertise, Authoritativeness, Trustworthiness, Google's quality framework). Stock photography that appears on four thousand other domains carries no provenance at all.
Visual classifiers weigh sharpness, lighting, and origin when deciding what an image depicts and whether to trust it. Accurate EXIF metadata (the technical data your camera writes into the file) reinforces that the asset is yours.
MaximusLabs AI pairs every original diagram with ImageObject markup and roughly 150 words of surrounding context, because an unexplained image is an uncitable one.
๐ The Facet Data Problem Nobody Fixes
Here is the failure I see most often on ecommerce and product pages. All the good visual attributes live inside JavaScript filters.
The closure type, the fabric, the material, the neck style, and the fit. A human clicks a filter and sees them. An AI agent cannot click a filter, so for retrieval purposes those attributes do not exist.
| Where the attribute lives | Reachable by RAG? | Fix |
| JavaScript facet filter | โ No | Surface in text |
| Product spec table in HTML | โ Yes | Keep as-is |
| FAQ block on the page | โ Yes | Add attribute Q&As |
| Image alt text | โ Yes | Use the four-part formula |
| PDF spec sheet only | โ Rarely | Publish as HTML |
Bring the facet vocabulary into your FAQs and your H3 headers. That is the only way it becomes reachable, and it is a recurring fix in our GEO work for e-commerce.
๐งช The Monday Test
Google reports Lens now handles more than 20 billion visual searches monthly, and 1 in 4 of those carries commercial intent. That makes this a bottom-of-funnel exercise, not a brand-awareness one.
So run the test rather than trusting the audit. Photograph your five highest-revenue products or product screens, publish them at 1200px or wider with ImageObject schema, then point Google Lens at each one and see what comes back.
MaximusLabs AI's read is that the standard advice gets this backwards. The category tells you to fill every alt tag on the site. Our engagements suggest ten deeply optimised images on revenue pages beat a thousand mechanically tagged ones, though I hold that loosely since sample sizes are still small.
MaximusLabs AI treats schema optimisation across Article, FAQ, and Product markup as a core service line, and every client image ships with its schema and context block attached before publish.
Q6. What Are Google's Hard Gates for Getting a Video Indexed and Cited?
Google states the watch page must itself be indexed and already performing well in Search before its video is considered for indexing. The video must be embedded and visible, with a valid thumbnail at a stable URL. No valid thumbnail means no indexing. Add VideoObject markup with Clip or SeekToAction for key moments, label hasPart chapters as user questions, and submit a video sitemap.
๐ง The Dependency Nobody Publishes
Almost every video SEO guide starts with schema. Google's own documentation starts somewhere else entirely.
The watch page has to be indexed and already performing well in Search before the video on it is considered. Video SEO is downstream of page authority, not a parallel track.
That reorders your whole plan. Commissioning video for a weak page is spending production budget on an asset Google will not evaluate.
โ The Five Gates, In Order

Google's Search Central team published these as ranked priorities. Treat each as pass or fail.
Publicly accessible watch page. Indexed, not gated, and not behind login.
VideoObject structured data. Describes the video to the system.
Valid thumbnail at a stable URL. No thumbnail, no indexing.
Video sitemap submitted. Tells Google the asset exists.
Fetchable content files. Robots rules must not block the video file itself.
Miss gate three and gates one, two, four, and five stop mattering. This is binary, not weighted, and it is the kind of gate our technical GEO implementation work clears first.
โฐ Chapters Phrased as Questions
Once the gates pass, the optimisation moves inside the video. Clip and SeekToAction structured data let you define key moments so the system can address a specific timestamp.
Label those hasPart chapters as the exact questions your buyers ask. Not "Section 2: Implementation." Instead, "How long does implementation actually take?"
The chapter title becomes the retrievable string. Phrase it the way a prompt is phrased, which is the same principle behind our query fan-out generator.
๐ธ The Nine-Month Problem
The technical work here is small. Schema, a thumbnail, a sitemap entry, and chapter labels are days of effort, sometimes hours.
What kills it is process. You walk into an enterprise, propose a two-week fix, and get told engineering has scoped it at nine months. I have watched that exact conversation more times than I can count, and it is why velocity matters more than sophistication in this discipline.
MaximusLabs AI compresses technical remediation into a week-one sprint precisely because of that gap, and our first optimised asset typically ships by day four rather than after a quarterly planning cycle.
๐ What to Do Before Commissioning Anything
Run the URL Inspection tool in Search Console against your top ten video watch pages. Confirm each is indexed, carries VideoObject markup, and has a thumbnail at a URL that will not change.
| Check | Tool | Pass condition |
| Watch page indexed | URL Inspection | "URL is on Google" |
| VideoObject present | Rich Results Test | No errors |
| Thumbnail stable | Manual | Static URL, no CDN churn |
| Video sitemap | Search Console | Submitted, no errors |
| Robots access | robots.txt tester | Video files fetchable |
Fix the failures first. New production waits until the existing library is eligible.
MaximusLabs AI audits video indexing eligibility before any production spend, because a video sitting on an unindexed watch page is budget Google will never look at.
Q7. Why Is YouTube Often a Bigger Retrieval Lever Than Your Own Domain?
MaximusLabs AI's client work consistently shows YouTube functioning as a first-class retrieval source rather than a social add-on, with Gemini treating it as a primary corpus and YouTube surfacing as one of the most-cited domains in Google AI Overviews. A chaptered, transcribed video hosted inside a corpus the model already trusts can outperform the identical asset embedded only on your own domain.
๐บ The Claim Most Strategy Decks Get Wrong
YouTube sits in the "social" line item on almost every marketing budget I review. It belongs in the search line item.
Gemini treats YouTube as a first-class source. In practice, that means a video hosted there is not competing with your blog post. It is competing to be the thing the model quotes instead of your blog post, a dynamic we unpack in our Google AI and Gemini optimization work.
"More and more citations are coming from YouTube rather than websites."
Practitioner, r/SEO Reddit Thread
๐ง Why the Corpus Advantage Exists
Two mechanics drive this. First, transcripts on YouTube are already generated, indexed, and structured, which removes the text-key problem before you even start.
Second, corpus trust. The model has ingested enormous volumes from that domain and treats its structure as reliable. Your domain has to earn that from zero.
MaximusLabs AI runs YouTube strategy inside Search Everywhere Optimization for exactly this reason, and the same logic drives our Reddit and forum AEO work.
โ ๏ธ The Nuance I Will Not Skip
This does not mean abandon on-domain video. Google gates on-domain video indexing separately, through the watch-page and thumbnail requirements covered earlier.
The two surfaces run on different rules. Winning one does not win the other, and a brand that publishes only to YouTube surrenders the on-page multimodal cluster entirely.
I also want to be honest about the limits of the evidence here. Citation-share data by platform is still thin and moves fast. My read is directionally right, but anyone claiming precision on the numbers is overselling.
โ Dual-Publish With Chapter Parity
The practical move is publish twice, with the same skeleton. Upload to YouTube with chapters, then embed that video on a matching on-domain page with identical chapter labels and a full transcript.
| Surface | What it wins | What it needs |
| YouTube | Gemini and AIO citations | Chapters, clean transcript |
| Own domain | Page-level topical depth | Indexed watch page, VideoObject |
| Both, in parity | Reinforcing entity signals | Identical chapter phrasing |
Parity matters more than volume. When the chapter labels match across both surfaces, you give the retrieval system two consistent paths to the same fact, which is the core of citation consistency.
MaximusLabs AI's read is that the standard advice gets this backwards: agencies optimise the website and ignore the wider web, when the model is spending most of its time on platforms it already trusts more than yours.
Q8. How Do You Make Podcasts and Audio Citable?
MaximusLabs AI treats audio as the largest untapped citation surface, and the fix is procedural: give every episode an indexable page with a full transcript, named and credentialed speakers, timestamped sections, and PodcastEpisode schema. Chunk key passages into 40 to 60 word self-contained blocks so speakable schema and sentence-level citation systems can lift a clean answer. Audio with no on-domain transcript is invisible to text retrieval.
๐๏ธ The Biggest Blind Spot on the SERP
Scan the top-ranking multimodal guides and count how many give audio real treatment. Almost none do.
That gap is the opportunity. Every serious B2B brand now runs a podcast or appears on someone else's, and nearly all of that audio sits on a hosting platform with no indexable text anywhere.
The model cannot listen. If the words are not published as text on a page you control, the episode contributes nothing to retrieval, and the same constraint shapes GEO and voice search.
๐งฑ The Six-Step Episode Build
Dedicated page per episode. One indexable URL, not a rolling feed.
Full transcript in HTML. Not a PDF, not a modal, and not lazy-loaded.
Named speakers with credentials. Person schema where possible.
Timestamped section headers. Phrased as questions.
PodcastEpisode schema. Plus AudioObject on the file.
40 to 60 word answer chunks. Placed at each section head.
That final step is the one people skip. Speakable schema and sentence-level citation systems retrieve at roughly that length, so writing to it is a deliberate design choice.
MaximusLabs AI's 40 to 80 word answer nugget standard maps almost exactly onto this chunking rule, which is why the same discipline transfers cleanly from articles to audio.
๐ The Oxford Moment
A practitioner I follow watched Perplexity summarise their article and describe the team as Oxford researchers. Nobody on the team went to Oxford.
The engine had pulled the association from mentions elsewhere on the web, not from the About page. That is the whole lesson in one glitch: the agent weights what others say about you above what you say about yourself.
๐ฏ Why Guest Appearances Beat Owned Episodes
This reframes podcast strategy entirely. A transcript of you on someone else's show is earned media, and earned media carries more retrieval weight than a self-published claim.
| Asset | Retrieval weight | Effort |
| Guest appearance transcript | High (earned) | Low, pitch only |
| Own episode with transcript | Medium (owned) | Medium |
| Own episode, audio only | Near zero | Wasted |
| About page claim | Low (self-asserted) | Low |
Chase the transcripts, not just the downloads. Ask every host to publish a full text version, and republish your own segment on-domain with proper attribution using AI citation acquisition tactics.
MaximusLabs AI prioritises podcast and third-party transcripts inside Search Everywhere Optimization, because earned mentions consistently carry more retrieval weight than anything a brand publishes about itself.
Q9. How Do You Keep Your Brand and Your Modalities Telling the Same Story?
Establish one unambiguous entity, then keep every modality agreeing with it. Publish Organization, Person, and WebSite schema with a logo meeting Google's 112x112px square minimum, and close the sameAs loop so a crawler can traverse website to Wikidata to LinkedIn to Crunchbase to G2 and back. When narration, on-screen text, alt text, and page copy contradict each other, retrieval confidence drops and the model skips you.
๐งฉ The Failure Mode Looks Like a Compliment
An engine invents a credential for you and states it confidently. Somebody screenshots it and sends it to your founder.
That is not a bug in the model. It is the model filling a gap you left open, using scraps from across the web because your own entity signals were too thin to override them.
๐ Closing the sameAs Loop

Entity work is a traversal problem, not a tagging problem. The goal is a closed circuit the crawler can walk, which is the core idea behind GEO knowledge graphs.
Website to Wikidata, to LinkedIn, to Crunchbase, to G2, and back to your website. Each hop confirms the last one.
| Node | What it confirms | Common gap |
| Website schema | Canonical name, logo | Logo under 112x112px |
| Wikidata | Machine-readable identity | Entry missing entirely |
| People and org linkage | Name variant mismatch | |
| Crunchbase | Funding, founding facts | Stale data |
| G2 or Capterra | Category, peer proof | No sameAs back-link |
MaximusLabs AI builds G2, Capterra, and Gartner Peer Insights presence as entity anchors rather than social proof, because those profiles are nodes a crawler actually traverses.
โ ๏ธ Operationalising Cross-Modal Consistency
Cross-modal consistency gets named in every guide and defined in almost none. Here is what the audit actually looks like.
Take one asset. Pull four text surfaces: the video transcript, the on-screen text, the image alt text, and the page copy. Put them side by side and read for contradictions.
Common contradictions I find:
The transcript says "under two weeks" while the page says "30 days."
The product name has a hyphen in the alt text and not on the page.
The narration cites a 2024 figure and the copy cites a 2026 one.
On-screen text abbreviates a term the copy spells out.
Every mismatch is a confidence penalty. The model has two versions of a fact and no way to pick, so it goes elsewhere, which is why citation consistency is a retrieval issue and not a style issue.
๐๏ธ What Does an Agent Feel When It Sees Your Logo?
That question sounds strange until you sit with it. A model does not feel anything, but it does form a representation of your brand from thousands of scattered signals.
If those signals are consistent, the representation is sharp. If they conflict, it is blurry, and blurry entities do not get recommended.
MaximusLabs AI's read is that the category treats this as schema hygiene when it is really brand work. Brand is the moat here, not the hack, and I hold that view strongly even though it is harder to sell than a technical fix.
๐ The Monday Audit
Search your brand name across the five nodes above and write down every variant you find. Legal name, trading name, spacing, capitalisation, and the abbreviation your sales team uses.
Pick one. Enforce it everywhere, including video narration scripts and podcast intros.
"AI Citation Tracking feels like snake oil. The results vary every single time."
Practitioner, r/SEO Reddit Thread
That scepticism is fair, and entity consistency is one of the few levers that reduces the variance rather than just measuring it.
MaximusLabs AI runs review-platform optimisation across G2, Capterra, and Gartner inside its online reputation management service line, and treats each profile as a machine-readable identity claim rather than a badge for the homepage.
Q10. How Do You Score and Measure Multimodal Citation Readiness?
MaximusLabs AI scores assets 0 to 4 across four pillars, image extractability, video indexing eligibility, audio transcript coverage, and cross-modal consistency, where 4 means citable without human interpretation and 0 means invisible. Measurement then runs four layers: Search Console image and video indexing, GA4 media engagement events, citation appearances by modality across ChatGPT, Perplexity, Gemini, and AI Overviews, and assisted conversions.
โ Why Flat Checklists Fail
Every multimodal checklist on the web is unweighted. Twenty boxes, all equal, no thresholds.
That design makes prioritisation impossible. A missing video thumbnail (a hard indexing gate) sits beside a nice-to-have filename tweak, and the intern fixes the filename.
โญ The 0 to 4 Rubric
Microsoft's internal grading systems use a 0 to 4 integer scale to score results against confirmed intent, where 4 is ideal and 0 is irrelevant. That granularity works well here.
| Score | Meaning | Example |
| 4 | Citable as-is, no interpretation needed | Chaptered video, reviewed VTT, VideoObject |
| 3 | Citable with minor cleanup | Good transcript, weak chapter labels |
| 2 | Present but ambiguous | Alt text exists, generic phrasing |
| 1 | Technically indexed, semantically empty | "IMG_4471.jpg", no context |
| 0 | Invisible | Audio-only, no transcript |
Anything scoring 2 or below on a bottom-of-funnel page is the queue. Everything else waits.
โ The Ten-Step Runnable Audit
Inventory every image, video, and audio asset.
Score resolution and file format.
Rewrite alt text using the four-part formula.
Verify ImageObject, VideoObject, and PodcastEpisode schema.
Confirm video watch pages are indexed.
Replace auto-captions with reviewed VTT.
Add chaptered timestamps phrased as user questions.
Test key queries directly in Google Lens.
Submit image and video sitemaps.
Track citations and impressions in Search Console.
MaximusLabs AI applies the same scoring discipline it uses on articles, where a 10-dimension scorecard sets an 80 out of 100 floor for hub pages, to every media asset before publish. Our AI content optimizer runs the same checks on text.
๐ The Four Measurement Layers
| Layer | Where it lives | What it tells you |
| Indexing | Search Console image and video reports | Is the asset eligible? |
| Engagement | GA4 media events | Do humans use it? |
| Citation | Prompt panels across four engines | Does the model quote it? |
| Revenue | Assisted conversions | Does it pay? |
Layer three is the one nobody builds. Run a fixed set of prompts weekly, log which modality got cited, and chart share of voice across question variants rather than rank, using AI search visibility and brand mention tracking.
โ ๏ธ The Honest Part About Attribution
Practitioners are right to be suspicious of this measurement layer.
"Getting cited by AI search but no traffic. Anyone successfully measuring GEO impact?"
Practitioner, r/SEO Reddit Thread
AI citation attribution remains genuinely imperfect. Results shift between runs, referral data is patchy, and anyone selling you a clean number is smoothing over real variance. We cover the same limits in our work on GEO revenue attribution.
MaximusLabs AI's data points toward citation share being the most stable leading indicator available, though I might be reading it more confidently than the sample size deserves. What I will defend is the direction: with AI Mode sessions running 93% zero-click, sessions alone cannot be the scoreboard.
MaximusLabs AI tracks share of voice across thousands of question variants rather than single rankings, because inside a generated answer there is no position two to measure against.
Q11. What Should You Fix First With One Quarter and One Designer?
MaximusLabs AI sequences multimodal work by commercial intent rather than pillar order. Weeks 1 to 4: original images with ImageObject schema on bottom-of-funnel product, pricing, and comparison pages, where 1 in 4 visual searches already carries purchase intent. Weeks 5 to 8: reviewed VTT captions and question-labelled chapters on your top 10 videos. Weeks 9 to 12: episode transcripts, entity sameAs closure, and CMS publish gates that block non-compliant assets.
๐ธ The Complication: Four Equal Pillars Is a Plan Nobody Finishes
Give a team of two an evenly weighted four-pillar plan and watch what happens. They start on images, drift into video, and abandon both by week seven.
Completeness is the enemy here. The constraint is one designer and eighty working days, and that constraint should drive the order.
๐ฐ The Resolution: Weight by Commercial Intent
Google reports 1 in 4 Lens searches carries commercial intent. That single figure decides the sequence, and it anchors the GEO strategy framework we run with constrained teams.
| Window | Focus | Why now |
| Weeks 1-4 | BOFU images plus ImageObject schema | Visual search skews commercial |
| Weeks 5-8 | Top 10 videos: VTT, chapters, indexing gates | Highest MOFU retrieval value |
| Weeks 9-12 | Audio transcripts, sameAs closure, publish gates | Authority and durability |
Notice audio comes last. It is the biggest gap on the SERP, and it is still the right thing to defer when cash and hours are finite.
๐ Governance: The Part That Stops the Decay
Here is what kills every audit I have ever seen. The work gets done, the quarter ends, and six months later the new intern publishes forty images with no alt text.
Fix it at the template, not the person. Three enforcement moves:
Naming conventions, enforced at upload, so no asset enters the CMS as "final_v3.png".
Transcript ownership, assigned to a named role, not "the content team".
Publish gates, where the CMS refuses to publish a post with a missing alt attribute or absent VideoObject markup.
MaximusLabs AI implements schema at the template level during its week-one technical audit sprint, which means compliance survives staff turnover instead of depending on it.
๐ป The Ghost Kitchen Problem
A useful way to think about what comes next. Your website is the dining room, designed, styled, and photographed for humans.
Agentic commerce is the kitchen. The delivery driver, an AI agent, never sees the dining room. It only needs the data feed to fulfil the order.
I tested this myself. I asked Gemini to buy snowboard pants and complete checkout end to end. It did not work.
โฐ What That Failure Actually Signals
The failure was not the model's. It was the storefronts, which exposed nothing an agent could act on.
A beautiful site with no clean data layer is a technical liability now, not just a missed opportunity. That gap will close faster than most roadmaps assume, as our state of agentic commerce 2026 research sets out.
MaximusLabs AI starts clients BOFU-first and ships the first optimised asset within four days, because sequencing beats completeness when the quarter is already half gone.
๐ค What I Am Sitting With
My open question is whether cross-modal consistency becomes a formal ranking input or stays an implicit confidence signal. If it formalises, entity work stops being defensive and becomes the whole game.
I do not know the answer yet. If you are running these tests inside your own stack, I would genuinely like to compare notes: krishna@maximuslabs.ai.
Frequently asked questions
What is multimodal search optimization, and how is it different from traditional image SEO?
Multimodal search optimization is the practice of structuring images, video, audio, and text with descriptive metadata and schema so search engines and generative AI can retrieve, understand, and cite them. It spans four pillars: image, video, audio, and cross-modal consistency. The difference from image SEO is the outcome you are competing for. Image SEO was built for a results page with slots. You optimized to occupy a slot. Multimodal GEO is built for a synthesized answer that has no slots at all. The model is deciding whether your media contains a fact clean enough to repeat. That reframe changes the atomic unit. It is no longer the page. It is the roughly 150 character excerpt, or the specific video timestamp, that an agent pulls. If that fragment cannot stand alone, the brand does not rank lower. It gets skipped. MaximusLabs AI engineers every section to a 40 to 80 word answer nugget standard for exactly this reason, because an asset that cannot survive extraction cannot be cited, and an uncited asset is invisible regardless of how good it looks. If you want the underlying discipline, our multimodal GEO framework explains how media becomes an evidence object rather than decoration.
How do AI engines actually see your images, video, and audio?
They do not watch or listen. They read the text representation of the asset, and if no clean text key is pushed into the index, the model cannot retrieve your media, cannot verify it, and cites the competitor who supplied one. The text keys that matter: Human-reviewed VTT captions , treated as the semantic source of truth for a video. Auto-generated captions , which are a draft, not an asset, because they mangle product names and invent words. On-screen text , lifted through OCR, which means slide titles and product labels become retrievable strings. Alt text, filenames, and EXIF data , which reinforce subject and provenance. ImageObject and VideoObject schema , which confirm context. Think of every media asset as a locked box. The text representation is the key you hand the retrieval system. No key, no access, regardless of how much the shoot cost. MaximusLabs AI runs a JavaScript-off render check on every client template, because asynchronously loaded reviews, specs, and media are invisible to the crawlers doing the summarising. What you see in the browser is not what the bot got. Our technical audit process treats HTML-rendered critical content as a service line, since extractability beats production value every time.
What are the file-level specs for an AI-ready image, and do Core Web Vitals still matter?
An AI-ready image is original, 1200px or wider, served as WebP or AVIF under roughly 200KB, with a keyword-descriptive filename, explicit width and height attributes, async decoding, and lazy loading everywhere except the LCP element. Width: 1200px minimum, which is the Google Discover eligibility floor. Below it the asset is not considered. Format: WebP or AVIF, roughly 25 to 50 percent smaller than JPEG or PNG at equivalent quality. Hero images: never lazy-load them. Preload the LCP image and set fetchpriority high. Origin: original photography, since stock that appears on four thousand domains carries no provenance. On Core Web Vitals, the honest verdict is that they are a floor, not a strategy. Unset dimensions genuinely break layout, and a page that renders badly for a bot renders badly for retrieval, but no CLS score has ever earned a citation on its own. MaximusLabs AI compresses technical remediation into a week-one sprint and then redirects the remaining budget to extractability: alt text quality, schema coverage, transcripts, and surrounding context. You can pressure-test your own baseline with our AI crawlability checker before committing engineering hours.
What does Google require before it will index and cite a video?
Google gates video indexing behind five hard requirements, and they are binary rather than weighted. Miss the third and the other four stop mattering. A publicly accessible watch page that is indexed, not gated, and not behind a login. VideoObject structured data describing the video to the system. A valid thumbnail at a stable URL. No thumbnail, no indexing. A submitted video sitemap telling Google the asset exists. Fetchable content files , with robots rules that do not block the video file itself. The dependency almost nobody publishes is that the watch page must already be indexed and performing well in Search before the video on it is considered. Video SEO is downstream of page authority, not a parallel track, which means commissioning video for a weak page is spending production budget on an asset Google will not evaluate. Once the gates pass, optimisation moves inside the video. Use Clip and SeekToAction markup for key moments and label hasPart chapters as the exact questions buyers ask, not "Section 2: Implementation." MaximusLabs AI audits video indexing eligibility before any production spend, and our technical GEO implementation work handles the schema layer at template level.
Why is YouTube often a bigger retrieval lever than your own domain?
Because the model already trusts the corpus. YouTube functions as a first-class retrieval source rather than a social add-on, with Gemini treating it as a primary corpus and YouTube surfacing as one of the most-cited domains in Google AI Overviews. Two mechanics drive this: Transcripts already exist. They are generated, indexed, and structured on the platform, which removes the text-key problem before you start. Corpus trust. The model has ingested enormous volumes from that domain and treats its structure as reliable. Your domain has to earn that from zero. This does not mean abandoning on-domain video. Google gates on-domain indexing separately through the watch-page and thumbnail requirements, so the two surfaces run on different rules and winning one does not win the other. The practical move is to publish twice with the same skeleton: upload to YouTube with chapters, then embed that video on a matching on-domain page with identical chapter labels and a full transcript. MaximusLabs AI runs YouTube strategy inside Search Everywhere Optimization for this reason, and the same logic drives our Reddit and review-platform work. Citation-share data by platform is still thin and moves fast, so treat the direction as reliable and the precise numbers as provisional.
How do you make podcasts and other audio content citable by AI engines?
Audio is the largest untapped citation surface on the SERP, and the fix is procedural rather than creative. The model cannot listen, so if the words are not published as text on a page you control, the episode contributes nothing to retrieval. The six-step episode build: A dedicated page per episode. One indexable URL, not a rolling feed. The full transcript in HTML. Not a PDF, not a modal, not lazy-loaded. Named speakers with credentials , using Person schema where possible. Timestamped section headers , phrased as questions. PodcastEpisode schema , plus AudioObject on the file. Answer chunks of 40 to 60 words placed at each section head. That final step is the one people skip. Speakable schema and sentence-level citation systems retrieve at roughly that length, so writing to it is a deliberate design choice rather than a stylistic one. MaximusLabs AI treats audio as the biggest blind spot in most content plans, and our 40 to 80 word answer nugget standard maps almost exactly onto this chunking rule, which is why the discipline transfers cleanly from articles to episodes. The same structural logic underpins our content formatting standards for AI search .
How do you measure multimodal citation readiness and sequence the work with one quarter and one designer?
Score first, then sequence by commercial intent rather than pillar order. MaximusLabs AI scores assets 0 to 4 across four pillars: image extractability, video indexing eligibility, audio transcript coverage, and cross-modal consistency. A 4 means citable without human interpretation. A 0 means invisible, such as audio with no transcript. Anything scoring 2 or below on a bottom-of-funnel page enters the queue, and everything else waits. Measurement then runs four layers: Indexing: Search Console image and video reports tell you whether the asset is eligible. Engagement: GA4 media events tell you whether humans use it. Citation: a fixed weekly prompt set across ChatGPT, Perplexity, Gemini, and AI Overviews tells you whether the model quotes it. Revenue: assisted conversions tell you whether it pays. For sequencing on eighty working days: weeks 1 to 4 on BOFU images with ImageObject schema, weeks 5 to 8 on reviewed VTT captions and question-labelled chapters for your top ten videos, weeks 9 to 12 on episode transcripts, entity sameAs closure, and CMS publish gates. Attribution here remains genuinely imperfect and results shift between runs, which is why our GEO measurement approach reports citation share by modality instead of sessions.