The digital marketing industry is currently undergoing a fundamental shift as artificial intelligence fundamentally alters the mechanics of information retrieval. For two decades, search engine optimization (SEO) has relied on relatively stable metrics: keyword rankings, impressions, and click-through rates. However, the rise of AI-driven search engines and "agentic" web behavior has rendered these traditional benchmarks increasingly obsolete. As brands rush to measure their "AI visibility," industry experts warn that the market is currently flooded with "vanity metrics" that provide a false sense of progress while failing to capture actual business value.
The transition from traditional search to AI-integrated platforms like ChatGPT, Perplexity, and Google’s AI Overviews has created a measurement gap. While new tools promise to track brand mentions within AI responses, these metrics often mirror the "rank tracking" of the early 2000s—a modality that does not fit the non-deterministic nature of large language models (LLMs). The challenge for modern marketing teams is distinguishing between mere citations and actual recommendations, while navigating a landscape where machines, rather than humans, are responsible for a growing share of search activity.
The Problem with Prompt Tracking and the Rise of "Machine Noise"
The dominant method currently used to measure AI search performance is prompt tracking. In this model, software tools input a predetermined list of queries into various AI models and report how often a specific brand appears in the output. While this provides a familiar dashboard for executives, technical consultants argue it is the wrong instrument for the current era. Jono Alderson, a prominent technical SEO consultant, suggests that the industry is simply "copy-pasting" old habits into a new environment. According to Alderson, the goal should not be to track prompts—which are often based on guesswork rather than real user data—but to influence how the machine perceives the brand as an entity.
Furthermore, the data used to ground these prompts is becoming increasingly corrupted by AI behavior itself. In late 2024 and early 2025, investigations by analytics consultants including Jason Packer and reports from outlets such as Ars Technica revealed a significant "leak" in search data. Private ChatGPT prompts were appearing within Google Search Console reports, a phenomenon traced to "query fan-out." This occurs when an AI system hits Google multiple times to ground a single user answer, creating a "crocodile mouth" pattern in analytics: impressions spike as machines crawl the web for information, but clicks remain flat or decline because no human ever sees the search results.
This distortion means that search-trend and keyword-volume data are no longer reliable indicators of human demand. When a brand sees its impressions rise in an AI-heavy environment, it may simply be a reflection of machines consuming its data to provide answers elsewhere, rather than a sign of growing consumer interest.
Citations vs. Recommendations: The Consensus Gap
One of the most critical distinctions in the new search landscape is the difference between being cited and being recommended. In traditional SEO, a link is generally a positive signal. In AI search, a citation is merely a footnote—a reference to a source. A recommendation, conversely, is an active endorsement by the model.
Recent data highlights a staggering disconnect between these two states. An analysis by SEO expert Lily Ray, tracking 100 business software queries over a three-month period in 2024, found that when a brand’s own self-promotional content was cited as a source in an AI Overview, that brand was excluded from the actual recommendation 69% of the time. Effectively, Google’s AI was reading the brand’s data and then using it to recommend competitors.
Supporting this, research from Visibility Labs involving 20,000 ChatGPT responses showed that product recommendations changed by over 80% once the "search" feature was activated. The correlation between being a cited source and being the recommended choice was a negligible 0.4%. Additionally, Kevin Indig’s analysis of 3.7 million citations revealed a "consensus gap," where 91% of cited URLs appeared in only one engine. This suggests that a brand’s visibility in one AI model does not necessarily translate to others, making a "broad" AI visibility score almost impossible to achieve through traditional means.
Alisa Scharf, Chief AI Officer at Seer Interactive, characterizes citations as a "leading indicator" at best, akin to ranking on the second or third page of Google. The hierarchy of value in the AI era begins with a citation, moves to a brand mention, and culminates in a specific recommendation. Currently, most tracking tools treat these as identical, leading brands to overinvest in "footnote visibility" that does not drive conversion.
The Stochastic Challenge: Why Single-Shot Measurement Fails
Unlike traditional search engines, which provide relatively stable results for a given query, LLMs are stochastic—they are probabilistic and non-deterministic. This means the same prompt can yield different results every time it is entered.
Rand Fishkin, founder of SparkToro, conducted research to quantify this volatility. His findings suggest that an AI answer is not a fixed data point but one of millions of potential variations. Fishkin noted that to find two identical lists of brand recommendations in the same order from a model like Claude or ChatGPT, a user would need to ask the same prompt an average of 1,500 times.
This volatility makes "single-shot" rank tracking—checking a rank once a day—statistically worthless. Instead, experts argue that AI visibility must be measured like a political poll, using a high volume of varied prompts to establish a "statistical presence" with a margin of error. "Percent of visibility" is the only honest metric in this context. It functions less like a ranking and more like a 20th-century brand awareness survey, measuring how often a brand exists within the "answer space" of a given category.
From Rankings to Brand Accuracy: A New Framework
To navigate this complexity, the focus of digital strategy is shifting toward "Brand Accuracy" and "Entity Confidence." Before a brand can worry about recommendation share, it must ensure that AI models hold accurate facts about its business.
A "Brand Accuracy Audit," as proposed by industry leaders, involves testing models on objective criteria: founding dates, headquarters locations, product lists, and key competitors. If a model consistently gets these basic facts wrong, any attempt at optimization is built on a flawed foundation. Duane Forrester, a former executive at Bing and a key figure in the development of Schema.org, argues that the goal for brands is to become the "canonical source" of knowledge for their niche.
The rationale is economic: it costs "cycles and tokens" for an AI model to build trust in a new source. If a model has already established a high confidence score for a specific brand as a trusted authority, it is unlikely to switch to a less certain source, even if that source has better "SEO."
Legal Implications and the Confidence Threshold
The necessity for brand accuracy is being reinforced by the legal system. In a landmark case in Germany, a court held Google liable for false statements generated by its AI Overviews regarding a business, ruling that the AI’s output constitutes Google’s own speech.
This legal precedent creates a powerful incentive for AI platforms to implement a "confidence threshold." It is speculated that systems will increasingly only surface entities they are highly certain about to avoid litigation and factual errors. In this environment, the most important metric is not "how high do I rank," but "how certain is the machine that I am who I say I am?"
Chronology of the AI Search Evolution
- Late 2022: Launch of ChatGPT initiates the "AI Search" era, though without live web grounding.
- Early 2023: Microsoft integrates GPT-4 into Bing; Google announces Search Generative Experience (SGE).
- Late 2023: SEOs begin noticing the "Crocodile Mouth" effect in Search Console, where AI bots inflate impression data.
- Mid-2024: Google rolls out AI Overviews (formerly SGE) to the general public; initial reports show a 61% drop in CTR for cited pages compared to traditional results.
- Late 2024: Major leaks show private user prompts appearing in public search data; German courts establish platform liability for AI hallucinations.
- 2025: The "Agentic Web" becomes the primary focus, with measurement shifting from clicks to "recommendation share" and "brand accuracy."
Conclusion: The Future of Digital Attribution
The search industry is relearning a lesson it first encountered two decades ago: that impressions and clicks are vanity numbers if they do not lead to revenue. However, the AI era adds a layer of complexity by removing the human from the middle of the transaction.
Wil Reynolds, founder of Seer Interactive, suggests that the current obsession with AI visibility is a "regurgitation" of the early days of SEO. The industry spent years moving from rankings to traffic, and then from traffic to business outcomes. AI visibility follows the same trajectory. If a brand is visible in an AI response but no action is taken, the visibility is functionally useless.
Ultimately, the shift toward AI search requires a return to fundamental branding. Brands must communicate clearly and consistently across all platforms—schema, social profiles, and third-party mentions—to ensure AI systems form an accurate picture of their entity. In the agentic web, visibility is a byproduct of certainty. Those who focus on being the most "trusted source" rather than the "highest ranker" will be the ones who survive the transition from the keyword era to the entity era.




