For years, digital marketers have operated in a state of strategic ambiguity regarding artificial intelligence. While professionals widely recognize that AI models learn about corporate brands by ingesting website content and subsequently drive visitors through chat-based citations, the exact mechanics of this relationship have remained largely opaque. Questions regarding which specific web pages attract the highest volume of AI bot attention, and which URLs actually convert that attention into human referral traffic, have largely gone unanswered.

To bridge this crucial information gap, researchers have recently completed a comprehensive empirical study examining 74 distinct marketing accounts. By analyzing upstream bot-tracking metrics—specifically utilizing telemetry data from Cloudflare platforms like the "Most Crawled Paths" and "AI Referral Traffic" reports—analysts have mapped out how artificial intelligence systems interact with corporate and publishing websites. The findings challenge long-held assumptions about content strategy, search optimization, and site architecture in an increasingly AI-driven digital economy.
Homepages Command Disproportionate AI Attention
The investigation’s primary revelation centers on the sheer concentration of artificial intelligence crawls directed toward corporate homepages. According to the dataset, AIs visit homepages roughly 15 times more frequently than any other type of individual web page.

When evaluating URL distribution, standard statistical expectations dictate that if bot requests were evenly dispersed, a page type representing 5 percent of a site’s total URLs would capture approximately 5 percent of the crawl volume. However, the empirical data completely shatters this baseline. Even when researchers aggregated alternative page classifications—including product and service pages, blog articles, "about" sections, contact pages, pricing lists, and case studies—the combined total failed to match the proportional weight assigned by AI models to the homepage alone.
This disparity suggests a fundamental shift in how digital strategists must view brand real estate. While content marketers have historically obsessed over the question of what kind of content to publish to capture search visibility, the Cloudflare data indicates that where information is hosted holds paramount importance for AI ingestion. Industry experts note that many publishers have historically undersold the strategic value of their homepages, frequently offloading core brand messaging to secondary support or resource sections. If a brand seeks to ensure that large language models accurately understand its foundational identity, essential messaging must be anchored directly on the homepage.

Website Scale and the Correlation with Bot Interest
Expanding beyond individual page types, the study explored whether the overall size of a website influences bot behavior. The findings reveal a robust positive correlation (0.86) between total page count and overall AI attention. Simply put, larger websites command a higher volume of total bot requests.
This relationship grows in a nearly proportional manner relative to page count, demonstrating neither a sudden drop-off nor a compounding advantage at extreme scales. Much like traditional search engine crawlers, automated AI bots allocate resources across expanded URL footprints, recognizing that a greater volume of pages offers more entry points to match against diverse user queries and prompts.

Nevertheless, the data highlights notable outliers. Several smaller websites within the dataset—some maintaining a lean footprint of roughly 50 pages—attracted bot attention levels comparable to sprawling sites containing upwards of 1,000 URLs. Analysts attribute this over-performance to superior technical optimization, higher-authority brand recognition baked into foundational model training data, or exceptionally disciplined content architecture.
The Conversion Gap: Crawls Versus Referrals
A critical distinction drawn in the report separates raw AI crawls from actual referral traffic. While a crawl signifies an automated bot reading a page to update a model’s knowledge base, a referral represents a human user clicking an active citation link embedded within an AI-generated conversational response.

By examining a subset of 27 websites equipped with advanced Cloudflare tracking tiers, researchers discovered pronounced efficiency variations across page categories. Service and product pages—defined as decision-shaping promotional assets—vastly outperform informational articles in generating actual human visits. On a per-page basis, service and product pages earn roughly three times more total AI-referral traffic than a standard blog article.
Conversely, informational resources suffer from what researchers term the "Dark Library Effect." Articles and blog posts are heavily crawled and frequently summarized by AI systems to answer user inquiries directly, yet they rarely generate outbound clicks. This phenomenon mirrors the zero-click search dynamics that have disrupted traditional search engine optimization (SEO) for years. Consequently, marketing analysts caution against applying rigid search traffic metrics to content marketing assets designed primarily for top-of-funnel brand awareness, educational support, or thought leadership.

URL Depth and the Architecture Tax
The analysis also evaluated URL structure and page depth, tracking how far down folder hierarchies AI bots are willing to travel. The data indicates that AI crawlers are technically proficient at discovering pages situated three or four folders deep within a site’s directory.
However, a steep drop-off occurs regarding referral traffic. While bots index deep pages, they rarely recommend them to end users. A page buried three folders deep typically earns only a fraction of the referral traffic predicted by its footprint, while pages four levels deep experience near-zero referral performance.

This trend reflects standard bot crawl budget limitations, wherein automated systems allocate finite computational resources efficiently. For web developers and digital architects, this introduces an implicit "architecture tax" on convoluted site designs. While experts advise against haphazardly restructuring existing websites solely to appease AI referral patterns, the insight offers valuable guidance for upcoming site migrations and new architectural builds.
Diagnostic Tools and Implications for Digital Strategy
To help organizations evaluate their own digital infrastructure, the researchers released open-source diagnostic prompts designed to parse raw CSV exports from Cloudflare’s AI Crawl Control metrics. By integrating this data into analytical workflows, digital teams can quickly audit whether their proprietary content is effectively channeled toward lead generation or trapped within inaccessible folder structures.

Furthermore, ancillary evaluations of domain accessibility reveal that many web publishers continue to experiment with blocking AI crawlers entirely via robots.txt protocols or security layers. However, industry analysts emphasize a fundamental paradox: blocking automated bots prevents models from understanding a brand, ultimately undermining the capacity of AI systems to cite and recommend products or services to prospective buyers.
Ultimately, the empirical findings underscore that while public internet data trains foundation models, corporate websites remain the single domain under direct brand control. As artificial intelligence increasingly mediates the discovery phase of the buyer journey, aligning site architecture, homepage prominence, and high-value service content will remain vital for organizations seeking to secure visibility in automated recommendation ecosystems.




