
Google's web crawler has achieved unprecedented dominance in AI data collection, according to Cloudflare CEO Matthew Prince. As reported by Business Standard, Googlebot now sees 3.2 times more of the web than OpenAI's version and 4.8 times more than Microsoft Corp.'s, giving the company a significant advantage in AI training data. This dominance is particularly concerning given that Google still controls 90% of the search market, making it a de facto gateway to the web. The company's strategy of using its search infrastructure to strengthen AI products has largely flown under the radar until recently, with the line between content collected for search and content used to generate AI answers becoming increasingly blurred as the bot helps feed both search and Google's flagship AI model, Gemini.
The impact of AI agents on internet traffic has reached a critical milestone, with Cloudflare data showing AI agents generated more than 57% of web traffic in 2026 versus about 42% for humans. This marks the first time in history that machine activity has surpassed human activity on the internet. According to Business Standard, Prince warns that this divergence will become more extreme in the coming years, potentially making humans a "rounding error" on the internet. The trend echoes the "dead internet theory" where machines increasingly create and consume online content without human interaction, with agents still largely the preserve of software developers but Mark Zuckerberg announcing last week that Meta would soon launch consumer versions for Facebook Messenger, WhatsApp and Instagram.
The UK's Competition and Markets Authority has taken decisive action against Google's practices, ordering the company in June to give website owners clear choice to block their content from being used for AI products while remaining in search results. As reported by Business Standard, the authority prohibits Google from punishing websites that choose to block their content by pushing them lower in search rankings. Additionally, Cloudflare has given Google an ultimatum to block mixed-purpose crawlers for ad-supported customers by September 15, potentially affecting millions of websites using the infrastructure firm. According to internal slides from 2024 disclosed as part of a US antitrust court case, Google considered giving websites the choice to opt out of AI data use before deciding against it because the company was "evolving into a space for monetization."
According to Business Standard, there are "rumblings" that Google wouldn't just follow the UK standard for content blocking but apply it globally. A Google spokesperson confirmed the company is testing a new setting that lets websites opt out of AI answers without hurting their search rankings, with plans to roll out the feature worldwide once UK testing is complete. This represents a significant shift from the status quo where major tech firms operate by their own rules, though the company's crawlers remain technically bundled for both search and AI purposes. The current system makes it difficult for publishers to block AI training without losing the search traffic they rely on, with websites having to trust Google to respect their wishes when the company's internal slides from 2024 showed revenue considerations can influence decisions.