The internet faces a mounting threat that operates largely unseen: artificial intelligence systems are degrading the web's core function of helping people find and share information. Over 60% of adult internet users rely on the web primarily for information discovery, yet the infrastructure supporting this mission is buckling under pressure from AI extraction.
Large language models powering services like OpenAI's ChatGPT and Anthropic's Claude depend on massive quantities of scraped web content. The automated bots performing this extraction traverse websites, download their content, and follow links to other sites in a recursive process that strains infrastructure while returning little value to the sources being mined. As Dr. Zachary McDowell, an academic expert on the topic, told us, 'These AI systems are not web crawlers really, more miners. Crawlers are like cartographers, while AI systems are like strip miners.'
The consequences are tangible: websites slow to a crawl under bot traffic, publishers erect paywalls to protect their work, and openly accessible repositories like Wikipedia and libraries struggle with mounting technical costs. The damage extends beyond individual sites to the internet's fundamental ability to serve as infrastructure for democratic communication and access to reliable information.
Why this matters now
These disruptions strike at what makes the internet indispensable, particularly during moments when trustworthy information is most needed. Rather than functioning as a platform for human communication, the internet increasingly serves as raw material for commercial software production.
The degradation unfolds in multiple ways. First, AI bot activity imposes severe technical strain, increasing the time and expense required for human users to access websites. Second, large language models trained on scraped content now answer user queries directly, eliminating the need to visit original sources and severing the traffic that sustained content creators, including news organizations.
The numbers reflect this shift. Globally, monthly search engine traffic has declined by between 15% and 35% depending on estimates, with outlets like the Wall Street Journal experiencing drops exceeding 25%. This decline persists despite serious accuracy concerns—research from the Dutch Consumer Protection Association found that nearly a third of Google's AI summaries contained errors.
Haunted infrastructure: the internet's broken compact
For decades, web crawlers operated within a sustainable economic framework. Search engines deployed bots to index content, and in exchange, they directed traffic back to the sites they crawled. Publishers accepted indexing in hopes of landing on search results' first page, where visibility could drive visitors and advertising revenue. The arrangement was imperfect but reciprocal.
AI crawlers represent a fundamentally different proposition—extractive rather than symbiotic. They harvest content, news articles, blog posts, and open-source code while providing little in return. When search engines like Google begin offering AI-generated summaries instead of directing users to underlying sources, the traffic that once sustained publishers evaporates.
Simultaneously, infrastructure costs escalate dramatically. AI bot traffic strains websites far more severely than traditional search indexing crawlers, increasing hosting fees, bandwidth charges, and server demands. This burden falls hardest on those least equipped to bear it: activist websites, independent news outlets, and other organizations with limited technical resources. Studies show that half of users abandon sites requiring more than 2 seconds to load—a threshold AI bots directly undermine through increased load times.
How crawlers work (and why AI crawlers are worse)
Web crawlers are autonomous software that visit websites, download content, and follow links to other sites in an iterative process. The analogy is straightforward: imagine entering a library, photocopying every book on the shelf, then tracking down every book mentioned in those copies, photocopying those as well, and continuing indefinitely through chains of references.

The recursive nature of crawling—visiting multiple webpages in rapid succession—creates the greatest expense for websites. Internet infrastructure relies on caching systems that store frequently requested information locally for quick, cost-effective delivery. Most human requests are satisfied from these local caches. However, crawler traffic typically accesses less-visited pages, forcing requests to main data centers where all website information resides. This process is slower and more expensive, and it cascades: each request from a main database updates the local cache, displacing older information.
The library analogy extends further: imagine malicious actors flooding the system with simultaneous requests for thousands of obscure materials—ten thousand times normal demand. The retrieval system would overwhelm, slowing access for legitimate patrons. Worse, these actors are not reading individual books but extracting entire collections to feed massive commercial systems that will never direct anyone back to the library. This extractive burden degrades the internet's ability to function as infrastructure for freedom of expression and access to reliable information.
The human cost of bleeding the web

AI companies are making it harder for people to access free knowledge in order to train systems that may eventually compete with those very sources. The financial and technical strain they impose on the internet is substantial and unevenly distributed.
Technical interventions exist—robots.txt files allow websites to communicate crawling boundaries—but many AI bots disregard these standard mechanisms and access content they have been explicitly forbidden from taking. For large publishers, the resulting costs are unwelcome but manageable. For independent journalists, nonprofit media outlets, and civil society organizations, they can prove unsustainable.
When local news sites, nonprofit outlets, and community organizing platforms shut down or restrict access due to infrastructure costs, communities lose critical resources for democratic participation. Archives go offline. Knowledge becomes inaccessible. People lose their ability to access information and share knowledge—the very functions that make the internet essential.
What can we do now?
The question is no longer whether AI crawlers are disrupting the internet—they clearly are—but what stakeholders can do about it. Technical solutions exist, though their effectiveness depends on whether AI companies choose to respect them. Market-based interventions that monetize content access risk centralizing control in unaccountable tech companies.
The priority must be protecting people's ability to access and share information freely. With that in mind, concrete actions exist for different stakeholders:
For technologists and developers
- Participate in internet governance debates: The Internet Engineering Task Force (IETF) and other standards bodies are actively developing protocols for AI crawler and monetization management. Technical expertise grounded in human rights perspectives is needed.
- Discuss and develop open standards: Rather than allowing fast-moving commercial gatekeepers to control AI crawler content monetization, support the development of open, transparent standards through multi-stakeholder processes at bodies like the IETF and W3C.
For policymakers and advocates
- Government interventions: Competition regulators should investigate the mass extraction of content without adequate compensation by major AI players, which entrenches their market position. Governments should provide creative funding mechanisms, as demonstrated by the Digital Infrastructure Insights Fund and the Sovereign Tech Agency.
- Support public interest infrastructure: Contribute technical expertise and funding to organizations like Wikimedia, the Internet Archive, and nonprofit news organizations bearing disproportionate costs from AI crawling while developing alternative, non-centralizing architectures.
For content creators and website owners
- Assess your situation: Check server logs to identify AI bot traffic. While many AI crawlers ignore robots.txt, some respect it. Add specific AI crawler user agents to your disallow list, understanding this is a request rather than enforcement.
- Consider your technical options: Services like Cloudflare offer tools to block or manage AI crawlers, though consolidating control over crawling permissions in a single intermediary raises concerns about centralized power over information flows.
For everyone
- Demand accountability: Engage in political debates about AI, the internet, and related technologies at local and national levels. This is fundamentally a question of how the internet should function and whom it should serve.
- Spread awareness: Most people do not understand that AI crawlers are degrading their ability to access information online. Sharing this knowledge increases pressure for change.
Defeating the monsters: what comes next
The scariest story this Halloween is not fiction—it is unfolding right now. Every time someone tries to access information online and encounters slowdowns caused by AI bots serving corporate interests rather than human needs, the threat becomes real. This extraction operates in the shadows of the internet, visible only in server logs and bandwidth reports, in the gradual degradation of sites people depend on.
Unlike seasonal frights that fade with November, this extraction accelerates daily. The monsters are not fantastical but corporate AI crawlers treating collective knowledge infrastructure as free raw material. They will not be defeated by daylight, garlic, or silver bullets, but by collective action, technical standards, and policy frameworks that prioritize human access to information over AI production.
Source: Tech Policy Press



