A Brief Bit on BHL Battling a Barrage of Bots

Summary

Over the past year, BHL has been forced to contend with the issue of AI Bot (short for robot) Crawlers and the effect of their unreasonable requests for BHL’s content. I say unreasonable because while BHL’s content is made freely available and is probably beneficial to modern Large Language Models (LLMs), the behavior of these bots and their disruptive effect on BHL has proven that they are bad partners in the exchange of information.

BHL is not the only organization to contend with this in recent months. Other cultural heritage organizations have reported the same or similar experiences, most notably this detailed report from Michael Weinberg at the GLAM-E Lab.

By the time this and other articles had been posted in mid-2025, BHL had already crested the first wave of interruptions.

A bit of background

When a request is made to a website, all actors are expected to identify themselves. Your browser informs the website that it is Firefox or Chrome. Many of the bots that come to a site are there to fill the index for a search engine. They have clever names like “GoogleBot” or “BingBot”. Furthermore, many LLM bots identify themselves, too. “ClaudeBot” or “GPTBot”, for example.

On the other side of the equation we have the “robots.txt” file which is meant to create an informal agreement between website and Bot to set rules on how the bots should behave. It contains directives like “you can go to these these URLs”, “please don’t go to these other URLs” or “please don’t make more than 10 requests per second”. These directives can be set individually for different bots to help guide them to what we think is the most useful content for them.

Some LLM bots ignore all of this.

The disruptive effect

In fact, the Grok bot sneaks its way into web servers by lying about who it is. No one has seen a “GrokBot” visiting their sites. This behavior is recently well documented. This is bad behavior, and it’s not something a handful of humans can fight alone.

We found that this is what was happening at BHL: The traffic that we saw did not identify themselves and they also came from many thousands of locations around the internet, making it difficult to control or block them. The numbers show it:

A line graph showing overall traffic to BHL with a step in the average traffic in the middle of the chart.

This chart shows traffic to BHL’s “/search” URL for approximately the past year. The peaks in June 2025 and February 2026 exceeded BHL’s capacity.

Beginning in April 2025, we saw a gradual increase in traffic that peaked on 19 June 2025 when it dropped off completely. This was the first wave of disruptions which we saw again later in 2025 and into 2026.

During this time, the Smithsonian was the host for the BHL servers and making changes in their ecosystem was a challenge, but the firewalls (and the team) at the Smithsonian are amazing and played a large part in keeping BHL operational. Regardless, we believe that enough traffic was allowed to get through the firewalls to still cause disruptions and unexpected errors trying to get to BHL.

Bar chart showing tall bars for many days over the past month.

The Smithsonian’s firewalls were an important part of blocking traffic to BHL in February 2026. This chart shows requests per hour that were blocked by the firewalls for one week in February.

Our Response

Starting in June 2025 and continuing through September, Mike Lichtenberg, BHL Lead Software Engineer, made some small but effective changes to how BHL works internally. This allowed BHL to serve a new, higher level of traffic on the same technical infrastructure. Another change we made was to redirect certain requests to the Internet Archive, which was already protected by Cloudflare. This freed up capacity at BHL to serve more requests.

A line graph showing overall traffic to BHL with a step in the average traffic in the middle of the chart.

This chart shows per-day traffic to BHL. On the left side before August 2025, our regular traffic was around 2.5 million hits per day. On the right side, after Mike’s improvements the new baseline is closer to 5 million hits per day.

Unfortunately this was not enough. In January 2026, the bots returned, and worse than before.

The Cloudflare Cavalry

After a few weeks, it became clear that our efforts were no longer sufficient. Around 19 June 2025, BHL had responded to 870,000 search requests. This far exceeded our usual 22,000 searches per day.

On 10 February 2026, BHL received a peak of 943,000 search requests.

We’d been encouraged to pursue Cloudflare for some time, but it wasn’t until mid-February that our technical infrastructure was ready for the switch. On 20 February we officially started using Cloudflare and after a few missteps, everything was in place by 24 February.

But what is Cloudflare, you ask? It’s a service that sits between a web site and the rest of the internet and it looks at requests to that website to decide if they should be allowed through. As a very large player on the internet, Cloudflare has access to mountains of data on bot-like behavior to accurately analyze traffic in real-time.

After activating Cloudflare, we immediately saw a reduction in traffic and much more stability at BHL. In the following weeks, we made other changes, such as adding the ubiquitous “are you human?” checkbox that has become common across the internet.

A line chart showing a 60% drop in traffic around February 23 2026

This chart shows traffic to BHL’s servers dropped dramatically when Cloudflare was fully activated.

Cloudflare is not simply a gatekeeper preventing unwanted traffic from coming to BHL. They also cache regularly requested parts of BHL (images, scripts, and styles) to deliver them to you from a location that is geographically closer to you. This means BHL will often seem to load faster than it did before. This also accounts for the drop in traffic to BHL as indicated in the chart above.

A line chart with three lines indicating that 2/3 of traffic to BHL is blocked by Cloudflare.

Cloudflare is now blocking a large amount of automated traffic coming to BHL. 

This chart shows a bit more detail of how Cloudflare is affecting requests to BHL. The top line is requests blocked by Cloudflare, the middle line is requests served by BHL’s own infrastructure, and the bottom line is requests served by  Cloudflare’s caches.

In Closing

It’s been a long journey, but BHL is much safer from the army of bots. This doesn’t mean that our work is done. We’ve had to make special efforts to not block access to BHL’s APIs and data feeds, which are often used by software or code…and which often behave like a bot! These settings have been in place for a few weeks now.

Additionally, we are investigating whether to allow the LLM bots to resume accessing our high-quality content, but we must be able to limit their flow of traffic to prevent overwhelming BHL’s servers.

Rest assured that BHL’s Technical Team is committed to keeping BHL up and running as we always have, but now we have a powerful new tool to help us in those efforts.


Discover more from Biodiversity Heritage Library

Subscribe to get the latest posts sent to your email.

Avatar for Joel Richard
Written by

Joel is the Technical Coordinator for BHL. When he's not serving as the Head of Web and IT for the Smithsonian Libraries and Archives, he's also working on BHL's Macaw software.