Biodiversity Heritage Library - Program news and collection highlights from BHL
  • Home
  • News
  • Featured Books
    • All Featured Books
    • Book of the Month Series
    • BHL at 20
  • User Stories
  • Campaigns
    • Fossil Stories
    • Garden Stories
    • Monsters Are Real
    • Page Frights
    • Her Natural History
    • Earth Optimism 2020
  • Tech Blog
  • Visit BHL
Home
News
Featured Books
    All Featured Books
    Book of the Month Series
    BHL at 20
User Stories
Campaigns
    Fossil Stories
    Garden Stories
    Monsters Are Real
    Page Frights
    Her Natural History
    Earth Optimism 2020
Tech Blog
Visit BHL
  • Home
  • News
  • Featured Books
    • All Featured Books
    • Book of the Month Series
    • BHL at 20
  • User Stories
  • Campaigns
    • Fossil Stories
    • Garden Stories
    • Monsters Are Real
    • Page Frights
    • Her Natural History
    • Earth Optimism 2020
  • Tech Blog
  • Visit BHL
Biodiversity Heritage Library - Program news and collection highlights from BHL
Blog Reel, Featured Books

The Treasure Between the Covers: Making BHL’s Articles Discoverable and Citable

A Gif zooming into a page of many tiny images of scanned pages

This post is part of BHL at 20: Treasures from the Biodiversity Heritage Library, a series contributed by members of the BHL community, highlighting remarkable works from across the collection in celebration of its 20th anniversary.

One of the greatest challenges for a digital library, especially one as large as the Biodiversity Heritage Library, is simply finding the content you are after. Recently, I made a website to give a sense of this challenge. The website features about 200,000 pages of content from BHL Australia, a small fraction of what is in BHL overall, but already it’s something of a challenge to find specific items you might be after. If you were looking for a particular article, how would you find it?

A Gif zooming into a page of many tiny images of scanned pages

Interactive browser of BHL Australia

Articles, articles, articles

For most scientists, the article is the fundamental unit of research, not the journal title, and not a journal volume. The article is what we download as a PDF, what we store in our reference managers, and what gets cited. At the outset BHL did not have articles, so over a decade ago I set about developing a tool to find those. This tool became BioStor, which was described in a paper in 2011 (Extracting scientific articles from a large digital archive: BioStor and the Biodiversity Heritage Library). The basic idea behind BioStor is to take information about an article, such as journal, volume, pages, and year, and then try and find that article in BHL.

A diagram with text and large blue arrows showing the mapping between BHL and articles

Mapping journal, volume, and pagination from an article to BHL.

In principle this seems straightforward, but often the vagaries of metadata complicate the task. The image below shows some of the issues encountered with the journal Ibis. The source of metadata for the articles was CrossRef, via a commercial publisher (Wiley). You might expect this data to be high quality, but it contains errors such as bad character encoding. To further complicate things, Wiley decided to renumber all the volumes of the journal, so that the original volume information we see in BHL (such as series 2, volume 1) bears little relation to what is in CrossRef (volume 7, issue 1).

Four examples of citations, with some text highlighted in orange. The Crossref and BHL logos are on the right.

Matching CrossRef metadata for an article in Ibis to BHL.

The list of metadata messes like this is almost endless. There are journals that have more than one numbering system for the same volumes (e.g., Annali del Museo civico di storia naturale Giacomo Doria where the same item is both series 3, volume 7 and volume 47), there are multiple abbreviations for the same journal, and there are journals with multiple names (e.g., title 8097 is “Annuaire du Musée zoologique de l’Académie des sciences de St. Pétersbourg”, “ЕЖЕГОДНИКЬ ЗООЛОГИЧЕСКАГО МУЗЕЯ ИМПЕРАТОРСКОЙ АКАДЕМІЙ НАУКЬ”, and “Ezhegodnik Zoologicheskago muzeia …”).

A further complication is that a scanned item in BHL may contain several issues or volumes, each with its own set of overlapping page numbers, which means we have to decide which page “1” is the page 1 that matches the article we are searching for. Once we solve all that, we encounter further problems. Perhaps the most challenging is pagination. In most modern articles, the page range, e.g. 1–5, completely encompasses the article, including figures, charts, illustrations, etc. But for the older literature this is often not the case. Typesetting text and reproducing plates were different processes, and hence the plates might be disconnected from the article (often appearing at the end of a volume). This means that extracting, say pages 1–5, from BHL is no guarantee that you have the whole article.

A good deal of code in BioStor is trying to make sense of matching article metadata to BHL items, finding the correct page to match to, and extracting the set of pages that correspond to that article, as well as providing tools to manually correct metadata and add missing pages (for example, the plates mentioned above). Hence the process of finding articles is, at best, semiautomated.

BioStor old and new

The original BioStor website dates back to 2009, and looked something like this:

A screenshot from a website showing an image viewer with a yellow page from an book and fields of metadata.

Original BioStor website.

This site could display individual articles, and you could edit metadata. For a variety of reasons, it was no longer feasible to host this at the university where I was based, so I split the website into two versions. The original site now runs only on my laptop, and I use it to process files and locate articles in BHL. The new version runs in the cloud and features a cleaner interface, along with much better search. Below is the same article in the current BioStor.

A screenshot of a website showing a search bar, bibliographic data, and page images.

Current BioStor website.

Once articles are discovered using the old BioStor, they get pushed to the public version of BioStor at https://biostor.org. This website is also the point of contact between BioStor and BHL: each day BHL runs an automated process which asks BioStor whether it has any new articles, and, if the answer is yes, it fetches those and adds them to BHL (in BHL articles are referred to as either “parts” or “segments”). The end result is that articles defined in BioStor now become visible in the Table of Contents in BHL.

A screenshot of the BHL website showing an image viewer with a yellow page of a journal, bibliographic metadata and a highlighted article in a contents page.

BioStor article displayed in BHL.

One advantage of having a separate project such as BioStor is that I can use it to experiment with different ways to view BHL content. For example, BioStor looks for geographic coordinates (latitude and longitude) in the OCR text for each article. Any pairs of coordinates that it finds get stored in a map, which you can browse. In the diagram below we have selected a small region in the centre of the map, on the right you can see a list of articles about that area.

A map of the island of Sulawesi with small red dots sprinkled across it. There is a pink rectangle over a cluster of red dots.

Maps showing localities on the island of Sulawesi that are mentioned in BioStor articles.

Identifiers

BioStor has been running since 2009. In that time it has contributed over 260,000 articles to BHL, making it the single largest source of BHL “parts”. Having articles is nice, but even better is having articles with persistent, citable identifiers, such as DOIs. The Persistent Identifier Working Group has been working to add DOIs to BHL content, especially “parts”. This work has focused on two kinds of DOIs. The first are existing DOIs minted, for example, by commercial publishers. BioStor adds a lot of articles using CrossRef metadata, so we get these “for free” (there are other sources of DOIs that BioStor uses, but that is another story). Why does it matter to have external DOIs for BHL content? Well, many of these articles are free in BHL but behind a paywall on the publisher’s website. Services such as Unpaywall can link existing DOIs to free versions of the corresponding article, and BHL is one of Unpaywall’s providers.

But the more exciting (and onerous) task is minting new DOIs for articles in BHL, so that BHL is the version of record for that content. This has several implications. It means that BHL is effectively a publisher, and has the responsibility to maintain access to this content in perpetuity. It also changes the way we think about adding articles. For example, most of my work with BioStor has been opportunistic – I’m working on a taxonomic database, I see that there are some papers that should be in BHL, find them, then add them to BioStor so future BHL users can find those articles. But once we start creating DOIs, the goal is quite different: you want to get every article in the journal that is in BHL, and mint DOIs for all of them. While this appeals to a completionist mindset, it does mean getting metadata for every article before you can add DOIs.

Luckily, the hard work in minting DOIs has a striking payback, we can see how many times articles in BHL are cited in the scientific literature. The last time the results were analyzed, BHL articles had been cited some 74,446 times! Without BHL these publications would appear as simple text strings in the literature cited, now they are first-class digital citizens with clickable DOI links.

Metadata matters

By now it is obvious that the way BioStor finds articles depends on having good quality metadata for articles (or chapters), which it then attempts to locate in BHL. The lack of freely accessible metadata is a major impediment to increasing the rate at which articles are added. In the past I have made extensive use of taxonomic databases as a source of bibliographic data (see my BioNames project, for example). Yet the quality of citations in these databases is often poor. I have also made extensive use of sources such as CrossRef, which covers articles that have been assigned DOIs by that agency, and also data provided by volunteers, such as those working with Nicole Kearney (thank you Bob Griffith and Heidi Griffith!). Another major source of data has come from scraping the web, a time-consuming process that is becoming increasingly difficult to do as the web becomes increasingly closed under the onslaught of AI bots (see also Joel Richard’s blog post A Brief Bit on BHL Battling a Barrage of Bots).

There is a clear need for a free and open bibliographic database. The nearest we have is OpenAlex, whose tagline is “All the world’s research, connected and open.” Sadly this is still more of an aspirational goal rather than a fact: a lot of taxonomic literature is not in OpenAlex. Perhaps it is time, therefore, to revive “CiteBank”, which was an early BHL project to collect bibliographic metadata. If we had a comprehensive database of the taxonomic and related literature, locating articles in BHL would be a much easier task.

Machines reading

BioStor’s method of finding articles works, but it is not the only way we could locate articles. Instead of relying on external sources of metadata, what if we could simply have a computer read the volume and extract the articles automatically? Early attempts to do this for BHL content were not particularly successful, see for example A metadata generation system for scanned scientific volumes. But the advent of large language models (LLMs) and AI chatbots has dramatically changed the way we can tackle finding articles in BHL. In my own work I routinely use AI to extract articles in bulk from a scanned volume. Typically the approach involves finding tables of contents in the scanned volume, using AI to parse that into structured data, then finding the corresponding pages in the volume, checking that they match the table of contents, and then using AI to extract bibliographic data (e.g., article title, authors, etc.). The result of this process is a data file that gets fed into BioStor, so that articles get found and added to BHL in the usual way. It is not bulletproof, and AI can quite happily make mistakes, but in my experience it works well.

But the holy grail would be to simply point an AI at a volume and it would identify and extract all the articles, find any stray plates, and present the results to BHL. Given the spectacular advances in OCR text and understanding document layout in recent years, perhaps there will be a point where BioStor can gracefully retire from the scene. Its hundreds upon hundreds of lines of regular expressions and special-case hacks quietly gathering dust in a GitHub repo while machines of loving grace read BHL for us.

From the Biodiversity Heritage Library

As BHL celebrates twenty years of open biodiversity knowledge, this post reminds us that access depends not only on digitised pages, but on the tools, metadata, identifiers, and infrastructure that make them discoverable and citable. With your support, BHL can continue strengthening the systems that connect biodiversity literature to the researchers, communities, and future discoveries that depend on it.Orange button with a heart icon

June 16, 2026by nkearney
BHL News, Blog Reel, Tech Updates

BHL Servers are Moving!

An old map of the northeast United States with an arching arrow from Washington DC and Chicago

The big day is here! BHL’s Tech Team is starting the three-week process of moving BHL’s technical infrastructure from the Smithsonian data center just outside of Washington, DC, to the Field Museum in Chicago, Illinois.

First, though, we are deeply grateful to the Field Museum for welcoming BHL’s technology and helping ensure that biodiversity knowledge remains available to the global community. The support from their administration and technical team is invaluable.

We’ve been testing and configuring the network at the Field Museum and BHL is already quietly running on their network, so we are confident the process will go smoothly when we ship and install the production database and full-text-search servers.

“How will this affect BHL?”, you say. Great question! As much as possible, the BHL website will remain online and we do not expect any time where BHL is unavailable. The shipments to Chicago will take place in two parts to ensure BHL is always available, with some limitations. The first shipment on 11 June 2026 will have no impact on the regular functions of the site.

The second shipment on 25 June 2026 will cause the usual full-text search function to be temporarily unavailable while the full-text-search server is in transit. This means users will still be able to search titles, authors, subjects, and other catalog metadata, but users will not be able to search the full text of books and articles. This will last for 5-6 days over a weekend.

The 11th of June is a warm up, but the big day is the 25th of June when we officially begin serving BHL from the Field Museum.

We’ve set up an informational page with more details and will update it as things progress. We’ve also provided a map to illustrate where BHL’s servers are during transit.

An old map of the northeast united states with a winding dashed line between DC and Chicago.

A map of the route the servers will take to get to Chicago.
Please note that the path is neither accurate nor to scale. (Map Source)

But will the BHL servers feel at home in Chicago? To help them acclimate, we looked to historical maps of rainfall, soils, and vegetation to compare BHL’s former and future environments.


Vintage map showing rainfall in the northeastern part of the United States.

From: Atlas of American agriculture, no.5 (Source)

In Chicago, BHL will be receiving slightly less rainfall (30-35 ” or 75-88 cm per year) than on the U.S. East Coast (35-40″ or 85-100 cm). This, however, is mitigated by its proximity to Lake Michigan, which is one of the largest freshwater lakes in the world.

This looks like sufficient water for BHL, but these are historical values and may not reflect the effect of global climate change. In fact, it’s reported that the north-central part of the United States is in a moderate drought.


Vintage map showing soil types in the northeastern part of the United States.

From: Atlas of American agriculture, no.8 (Source)

What about soil type? BHL is leaving an area of Gray-Brown Podzolic soils to the much less prodigious Soils of the Northern Prairie. Looking closely at this map, we can see that BHL’s familiar soils are in close proximity to its new soil environment in Chicago, so we hope that this will help BHL rapidly make the transition to its new home.


Vintage map showing types of vegetation in the northeastern part of the United States.

From: Atlas of American agriculture, no.6 (Source)

Lastly, we look at vegetation. Historically the eastern seaboard of the United States is abundant in Oak-Pine and Chestnut-Oak-Yellow Poplar forests. The move to Chicago brings a potentially dramatic change as BHL moves from forests to the Tall Grasses of the Prairie Grasslands. Much like the new soil type, we see that BHL is still in close proximity to Oak-Hickory forests and we hope that the presence of the familiar Oaks will also help make the transition.


With familiar oaks nearby, sufficient water, and a welcoming new home at the Field Museum, we think BHL is well positioned to put down roots in Chicago.

 

June 9, 2026by Joel Richard
Blog Reel, Featured Books

Collated and Perfect: Pliny the Elder’s Natural History at the Natural History Museum in London

A page from an historic book in Latin with elaborately decorated margins and a large decorated "L".

This post is part of BHL at 20: Treasures from the Biodiversity Heritage Library, a series contributed by members of the BHL community, highlighting remarkable works from across the collection in celebration of its 20th anniversary.

One afternoon in AD 79 a family had gathered at a villa overlooking the Bay of Naples. Outside, a mother and her seventeen-year-old son observed a strange cloud which had appeared in the distance. The mother, Plinia, saw it first, but it would be her son, Pliny the Younger, who recorded for posterity the sight on the horizon and the terrible events that had long since started but were now nearing their catastrophic conclusion. The teenager rushed inside where his uncle, who was also Admiral of the Roman Fleet in the Bay, Pliny the Elder, was reading. The younger Pliny described a cloud in the shape of a pine tree, its branches supported by a thick trunk. From the perspective of a post-twentieth century world, we might now call it a mushroom cloud. But to Pliny the Elder, ever inquisitive, it was like nothing ever seen before and he hastened to investigate.

Anyone with a passing knowledge of Roman history will know that what that confusingly named family had observed was the eruption of Mount Vesuvius. Very soon, Pompeii and Herculaneum would be lost, covered in a pyroclastic flow. What had started as the attractive prospect of observing a volcanic eruption first hand quickly turned to the realisation that a rescue mission was required. Events for the elder Pliny would end at Stabiae. Already suffering from troubled breathing, he was overcome by fumes, or perhaps suffered a heart attack, and died.

Much of what we know of Pliny the Elder comes to us from two letters written by his nephew to the historian Tacitus. Pliny the Elder’s writings were extensive, spanning the wars in Germania, where he served in the cavalry, and a treatise on throwing a lance from horseback to a Latin grammar. All are now lost. Yet one great monument to his learning survives, the Historia Naturalis, here called the Natural History. Divided into 37 books, and likely published posthumously, the Natural History covers topics from astronomy to horticulture.

Page from a printed book in Latin, arranged in two columns with red and blue initials, handwritten page numbers, and marginal annotations.

Book one’s summarium allowed readers to navigate the Natural History. In the NHM’s copy, annotation identifies where in the main work information can be found. Contributed in BHL from the Natural History Museum, London.

The Natural History is sometimes referred to as the first encyclopaedia, yet some recent scholarship has viewed this classification as overly simplistic. There was no ancient tradition of encyclopaedias to which Pliny himself could identify and whilst the author celebrated his accumulation of 20,000 distinct pieces of information, he also intended his work to stand as a coherent whole. By bringing together as much known information as possible, and not just about natural history, Pliny’s work can be seen as an Imperial project. He is laying claim to knowledge for Rome.

Opening page of a printed book, with a large decorated initial, a colorful interlaced border, printed Latin text, and a small painted roundel at the bottom of the page.

The decorated page at the beginning of Book One immediately identifies the work as a luxury product for an elite clientele. Contributed in BHL from the Natural History Museum, London.

From the time of its publication in manuscript form, Pliny’s Natural History was extraordinarily influential. It circulated widely and prior to its first printing some 200 manuscript copies were known to survive. Thus, when the Venetian Senate granted Johann de Speyer a commercial printing licence in 1469, just a decade and a half after Gutenberg used moveable type to publish his Bible, the Natural History was one of the first titles he printed. The copy in the Library of the Natural History Museum in London is one of those 1469 first editions.

Page from a printed book in Latin, showing a block of text with wide margins and handwritten annotations in brown ink.

Annotations to the margins, here to Book 33, demonstrate use of the volume. Contributed in BHL from the Natural History Museum, London.

Printing in the fifteenth century was, as now, a commercial endeavour and Johann de Speyer knowingly produced a product that would appeal to a wealthy clientele. His edition featured wide margins and decorated letters which echoed manuscript production and spoke of luxury and prestige. These margins also allowed the reader to make annotations, often as a means of navigating the text. In the NHM’s copy, those signs of use are present and at least two hands can be identified in the marginal annotations.

BHL contains a number of incunabula of Pliny (incunabula is the term used to indicate books printed before 1501). There is the 1469 edition from Venice contributed by the NHM and the 1472 edition, also from Venice, held by Boston Public Library. McGill University Library has contributed the 1473 (Rome), 1481 (Parma) and 1497 (Venice) editions. The NHM also holds the 1472 Venice and 1481 Parma editions as well as later copies such as the first English translation produced by Philemon Holland in 1601. Yet it is the de Speyer’s 1469 edition that is the oldest book in our library.

Blank page from an early printed book with several handwritten ownership and bibliographic notes.

The flyleaf gives provenance and condition information about the volume. Contributed in BHL from the Natural History Museum, London.

At the recent BHL Annual Meeting in London, delegates were given the opportunity to view the NHM’s 1469 Pliny. It has been suggested that only 100 copies of this edition were produced, and it is a wonderful survival of arguably the first printed book on natural history. As noted on the flyleaf, the volume was ‘collated’ and identified as ‘perfect, which is almost a miracle’. In actual fact, two facsimile folios were added at a later date, though this only makes for a more interesting object.

So much has been written about Pliny the Elder and his Natural History. My text lays no claim to originality. Indeed, even the NHM’s copy is relatively well known, having once been celebrated as the oldest item on BHL. As an earlier blog announced that baton has now been passed to the Circa Instans at the LuEsther T. Mertz Library of the New York Botanical Garden, evidence, were it needed, that BHL continues to thrive, with new material added all the time. It’s an endeavour we can be proud of, and one of which Pliny the Elder would certainly have approved.

From the Biodiversity Heritage Library

For 20 years, BHL has made treasures like this freely accessible to the world. With your support, we can ensure these treasures – and the knowledge they hold – remain open for generations to come.

Orange button with a heart icon

Further reading

The literature on Pliny the Elder and his Natural History is extensive, but the following were useful when preparing this blog.

Chibnall, M., ‘Pliny’s Natural History and the Middle Ages’ in T.A. Dorey (ed.), Empire and Aftermath, Silver Latin II (London, 1975).

Di Tommaso, L., ‘Pliny the Elder: Historia Naturalis’ in J. Magee (ed.), Rare Treasures from the Library of the Natural History Museum (London, 2017), pp. 6-11.

Doody, A., Pliny’s Encyclopaedia: The reception of the Natural History (Cambridge, 2010).

Dunn, D., In the Shadow of Vesuvius: A life of Pliny (London, 2019).

Gudger, E.W., ‘Pliny’s Historia Naturalis: The most popular natural history ever published’, Isis, vol. 6, no. 3 (1924), pp. 269-81.

June 3, 2026by nkearney

Help Support BHL

BHL's existence depends on the financial support of its patrons. Help us keep this free resource alive!

search

About BHL

The Biodiversity Heritage Library (BHL) is the world’s largest open access digital library for biodiversity literature and archives. BHL operates as a worldwide consortium of natural history, botanical, research, and national libraries working together to digitize the natural history literature held in their collections and make it freely available for open access as part of a global “biodiversity community.”

Join Our Mailing List

Sign up to receive the latest news, content highlights, and promotions.

Subscribe Now

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 319 other subscribers

Subscribe to Blog Via RSS

Subscribe to the blog RSS feed to stay up-to-date on all the latest BHL posts.

Access RSS Feed

Inspiring Discovery through Free Access to Biodiversity Knowledge.

The Biodiversity Heritage Library makes it easier than ever for you to access the information you need to study and explore life on Earth…for free, anytime, anywhere.

 

64+ Million Pages of
Biodiversity Literature Online.

EXPLORE

Tools and Services
to Transform Research.

EXPLORE

300,000+
Illustrations on Flickr.

EXPLORE

ABOUT | HARMFUL CONTENT | PRIVACY | SITE MAP | TERMS OF USE