Biodiversity Heritage Library - Program news and collection highlights from BHL
  • Home
  • News
  • Featured Books
    • All Featured Books
    • Book of the Month Series
    • BHL at 20
  • User Stories
  • Campaigns
    • Fossil Stories
    • Garden Stories
    • Monsters Are Real
    • Page Frights
    • Her Natural History
    • Earth Optimism 2020
  • Tech Blog
  • Visit BHL
Home
News
Featured Books
    All Featured Books
    Book of the Month Series
    BHL at 20
User Stories
Campaigns
    Fossil Stories
    Garden Stories
    Monsters Are Real
    Page Frights
    Her Natural History
    Earth Optimism 2020
Tech Blog
Visit BHL
  • Home
  • News
  • Featured Books
    • All Featured Books
    • Book of the Month Series
    • BHL at 20
  • User Stories
  • Campaigns
    • Fossil Stories
    • Garden Stories
    • Monsters Are Real
    • Page Frights
    • Her Natural History
    • Earth Optimism 2020
  • Tech Blog
  • Visit BHL
Biodiversity Heritage Library - Program news and collection highlights from BHL
Blog Reel, Featured Books

The Treasure Between the Covers: Making BHL’s Articles Discoverable and Citable

A Gif zooming into a page of many tiny images of scanned pages

This post is part of BHL at 20: Treasures from the Biodiversity Heritage Library, a series contributed by members of the BHL community, highlighting remarkable works from across the collection in celebration of its 20th anniversary.

One of the greatest challenges for a digital library, especially one as large as the Biodiversity Heritage Library, is simply finding the content you are after. Recently, I made a website to give a sense of this challenge. The website features about 200,000 pages of content from BHL Australia, a small fraction of what is in BHL overall, but already it’s something of a challenge to find specific items you might be after. If you were looking for a particular article, how would you find it?

A Gif zooming into a page of many tiny images of scanned pages

Interactive browser of BHL Australia

Articles, articles, articles

For most scientists, the article is the fundamental unit of research, not the journal title, and not a journal volume. The article is what we download as a PDF, what we store in our reference managers, and what gets cited. At the outset BHL did not have articles, so over a decade ago I set about developing a tool to find those. This tool became BioStor, which was described in a paper in 2011 (Extracting scientific articles from a large digital archive: BioStor and the Biodiversity Heritage Library). The basic idea behind BioStor is to take information about an article, such as journal, volume, pages, and year, and then try and find that article in BHL.

A diagram with text and large blue arrows showing the mapping between BHL and articles

Mapping journal, volume, and pagination from an article to BHL.

In principle this seems straightforward, but often the vagaries of metadata complicate the task. The image below shows some of the issues encountered with the journal Ibis. The source of metadata for the articles was CrossRef, via a commercial publisher (Wiley). You might expect this data to be high quality, but it contains errors such as bad character encoding. To further complicate things, Wiley decided to renumber all the volumes of the journal, so that the original volume information we see in BHL (such as series 2, volume 1) bears little relation to what is in CrossRef (volume 7, issue 1).

Four examples of citations, with some text highlighted in orange. The Crossref and BHL logos are on the right.

Matching CrossRef metadata for an article in Ibis to BHL.

The list of metadata messes like this is almost endless. There are journals that have more than one numbering system for the same volumes (e.g., Annali del Museo civico di storia naturale Giacomo Doria where the same item is both series 3, volume 7 and volume 47), there are multiple abbreviations for the same journal, and there are journals with multiple names (e.g., title 8097 is “Annuaire du Musée zoologique de l’Académie des sciences de St. Pétersbourg”, “ЕЖЕГОДНИКЬ ЗООЛОГИЧЕСКАГО МУЗЕЯ ИМПЕРАТОРСКОЙ АКАДЕМІЙ НАУКЬ”, and “Ezhegodnik Zoologicheskago muzeia …”).

A further complication is that a scanned item in BHL may contain several issues or volumes, each with its own set of overlapping page numbers, which means we have to decide which page “1” is the page 1 that matches the article we are searching for. Once we solve all that, we encounter further problems. Perhaps the most challenging is pagination. In most modern articles, the page range, e.g. 1–5, completely encompasses the article, including figures, charts, illustrations, etc. But for the older literature this is often not the case. Typesetting text and reproducing plates were different processes, and hence the plates might be disconnected from the article (often appearing at the end of a volume). This means that extracting, say pages 1–5, from BHL is no guarantee that you have the whole article.

A good deal of code in BioStor is trying to make sense of matching article metadata to BHL items, finding the correct page to match to, and extracting the set of pages that correspond to that article, as well as providing tools to manually correct metadata and add missing pages (for example, the plates mentioned above). Hence the process of finding articles is, at best, semiautomated.

BioStor old and new

The original BioStor website dates back to 2009, and looked something like this:

A screenshot from a website showing an image viewer with a yellow page from an book and fields of metadata.

Original BioStor website.

This site could display individual articles, and you could edit metadata. For a variety of reasons, it was no longer feasible to host this at the university where I was based, so I split the website into two versions. The original site now runs only on my laptop, and I use it to process files and locate articles in BHL. The new version runs in the cloud and features a cleaner interface, along with much better search. Below is the same article in the current BioStor.

A screenshot of a website showing a search bar, bibliographic data, and page images.

Current BioStor website.

Once articles are discovered using the old BioStor, they get pushed to the public version of BioStor at https://biostor.org. This website is also the point of contact between BioStor and BHL: each day BHL runs an automated process which asks BioStor whether it has any new articles, and, if the answer is yes, it fetches those and adds them to BHL (in BHL articles are referred to as either “parts” or “segments”). The end result is that articles defined in BioStor now become visible in the Table of Contents in BHL.

A screenshot of the BHL website showing an image viewer with a yellow page of a journal, bibliographic metadata and a highlighted article in a contents page.

BioStor article displayed in BHL.

One advantage of having a separate project such as BioStor is that I can use it to experiment with different ways to view BHL content. For example, BioStor looks for geographic coordinates (latitude and longitude) in the OCR text for each article. Any pairs of coordinates that it finds get stored in a map, which you can browse. In the diagram below we have selected a small region in the centre of the map, on the right you can see a list of articles about that area.

A map of the island of Sulawesi with small red dots sprinkled across it. There is a pink rectangle over a cluster of red dots.

Maps showing localities on the island of Sulawesi that are mentioned in BioStor articles.

Identifiers

BioStor has been running since 2009. In that time it has contributed over 260,000 articles to BHL, making it the single largest source of BHL “parts”. Having articles is nice, but even better is having articles with persistent, citable identifiers, such as DOIs. The Persistent Identifier Working Group has been working to add DOIs to BHL content, especially “parts”. This work has focused on two kinds of DOIs. The first are existing DOIs minted, for example, by commercial publishers. BioStor adds a lot of articles using CrossRef metadata, so we get these “for free” (there are other sources of DOIs that BioStor uses, but that is another story). Why does it matter to have external DOIs for BHL content? Well, many of these articles are free in BHL but behind a paywall on the publisher’s website. Services such as Unpaywall can link existing DOIs to free versions of the corresponding article, and BHL is one of Unpaywall’s providers.

But the more exciting (and onerous) task is minting new DOIs for articles in BHL, so that BHL is the version of record for that content. This has several implications. It means that BHL is effectively a publisher, and has the responsibility to maintain access to this content in perpetuity. It also changes the way we think about adding articles. For example, most of my work with BioStor has been opportunistic – I’m working on a taxonomic database, I see that there are some papers that should be in BHL, find them, then add them to BioStor so future BHL users can find those articles. But once we start creating DOIs, the goal is quite different: you want to get every article in the journal that is in BHL, and mint DOIs for all of them. While this appeals to a completionist mindset, it does mean getting metadata for every article before you can add DOIs.

Luckily, the hard work in minting DOIs has a striking payback, we can see how many times articles in BHL are cited in the scientific literature. The last time the results were analyzed, BHL articles had been cited some 74,446 times! Without BHL these publications would appear as simple text strings in the literature cited, now they are first-class digital citizens with clickable DOI links.

Metadata matters

By now it is obvious that the way BioStor finds articles depends on having good quality metadata for articles (or chapters), which it then attempts to locate in BHL. The lack of freely accessible metadata is a major impediment to increasing the rate at which articles are added. In the past I have made extensive use of taxonomic databases as a source of bibliographic data (see my BioNames project, for example). Yet the quality of citations in these databases is often poor. I have also made extensive use of sources such as CrossRef, which covers articles that have been assigned DOIs by that agency, and also data provided by volunteers, such as those working with Nicole Kearney (thank you Bob Griffith and Heidi Griffith!). Another major source of data has come from scraping the web, a time-consuming process that is becoming increasingly difficult to do as the web becomes increasingly closed under the onslaught of AI bots (see also Joel Richard’s blog post A Brief Bit on BHL Battling a Barrage of Bots).

There is a clear need for a free and open bibliographic database. The nearest we have is OpenAlex, whose tagline is “All the world’s research, connected and open.” Sadly this is still more of an aspirational goal rather than a fact: a lot of taxonomic literature is not in OpenAlex. Perhaps it is time, therefore, to revive “CiteBank”, which was an early BHL project to collect bibliographic metadata. If we had a comprehensive database of the taxonomic and related literature, locating articles in BHL would be a much easier task.

Machines reading

BioStor’s method of finding articles works, but it is not the only way we could locate articles. Instead of relying on external sources of metadata, what if we could simply have a computer read the volume and extract the articles automatically? Early attempts to do this for BHL content were not particularly successful, see for example A metadata generation system for scanned scientific volumes. But the advent of large language models (LLMs) and AI chatbots has dramatically changed the way we can tackle finding articles in BHL. In my own work I routinely use AI to extract articles in bulk from a scanned volume. Typically the approach involves finding tables of contents in the scanned volume, using AI to parse that into structured data, then finding the corresponding pages in the volume, checking that they match the table of contents, and then using AI to extract bibliographic data (e.g., article title, authors, etc.). The result of this process is a data file that gets fed into BioStor, so that articles get found and added to BHL in the usual way. It is not bulletproof, and AI can quite happily make mistakes, but in my experience it works well.

But the holy grail would be to simply point an AI at a volume and it would identify and extract all the articles, find any stray plates, and present the results to BHL. Given the spectacular advances in OCR text and understanding document layout in recent years, perhaps there will be a point where BioStor can gracefully retire from the scene. Its hundreds upon hundreds of lines of regular expressions and special-case hacks quietly gathering dust in a GitHub repo while machines of loving grace read BHL for us.

From the Biodiversity Heritage Library

As BHL celebrates twenty years of open biodiversity knowledge, this post reminds us that access depends not only on digitised pages, but on the tools, metadata, identifiers, and infrastructure that make them discoverable and citable. With your support, BHL can continue strengthening the systems that connect biodiversity literature to the researchers, communities, and future discoveries that depend on it.Orange button with a heart icon

June 16, 2026by nkearney
BHL News, Blog Reel, Tech Updates

What Is BHL’s New Persistent Identifier Working Group DOI’ng?

Graphic showing the members of BHL's Persistent Identifier Working Group

In October 2020, BHL launched a new working group with a momentous goal: to make the content on BHL persistently discoverable, citable and trackable using DOIs (Digital Object Identifiers).

Graphic showing the members of BHL's Persistent Identifier Working Group

The members of BHL’s new Persistent Identifier Working Group (PIWG).

A DOI is like an electronic fingerprint in the form of a unique and permanent alphanumeric string that provides a persistent link to a piece of content online. Modern publications receive a DOI at the point of publication. This DOI becomes a key part of a publication’s bibliographic metadata that should be included in any mention or citation of that publication. Reference lists in modern publications are filled with DOIs, which allows readers to click from publication to publication in (in theory) a never-ending chain of knowledge.

This reciprocal linking of DOIs has created a great linked network of scholarly research, but that network is missing the historic literature. The vast majority of historic publications lack DOIs. This means they appear in reference lists as unlinked citations. In our increasingly online world, readers are far more likely to read (and thus cite) publications they can click through to (particularly when libraries are inaccessible during a global pandemic). The upshot of this is that the millions of pages of historic literature on BHL—the foundation of our understanding of biodiversity—is in danger of falling into obscurity.

BHL has been retrospectively minting DOIs for historic publications since 2011, but the focus has primarily been on monographs. BHL’s new Persistent Identifier Working Group (PIWG) is (at least initially) focusing on journal articles. Minting DOIs for articles on BHL is a far more complex and time-consuming task than minting DOIs for monographs. This is because article DOIs need article data: every journal volume uploaded onto BHL must be accompanied by journal and volume data, but there is no requirement that contributors provide article data.

Thankfully, there have been considerable efforts to add article data to BHL (and thereby make it possible to search for the titles and authors of these articles both within BHL and via external engines). A huge proportion of this article data has been contributed to BHL by Roderic Page via BioStor: 75% of the 300,753 articles indexed in BHL as of 4 May 2021 were “defined” by BioStor. It is very difficult to determine how many articles are actually on BHL (hidden within all those journal volumes). But, while we don’t know what proportion of BHL’s journal content still needs to be made discoverable, we know there is still a huge amount of work to do.

COVID-19 provided an unexpected opportunity to make a considerable dent in this work. With no access to scanners or library materials, a number of BHL contributors, including Harvard University Libraries, Muséum National d’Histoire Naturelle and BHL Australia, pivoted from making new content accessible to making their existing content on BHL more discoverable. For example, BHL Australia’s digitisation volunteers gathered, gap filled and checked article-level metadata for over 30,000 articles in 2020.

Once an article has been defined, i.e. it exists as a publication unit in BHL and has its own article landing page (and we’ve checked that the article does not already have a DOI), we can assign a DOI to it. Articles that have recently been assigned BHL DOIs include some very old publications, such as the first scientific description of the Duck-billed Platypus, published in 1799 (https://doi.org/10.5962/p.304567), and the species descriptions from A specimen of the botany of New Holland, the first publication dedicated to Australian flora (1793-5), e.g. https://doi.org/10.5962/p.312432. The PIWG has also started assigning DOIs to in-copyright publications (with permission from the rights holders). These include articles from the Bulletin of the British Museum, e.g. https://doi.org/10.5962/p.310418, and the Bulletin of the African Bird Club, e.g. https://doi.org/10.5962/p.308885.

Screenshot of the landing page in BHL for the description of the The Duck-Billed Platypus, Platypus anatinus.

Shaw, George (1799), The Duck-Billed Platypus, Platypus anatinus, The Naturalist’s Miscellany: https://doi.org/10.5962/p.304567 (illustration by Frederick Polydore Nodder).

If an article on BHL has an existing non-BHL DOI, we add this key piece of bibliographic metadata to the BHL landing page for the article. This ensures that BHL users can link to the definitive version of the article (the one the DOI resolves to), and more importantly, that other parties (and their algorithms) can find our versions from elsewhere. This is particularly important when commercial websites lock their DOI’d versions of public domain articles behind paywalls. Having their DOIs on our freely accessible versions ensures that services like Unpaywall can find them. To learn more about how this works, see our blog post: BHL Journal Articles Are Now Discoverable via Unpaywall.

DOIs not only improve discoverability and enable persistent linking to our historic content; they also allow us to track how BHL content is being used. In the six months following the minting of its new DOI (Oct 2020 to April 2021), the 1799 Platypus description was tweeted by 219 Twitter accounts, referenced in six Wikipedia pages, picked up by one news outlet and cited in one academic paper (data from Altmetric, April 2021). We know this because the article has a DOI.

Screenshot of the Altmetric dashboard for the first scientific description of the Duck-billed Platypus

Altmetric’s overview of attention for the first scientific description of the Duck-billed Platypus (Shaw 1799): https://www.altmetric.com/details/91788579.

The PIWG has spent the past six months creating, refining and testing tools that will allow BHL contributors to do this work themselves. We have also been producing documentation that explains a) how to use the new tools, and b) why this work is so important. These tools will facilitate every step in the article discoverability and DOI assignment process including: downloading existing article data for a given journal title to allow for correction and gap-filling (in development); bulk uploading of article data for new articles (available now); and adding articles and titles to BHL’s (new) DOI Assignment Queue (available now). Our dream is that, whenever anyone uploads a journal volume to BHL, they also provide the data for the articles it contains (and thus take responsibility for making that content discoverable).

The Persistent Identifier Working Group (PIWG) is fueled by the technical expertise, metadata dexterity and incredible passion of:

  • Nicole Kearney, Manager BHL Australia (Chair)
  • Mike Lichtenberg, BHL Lead Developer
  • Susan Lynch, Systems, Digitization & Web Services Librarian, The New York Botanical Garden
  • Bess Missell, Metadata Librarian, Smithsonian Libraries and Archives
  • Roderic Page, Professor of Taxonomy, University of Glasgow
  • Joel Richard, BHL Technical Coordinator | Head of Web Services & IT, Smithsonian Libraries, Smithsonian Libraries and Archives
  • Diane Rielinger, Digital Projects Librarian, Botany Libraries, Harvard University Herbaria
  • Colleen Funkhouser, BHL Program Manager

The specific goals of the group are:

  • To add article-level metadata to journal articles on BHL
  • To add existing DOIs to (new and existing) article landing pages on BHL (particularly for those articles where the DOI’d version is behind a paywall elsewhere)
  • To assign BHL DOIs to articles that lack DOIs

Want to know more about BHL’s Persistent Identifier Working Group? See:

  • Discovering the Platypus: From its scientific description to its DOI, Biodiversity Information Science and Standards (TDWG) Conference, 6 October 2020: https://youtu.be/4UVSEoWsSrw?t=1285
  • #RetroPIDs: making historic Platypus Infinitely Discoverable (PID), PIDapalooza: the Festival of Persistent Identifiers, 28 January 2021: https://youtu.be/CSeQNe5KR5U

For the latest news about BHL’s DOI work, check out #RetroPIDs on Twitter.

May 10, 2021by michelle.underhill
BHL News, Blog Reel

BHL Australia Turns 10!

map of australia with BHL Australia contributors highlighted

Ten years ago in June 2010, the Atlas of Living Australia and Museums Victoria signed an agreement with the Biodiversity Heritage Library – and BHL Australia was born.

Signing of BHL agreement.jpg

Dr. Mark Lonsdale (left), the then Chief of the Commonwealth Scientific and Industrial Research Organisation’s (CSIRO’s) Entomology Division, and Martin Kalfatovic (right), BHL Program Director, signing the Relationship Agreement between the Atlas of Living Australia, Museums Victoria and BHL. Image sourced from Martin Kalfatovic.

BHL Australia’s mission is to make Australia’s biodiversity literature freely accessible and discoverable. Ten years ago, we started with a single contributing organisation, Museums Victoria, and a team of five incredibly dedicated volunteers.

Over the past 10 years, BHL Australia has grown considerably. Our operation is still hosted by Museums Victoria (at the Melbourne Museum), but we now digitise literature (and ingest born-digital material) on behalf of 27 organisations across the country.

map of australia with BHL Australia contributors highlighted

We are now a truly national project, representing Australia’s state and territory museums, herbaria, royal societies and field naturalists clubs, as well as government agencies and natural history publishers. Together these organisations have contributed more than 350,000 pages from over 2,400 volumes.

Ely with scanner.jpg

Dr. Elycia Wallis, Project Lead for BHL Australia at the Atlas of Living Australia, with our Zeutschel 16000 Book Scanner, purchased in 2017. Photo Credit: Nicole Kearney.

These volumes include treasures such as George Shaw’s The Naturalist’s Miscellany (1789-1813), Helena Forde and Harriet Scott’s Australian lepidoptera and their transformations, drawn from the life (1890-1898) and John Gould’s A synopsis of the birds of Australia, and the adjacent Islands (1837).

BHL Australia has also uploaded an extensive list of journals onto BHL (they may not be as pretty, but they’re just as important). To peruse all 2,400+ volumes, see our full BHL Australia Collection.

Shaw's Platypus resized.jpg

The Naturalist’s Miscellany includes the first published scientific description and illustration of the Duck-billed Platypus (Ornithorhynchus anatinus). Shaw, George. The Naturalist’s Miscellany. Volume 10. 1799. Contributed to  BHL by Museums Victoria. DOI: https://doi.org/10.5962/p.304567. 

Our volunteer team has also grown since 2010; BHL Australia now has 15 amazing volunteers who do the majority of our scanning, cropping, image processing and metadata addition work, as well as three science communication volunteers. (You may have seen their hugely successful takeover of the BHL Instagram account during #BirdWeek last year.)

Tiziana cropped.PNG

BHL Australia volunteer, Tiziana Tizian, cropping page images in preparation for upload onto BHL. Photo Credit: Nicole Kearney.

Of course, like so many other digitisation operations around the globe, BHL Australia is now in lockdown (and has been since mid-March). But out of adversity comes opportunity. While in lockdown, we’ve switched our focus from physical to born-digital material. We’ve welcomed new contributors and have uploaded journal volumes published as recently as 2020. We’ve also started a major project (in collaboration with BHL superuser Rod Page) to upload article metadata for every Australian journal on BHL (we’re pretty obsessed with discoverability).

Records of WAM.png

The Records of the Western Australian Museum is now complete on BHL from 1910 to 2019 as a result of our efforts (during our COVID-19 lockdown) to upload born-digital material from Australia’s journals.

And, like so many others, we’ve also had to postpone our celebrations. We had planned to invite all our Australian contributors and volunteers to a big BHL birthday bash. That’s on hold for now, but in the meantime, here are our BHL staff — Cerise, Chris, Veronica and myself — waving our thanks to all those who support BHL Australia. To our volunteers, our contributors, the Atlas of Living Australia and our BHL community around the world — thank you for a wonderful 10 years!

BHL staff waving for volunteers

The BHL Australia team: Manager Nicole Kearney (top right); Digitisation Coordinators Veronica Scholes (top left) and Cerise Howard (bottom right); and Technician Chris Healey (bottom left).

June 30, 2020by Joel Richard
BHL News, Blog Reel, Tech Updates

BHL Journal Articles Are Now Discoverable via Unpaywall

Earlier this week, Rod Page and I received an email from Richard Orr, the Lead Developer at Unpaywall, telling us that he had created a work-around that would finally enable the Unpaywall extension to discover content in BHL. And I (Nicole) literally spent the rest of the day jumping for joy. 

Let us explain: 

Firstly, what’s a DOI?

DOIs (Digital Object Identifiers) are used throughout the scholarly research community to uniquely identify academic articles. They help readers locate the definitive version of a published article, and they make linking together the academic literature much easier – look at any recent paper and you’ll see that most of the references cited have DOIs.

DOIs do two things: 1) they uniquely identify an article, and 2) they point to the online location of the definitive version of that article (typically hosted by the article’s publisher). 

Many articles are not free to read: a significant proportion of both recently-published and historic articles are locked behind paywalls. However, with the rise of open access, it’s increasingly the case that there may be a free version of an article available somewhere online. BHL, for example, has scanned and made available tens of thousands of articles that also exist on commercial publishers’ websites. DOIs always direct users to the definitive version of an article. But if the definitive version is behind a paywall (as they so often are), how do we tell users that BHL has a free version available? Enter Unpaywall.

What’s Unpaywall?

Unpaywall finds (legally) open access versions of paywalled literature. Since its launch in 2016, Unpaywall has become an indispensable tool for scientists  (see “How Unpaywall is transforming open science”). Unpaywall’s free browser extension (downloadable via their website) displays a discrete padlock symbol on the side of your browser whenever you are on a paywalled paper. If Unpaywall is able to locate a freely-accessible copy of the article elsewhere, the padlock symbol appears green and clicking on it will take you directly to the open access version. To discover whether an article is free, Unpaywall scans a database of millions of articles compiled from over 50,000 sources. Until this week, BHL wasn’t one of them. 

Screenshot of Unpaywall homepage

Why couldn’t Unpaywall link to BHL?

BHL contains hundreds of thousands of journal articles. More than a quarter of a million of these articles have been indexed (the vast majority by Rod Page via Biostor), which means that they now have article-level landing pages containing their article-level metadata. Tens of thousands of these article landing pages now have DOIs. This should have made them discoverable, but Unpaywall still couldn’t find them. 

In June, we contacted Unpaywall to find out why. It turns out that the reason BHL content has never been picked up by Unpaywall is because of the way BHL uploads and presents its journal content. Most providers of online journals present each article neatly packaged as an individual PDF. BHL, however, is first and foremost a virtual library. We upload complete volumes of journals made up of individual page images. Our article landing pages don’t link to PDFs; they link to the page in the volume upon which the article starts. Unpaywall looks for a PDF link to confirm that the document is actually available. Thus, the 57 million pages of open access content on BHL was excluded from Unpaywall’s database.  

After we explained to Unpaywall how significant BHL’s content was, Richard Orr, Unpaywall’s Lead Developer, very kindly agreed to create a work-around, specifically for BHL, that would enable Unpaywall to link to BHL article landing pages. The result of this work-around is that (as of this week) 43,000 journal articles on the BHL website are now discoverable via Unpaywall. 

To demonstrate how useful this is, here are two examples:

The first description of the iconic (sadly now-extinct) Thylacine (or Tasmanian Tiger) was published in the Transactions of the Linnean Society of London in 1808. This article is well and truly out of copyright and yet the definitive version of this article is behind a paywall on the Wiley Online website: https://doi.org/10.1111/j.1096-3642.1818.tb00336.x. Downloading the PDF of this out-of-copyright article from the Wiley website will cost you $42 USD.

Screenshot of an article landing page on Wiley Online

The definitive DOI version of this article is behind a paywall on the Wiley Online website, but it is freely available on the BHL website. With the Unpaywall extension, you can now easily navigate to the free version on BHL.

Now that BHL’s content is discoverable via Unpaywall, anyone directed to the Wiley version via the article’s DOI can discover the free version on BHL (via the Unpaywall extension).

Screenshot of an article in BHL

Article freely available via the Biodiversity Heritage Library.

It’s important to note that making BHL discoverable via Unpaywall doesn’t just enhance access to legacy literature such as the Thylacine paper; it also applies to much more recent research. For example, “The generic relationships of the new endemic Australian ant spider genus Notasteron (Araneae, Zodariidae)” was published in The Journal of Arachnology in 2005. This article has the DOI (https://doi.org/10.1636/04-56.1) and is behind a paywall on BioOne. With Unpaywall’s extension in your browser you can now discover the free version hosted by BHL.

How to make even more BHL content discoverable via Unpaywall

43,000 may seem like a large number, but it’s actually only a tiny fraction of the articles freely available on BHL. For the Unpaywall extension to be able to locate all the journal content on BHL, we need to unlock the rest of the journal articles on BHL by 1) adding more article-level metadata, and 2) ensuring that, for every article on BHL that has an existing DOI, we include that DOI in the article-level metadata. That’s our next task…

We would like to thank Unpaywall for providing access to an ever-increasing number of open access scholarly articles (23,943,966 at 14/8/19) and to particularly thank their Lead Developer, Richard Orr, for making it possible for BHL (a square peg) to fit into their open database (a round hole).  

August 16, 2019by Joel Richard

Help Support BHL

BHL's existence depends on the financial support of its patrons. Help us keep this free resource alive!

search

About BHL

The Biodiversity Heritage Library (BHL) is the world’s largest open access digital library for biodiversity literature and archives. BHL operates as a worldwide consortium of natural history, botanical, research, and national libraries working together to digitize the natural history literature held in their collections and make it freely available for open access as part of a global “biodiversity community.”

Join Our Mailing List

Sign up to receive the latest news, content highlights, and promotions.

Subscribe Now

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 319 other subscribers

Subscribe to Blog Via RSS

Subscribe to the blog RSS feed to stay up-to-date on all the latest BHL posts.

Access RSS Feed

Inspiring Discovery through Free Access to Biodiversity Knowledge.

The Biodiversity Heritage Library makes it easier than ever for you to access the information you need to study and explore life on Earth…for free, anytime, anywhere.

 

64+ Million Pages of
Biodiversity Literature Online.

EXPLORE

Tools and Services
to Transform Research.

EXPLORE

300,000+
Illustrations on Flickr.

EXPLORE

ABOUT | HARMFUL CONTENT | PRIVACY | SITE MAP | TERMS OF USE