Biodiversity Heritage Library - Program news and collection highlights from BHL
  • Home
  • News
  • Featured Books
    • All Featured Books
    • Book of the Month Series
    • BHL at 20
  • User Stories
  • Campaigns
    • Fossil Stories
    • Garden Stories
    • Monsters Are Real
    • Page Frights
    • Her Natural History
    • Earth Optimism 2020
  • Tech Blog
  • Visit BHL
Home
News
Featured Books
    All Featured Books
    Book of the Month Series
    BHL at 20
User Stories
Campaigns
    Fossil Stories
    Garden Stories
    Monsters Are Real
    Page Frights
    Her Natural History
    Earth Optimism 2020
Tech Blog
Visit BHL
  • Home
  • News
  • Featured Books
    • All Featured Books
    • Book of the Month Series
    • BHL at 20
  • User Stories
  • Campaigns
    • Fossil Stories
    • Garden Stories
    • Monsters Are Real
    • Page Frights
    • Her Natural History
    • Earth Optimism 2020
  • Tech Blog
  • Visit BHL
Biodiversity Heritage Library - Program news and collection highlights from BHL
BHL News, Blog Reel, Tech Updates

New Article PDF Content Available

A sample of a printed page of a book with highlighted text superimposed over the printed text

The BHL Tech Team is pleased to announce a new form of content available in BHL: Article PDFs. While this may not sound like anything new, after all, we have had a tool to download PDF content for some time, this update changes both how the PDFs are created and maintained, and how BHL is viewed by content aggregators on the internet, most notably Unpaywall.

Screenshot of the Download PDF icon.

The new Download PDF icon

How to use it? While browsing an article, you will now see a Download PDF icon below the View Article link on the right side of the page. Clicking the link will immediately download the PDF to your computer (or view it in your web browser, depending on your settings.)

The benefits of the immediate download are:

  • No waiting.
  • No selecting pages.
  • The PDF contains embedded, searchable, copy-paste-able text.†
  • The PDF contains rich XMP-based metadata about the article.

An important change to note is that when viewing an article within an item at BHL, the Download Contents > Download Article link will now direct the visitor’s browser to the new PDFs for immediate download. This is a departure from what we had before in that the pages of the article were pre-selected for download and the visitor was then required to complete the process and wait for the PDF to be generated. We expect the new PDFs to be an improvement for our visitors who come to download articles. View the How do I download a PDF of an article? FAQ for simple download instructions.

Visitors to BHL are still able to manually create PDFs using the Download Contents > Select Pages to Download feature. This feature has not been removed, but it still means that it takes some time to create those PDFs and email the person when the PDF is ready. This option is useful for articles that have not been indexed in BHL, and therefore do not have a Download Article link. View the How do I generate a custom PDF of selected pages from the book? FAQ for complete instructions.

The most important feature of the new Article PDFs is the embedded text† within the document. The select-able text is an invisible text layer in the PDF, but it appears when you select or search for text within the document:

A sample of a printed page of a book with highlighted text superimposed over the printed text

An example of select-able text in an Article PDF.

While the appearance of the text may look… less than ideal, rest assured that the text can be copied out intact and used in another program. Example:

It is perhaps needless for me here to reiterate the great importance
of arriving at a final decision as to the real nature of
the haloliranic forms, for it will be obvious that if they have
nothing to do with the normal fresh-water series, and are to
be regarded as the remnant of an ancient sea, our views
respecting the past history of the African interior must be
greatly changed.

Other, less visible benefits to the PDFs are that they are directly linked from the citation_pdf_url meta-tag on the web page which makes them more findable by Google Scholar, Unpaywall, and potentially other aggregators.

For the technical-minded, the PDFs (many tens of thousands of them) are created in advance and stored on BHL’s servers. Changes to data within BHL will cause the PDF to be updated automatically, usually within several hours.

We hope that this is a welcome addition to BHL.

 

† – Please note that the text is only as good as the OCR that was generated for the text on the page. While the OCR text is probably very good for the prose sections of an article, titles, tables, and other special content may not appear as expected.

March 14, 2022by Sheila Rabun
BHL News, Blog Reel, Tech Updates

What Is BHL’s New Persistent Identifier Working Group DOI’ng?

Graphic showing the members of BHL's Persistent Identifier Working Group

In October 2020, BHL launched a new working group with a momentous goal: to make the content on BHL persistently discoverable, citable and trackable using DOIs (Digital Object Identifiers).

Graphic showing the members of BHL's Persistent Identifier Working Group

The members of BHL’s new Persistent Identifier Working Group (PIWG).

A DOI is like an electronic fingerprint in the form of a unique and permanent alphanumeric string that provides a persistent link to a piece of content online. Modern publications receive a DOI at the point of publication. This DOI becomes a key part of a publication’s bibliographic metadata that should be included in any mention or citation of that publication. Reference lists in modern publications are filled with DOIs, which allows readers to click from publication to publication in (in theory) a never-ending chain of knowledge.

This reciprocal linking of DOIs has created a great linked network of scholarly research, but that network is missing the historic literature. The vast majority of historic publications lack DOIs. This means they appear in reference lists as unlinked citations. In our increasingly online world, readers are far more likely to read (and thus cite) publications they can click through to (particularly when libraries are inaccessible during a global pandemic). The upshot of this is that the millions of pages of historic literature on BHL—the foundation of our understanding of biodiversity—is in danger of falling into obscurity.

BHL has been retrospectively minting DOIs for historic publications since 2011, but the focus has primarily been on monographs. BHL’s new Persistent Identifier Working Group (PIWG) is (at least initially) focusing on journal articles. Minting DOIs for articles on BHL is a far more complex and time-consuming task than minting DOIs for monographs. This is because article DOIs need article data: every journal volume uploaded onto BHL must be accompanied by journal and volume data, but there is no requirement that contributors provide article data.

Thankfully, there have been considerable efforts to add article data to BHL (and thereby make it possible to search for the titles and authors of these articles both within BHL and via external engines). A huge proportion of this article data has been contributed to BHL by Roderic Page via BioStor: 75% of the 300,753 articles indexed in BHL as of 4 May 2021 were “defined” by BioStor. It is very difficult to determine how many articles are actually on BHL (hidden within all those journal volumes). But, while we don’t know what proportion of BHL’s journal content still needs to be made discoverable, we know there is still a huge amount of work to do.

COVID-19 provided an unexpected opportunity to make a considerable dent in this work. With no access to scanners or library materials, a number of BHL contributors, including Harvard University Libraries, Muséum National d’Histoire Naturelle and BHL Australia, pivoted from making new content accessible to making their existing content on BHL more discoverable. For example, BHL Australia’s digitisation volunteers gathered, gap filled and checked article-level metadata for over 30,000 articles in 2020.

Once an article has been defined, i.e. it exists as a publication unit in BHL and has its own article landing page (and we’ve checked that the article does not already have a DOI), we can assign a DOI to it. Articles that have recently been assigned BHL DOIs include some very old publications, such as the first scientific description of the Duck-billed Platypus, published in 1799 (https://doi.org/10.5962/p.304567), and the species descriptions from A specimen of the botany of New Holland, the first publication dedicated to Australian flora (1793-5), e.g. https://doi.org/10.5962/p.312432. The PIWG has also started assigning DOIs to in-copyright publications (with permission from the rights holders). These include articles from the Bulletin of the British Museum, e.g. https://doi.org/10.5962/p.310418, and the Bulletin of the African Bird Club, e.g. https://doi.org/10.5962/p.308885.

Screenshot of the landing page in BHL for the description of the The Duck-Billed Platypus, Platypus anatinus.

Shaw, George (1799), The Duck-Billed Platypus, Platypus anatinus, The Naturalist’s Miscellany: https://doi.org/10.5962/p.304567 (illustration by Frederick Polydore Nodder).

If an article on BHL has an existing non-BHL DOI, we add this key piece of bibliographic metadata to the BHL landing page for the article. This ensures that BHL users can link to the definitive version of the article (the one the DOI resolves to), and more importantly, that other parties (and their algorithms) can find our versions from elsewhere. This is particularly important when commercial websites lock their DOI’d versions of public domain articles behind paywalls. Having their DOIs on our freely accessible versions ensures that services like Unpaywall can find them. To learn more about how this works, see our blog post: BHL Journal Articles Are Now Discoverable via Unpaywall.

DOIs not only improve discoverability and enable persistent linking to our historic content; they also allow us to track how BHL content is being used. In the six months following the minting of its new DOI (Oct 2020 to April 2021), the 1799 Platypus description was tweeted by 219 Twitter accounts, referenced in six Wikipedia pages, picked up by one news outlet and cited in one academic paper (data from Altmetric, April 2021). We know this because the article has a DOI.

Screenshot of the Altmetric dashboard for the first scientific description of the Duck-billed Platypus

Altmetric’s overview of attention for the first scientific description of the Duck-billed Platypus (Shaw 1799): https://www.altmetric.com/details/91788579.

The PIWG has spent the past six months creating, refining and testing tools that will allow BHL contributors to do this work themselves. We have also been producing documentation that explains a) how to use the new tools, and b) why this work is so important. These tools will facilitate every step in the article discoverability and DOI assignment process including: downloading existing article data for a given journal title to allow for correction and gap-filling (in development); bulk uploading of article data for new articles (available now); and adding articles and titles to BHL’s (new) DOI Assignment Queue (available now). Our dream is that, whenever anyone uploads a journal volume to BHL, they also provide the data for the articles it contains (and thus take responsibility for making that content discoverable).

The Persistent Identifier Working Group (PIWG) is fueled by the technical expertise, metadata dexterity and incredible passion of:

  • Nicole Kearney, Manager BHL Australia (Chair)
  • Mike Lichtenberg, BHL Lead Developer
  • Susan Lynch, Systems, Digitization & Web Services Librarian, The New York Botanical Garden
  • Bess Missell, Metadata Librarian, Smithsonian Libraries and Archives
  • Roderic Page, Professor of Taxonomy, University of Glasgow
  • Joel Richard, BHL Technical Coordinator | Head of Web Services & IT, Smithsonian Libraries, Smithsonian Libraries and Archives
  • Diane Rielinger, Digital Projects Librarian, Botany Libraries, Harvard University Herbaria
  • Colleen Funkhouser, BHL Program Manager

The specific goals of the group are:

  • To add article-level metadata to journal articles on BHL
  • To add existing DOIs to (new and existing) article landing pages on BHL (particularly for those articles where the DOI’d version is behind a paywall elsewhere)
  • To assign BHL DOIs to articles that lack DOIs

Want to know more about BHL’s Persistent Identifier Working Group? See:

  • Discovering the Platypus: From its scientific description to its DOI, Biodiversity Information Science and Standards (TDWG) Conference, 6 October 2020: https://youtu.be/4UVSEoWsSrw?t=1285
  • #RetroPIDs: making historic Platypus Infinitely Discoverable (PID), PIDapalooza: the Festival of Persistent Identifiers, 28 January 2021: https://youtu.be/CSeQNe5KR5U

For the latest news about BHL’s DOI work, check out #RetroPIDs on Twitter.

May 10, 2021by michelle.underhill
BHL News, Blog Reel, Tech Updates

BHL Journal Articles Are Now Discoverable via Unpaywall

Earlier this week, Rod Page and I received an email from Richard Orr, the Lead Developer at Unpaywall, telling us that he had created a work-around that would finally enable the Unpaywall extension to discover content in BHL. And I (Nicole) literally spent the rest of the day jumping for joy. 

Let us explain: 

Firstly, what’s a DOI?

DOIs (Digital Object Identifiers) are used throughout the scholarly research community to uniquely identify academic articles. They help readers locate the definitive version of a published article, and they make linking together the academic literature much easier – look at any recent paper and you’ll see that most of the references cited have DOIs.

DOIs do two things: 1) they uniquely identify an article, and 2) they point to the online location of the definitive version of that article (typically hosted by the article’s publisher). 

Many articles are not free to read: a significant proportion of both recently-published and historic articles are locked behind paywalls. However, with the rise of open access, it’s increasingly the case that there may be a free version of an article available somewhere online. BHL, for example, has scanned and made available tens of thousands of articles that also exist on commercial publishers’ websites. DOIs always direct users to the definitive version of an article. But if the definitive version is behind a paywall (as they so often are), how do we tell users that BHL has a free version available? Enter Unpaywall.

What’s Unpaywall?

Unpaywall finds (legally) open access versions of paywalled literature. Since its launch in 2016, Unpaywall has become an indispensable tool for scientists  (see “How Unpaywall is transforming open science”). Unpaywall’s free browser extension (downloadable via their website) displays a discrete padlock symbol on the side of your browser whenever you are on a paywalled paper. If Unpaywall is able to locate a freely-accessible copy of the article elsewhere, the padlock symbol appears green and clicking on it will take you directly to the open access version. To discover whether an article is free, Unpaywall scans a database of millions of articles compiled from over 50,000 sources. Until this week, BHL wasn’t one of them. 

Screenshot of Unpaywall homepage

Why couldn’t Unpaywall link to BHL?

BHL contains hundreds of thousands of journal articles. More than a quarter of a million of these articles have been indexed (the vast majority by Rod Page via Biostor), which means that they now have article-level landing pages containing their article-level metadata. Tens of thousands of these article landing pages now have DOIs. This should have made them discoverable, but Unpaywall still couldn’t find them. 

In June, we contacted Unpaywall to find out why. It turns out that the reason BHL content has never been picked up by Unpaywall is because of the way BHL uploads and presents its journal content. Most providers of online journals present each article neatly packaged as an individual PDF. BHL, however, is first and foremost a virtual library. We upload complete volumes of journals made up of individual page images. Our article landing pages don’t link to PDFs; they link to the page in the volume upon which the article starts. Unpaywall looks for a PDF link to confirm that the document is actually available. Thus, the 57 million pages of open access content on BHL was excluded from Unpaywall’s database.  

After we explained to Unpaywall how significant BHL’s content was, Richard Orr, Unpaywall’s Lead Developer, very kindly agreed to create a work-around, specifically for BHL, that would enable Unpaywall to link to BHL article landing pages. The result of this work-around is that (as of this week) 43,000 journal articles on the BHL website are now discoverable via Unpaywall. 

To demonstrate how useful this is, here are two examples:

The first description of the iconic (sadly now-extinct) Thylacine (or Tasmanian Tiger) was published in the Transactions of the Linnean Society of London in 1808. This article is well and truly out of copyright and yet the definitive version of this article is behind a paywall on the Wiley Online website: https://doi.org/10.1111/j.1096-3642.1818.tb00336.x. Downloading the PDF of this out-of-copyright article from the Wiley website will cost you $42 USD.

Screenshot of an article landing page on Wiley Online

The definitive DOI version of this article is behind a paywall on the Wiley Online website, but it is freely available on the BHL website. With the Unpaywall extension, you can now easily navigate to the free version on BHL.

Now that BHL’s content is discoverable via Unpaywall, anyone directed to the Wiley version via the article’s DOI can discover the free version on BHL (via the Unpaywall extension).

Screenshot of an article in BHL

Article freely available via the Biodiversity Heritage Library.

It’s important to note that making BHL discoverable via Unpaywall doesn’t just enhance access to legacy literature such as the Thylacine paper; it also applies to much more recent research. For example, “The generic relationships of the new endemic Australian ant spider genus Notasteron (Araneae, Zodariidae)” was published in The Journal of Arachnology in 2005. This article has the DOI (https://doi.org/10.1636/04-56.1) and is behind a paywall on BioOne. With Unpaywall’s extension in your browser you can now discover the free version hosted by BHL.

How to make even more BHL content discoverable via Unpaywall

43,000 may seem like a large number, but it’s actually only a tiny fraction of the articles freely available on BHL. For the Unpaywall extension to be able to locate all the journal content on BHL, we need to unlock the rest of the journal articles on BHL by 1) adding more article-level metadata, and 2) ensuring that, for every article on BHL that has an existing DOI, we include that DOI in the article-level metadata. That’s our next task…

We would like to thank Unpaywall for providing access to an ever-increasing number of open access scholarly articles (23,943,966 at 14/8/19) and to particularly thank their Lead Developer, Richard Orr, for making it possible for BHL (a square peg) to fit into their open database (a round hole).  

August 16, 2019by Joel Richard

Help Support BHL

BHL's existence depends on the financial support of its patrons. Help us keep this free resource alive!

search

About BHL

The Biodiversity Heritage Library (BHL) is the world’s largest open access digital library for biodiversity literature and archives. BHL operates as a worldwide consortium of natural history, botanical, research, and national libraries working together to digitize the natural history literature held in their collections and make it freely available for open access as part of a global “biodiversity community.”

Join Our Mailing List

Sign up to receive the latest news, content highlights, and promotions.

Subscribe Now

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 319 other subscribers

Subscribe to Blog Via RSS

Subscribe to the blog RSS feed to stay up-to-date on all the latest BHL posts.

Access RSS Feed

Inspiring Discovery through Free Access to Biodiversity Knowledge.

The Biodiversity Heritage Library makes it easier than ever for you to access the information you need to study and explore life on Earth…for free, anytime, anywhere.

 

64+ Million Pages of
Biodiversity Literature Online.

EXPLORE

Tools and Services
to Transform Research.

EXPLORE

300,000+
Illustrations on Flickr.

EXPLORE

ABOUT | HARMFUL CONTENT | PRIVACY | SITE MAP | TERMS OF USE