Biodiversity Heritage Library - Program news and collection highlights from BHL
  • Home
  • News
  • Featured Books
    • All Featured Books
    • Book of the Month Series
    • BHL at 20
  • User Stories
  • Campaigns
    • Fossil Stories
    • Garden Stories
    • Monsters Are Real
    • Page Frights
    • Her Natural History
    • Earth Optimism 2020
  • Tech Blog
  • Visit BHL
Home
News
Featured Books
    All Featured Books
    Book of the Month Series
    BHL at 20
User Stories
Campaigns
    Fossil Stories
    Garden Stories
    Monsters Are Real
    Page Frights
    Her Natural History
    Earth Optimism 2020
Tech Blog
Visit BHL
  • Home
  • News
  • Featured Books
    • All Featured Books
    • Book of the Month Series
    • BHL at 20
  • User Stories
  • Campaigns
    • Fossil Stories
    • Garden Stories
    • Monsters Are Real
    • Page Frights
    • Her Natural History
    • Earth Optimism 2020
  • Tech Blog
  • Visit BHL
Biodiversity Heritage Library - Program news and collection highlights from BHL
Blog Reel, Featured Books

The Treasure Between the Covers: Making BHL’s Articles Discoverable and Citable

A Gif zooming into a page of many tiny images of scanned pages

This post is part of BHL at 20: Treasures from the Biodiversity Heritage Library, a series contributed by members of the BHL community, highlighting remarkable works from across the collection in celebration of its 20th anniversary.

One of the greatest challenges for a digital library, especially one as large as the Biodiversity Heritage Library, is simply finding the content you are after. Recently, I made a website to give a sense of this challenge. The website features about 200,000 pages of content from BHL Australia, a small fraction of what is in BHL overall, but already it’s something of a challenge to find specific items you might be after. If you were looking for a particular article, how would you find it?

A Gif zooming into a page of many tiny images of scanned pages

Interactive browser of BHL Australia

Articles, articles, articles

For most scientists, the article is the fundamental unit of research, not the journal title, and not a journal volume. The article is what we download as a PDF, what we store in our reference managers, and what gets cited. At the outset BHL did not have articles, so over a decade ago I set about developing a tool to find those. This tool became BioStor, which was described in a paper in 2011 (Extracting scientific articles from a large digital archive: BioStor and the Biodiversity Heritage Library). The basic idea behind BioStor is to take information about an article, such as journal, volume, pages, and year, and then try and find that article in BHL.

A diagram with text and large blue arrows showing the mapping between BHL and articles

Mapping journal, volume, and pagination from an article to BHL.

In principle this seems straightforward, but often the vagaries of metadata complicate the task. The image below shows some of the issues encountered with the journal Ibis. The source of metadata for the articles was CrossRef, via a commercial publisher (Wiley). You might expect this data to be high quality, but it contains errors such as bad character encoding. To further complicate things, Wiley decided to renumber all the volumes of the journal, so that the original volume information we see in BHL (such as series 2, volume 1) bears little relation to what is in CrossRef (volume 7, issue 1).

Four examples of citations, with some text highlighted in orange. The Crossref and BHL logos are on the right.

Matching CrossRef metadata for an article in Ibis to BHL.

The list of metadata messes like this is almost endless. There are journals that have more than one numbering system for the same volumes (e.g., Annali del Museo civico di storia naturale Giacomo Doria where the same item is both series 3, volume 7 and volume 47), there are multiple abbreviations for the same journal, and there are journals with multiple names (e.g., title 8097 is “Annuaire du Musée zoologique de l’Académie des sciences de St. Pétersbourg”, “ЕЖЕГОДНИКЬ ЗООЛОГИЧЕСКАГО МУЗЕЯ ИМПЕРАТОРСКОЙ АКАДЕМІЙ НАУКЬ”, and “Ezhegodnik Zoologicheskago muzeia …”).

A further complication is that a scanned item in BHL may contain several issues or volumes, each with its own set of overlapping page numbers, which means we have to decide which page “1” is the page 1 that matches the article we are searching for. Once we solve all that, we encounter further problems. Perhaps the most challenging is pagination. In most modern articles, the page range, e.g. 1–5, completely encompasses the article, including figures, charts, illustrations, etc. But for the older literature this is often not the case. Typesetting text and reproducing plates were different processes, and hence the plates might be disconnected from the article (often appearing at the end of a volume). This means that extracting, say pages 1–5, from BHL is no guarantee that you have the whole article.

A good deal of code in BioStor is trying to make sense of matching article metadata to BHL items, finding the correct page to match to, and extracting the set of pages that correspond to that article, as well as providing tools to manually correct metadata and add missing pages (for example, the plates mentioned above). Hence the process of finding articles is, at best, semiautomated.

BioStor old and new

The original BioStor website dates back to 2009, and looked something like this:

A screenshot from a website showing an image viewer with a yellow page from an book and fields of metadata.

Original BioStor website.

This site could display individual articles, and you could edit metadata. For a variety of reasons, it was no longer feasible to host this at the university where I was based, so I split the website into two versions. The original site now runs only on my laptop, and I use it to process files and locate articles in BHL. The new version runs in the cloud and features a cleaner interface, along with much better search. Below is the same article in the current BioStor.

A screenshot of a website showing a search bar, bibliographic data, and page images.

Current BioStor website.

Once articles are discovered using the old BioStor, they get pushed to the public version of BioStor at https://biostor.org. This website is also the point of contact between BioStor and BHL: each day BHL runs an automated process which asks BioStor whether it has any new articles, and, if the answer is yes, it fetches those and adds them to BHL (in BHL articles are referred to as either “parts” or “segments”). The end result is that articles defined in BioStor now become visible in the Table of Contents in BHL.

A screenshot of the BHL website showing an image viewer with a yellow page of a journal, bibliographic metadata and a highlighted article in a contents page.

BioStor article displayed in BHL.

One advantage of having a separate project such as BioStor is that I can use it to experiment with different ways to view BHL content. For example, BioStor looks for geographic coordinates (latitude and longitude) in the OCR text for each article. Any pairs of coordinates that it finds get stored in a map, which you can browse. In the diagram below we have selected a small region in the centre of the map, on the right you can see a list of articles about that area.

A map of the island of Sulawesi with small red dots sprinkled across it. There is a pink rectangle over a cluster of red dots.

Maps showing localities on the island of Sulawesi that are mentioned in BioStor articles.

Identifiers

BioStor has been running since 2009. In that time it has contributed over 260,000 articles to BHL, making it the single largest source of BHL “parts”. Having articles is nice, but even better is having articles with persistent, citable identifiers, such as DOIs. The Persistent Identifier Working Group has been working to add DOIs to BHL content, especially “parts”. This work has focused on two kinds of DOIs. The first are existing DOIs minted, for example, by commercial publishers. BioStor adds a lot of articles using CrossRef metadata, so we get these “for free” (there are other sources of DOIs that BioStor uses, but that is another story). Why does it matter to have external DOIs for BHL content? Well, many of these articles are free in BHL but behind a paywall on the publisher’s website. Services such as Unpaywall can link existing DOIs to free versions of the corresponding article, and BHL is one of Unpaywall’s providers.

But the more exciting (and onerous) task is minting new DOIs for articles in BHL, so that BHL is the version of record for that content. This has several implications. It means that BHL is effectively a publisher, and has the responsibility to maintain access to this content in perpetuity. It also changes the way we think about adding articles. For example, most of my work with BioStor has been opportunistic – I’m working on a taxonomic database, I see that there are some papers that should be in BHL, find them, then add them to BioStor so future BHL users can find those articles. But once we start creating DOIs, the goal is quite different: you want to get every article in the journal that is in BHL, and mint DOIs for all of them. While this appeals to a completionist mindset, it does mean getting metadata for every article before you can add DOIs.

Luckily, the hard work in minting DOIs has a striking payback, we can see how many times articles in BHL are cited in the scientific literature. The last time the results were analyzed, BHL articles had been cited some 74,446 times! Without BHL these publications would appear as simple text strings in the literature cited, now they are first-class digital citizens with clickable DOI links.

Metadata matters

By now it is obvious that the way BioStor finds articles depends on having good quality metadata for articles (or chapters), which it then attempts to locate in BHL. The lack of freely accessible metadata is a major impediment to increasing the rate at which articles are added. In the past I have made extensive use of taxonomic databases as a source of bibliographic data (see my BioNames project, for example). Yet the quality of citations in these databases is often poor. I have also made extensive use of sources such as CrossRef, which covers articles that have been assigned DOIs by that agency, and also data provided by volunteers, such as those working with Nicole Kearney (thank you Bob Griffith and Heidi Griffith!). Another major source of data has come from scraping the web, a time-consuming process that is becoming increasingly difficult to do as the web becomes increasingly closed under the onslaught of AI bots (see also Joel Richard’s blog post A Brief Bit on BHL Battling a Barrage of Bots).

There is a clear need for a free and open bibliographic database. The nearest we have is OpenAlex, whose tagline is “All the world’s research, connected and open.” Sadly this is still more of an aspirational goal rather than a fact: a lot of taxonomic literature is not in OpenAlex. Perhaps it is time, therefore, to revive “CiteBank”, which was an early BHL project to collect bibliographic metadata. If we had a comprehensive database of the taxonomic and related literature, locating articles in BHL would be a much easier task.

Machines reading

BioStor’s method of finding articles works, but it is not the only way we could locate articles. Instead of relying on external sources of metadata, what if we could simply have a computer read the volume and extract the articles automatically? Early attempts to do this for BHL content were not particularly successful, see for example A metadata generation system for scanned scientific volumes. But the advent of large language models (LLMs) and AI chatbots has dramatically changed the way we can tackle finding articles in BHL. In my own work I routinely use AI to extract articles in bulk from a scanned volume. Typically the approach involves finding tables of contents in the scanned volume, using AI to parse that into structured data, then finding the corresponding pages in the volume, checking that they match the table of contents, and then using AI to extract bibliographic data (e.g., article title, authors, etc.). The result of this process is a data file that gets fed into BioStor, so that articles get found and added to BHL in the usual way. It is not bulletproof, and AI can quite happily make mistakes, but in my experience it works well.

But the holy grail would be to simply point an AI at a volume and it would identify and extract all the articles, find any stray plates, and present the results to BHL. Given the spectacular advances in OCR text and understanding document layout in recent years, perhaps there will be a point where BioStor can gracefully retire from the scene. Its hundreds upon hundreds of lines of regular expressions and special-case hacks quietly gathering dust in a GitHub repo while machines of loving grace read BHL for us.

From the Biodiversity Heritage Library

As BHL celebrates twenty years of open biodiversity knowledge, this post reminds us that access depends not only on digitised pages, but on the tools, metadata, identifiers, and infrastructure that make them discoverable and citable. With your support, BHL can continue strengthening the systems that connect biodiversity literature to the researchers, communities, and future discoveries that depend on it.Orange button with a heart icon

June 16, 2026by nkearney
BHL News, Blog Reel, Tech Updates

BHL Journal Articles Are Now Discoverable via Unpaywall

Earlier this week, Rod Page and I received an email from Richard Orr, the Lead Developer at Unpaywall, telling us that he had created a work-around that would finally enable the Unpaywall extension to discover content in BHL. And I (Nicole) literally spent the rest of the day jumping for joy. 

Let us explain: 

Firstly, what’s a DOI?

DOIs (Digital Object Identifiers) are used throughout the scholarly research community to uniquely identify academic articles. They help readers locate the definitive version of a published article, and they make linking together the academic literature much easier – look at any recent paper and you’ll see that most of the references cited have DOIs.

DOIs do two things: 1) they uniquely identify an article, and 2) they point to the online location of the definitive version of that article (typically hosted by the article’s publisher). 

Many articles are not free to read: a significant proportion of both recently-published and historic articles are locked behind paywalls. However, with the rise of open access, it’s increasingly the case that there may be a free version of an article available somewhere online. BHL, for example, has scanned and made available tens of thousands of articles that also exist on commercial publishers’ websites. DOIs always direct users to the definitive version of an article. But if the definitive version is behind a paywall (as they so often are), how do we tell users that BHL has a free version available? Enter Unpaywall.

What’s Unpaywall?

Unpaywall finds (legally) open access versions of paywalled literature. Since its launch in 2016, Unpaywall has become an indispensable tool for scientists  (see “How Unpaywall is transforming open science”). Unpaywall’s free browser extension (downloadable via their website) displays a discrete padlock symbol on the side of your browser whenever you are on a paywalled paper. If Unpaywall is able to locate a freely-accessible copy of the article elsewhere, the padlock symbol appears green and clicking on it will take you directly to the open access version. To discover whether an article is free, Unpaywall scans a database of millions of articles compiled from over 50,000 sources. Until this week, BHL wasn’t one of them. 

Screenshot of Unpaywall homepage

Why couldn’t Unpaywall link to BHL?

BHL contains hundreds of thousands of journal articles. More than a quarter of a million of these articles have been indexed (the vast majority by Rod Page via Biostor), which means that they now have article-level landing pages containing their article-level metadata. Tens of thousands of these article landing pages now have DOIs. This should have made them discoverable, but Unpaywall still couldn’t find them. 

In June, we contacted Unpaywall to find out why. It turns out that the reason BHL content has never been picked up by Unpaywall is because of the way BHL uploads and presents its journal content. Most providers of online journals present each article neatly packaged as an individual PDF. BHL, however, is first and foremost a virtual library. We upload complete volumes of journals made up of individual page images. Our article landing pages don’t link to PDFs; they link to the page in the volume upon which the article starts. Unpaywall looks for a PDF link to confirm that the document is actually available. Thus, the 57 million pages of open access content on BHL was excluded from Unpaywall’s database.  

After we explained to Unpaywall how significant BHL’s content was, Richard Orr, Unpaywall’s Lead Developer, very kindly agreed to create a work-around, specifically for BHL, that would enable Unpaywall to link to BHL article landing pages. The result of this work-around is that (as of this week) 43,000 journal articles on the BHL website are now discoverable via Unpaywall. 

To demonstrate how useful this is, here are two examples:

The first description of the iconic (sadly now-extinct) Thylacine (or Tasmanian Tiger) was published in the Transactions of the Linnean Society of London in 1808. This article is well and truly out of copyright and yet the definitive version of this article is behind a paywall on the Wiley Online website: https://doi.org/10.1111/j.1096-3642.1818.tb00336.x. Downloading the PDF of this out-of-copyright article from the Wiley website will cost you $42 USD.

Screenshot of an article landing page on Wiley Online

The definitive DOI version of this article is behind a paywall on the Wiley Online website, but it is freely available on the BHL website. With the Unpaywall extension, you can now easily navigate to the free version on BHL.

Now that BHL’s content is discoverable via Unpaywall, anyone directed to the Wiley version via the article’s DOI can discover the free version on BHL (via the Unpaywall extension).

Screenshot of an article in BHL

Article freely available via the Biodiversity Heritage Library.

It’s important to note that making BHL discoverable via Unpaywall doesn’t just enhance access to legacy literature such as the Thylacine paper; it also applies to much more recent research. For example, “The generic relationships of the new endemic Australian ant spider genus Notasteron (Araneae, Zodariidae)” was published in The Journal of Arachnology in 2005. This article has the DOI (https://doi.org/10.1636/04-56.1) and is behind a paywall on BioOne. With Unpaywall’s extension in your browser you can now discover the free version hosted by BHL.

How to make even more BHL content discoverable via Unpaywall

43,000 may seem like a large number, but it’s actually only a tiny fraction of the articles freely available on BHL. For the Unpaywall extension to be able to locate all the journal content on BHL, we need to unlock the rest of the journal articles on BHL by 1) adding more article-level metadata, and 2) ensuring that, for every article on BHL that has an existing DOI, we include that DOI in the article-level metadata. That’s our next task…

We would like to thank Unpaywall for providing access to an ever-increasing number of open access scholarly articles (23,943,966 at 14/8/19) and to particularly thank their Lead Developer, Richard Orr, for making it possible for BHL (a square peg) to fit into their open database (a round hole).  

August 16, 2019by Joel Richard
Blog Reel, Featured Books

Kirtlandia and the Cleveland Museum of Natural History

Origins of the Cleveland Museum of Natural History

The story of the Cleveland Museum of Natural History (CMNH) begins in the1830s, when a small group of men filled a two-room wooden building in Public Square, downtown Cleveland, with mounted animals. This building was known as the “Ark,” and the men who gathered there, united in their passion for natural history, were called “Arkites.”

The Arkites were led by William Case, who would later become mayor of Cleveland. He, his brother, and his father had used the Ark as a place to retreat from work, and in the absence of any other museums in the city, it became a hub for all kinds of collection and research.

View Full Size Image
Case Hall, engraving by William Payne,
courtesy of Special Collections,
Cleveland State University Library

In 1876, the Ark was relocated to Case Hall. It shared the space with other organizations, including the Kirtland Society (formerly the Cleveland Academy of Natural Sciences), named after renowned naturalist and fellow Arkite Jared Potter Kirtland, who died the following year. Case Hall remained the home of the Ark until 1916, when it was demolished to make way for the U.S. Post Office, Court House, and Custom House.

View Full Size Image
Kirtland’s Warbler, namesake of Jared Potter Kirtland,
from C.J. Maynard’s The Birds of Eastern North America, 1896,
Plate XXI. Digitized by Smithsonian Libraries.

The collections of both the Ark and the Kirtland Society found a new home in the Cleveland Museum of Natural History, founded in 1920. CMNH relocated several times as its collections grew, settling eventually at Wade Park, where it is today. William Case’s bird collection can still be viewed there, as well as many of the specimens provided by the Kirtland Society. Other exhibits include Balto, the hero dog of Nome, Alaska; “Dunk,” a large specimen of Dunkleosteus terrelli; and, most famous of all, “Lucy,” discovered in 1974 by former CMNH curator Donald Johanson.

The collections and physical space of the museum continue to grow today; this spring, CMNH unveiled the Ralph Perkins II Wildlife Center & Woods Garden, where visitors can view Ohio flora and fauna in their native habitat.

Kirtlandia

In 1972, CMNH began publishing Kirtlandia, a journal of original, peer-reviewed research by Museum staff. Wendy Wasman, Librarian and Archivist of the Harold T. Clark Library at CMNH, notes that Kirtlandia has a “long history of publishing cutting edge research in the natural sciences…Over the years, there have been articles on dinosaurs, fossil sharks, archaeology, botany, herpetology, mussels, moths, and even an entire issue devoted to paleontological research of the Kenya Rift Valley.”

View Full Size Image
Partial skeleton of “Lucy,”from Johanson, et al.,
“A New Species of the Genus Australopithecus…“,
Kirtlandia
, no. 28, 1978. Digitized by Smithsonian Libraries.

Early on, Kirtlandia issues focused on a single topic. In the late 1970s, however, issues began to contain multiple articles, and the length of those articles increased. They also featured more photographs and diagrams, though they retained their simple, sparse design.

Kirtlandia is supported by the Kirtlandia Society, founded in 1976 to advance research and education at the museum. Wasman says that because the CMNH library (named the Harold Terry Clark Library in 1972) had an active publications exchange program from the beginning, Kirtlandia can now be found in over 200 university and museum libraries worldwide, and that while publication is currently in hiatus, BHL has given it an even wider reach.

Kirtlandia was digitized by the Smithsonian Libraries as a part of the IMLS-funded Expanding Access to Biodiversity Literature (EABL) project. Susan Lynch, an EABL team member at the New York Botanical Garden, worked with Rod Page and BioStor, using metadata provided by Wendy Wasman, to define all of the articles in Kirtlandia. This allows users to search for and navigate to individual articles in the journal without having to scroll through an entire volume or set of volumes.

In search results, Kirtlandia articles can be found under the Articles/Chapters/Treatments tab:
View Full Size Image

When browsing from the title page, articles can be found by clicking View Identified Parts:

View Full Size Image

Thank you to the Cleveland Museum of Natural History and the Harold T. Clark Library for giving us permission to make Kirtlandia available in BHL!

References

“ARK.” The Encyclopedia of Cleveland History. Last modified July 10, 1997. Accessed November 16, 2016. http://ech.case.edu/cgi/article.pl?id=A16.

“Case Hall.” The Encyclopedia of Cleveland History. Last modified November 9, 2005. Accessed November 16, 2016. http://ech.case.edu/cgi/article.pl?id=CH.

“History.” Cleveland Museum of Natural History. Accessed November 16, 2016. https://www.cmnh.org/about-the-museum/history.

“Kirtlandia Society.” Cleveland Museum of Natural History. Accessed November 16, 2016.
https://www.cmnh.org/kirtlandiasociety.

Splain, Emily. “Cleveland Museum of Natural History.” Cleveland Historical. Accessed November 16, 2016. https://clevelandhistorical.org/items/show/41.

November 17, 2016by eomeara
BHL News, Blog Reel

Announcing the New Biodiversity Heritage Library!

View Full Size Image
The Homepage of the New BHL! Click image to enlarge.

Today the Biodiversity Heritage Library released a new user interface, including an updated website design, improved book navigation, and article-level access to collections. The new interface was informed by usability studies and is based on the design and functionality of the BHL-Australia portal.

Current Improvements Include:

  • Updated Design: The website’s design has been upgraded to reflect the celebrated aesthetics of the BHL-Australia portal. 
  • Article and Chapter Access: The ability to search BHL by article or chapter titles has been implemented. To date, over 81,000 articles and chapters have been indexed and are searchable within BHL. Additional articles and chapters will become available as the collections continue to be indexed.
  • Open Data Enhancements: BHL’s APIs, OpenURL interface, and Data Exports have been modified to include available article and chapter information.
  • Book Viewer Updates: The BHL book viewer has been updated, allowing users to view multiple columns of pages on screen at once and more easily navigate to a specific page within a book. Users can also view OCR text alongside page images, and, where the books have been indexed, users can navigate directly to the articles or chapters within using a new Table of Contents feature.
  • PDF Creation Improvements: The custom PDF creation process has been improved, allowing users to select pages for their PDF while in the book-viewer mode and more easily review the PDF before creation. Learn more about the new creation process in our Guide!

 

View Full Size Image
New and Improved BHL Book Viewer, with option to view multiple columns of pages at once and view OCR text alongside page images. Select books also have a Table of Contents feature which displays the articles/chapters identified within the text, with the ability for users to click on each to navigate directly to the corresponding part. Click image to enlarge.
View Full Size Image
New custom-PDF creation process, with ability to select pages for your PDF while viewing them and review your PDF before generation. Click image to enlarge.

Upcoming Improvements Include:

  • Improved Taxon Name Finding Algorithms: BHL will soon implement a new algorithm capable of identifying previously undiscovered taxon names throughout the BHL corpus. Test applications of this algorithm have already resulted in an increase of nearly 50 million names instances in BHL, translating to over 20 million unique names identified. These newly-identified names are currently available in BHL.

These developments follows BHL’s December, 2012, milestone achievement of providing access to over 40 million pages and over 110,000 volumes of freely-available biodiversity literature.

Explore the changes to BHL in-depth in our Guide to the New BHL.

We’d like to send a special thanks to everyone who made the new BHL possible. To start, thanks to the BHL-Australia team for their contributions to the process. First, to those who designed and developed the original BHL-Australia portal on which our new website is based, we thank Simon O’Shea (Designer) and Michael Mason (Developer). Secondly, to the BHL-Australia staff that worked with the US staff to merge the two UIs, we thank Simon Sherrin and Ajay Ranipeta (Developers) and Simone Downey (Designer – design based on original design by Simon O’Shea). And finally, a special thanks to Ely Wallis, BHL-Australia Director and Chair of the Global BHL Executive Committee, who selflessly supported the dedication of her staff’s time to this process.

Secondly, thanks to all of the members of the BHL TAG (Technical Advisory Group), including William Ulate (BHL Technical Director, Missouri Botanical Garden), Joe deVeer (Harvard-MCZ), John Mignault (The New York Botanical Garden), Joel Richard (Smithsonian Libraries), Jenna Nolt (United States Geological Survey), Francis Webb (Cornell University), Keri Thompson (Smithsonian Libraries), and Chris Freeland (Washington University). Furthermore, thanks to the Missouri Botanical Garden for hosting the BHL Technical Team, whose hard work made this vision a reality!

Thirdly, we would like to thank everyone who helped alpha and beta test the new site to ensure that our users’ experience would be the best that it could possibly be. Specifically, thanks to Rod Page, Pat LaFollette, and Francisco Welter-Schultes, BHL superstar users, for the valuable user-perspective input they provided.

Finally, we’d like to give a standing ovation to BHL’s Lead Developer, Mike Lichtenberg (Missouri Botanical Garden), who has dedicated countless hours over the past year to creating, with the support of all mentioned above, the new Biodiversity Heritage Library!

Big News Across the Pond: Launching BHL-Europe

Coinciding with the launch of our new portal, the BHL-Europe portal is also officially launching today. Providing access to material scanned from 92 content providers in Europe and the United States (including a subset of the BHL-US/UK corpus), the BHL-Europe collection currently comprises over 6,000 items, constituting over 1 million pages, of open access biodiversity literature, with more content being added daily. BHL-Europe’s content is also available through the Europeana portal. And, while you’re exploring BHL-Europe, be sure to check out Biodiversity Library Exhibitions (BLE), online exhibitions from BHL-Europe featuring books, images, stories and factoids about such topics as expeditions, spices, and poisonous nature.

Congratulations to all of our colleagues at BHL-Europe on the exciting launch of your new portal! Click here to learn more about the BHL-Europe launch.

We hope you’ll be as excited about all of these changes as we are! Visit the new and improved Biodiversity Heritage Library today! Explore the improvements in the “Guide to the New BHL.” Also, be sure to check out the new BHL-Europe portal and their online exhibitions! Tell us what you think of our improvements by sending us feedback, writing to [email protected], or leaving a comment on this post.

March 18, 2013by michelle.underhill
BHL News, Blog Reel

Coming Soon! A New and Improved Biodiversity Heritage Library

View Full Size Image
Sneak Peek! The soon-to-be-released BHL homepage!

The BHL team has been hard at work for the past year developing a new user interface (which will become live on March 18, 2013!) based on usability studies and the design and functionality of the BHL-Australia portal. BHL-Australia staff* partnered with BHL-US/UK staff to merge the user-praised aesthetics and book viewer of the Au portal with the functionality of the US/UK portal.

Besides a whole new look and feel, users will be able to navigate more easily within books, with the ability to view multiple columns of pages on screen at once, scroll quickly to a specific page within the book, and view the book’s OCR text alongside page images.

View Full Size Image
The new book viewer, which will include options to view multiple columns of pages on the screen at once and view OCR text alongside page images!
March 4, 2013by michelle.underhill
Blog Reel, User Stories

BHL and Our Users: Rod Page and BioStor

This week, we feature one of our users that has been extraordinarily active in not only using BHL content, but in creating applications that significantly enhance the information and knowledge that can be gleaned from our resources. The creator of BioStor and a huge player in the realm of biodiversity informatics, meet Dr. Roderic Page!

View Full Size Image

In the beginning…”meh”

I first became aware of the Biodiversity Heritage Library around 2007. To be honest, initially I was underwhelmed. BHL didn’t seem to have much literature, what it did have was mostly about plants (I’m a zoologist by background), the interface was a bit clunky, and most of the content was pre-1923, which to me simply echoed the impression that taxonomy is a science that is something of a backwater, obsessed with ancient documents and arcane terminology.

So at the start I wasn’t much of a fan. But as BHL grew it started to add more recent content, particularly for museum journals, as well as vital content such as the Bulletin of Zoological Nomenclature and I realised that it was going to be much more useful than I’d previously thought. So I started playing with ways to visualise content from BHL, such as timelines to plot search results over time, and sparklines to show how the relative frequency of different names for the same organism would change over time (similar to the nice visualisations Ryan Schenk has done recently)

But where are the articles?

These experiments were fun, but I keep coming up against what for me was the show stopper: BHL had no concept of a scientific article. Because it was a library project the basic unit in BHL was a scanned item, which could correspond to anything from a book, one or more volumes of a journal, or a single article. Whereas librarians deal with volumes on shelves, for most scientists the unit that matters is the article, and there was no easy way to find articles in BHL. To be fair, BHL was well aware of this mismatch between library practice and the expectations of scientists (see Chris Freeland’s post But where are the articles??).

I’d spent a lot of time developing a tool called bioGUID, which was designed to find articles online using just the journal name, volume, and starting page. It uses a range of web services to find the article, such as talking to CrossRef to see if it had a DOI (the ubiquitous identifier for modern articles), as well as searching other sources, such as JSTOR. I wanted something like this for BHL, where you could simply take those three things – journal, volume, starting page – and go straight to the article. A common way to provide this service is through a protocol called OpenURL, which takes the journal, volume, starting page for an article and looks for it online.

However, finding articles in BHL is a challenging task, not least because there is little standardisation in how library catalogues record bibliographic information. To give just one example, for the journal Proceedings of the Zoological Society of London here are some of the ways information about a volume is recorded.

  • Part 1- Part 4 (1833-38)
  • 1901, v. 1 (Jan.-Apr.)
  • Jan-Apr 1906
  • 1912 v. 2
  • 1923, pt. 1-2 (pp. 1-481)

So any tool to find articles has to deal with these issues. But after a few experiments I decided it would be possible to find lots of articles in BHL, especially if I had access to all the BHL data on my own computers. So, I grabbed a copy of the data and created BioStor.

BioStor

Below is a screen shot of BioStor, which at the moment has over 112,000 articles from BHL.View Full Size Image

There are two main ways to use BioStor. The first is as a website where you can browse or search for articles. You can search for articles about taxa by adding the taxon name to http://biostor.org/name/, for example http://biostor.org/name/Zonosaurus. In addition to displaying the article, BioStor displays the names found in the article as a tag cloud and a classification, and in some cases also shows a map with localities that have been automatically extracted from the text and displayed on the map, such as this example from A revision of the dwarf Zonosaurus Boulenger (Reptilia: Squamata: Cordylidae) from Madagascar, including descriptions of three new species:

View Full Size Image

The other way you can use BioStor is as an OpenURL resolver. Bibliographic software and websites such as EndNote, Zotero, and Mendeley all support OpenURL, so you can be looking at an article in one of those databases and automatically look for it in BioStor.

BioStor needs bibliographies

One thing I’ve glossed over is how BioStor has managed to find thousands of articles. Some have been added manually, but this rapidly gets tedious. For the majority of articles what I’ve done is take an existing bibliography for a journal, or a taxonomic group, and write a small computer programme (or “script”) to get BioStor to find the articles automatically. For example, I quickly added most of the articles in the journal Tijdschrift voor Entomologie because I had an EndNote file containing those references.

I spend a lot of time searching for bibliographies, downloading them or scrapping them from websites, converting them into a readable format, then using scripts to ask BioStor to locate the article in BHL. I’m somewhat taken aback by how hard it is to get these bibliographies. If taxonomists and/or journal editors made these available, we could add many more articles to BioStor. While one approach is to beg, borrow, or steal bibliographies, I’m hoping that the rise of online bibliography databases and associated social networks, especially Mendeley, will generate the bibliographies I need to efficiently find articles in BHL.

What’s next?

BioStor has some obvious limitations, notably the assumption that older literature works the same way as modern articles. Whereas today figures, tables, and text are all contained within the page range of an article, it’s not uncommon in older (pre-20th centruy) literature for figures and plates to be physically separate from the text. BioStor can’t really handle this, so one day I plan to add the ability to have discontinuous page ranges that will include these figures and plates.

BioStor’s data on articles is also now being fed back into BHL, meaning that you can now discover content at the article-level within the BHL portal itself. As of July 2015, BHL collections contained over 111,500 articles from BioStor.

What do I think of BHL now?

Despite my initial lack of enthusiasm, I now see BHL as one of the great resources of biodiversity informatics. There’s some extraordinary stuff in BHL, and it keeps growing. It’s also been great working with Chris Freeland, Phil Cryer, and Mike Lichtenberg, who have all been very helpful, even when I’ve written blog posts venting my frustration with BHL’s limitations. I think it’s definitely one of those cases where you only complain about the things you actually care about.

June 7, 2011by ulib-libraryjobs
BHL News, Blog Reel

BHL and Culturomics

On December 16, 2010 Science released a paper, “Quantitative Analysis of Culture Using Millions of Digitized Books” that describes data mining research using a vast textual archive created by the Google Books. The abstract reads, “We constructed a corpus of digitized texts containing about 4% of all books ever printed. Analysis of this corpus enables us to investigate cultural trends quantitatively. We survey the vast terrain of “culturomics”, focusing on linguistic and cultural phenomena that were reflected in the English language between 1800 and 2000. We show how this approach can provide insights about fields as diverse as lexicography, the evolution of grammar, collective memory, the adoption of technology, the pursuit of fame, censorship, and historical epidemiology. ‘Culturomics’ extends the boundaries of rigorous quantitative inquiry to a wide array of new phenomena spanning the social sciences and the humanities.”

The paper and subsequent commentary has accelerated nascent efforts at macroscopic, algorithmic questioning of large historical textual data sets. Can similar methods be applied fruitfully to the BHL corpus?

Already, Rod Page, bioinformatician and developer of BioStor, has demonstrated suggestive evidence in the affirmative by tracking a small sample of species names for the same organism in BHL texts through time and plotting the number of citations. The graph may be a visual representation of scientific debate and usage. Many other uses are possible, including:

  • Co-occurrence of place names with species
  • Frequency of co-occurrence of species names esp. with key words such as host, prey, predator, symbiont etc.
  • Tracking trends in zoological and botanical research by tracking methodological terminology through time.
  • Identification of taxonomically significant “events” in the literature based on textual cues.

Much of the follow-on activity to the Science paper is occurring in the “Digging into the Data” program. Thus, on May 9, the BHL made its data available for researchers in the Digging into the Data program.

View Full Size Image

BHL Director, Tom Garnett, will be attending the conference, “Digging into the Data” in June where speakers, including the authors of the Science article, will address issues of and opportunities in data mining of large textual corpora. With suitable partners, it is possible that we can seek NSF or Google funding for the unique use case our increasing text corpus presents. The framework for a proposal would be a team of biologists and a team of computer scientists posing research questions for the BHL corpus that would be amenable to algorithmic investigation. Even if funding is not forthcoming, if third party researchers use the BHL corpus to produce scientifically or historically salient results, it will enhance the value and use of the BHL, which can lead to further collaborations.

May 16, 2011by ulib-libraryjobs

Help Support BHL

BHL's existence depends on the financial support of its patrons. Help us keep this free resource alive!

search

About BHL

The Biodiversity Heritage Library (BHL) is the world’s largest open access digital library for biodiversity literature and archives. BHL operates as a worldwide consortium of natural history, botanical, research, and national libraries working together to digitize the natural history literature held in their collections and make it freely available for open access as part of a global “biodiversity community.”

Follow BHL

  • Bluesky logo
  • Instagram logo
  • Facebook logo
  • Flickr logo
  • Twitter logo

Join Our Mailing List

Sign up to receive the latest news, content highlights, and promotions.

Subscribe Now

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 319 other subscribers

Subscribe to Blog Via RSS

Subscribe to the blog RSS feed to stay up-to-date on all the latest BHL posts.

Access RSS Feed

 

64+ Million Pages of
Biodiversity Literature Online.

EXPLORE

Tools and Services
to Transform Research.

EXPLORE

300,000+
Illustrations on Flickr.

EXPLORE

ABOUT | HARMFUL CONTENT | PRIVACY | SITE MAP | TERMS OF USE