Biodiversity Heritage Library - Program news and collection highlights from BHL
  • Home
  • News
  • Featured Books
    • All Featured Books
    • Book of the Month Series
    • BHL at 20
  • User Stories
  • Campaigns
    • Fossil Stories
    • Garden Stories
    • Monsters Are Real
    • Page Frights
    • Her Natural History
    • Earth Optimism 2020
  • Tech Blog
  • Visit BHL
Home
News
Featured Books
    All Featured Books
    Book of the Month Series
    BHL at 20
User Stories
Campaigns
    Fossil Stories
    Garden Stories
    Monsters Are Real
    Page Frights
    Her Natural History
    Earth Optimism 2020
Tech Blog
Visit BHL
  • Home
  • News
  • Featured Books
    • All Featured Books
    • Book of the Month Series
    • BHL at 20
  • User Stories
  • Campaigns
    • Fossil Stories
    • Garden Stories
    • Monsters Are Real
    • Page Frights
    • Her Natural History
    • Earth Optimism 2020
  • Tech Blog
  • Visit BHL
Biodiversity Heritage Library - Program news and collection highlights from BHL
BHL News, Blog Reel, Tech Updates

BHL Improves the Speed and Accuracy of its Taxonomic Name Finding Services with gnfinder

New and improved BHL name finding services

BHL has deployed a new taxonomic name finding tool to improve the speed and accuracy of identifying names throughout its 58+ million pages.

BHL is now using Global Names Architecture’s (GNA) gnfinder tool to locate taxonomic names in the BHL corpus. Prior to this deployment, BHL’s name finding services were based on an index of scientific names created by GNA developers six years ago by parsing every page in BHL one by one. This took 45 days to accomplish, and the cost of repeating this process made updating or improving the index infeasible.

The gnfinder tool uses fast, scalable programming languages to significantly reduce computational time. Using Open Source applications in Go and Scala, the tool detects candidate scientific names and compares them to millions of scientific name-strings aggregated by GNA for verification. The new process decreases the time needed for name detection and name verification from 35 days to 5 hours and from 7 days to 12 hours, respectively. As a result, the entire BHL corpus can now be indexed in less than a day, compared to the 45 days needed for the previous index. Additionally, by significantly reducing computational time, implementing iterative improvements to the index is now achievable.

The accuracy of the names identified has also been improved with this deployment. By eliminating questionable results and false positives from the previous index, gnfinder produces a more accurate index of names in BHL. More than 34 million unique names — representing more than 239 million total instances of taxonomic name strings — were identified across the BHL corpus as of 21 July 2020. Of these, approximately 11.7 million are “Verified Names”, meaning they are unique names that have been resolved against a name authority (NameBank, Catalogue of Life, etc).

The gnfinder tool was developed by Dmitry Mozzherin and Alexander Myltsev as part of GNA project work at the University of Illinois at Urbana-Champaign. Mozzherin shared more about the process of developing this tool at the Biodiversity Next conference in Leiden, The Netherlands in 2019. Learn more in the presentation slides.

You can learn more about how the BHL implementation of the gnfinder tool works in our FAQ.

We would like to thank our colleagues at Global Names Architecture — especially Dmitry Mozzherin, Alexander Myltsev, and David Patterson — for their work to develop these tools. Thanks also to Joel Richard (BHL Technical Coordinator and Head of Web Services and IT at Smithsonian Libraries) and Mike Lichtenberg (BHL Lead Developer) for their work to deploy gnfinder on the BHL website.

If you have questions about gnfinder or would like to provide feedback or suggestions, please contact Global Names Architecture via the Global Names BHL project on GitHub.

Global Names development on BHL indexing is supported by National Science Foundation grants #1356347 and #1645959 as well as the Species File Group at the University of Illinois.

July 21, 2020by michelle.underhill
BHL News, Blog Reel, User Stories

Rod Page Talks Bioinformatics, Linked Data, and Primary Literature at the Smithsonian

Portrait version of the Biodiversity Heritage Library logo.

On 22 September, Dr. Rod Page, Professor at the University of Glasgow and creator of BioStor, gave a presentation at the Smithsonian’s National Museum of Natural History Library about ideas for extracting and linking data across biodiversity repositories and primary literature.

The talk, entitled “The Sam Adams Talk,” covered topics including phylogenetics, geophylogeny visualizations, linking article and specimen data across repositories, annotations, the biodiversity knowledge graph, and data as source code. A lively Q&A; session between Rod and the audience followed the presentation.

You can view a recording of the presentation on YouTube.

Rod’s talk at the Smithsonian followed the Scientific Collections International Food Security Symposium at the National Agricultural Library, 19-21 September. At the invitation of BHL, Rod gave a presentation at the symposium about using publication data to support food security research. You can view Rod’s Food Security Symposium presentation, entitled “Unknown Knowns, Long Tails, and Long Data“, on his Slideshare.

Visit Rod’s blog iPhylo and follow him on Twitter at @rdmpage to see more of his thoughts on bioinformatics, linked data, and primary literature.

September 29, 2016by ulib-libraryjobs
Blog Reel

How’s your fern and bird coverage, BHL?

“Every one knows what a bird is,” asserts an early 20th century book that I found while browsing the Biodiversity Heritage Library (BHL).

As I’ve learned during my Professional Development Internship with Jacqueline Chapman at Smithsonian Libraries this summer, it’s not always that simple. Taxonomy is ever-changing, especially at the granular level needed by subject specialists around the world who use BHL to conduct research on organisms ranging from mosses to turtles to fungi.

BHL is a consortial digital library whose member libraries digitize works in natural history and botany based on both user requests and subject librarians’ selections. My project for this summer was to refine a collection assessment methodology for BHL using both taxonomic and bibliographic analyses. Along the way, I’ve learned valuable lessons in using library tools, troubleshooting in Python (a computer programming language), and understanding the thought processes of 19th century ornithologists and pteridologists.

View Full Size Image
Becca Greenstein

Last year, Jacqueline worked with Robin Everly, the Smithsonian’s Botany and Horticulture Librarian, to conduct a taxonomic and bibliographic analysis to assess the depth of the BHL’s fern and lycophyte literature. They presented their results at an international conference on ferns, Next Generation Pteridology, and had the unique ability to talk with many subject-specialist users from around the world. Jacqueline later shared this proof-of-concept with researchers at TDWG in Nairobi, Kenya.

For the bibliographic portion of the project, Fern Books and Related Items in English before 1900 was used to create a list that could be referenced to determine whether a book was available on BHL, and if not, if we had access to it. A year later, I furthered this analysis by seeing what has changed in the past year and making requests for partner libraries to scan items to add to the collection. I enjoyed gathering data for books with titles such as Greenhouse Ferns and the Romance of Plant Life, Rambles in Search of Ferns, and The Fern Paradise: A Plea for the Culture of Ferns (2nd edition in BHL).

As the bibliography used included all editions of a particular work, regardless of whether the content had changed, I decided to not digitize the 53 works on the list whose content was already in BHL in another edition of the same work. As you can see in the graphs below, the number of fern books on BHL from this list has increased by 36% over the past year. The 112 titles from the list that are not yet in BHL but that we have access to via partner libraries will be in BHL after they are digitized. We lack access to only 37 of the titles on the list that would add content to BHL, and it will be interesting to follow up with this study to see if current partners acquire new resources or if new partners that possess these materials join the BHL Consortium.

View Full Size Image
2015 Bibliographic Analysis: Graph presents percentage of books from the list generated using Fern Books and Related Items in English before 1900 that are in BHL, are not in BHL but are held by a BHL partner, and are neither in BHL nor held by a BHL partner.
View Full Size Image
2016 Bibliographic Analysis, showing that the BHL collection of fern books has increased from 2015 to 2016.

For the taxonomic portion of the project, BHL’s coverage of a particular taxonomic grouping using scientific names was analyzed. The digitized material on BHL is in the form of images, which the computer does not recognize as text. Using Optical Character Recognition (OCR), the images are converted to machine-readable text. Taxonomic Name Recognition (TNR) then searches the OCR to find scientific names using multiple recognized lists of scientific names.

To use this powerful analytical tool to analyze BHL’s literature on birds, I upgraded the Python 2 code used for last year’s analysis to Python 3, the newest version of the programming language. Using my code, I counted the number of mentions in BHL of each genus of birds that appear in Catalogue of Life, as determined by TNR, to identify potential gaps in the BHL collection.

Of the 2234 genera analyzed, 99.6% of them are mentioned in the BHL corpus, 131 individual genera had more than 10,000 mentions in BHL, and 88% of them had more than 100 mentions.

I conducted an in-depth analysis of the 37 genera with fewer than ten mentions in BHL to figure out possible reasons for the paucity of literature. I determined that this lack of literature could be attributed to such things as the more-recent description of some of the genera, such as within the past 20 years, to the locality of some genera, as in some birds being endemic to far-away (to 19th century European ornithologists) places like New Guinea and Mozambique, and to taxonomic changes to the genera over the years. I then looked for the first mention of each of the 37 genera in books and journal articles online and in print, in addition to submitting scan requests for the books we have access to that weren’t already in BHL. There was something surreal about trekking up to the Birds Library, which is tucked away on the sixth floor of the National Museum of Natural History, finding Ornithologische Berichte on the shelf (and no, I don’t speak German), and opening to page 118 to find Wilhelm Meise’s initial description of Stresemannia bougainvillea.

View Full Size Image
Meise’s initial description of Stresemannia bougainvillea is next to my thumb.

My internship lasted six weeks, but it did not feel like that long. I hope that BHL will use my code to analyze larger sets of data and/or data at a higher level (for example, how is BHL doing at collecting literature on Kingdom Animalia?).

Through conducting my project, I’ve learned that things you learn in library school really do apply to the real world, how an academic library at an institution without students functions, and the workflow behind digitizing materials that appear in BHL and on the Smithsonian Digital Library. I’ve learned that library tools we take for granted can be unreliable, but aren’t usually, and that getting help from people who do research on ferns and those who do speak German can be very beneficial. I hope to bring the things I’ve learned back to my final two semesters of library school, as well as into my hoped-for career as a science librarian after I graduate.

________________________

About the Author

Becca Greenstein is getting her Master’s in Library Science at UNC-Chapel Hill. For her Bachelor’s degree, she went to Carleton College, where she majored in Biology and minored in Chinese. After graduating from Carleton, she worked as a lab technician at the University of Minnesota before starting library school. After she graduates, she hopes to continue honing these skills while working in an academic or special library as a science librarian.

August 23, 2016by ulib-libraryjobs
Blog Reel, User Stories

BHL and Our Users: Rod Page and BioStor

This week, we feature one of our users that has been extraordinarily active in not only using BHL content, but in creating applications that significantly enhance the information and knowledge that can be gleaned from our resources. The creator of BioStor and a huge player in the realm of biodiversity informatics, meet Dr. Roderic Page!

View Full Size Image

In the beginning…”meh”

I first became aware of the Biodiversity Heritage Library around 2007. To be honest, initially I was underwhelmed. BHL didn’t seem to have much literature, what it did have was mostly about plants (I’m a zoologist by background), the interface was a bit clunky, and most of the content was pre-1923, which to me simply echoed the impression that taxonomy is a science that is something of a backwater, obsessed with ancient documents and arcane terminology.

So at the start I wasn’t much of a fan. But as BHL grew it started to add more recent content, particularly for museum journals, as well as vital content such as the Bulletin of Zoological Nomenclature and I realised that it was going to be much more useful than I’d previously thought. So I started playing with ways to visualise content from BHL, such as timelines to plot search results over time, and sparklines to show how the relative frequency of different names for the same organism would change over time (similar to the nice visualisations Ryan Schenk has done recently)

But where are the articles?

These experiments were fun, but I keep coming up against what for me was the show stopper: BHL had no concept of a scientific article. Because it was a library project the basic unit in BHL was a scanned item, which could correspond to anything from a book, one or more volumes of a journal, or a single article. Whereas librarians deal with volumes on shelves, for most scientists the unit that matters is the article, and there was no easy way to find articles in BHL. To be fair, BHL was well aware of this mismatch between library practice and the expectations of scientists (see Chris Freeland’s post But where are the articles??).

I’d spent a lot of time developing a tool called bioGUID, which was designed to find articles online using just the journal name, volume, and starting page. It uses a range of web services to find the article, such as talking to CrossRef to see if it had a DOI (the ubiquitous identifier for modern articles), as well as searching other sources, such as JSTOR. I wanted something like this for BHL, where you could simply take those three things – journal, volume, starting page – and go straight to the article. A common way to provide this service is through a protocol called OpenURL, which takes the journal, volume, starting page for an article and looks for it online.

However, finding articles in BHL is a challenging task, not least because there is little standardisation in how library catalogues record bibliographic information. To give just one example, for the journal Proceedings of the Zoological Society of London here are some of the ways information about a volume is recorded.

  • Part 1- Part 4 (1833-38)
  • 1901, v. 1 (Jan.-Apr.)
  • Jan-Apr 1906
  • 1912 v. 2
  • 1923, pt. 1-2 (pp. 1-481)

So any tool to find articles has to deal with these issues. But after a few experiments I decided it would be possible to find lots of articles in BHL, especially if I had access to all the BHL data on my own computers. So, I grabbed a copy of the data and created BioStor.

BioStor

Below is a screen shot of BioStor, which at the moment has over 112,000 articles from BHL.View Full Size Image

There are two main ways to use BioStor. The first is as a website where you can browse or search for articles. You can search for articles about taxa by adding the taxon name to http://biostor.org/name/, for example http://biostor.org/name/Zonosaurus. In addition to displaying the article, BioStor displays the names found in the article as a tag cloud and a classification, and in some cases also shows a map with localities that have been automatically extracted from the text and displayed on the map, such as this example from A revision of the dwarf Zonosaurus Boulenger (Reptilia: Squamata: Cordylidae) from Madagascar, including descriptions of three new species:

View Full Size Image

The other way you can use BioStor is as an OpenURL resolver. Bibliographic software and websites such as EndNote, Zotero, and Mendeley all support OpenURL, so you can be looking at an article in one of those databases and automatically look for it in BioStor.

BioStor needs bibliographies

One thing I’ve glossed over is how BioStor has managed to find thousands of articles. Some have been added manually, but this rapidly gets tedious. For the majority of articles what I’ve done is take an existing bibliography for a journal, or a taxonomic group, and write a small computer programme (or “script”) to get BioStor to find the articles automatically. For example, I quickly added most of the articles in the journal Tijdschrift voor Entomologie because I had an EndNote file containing those references.

I spend a lot of time searching for bibliographies, downloading them or scrapping them from websites, converting them into a readable format, then using scripts to ask BioStor to locate the article in BHL. I’m somewhat taken aback by how hard it is to get these bibliographies. If taxonomists and/or journal editors made these available, we could add many more articles to BioStor. While one approach is to beg, borrow, or steal bibliographies, I’m hoping that the rise of online bibliography databases and associated social networks, especially Mendeley, will generate the bibliographies I need to efficiently find articles in BHL.

What’s next?

BioStor has some obvious limitations, notably the assumption that older literature works the same way as modern articles. Whereas today figures, tables, and text are all contained within the page range of an article, it’s not uncommon in older (pre-20th centruy) literature for figures and plates to be physically separate from the text. BioStor can’t really handle this, so one day I plan to add the ability to have discontinuous page ranges that will include these figures and plates.

BioStor’s data on articles is also now being fed back into BHL, meaning that you can now discover content at the article-level within the BHL portal itself. As of July 2015, BHL collections contained over 111,500 articles from BioStor.

What do I think of BHL now?

Despite my initial lack of enthusiasm, I now see BHL as one of the great resources of biodiversity informatics. There’s some extraordinary stuff in BHL, and it keeps growing. It’s also been great working with Chris Freeland, Phil Cryer, and Mike Lichtenberg, who have all been very helpful, even when I’ve written blog posts venting my frustration with BHL’s limitations. I think it’s definitely one of those cases where you only complain about the things you actually care about.

June 7, 2011by ulib-libraryjobs

Help Support BHL

BHL's existence depends on the financial support of its patrons. Help us keep this free resource alive!

search

About BHL

The Biodiversity Heritage Library (BHL) is the world’s largest open access digital library for biodiversity literature and archives. BHL operates as a worldwide consortium of natural history, botanical, research, and national libraries working together to digitize the natural history literature held in their collections and make it freely available for open access as part of a global “biodiversity community.”

Join Our Mailing List

Sign up to receive the latest news, content highlights, and promotions.

Subscribe Now

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 319 other subscribers

Subscribe to Blog Via RSS

Subscribe to the blog RSS feed to stay up-to-date on all the latest BHL posts.

Access RSS Feed

Inspiring Discovery through Free Access to Biodiversity Knowledge.

The Biodiversity Heritage Library makes it easier than ever for you to access the information you need to study and explore life on Earth…for free, anytime, anywhere.

 

64+ Million Pages of
Biodiversity Literature Online.

EXPLORE

Tools and Services
to Transform Research.

EXPLORE

300,000+
Illustrations on Flickr.

EXPLORE

ABOUT | HARMFUL CONTENT | PRIVACY | SITE MAP | TERMS OF USE