Biodiversity Heritage Library - Program news and collection highlights from BHL
  • Home
  • News
  • Featured Books
    • All Featured Books
    • Book of the Month Series
    • BHL at 20
  • User Stories
  • Campaigns
    • Fossil Stories
    • Garden Stories
    • Monsters Are Real
    • Page Frights
    • Her Natural History
    • Earth Optimism 2020
  • Tech Blog
  • Visit BHL
Home
News
Featured Books
    All Featured Books
    Book of the Month Series
    BHL at 20
User Stories
Campaigns
    Fossil Stories
    Garden Stories
    Monsters Are Real
    Page Frights
    Her Natural History
    Earth Optimism 2020
Tech Blog
Visit BHL
  • Home
  • News
  • Featured Books
    • All Featured Books
    • Book of the Month Series
    • BHL at 20
  • User Stories
  • Campaigns
    • Fossil Stories
    • Garden Stories
    • Monsters Are Real
    • Page Frights
    • Her Natural History
    • Earth Optimism 2020
  • Tech Blog
  • Visit BHL
Biodiversity Heritage Library - Program news and collection highlights from BHL
BHL News, Blog Reel, Tech Updates

Revised BHL Architecture

Portrait version of the Biodiversity Heritage Library logo.

A revised diagram & description of the BHL hardware architecture is available at:

http://www.slideshare.net/chrisfreeland/bhl-architecture-july-2008/

July 22, 2008by [email protected]
BHL News, Blog Reel, Tech Updates

Updated Harvesting Process from the Internet Archive

Portrait version of the Biodiversity Heritage Library logo.

Note: This is a revision of our previous blog post that described our process for harvesting digitized books from the Internet Archive. Their query interface changed, and we’ve updated our process & documentation accordingly.

Disclaimer: BHL is not directly or indirectly involved with the development of this query interface. We scan books through Internet Archive and are consumers of their services & interfaces. We have provided this documentation to help inform others of our process. Questions or comments concerning the query interface, results returned, etc., should be directed to the Internet Archive.


Overview
The following steps are taken to download data from Internet Archive and host it on the Biodiversity Heritage Library. Diagrams of the process are available in PDF.

  1. Get item identifiers from Internet Archive for items in the “biodiversity” collection that have been recently added/updated.
  2. For each item identifier:
  • Get the list of files (XML and images) that are available for download.
  • Download the XML and image files
  • Download the scan data if it is not included with the other downloaded files
  • Extract the item metadata from the XML files and store it in the import database.
  • Extract the OCR text from the XML files and store it on the file system (one file per page).

For each “approved” item, clean up and transform the metadata into an “importable” format and store the results in the import database.Read all data that is ready for import and insert/update the appropriate data in the production database.Internet Archive Metadata Files
The following table lists the key XML files containing metadata for items hosted by Internet Archive. It is possible that one or more of these files may not exist for an item. However, most items that have been “approved” (i.e. marked as “complete” by Internet Archive) do include each of these files.

Filename

Description

*iaidentifier*_files.xml

List of files that exist for the given identifier

*iaidentifier*_dc.xml

Dublin Core metadata. In many cases the data include here overlaps with the data in the _meta.xml file.

*iaidentifier*_meta.xml

Dublin Core metadata, as well as metadata specific to the item on IA (scan date, scanning equipment, creation date, update date, status of the item, etc)

*iaidentifier*_metasource.xml

Identifies the source of the item… not much meaningful data here

*iaidentifier*_marc.xml

MARC data for the item.

*iaidentifier*_djvu.xml

The OCR for the item, formatted as XML.

*iaidentifier*_scandata.xml

Raw data about the scanned pages. In combination with the OCR text (_djvu.xml), the page numbers and page types can be inferred from this data. This file may not exist, though in most cases it does. For the most part, only materials added to IA prior to late summery 2007 are likely to be missing this file

*iaidentifier*_scandata.xml

Raw data about the scanned pages.If there is no *iaidentifier*_scandata file for an item, we look in scandata.zip (via an IA API) for this file, which contains the same information.

Internet Archive Services
Search for Items
Internet Archive items belong to one or more collections. To search a particular Internet Archive collection for items that have been updated between two dates, use the following query:

http://www.archive.org/advancedsearch.php
?q={0}+AND+oai_updatedate:[{1}+TO+{2}] &fl;[]=identifier&fl;[]=oai_updatedate&fmt;=xml&xmlsearch;=Search

where

{0} = name of the Internet Archive collection; in our case, “collection:biodiversity”
{1} = start date of range of items to retrieve (YYYY-MM-DD)
{2} = end date of range of items to retrieve (YYYY-MM-DD)

To limit the item search to a particular contributing institution, modify the query as follows:

http://www.archive.org/advancedsearch.php
?q={0}+AND+oai_updatedate:[{1}+TO+{2}]+AND+contributor:(MBLWHOI Library)
&fl;[]=identifier&fl;[]=oai_updatedate
&rows;=100000&fmt;=xml&xmlsearch;=Search

To limit the results of the query to a particular number of items, modify the query as follows:

http://www.archive.org/advancedsearch.php
?q={0}+AND+oai_updatedate:[{1}+TO+{2}] &fl;[]=identifier&fl;[]=oai_updatedate
&rows;=100000&fmt;=xml&xmlsearch;=Search

To search for one particular item, use:

http://www.archive.org/advancedsearch.php
?q={0}&fl;[]=identifier&fl;[]=oai_updatedate
&fmt;=xml&xmlsearch;=Search

where

{0} = an Internet Archive item identifier

Download Files
To download a particular file for an Internet Archive item, use the following query:

http://www.archive.org/download/{0}/{1}

where

{0} = an Internet Archive item identifier
{1} = the name of the file to be downloaded

Downloading Files Contained In ZIP Archives
In some cases, a file cannot be downloaded directly, and may instead need to be extracted from a ZIP archive located at Internet Archive. One example of this is the scandata.xml file, which in some cases must be extracted from the scandata.zip file. To do this, two queries must be made. First invoke this query to get the physical file locations (on IA servers) for the given item:

http://www.archive.org/services/find_file.php
?file={0}
&loconly;=1

where

{0} = and Internet Archive item identifier

Then, invoke the second query to extract the scandata.xml file from the scandata.zip file (using the physical file locations returned by the previous query):

http://{0}/zipview.php
?zip={1}/scandata.zip
&file;=scandata.xml

where

{0} = host address for the file
{1} = directory location for the file

Note that the second query can be generalized to extract the contents of other zip files hosted at Internet Archive. The format for the query is:

http://{0}/zipview.php
?zip={1}/{2}
&file;={3}.jpg

where

{0} = host address for the file
{1} = directory location for the file
{2} = name of the zip archive from which to extract a file
{3} = the name of the file to extract from the zip archive

Documentation written by Mike Lichtenberg.

June 13, 2008by [email protected]
BHL News, Blog Reel, Tech Updates

Better maps! More bibliographic detail!

Portrait version of the Biodiversity Heritage Library logo.

View Full Size Image

Following some excellent suggestions gathered at a recent Encyclopedia of Life meeting, we’ve made changes to our Google Maps browse interface. To recap, we take Library of Congress Subject Headings and geocode and map them using the Google Maps API (details here).

Now that we’re managing nearly 10,000 volumes the standard Google Maps interface was getting cluttered and clunky, so we’ve refined the interface to show smaller points, weight the results using color, and display links to the titles for a given subject heading within the map itself (as demonstrated for “Africa” above). To view the map in full, visit http://www.biodiversitylibrary.org/browse/map.

We made another change based on requests to view full bibliographic details for a scanned title. When we harvest scans from the Internet Archive, we copy the MARCXML for the title to our servers and siphon off just enough of the metadata to facilitate our browse & search capabilities – to pull in the contents of the entire MARCXML would unnecessarily bloat our database with info we don’t expect to search across or expose via browse. But, it’s important data to have in the display, so we’ve skinned the MARCXML using XSLTs provided by the Library of Congress. To view in action, click the “Brief|Detailed|MARC” links at http://www.biodiversitylibrary.org/bibliography/1583, or for any title in our collection.

Finally, we’ve enhanced the display for our Discovered Bibliographies to return results in a more performant way, providing more visual feedback to the user that processes are at work. To view the refined interface, visit the result for Pomatomus saltatrix at http://www.biodiversitylibrary.org/name/Pomatomus_saltatrix.

April 25, 2008by [email protected]
BHL News, Blog Reel, Tech Updates

Major updates to BHL Portal released

Portrait version of the Biodiversity Heritage Library logo.

BHL developers have released several significant updates to the BHL portal today. These updates include:

  • Display of materials scanned by Internet Archive. BHL now manages more than 2.8 million pages from 7,500 digitized scientific texts. To stay updated on new titles, view our Recent Additions and subscribe to our feeds.
  • Filtering by Contributing Library. When users select “Browse By:” functions, they can filter results using the “For:” dropdown to view, for example, Authors from the New York Botanical Garden, or a Map of titles scanned by Smithsonian Institution Libraries, or Titles from All Contributors.
  • Browse by Names. Users can view the most frequently found names from our taxonomic name finding tools, and view a bibliography of their occurrences:
    http://www.biodiversitylibrary.org/browse/names
  • Stable URLs. BHL produces stable URLs for bookmarking and persistent linking to main parts of content, as in:Subject: Insects
    http://www.biodiversitylibrary.org/subject/Insects

    Author: Darwin, Charles, (1809 – 1882)
    http://www.biodiversitylibrary.org/creator/93

    Title: The Journal of the Linnean Society
    http://www.biodiversitylibrary.org/title/350

    Page: Carl Linnaeus’ Species Plantarum. 2 : 971. 1753.
    http://www.biodiversitylibrary.org/page/358992

  • Download link to all source files. For books scanned by Internet Archive, you can click the “Download” link to grab all of the files available:
    http://www.biodiversitylibrary.org/title/1539
  • Link to Developer Tools. We’ve documented our Name services, URLs, and all kinds of technical information at:
    http://www.biodiversitylibrary.org/Tools.aspx
  • Feedback tracking. Users can submit feedback or comments on records using the Feedback link at the top of the portal.

For a complete list of bugs and enhancements included in this release, visit our issue tracking web site.

February 25, 2008by [email protected]
BHL News, Blog Reel, Tech Updates

Senior Programmer needed to assist BHL development

Portrait version of the Biodiversity Heritage Library logo.

The Missouri Botanical Garden (MOBOT), located in St. Louis, MO, is seeking to hire a Senior Programmer Analyst to work on several large biodiversity informatics projects, including the Biodiversity Heritage Library (BHL) online at www.biodiversitylibrary.org.

Primary responsibilities for this position include leading the development effort for MOBOT’s LAMP-based applications, complementing the existing .Net team. Up first on the development schedule is the instantiation of Fedora (www.fedora-commons.org) at MOBOT as a repository layer in our multi-platform, SOA-based infrastructure, then refactoring applications and building new ones to utilize Fedora. Future projects include enhancement of the BHL GUI and development of tools for managing digital library content.

Qualifications include a BS in Computer Science or related field, 5 years experience developing enterprise-level applications, and 2 years experience leading a development team. Experience managing data and applications in an open source environment (LAMP and its variants) required. Experience managing biodiversity and/or library datasets preferred, but not required.

To apply online, please visit:
http://www.mobot.org/jobs/mbgjobs.asp#H005

January 4, 2008by [email protected]
Blog Reel

Eggplant Leafroller Moth reared on Potatoes

Portrait version of the Biodiversity Heritage Library logo.

The following page from Biologia Centrali Americana, Insecta Lepidoptera-Heterocera v. 4 shows an interesting example of a proximity search we’d like to support with BHL Name Services – “find species x within n characters/words of species y.”

View Full Size Image

http://www.biodiversitylibrary.org/page/593637

Halfway through the entry for Lineodes integra you’ll see a character that looks like a crosshair, followed by “Solanum spp. 4-5,8, S. radula4-5, S. jasminifolium4-5, S. tuberosum (=Potato)8.” According to Wolfram Mey, the leading lepidopterist of the Museum of Natural History (MfN), Berlin:

The symbol means that the species has been reared from/on the particular plant. The symbol has been in use particularly by the old British authors, particularly Lord Walsingham, and is also used on the labels attached to the specimens. (translation by Dr. Michael Ohl)

What this tells us is Lineodes integra (Eggplant Leafroller Moth) is reared on a variety of Solanum species, including Solanum tuberosum (Potato). This example was uncovered during a Name search for Solanum tuberosum; the resulting bibliography included a link to this volume on insects from the Biologia Centrali-Americana, which seemed unusual given the search was for a plant species. This demonstrates why we’d want to facilitate proximity searches, so that users could find pages where both Lineodes integra and Solanum tuberosum occurred to aid in the discovery of predator-prey, plant-pollinator, or other coevolutionary relationships.

This example also suggests that our OCR algorithms are woefully inadequate to infer these kinds of relationships through automated means; the crosshair symbol was identified as ©.

December 18, 2007by [email protected]
BHL News, Blog Reel, Tech Updates

‘Discovered Bibliographies’ through Natural Language Processing algorithms

Portrait version of the Biodiversity Heritage Library logo.

“Names, especially those ascribed to organisms, serve as a primary entry point into the scientific, medical, and technical literature…”
– Garrity, Lyons, 2003, Future Proofing Biological Nomenclature

A characteristic of the Biodiversity Heritage Library (BHL) that distinguishes it from other mass digitization projects is the incorporation of service-based algorithms to identify scientific name strings throughout digitized content. These ‘taxonomically intelligent’ services, powered by uBio.org’s TaxonFinder and NameBank, have been incorporated into the BHL Portal to provide names-based interfaces into taxonomic literature.

To begin a search, visit http://www.biodiversitylibrary.org/NameSearch.aspx, or view an example ‘discovered bibliography’ for Tapirus bairdi (Baird’s Tapir), including an illustration, at http://www.biodiversitylibrary.org/name/Tapirus_bairdi. The ability to generate these ‘discovered bibliographies’ for taxa will enable users to data mine taxonomic literature for references and resources in ways not previously possible.

How it works
Each digitized page image in BHL has an accompanying OCR text file. As users navigate to a page, the uncorrected OCR file is sent to uBio’s TaxonFinder, which identifies text strings that match the characteristics of Latin binomials. Those potential name strings are then compared to the 10.7 million+ names in uBio’s NameBank, and the results, both matched and unmatched, are stored in the BHL database. BHL also has automated processes to reindex pages at regular intervals since NameBank is a growing repository.

What we’ve found
As of 20 Nov 2007 more than 6.8 million potential name strings have been identified throughout the BHL corpus, with more than 3.8 million matched to a corresponding NameBank identifier. There are more than 431,000 unique names within that 3.8 million set. Of those, more than 156,000 are known by a single occurrence. These results will be evaluated more thoroughly in the coming months to determine potential errors such as false positives and how to refine the TaxonFinder algorithm to reduce them.

Caveat: These results are generated from uncorrected OCR, which range in quality from pretty good (contemporary publications, such as modern issues of Rhodora) to downright terrible (18th century Latin texts, such as Species Plantarum). Again, further evaluation is required to determine the full scope of this problem.

Where we’re headed
To see a simple example of how this can be used from external sites, check out the ‘External Links’ at the bottom of the Wikipedia article for Mimosa pudica L., the sensitive plant:
http://en.wikipedia.org/wiki/Mimosa_pudica

Up next is development of a service layer on top of the names index so that other application providers can query & display ‘discovered bibliographies’ within their own applications. This service will be deployed in early 2008. These services are now available for use.

November 21, 2007by [email protected]
Page 3 of 4«1234»

Help Support BHL

BHL's existence depends on the financial support of its patrons. Help us keep this free resource alive!

search

About BHL

The Biodiversity Heritage Library (BHL) is the world’s largest open access digital library for biodiversity literature and archives. BHL operates as a worldwide consortium of natural history, botanical, research, and national libraries working together to digitize the natural history literature held in their collections and make it freely available for open access as part of a global “biodiversity community.”

Join Our Mailing List

Sign up to receive the latest news, content highlights, and promotions.

Subscribe Now

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 319 other subscribers

Subscribe to Blog Via RSS

Subscribe to the blog RSS feed to stay up-to-date on all the latest BHL posts.

Access RSS Feed

Inspiring Discovery through Free Access to Biodiversity Knowledge.

The Biodiversity Heritage Library makes it easier than ever for you to access the information you need to study and explore life on Earth…for free, anytime, anywhere.

 

ABOUT | HARMFUL CONTENT | PRIVACY | SITE MAP | TERMS OF USE