A revised diagram & description of the BHL hardware architecture is available at:
http://www.slideshare.net/chrisfreeland/bhl-architecture-july-2008/
A revised diagram & description of the BHL hardware architecture is available at:
http://www.slideshare.net/chrisfreeland/bhl-architecture-july-2008/
Note: This is a revision of our previous blog post that described our process for harvesting digitized books from the Internet Archive. Their query interface changed, and we’ve updated our process & documentation accordingly.
Disclaimer: BHL is not directly or indirectly involved with the development of this query interface. We scan books through Internet Archive and are consumers of their services & interfaces. We have provided this documentation to help inform others of our process. Questions or comments concerning the query interface, results returned, etc., should be directed to the Internet Archive.
Overview
The following steps are taken to download data from Internet Archive and host it on the Biodiversity Heritage Library. Diagrams of the process are available in PDF.
For each “approved” item, clean up and transform the metadata into an “importable” format and store the results in the import database.Read all data that is ready for import and insert/update the appropriate data in the production database.Internet Archive Metadata Files
The following table lists the key XML files containing metadata for items hosted by Internet Archive. It is possible that one or more of these files may not exist for an item. However, most items that have been “approved” (i.e. marked as “complete” by Internet Archive) do include each of these files.
|
Filename |
Description |
|
*iaidentifier*_files.xml |
List of files that exist for the given identifier |
|
*iaidentifier*_dc.xml |
Dublin Core metadata. In many cases the data include here overlaps with the data in the _meta.xml file. |
|
*iaidentifier*_meta.xml |
Dublin Core metadata, as well as metadata specific to the item on IA (scan date, scanning equipment, creation date, update date, status of the item, etc) |
|
*iaidentifier*_metasource.xml |
Identifies the source of the item… not much meaningful data here |
|
*iaidentifier*_marc.xml |
MARC data for the item. |
|
*iaidentifier*_djvu.xml |
The OCR for the item, formatted as XML. |
|
*iaidentifier*_scandata.xml |
Raw data about the scanned pages. In combination with the OCR text (_djvu.xml), the page numbers and page types can be inferred from this data. This file may not exist, though in most cases it does. For the most part, only materials added to IA prior to late summery 2007 are likely to be missing this file |
|
*iaidentifier*_scandata.xml |
Raw data about the scanned pages.If there is no *iaidentifier*_scandata file for an item, we look in scandata.zip (via an IA API) for this file, which contains the same information. |
Internet Archive Services
Search for Items
Internet Archive items belong to one or more collections. To search a particular Internet Archive collection for items that have been updated between two dates, use the following query:
http://www.archive.org/advancedsearch.php
?q={0}+AND+oai_updatedate:[{1}+TO+{2}] &fl;[]=identifier&fl;[]=oai_updatedate&fmt;=xml&xmlsearch;=Search
where
{0} = name of the Internet Archive collection; in our case, “collection:biodiversity”
{1} = start date of range of items to retrieve (YYYY-MM-DD)
{2} = end date of range of items to retrieve (YYYY-MM-DD)
To limit the item search to a particular contributing institution, modify the query as follows:
http://www.archive.org/advancedsearch.php
?q={0}+AND+oai_updatedate:[{1}+TO+{2}]+AND+contributor:(MBLWHOI Library)
&fl;[]=identifier&fl;[]=oai_updatedate
&rows;=100000&fmt;=xml&xmlsearch;=Search
To limit the results of the query to a particular number of items, modify the query as follows:
http://www.archive.org/advancedsearch.php
?q={0}+AND+oai_updatedate:[{1}+TO+{2}] &fl;[]=identifier&fl;[]=oai_updatedate
&rows;=100000&fmt;=xml&xmlsearch;=Search
To search for one particular item, use:
http://www.archive.org/advancedsearch.php
?q={0}&fl;[]=identifier&fl;[]=oai_updatedate
&fmt;=xml&xmlsearch;=Search
where
{0} = an Internet Archive item identifier
Download Files
To download a particular file for an Internet Archive item, use the following query:
http://www.archive.org/download/{0}/{1}
where
{0} = an Internet Archive item identifier
{1} = the name of the file to be downloaded
Downloading Files Contained In ZIP Archives
In some cases, a file cannot be downloaded directly, and may instead need to be extracted from a ZIP archive located at Internet Archive. One example of this is the scandata.xml file, which in some cases must be extracted from the scandata.zip file. To do this, two queries must be made. First invoke this query to get the physical file locations (on IA servers) for the given item:
http://www.archive.org/services/find_file.php
?file={0}
&loconly;=1where
{0} = and Internet Archive item identifier
Then, invoke the second query to extract the scandata.xml file from the scandata.zip file (using the physical file locations returned by the previous query):
http://{0}/zipview.php
?zip={1}/scandata.zip
&file;=scandata.xmlwhere
{0} = host address for the file
{1} = directory location for the file
Note that the second query can be generalized to extract the contents of other zip files hosted at Internet Archive. The format for the query is:
http://{0}/zipview.php
?zip={1}/{2}
&file;={3}.jpgwhere
{0} = host address for the file
{1} = directory location for the file
{2} = name of the zip archive from which to extract a file
{3} = the name of the file to extract from the zip archive
Documentation written by Mike Lichtenberg.
Following some excellent suggestions gathered at a recent Encyclopedia of Life meeting, we’ve made changes to our Google Maps browse interface. To recap, we take Library of Congress Subject Headings and geocode and map them using the Google Maps API (details here).
Now that we’re managing nearly 10,000 volumes the standard Google Maps interface was getting cluttered and clunky, so we’ve refined the interface to show smaller points, weight the results using color, and display links to the titles for a given subject heading within the map itself (as demonstrated for “Africa” above). To view the map in full, visit http://www.biodiversitylibrary.org/browse/map.
We made another change based on requests to view full bibliographic details for a scanned title. When we harvest scans from the Internet Archive, we copy the MARCXML for the title to our servers and siphon off just enough of the metadata to facilitate our browse & search capabilities – to pull in the contents of the entire MARCXML would unnecessarily bloat our database with info we don’t expect to search across or expose via browse. But, it’s important data to have in the display, so we’ve skinned the MARCXML using XSLTs provided by the Library of Congress. To view in action, click the “Brief|Detailed|MARC” links at http://www.biodiversitylibrary.org/bibliography/1583, or for any title in our collection.
Finally, we’ve enhanced the display for our Discovered Bibliographies to return results in a more performant way, providing more visual feedback to the user that processes are at work. To view the refined interface, visit the result for Pomatomus saltatrix at http://www.biodiversitylibrary.org/name/Pomatomus_saltatrix.
BHL developers have released several significant updates to the BHL portal today. These updates include:
Author: Darwin, Charles, (1809 – 1882)
http://www.biodiversitylibrary.org/creator/93
Title: The Journal of the Linnean Society
http://www.biodiversitylibrary.org/title/350
Page: Carl Linnaeus’ Species Plantarum. 2 : 971. 1753.
http://www.biodiversitylibrary.org/page/358992
For a complete list of bugs and enhancements included in this release, visit our issue tracking web site.
The Missouri Botanical Garden (MOBOT), located in St. Louis, MO, is seeking to hire a Senior Programmer Analyst to work on several large biodiversity informatics projects, including the Biodiversity Heritage Library (BHL) online at www.biodiversitylibrary.org.
Primary responsibilities for this position include leading the development effort for MOBOT’s LAMP-based applications, complementing the existing .Net team. Up first on the development schedule is the instantiation of Fedora (www.fedora-commons.org) at MOBOT as a repository layer in our multi-platform, SOA-based infrastructure, then refactoring applications and building new ones to utilize Fedora. Future projects include enhancement of the BHL GUI and development of tools for managing digital library content.
Qualifications include a BS in Computer Science or related field, 5 years experience developing enterprise-level applications, and 2 years experience leading a development team. Experience managing data and applications in an open source environment (LAMP and its variants) required. Experience managing biodiversity and/or library datasets preferred, but not required.
To apply online, please visit:
http://www.mobot.org/jobs/mbgjobs.asp#H005
The following page from Biologia Centrali Americana, Insecta Lepidoptera-Heterocera v. 4 shows an interesting example of a proximity search we’d like to support with BHL Name Services – “find species x within n characters/words of species y.”
http://www.biodiversitylibrary.org/page/593637
Halfway through the entry for Lineodes integra you’ll see a character that looks like a crosshair, followed by “Solanum spp. 4-5,8, S. radula4-5, S. jasminifolium4-5, S. tuberosum (=Potato)8.” According to Wolfram Mey, the leading lepidopterist of the Museum of Natural History (MfN), Berlin:
The symbol means that the species has been reared from/on the particular plant. The symbol has been in use particularly by the old British authors, particularly Lord Walsingham, and is also used on the labels attached to the specimens. (translation by Dr. Michael Ohl)
What this tells us is Lineodes integra (Eggplant Leafroller Moth) is reared on a variety of Solanum species, including Solanum tuberosum (Potato). This example was uncovered during a Name search for Solanum tuberosum; the resulting bibliography included a link to this volume on insects from the Biologia Centrali-Americana, which seemed unusual given the search was for a plant species. This demonstrates why we’d want to facilitate proximity searches, so that users could find pages where both Lineodes integra and Solanum tuberosum occurred to aid in the discovery of predator-prey, plant-pollinator, or other coevolutionary relationships.
This example also suggests that our OCR algorithms are woefully inadequate to infer these kinds of relationships through automated means; the crosshair symbol was identified as ©.
“Names, especially those ascribed to organisms, serve as a primary entry point into the scientific, medical, and technical literature…”
– Garrity, Lyons, 2003, Future Proofing Biological Nomenclature
A characteristic of the Biodiversity Heritage Library (BHL) that distinguishes it from other mass digitization projects is the incorporation of service-based algorithms to identify scientific name strings throughout digitized content. These ‘taxonomically intelligent’ services, powered by uBio.org’s TaxonFinder and NameBank, have been incorporated into the BHL Portal to provide names-based interfaces into taxonomic literature.
To begin a search, visit http://www.biodiversitylibrary.org/NameSearch.aspx, or view an example ‘discovered bibliography’ for Tapirus bairdi (Baird’s Tapir), including an illustration, at http://www.biodiversitylibrary.org/name/Tapirus_bairdi. The ability to generate these ‘discovered bibliographies’ for taxa will enable users to data mine taxonomic literature for references and resources in ways not previously possible.
How it works
Each digitized page image in BHL has an accompanying OCR text file. As users navigate to a page, the uncorrected OCR file is sent to uBio’s TaxonFinder, which identifies text strings that match the characteristics of Latin binomials. Those potential name strings are then compared to the 10.7 million+ names in uBio’s NameBank, and the results, both matched and unmatched, are stored in the BHL database. BHL also has automated processes to reindex pages at regular intervals since NameBank is a growing repository.
What we’ve found
As of 20 Nov 2007 more than 6.8 million potential name strings have been identified throughout the BHL corpus, with more than 3.8 million matched to a corresponding NameBank identifier. There are more than 431,000 unique names within that 3.8 million set. Of those, more than 156,000 are known by a single occurrence. These results will be evaluated more thoroughly in the coming months to determine potential errors such as false positives and how to refine the TaxonFinder algorithm to reduce them.
Caveat: These results are generated from uncorrected OCR, which range in quality from pretty good (contemporary publications, such as modern issues of Rhodora) to downright terrible (18th century Latin texts, such as Species Plantarum). Again, further evaluation is required to determine the full scope of this problem.
Where we’re headed
To see a simple example of how this can be used from external sites, check out the ‘External Links’ at the bottom of the Wikipedia article for Mimosa pudica L., the sensitive plant:
http://en.wikipedia.org/wiki/Mimosa_pudica
Up next is development of a service layer on top of the names index so that other application providers can query & display ‘discovered bibliographies’ within their own applications. This service will be deployed in early 2008. These services are now available for use.
The Biodiversity Heritage Library (BHL) is the world’s largest open access digital library for biodiversity literature and archives. BHL operates as a worldwide consortium of natural history, botanical, research, and national libraries working together to digitize the natural history literature held in their collections and make it freely available for open access as part of a global “biodiversity community.”
Sign up to receive the latest news, content highlights, and promotions.
Subscribe NowSubscribe to the blog RSS feed to stay up-to-date on all the latest BHL posts.
Access RSS Feed
Inspiring Discovery through Free Access to Biodiversity Knowledge.
The Biodiversity Heritage Library makes it easier than ever for you to access the information you need to study and explore life on Earth…for free, anytime, anywhere.
