Biodiversity Heritage Library - Program news and collection highlights from BHL
  • Home
  • News
  • Featured Books
    • All Featured Books
    • Book of the Month Series
    • BHL at 20
  • User Stories
  • Campaigns
    • Fossil Stories
    • Garden Stories
    • Monsters Are Real
    • Page Frights
    • Her Natural History
    • Earth Optimism 2020
  • Tech Blog
  • Visit BHL
Home
News
Featured Books
    All Featured Books
    Book of the Month Series
    BHL at 20
User Stories
Campaigns
    Fossil Stories
    Garden Stories
    Monsters Are Real
    Page Frights
    Her Natural History
    Earth Optimism 2020
Tech Blog
Visit BHL
  • Home
  • News
  • Featured Books
    • All Featured Books
    • Book of the Month Series
    • BHL at 20
  • User Stories
  • Campaigns
    • Fossil Stories
    • Garden Stories
    • Monsters Are Real
    • Page Frights
    • Her Natural History
    • Earth Optimism 2020
  • Tech Blog
  • Visit BHL
Biodiversity Heritage Library - Program news and collection highlights from BHL
BHL News, Blog Reel, Tech Updates

Export of titles & scientific names in BHL now available for download

Portrait version of the Biodiversity Heritage Library logo.

A series of files is now available for download that will enable libraries and other data providers to identify digitized titles available within BHL.

This suite of files also includes metadata about each volume scanned, as well as information about the millions of scientific names that have been identified throughout the BHL corpus and the pages on which those names occur.

Download files:

  • Documentation -updated 3/13/2009
  • Download .zip file of all tables (147MB)

NOTE: These files represent a first cut at how we want to make data providers and libraries aware of the content within BHL. Yes, we will build services, including an OpenURL resolver, but for now our partners have asked for a low-barrier export that they can manipulate for their own specific uses. The files above are automatically generated from the BHL database on a monthly basis. The datestamp on the files themselves indicate when they were last generated.

If you are interested only in the titles we have digitized, and the items (“books” or “volumes”) for each title, you only need to download the (significantly smaller) files for the following tables:

  • Documentation
  • Download contents of Title table as a tab-delimited text file. (4MB+)
  • Download contents of TitleIdentifier table as a tab-delimited text file. (400KB)
  • Download contents of Item table as a tab-delimited text file. (3MB+)

The full .zip download is not for the faint of heart! It’s a monster file because it includes the export of the 27 million 36 million occurrences of scientific names (updated 3/13/2009) identified in the BHL corpus through indexing by TaxonFinder.

Finally, we are considering this version a “warts and all” export. Merging the contents of multiple library catalogues and streamlining the digitization process to avoid duplication are the biggest challenges we face in building BHL, and to be frank our metadata is far from pristine in these early stages of our project. We are building functionality that allows librarians at BHL institutions to curate these digital books in ways that make sense to both scientists and librarians and that accommodate the variety of ways in which historic works have been catalogued over time. It’s a challenge we’ve just begun to tackle, and we look forward to any and all feedback you care to provide.

September 11, 2008by oneclickorders
BHL News, Blog Reel, Tech Updates

But where are the articles??

Portrait version of the Biodiversity Heritage Library logo.

Many researchers are used to searching or browsing for materials by article. Article level access to BHL content is a goal that we’re striving for, and one that we haven’t yet reached!

BHL is a mass scanning operation. Our member libraries are moving as quickly as possible through a range of materials – books, serials, etc. – in order to scan as much as possible during our relatively brief window of funding. Our goal is to scan & cache now, then add in advanced technology solutions for secondary post-processing as they are developed.

We’ve found that in scanning historic scientific monographs and journals, article identification is too labor intensive (and expensive) to do by hand. BHL staff, through connections formed by our scanning partner, Internet Archive, have been working with Penn State’s College of Information Sciences and Technology (the developers behind CiteSeer) to provide a test bed algorithm to extract article metadata from historic literature.

There are a number of challenges in our digitized historic literature that cause even the most scalable, sophisticated algorithms to return inaccurate results. These include:

  • uncorrected source OCR (accuracy is problematic)
  • multiple foreign languages, including Latin
  • irregular printing processes and type setting in historic literature
  • change of printing process and issue frequency during course of a journal run

Still, progress is being made:

  • Penn State’s algorithms have been demonstrated, as in this example
  • They need access to a wider testbed for improved machine learning, which is now available (7.4 million pages in BHL as of this writing)

But the work is far from finished. Next steps:

  • need to refine algorithms
  • need interfaces for human editing
    • too many possible inaccuracies upstream
    • possibly distribute this task via Mechanical Turk or other ‘clickworking’ network?
  • define workflow
    • first pass: algorithms; second pass: volunteers; editorial review?

But what about Google, you say? They’re scanning books en masse. They’re smart. Haven’t they solved the problem?? The quick answer is “No, not really.” Take a look at the “Contents” section for:
Zoologist: A Monthly Journal of Natural History, ser.4 v.12 1908

This is what Google can do…with all of their resources & grey matter. BHL is many things to many people, but we’re certainly not Google!

And here’s where we need your input:
So what can you, an enthusiastic supporter of BHL who wants access to articles in our collection, do to help? We’re glad you asked!

  1. Regardless of the means by which we actually get the article metadata, we’ve got to store that alongside all of our other content in BHL. We have released an UPDATED (9/3/2008) data model supporting articles for review & comment, guided by our research into NLM, OpenURL, and OAI-ORE, and with help from the expert developers behind the Public Library of Science.
  2. We need to hear how you, enthusiastic BHL supporter, expect to access and use article-based content. We’re looking for information about the sites you use and like with similar content, as well as your general expectations for delivery of articles in BHL. We know that’s a wide open question; here’s your chance to bend our ear(s).

Please use the “Comment” feature below to drop off your suggestions and ideas for both topics. Or, if you’d prefer to keep your opinions private, e-mail them to chris (dot) freeland (at) mobot (dot) org.

Looking forward to the feedback, and to providing this important method of access to the wealth of content in BHL.

August 19, 2008by oneclickorders
BHL News, Blog Reel, Tech Updates

Revised BHL Data Model

Portrait version of the Biodiversity Heritage Library logo.

The latest revision of the BHL Data Model is now available for review at:
http://www.biodiversitylibrary.org/documents/BHLDataModel_20080805.pdf

August 5, 2008by [email protected]
BHL News, Blog Reel, Tech Updates

Revised BHL Architecture

Portrait version of the Biodiversity Heritage Library logo.

A revised diagram & description of the BHL hardware architecture is available at:

http://www.slideshare.net/chrisfreeland/bhl-architecture-july-2008/

July 22, 2008by [email protected]
BHL News, Blog Reel, Tech Updates

Updated Harvesting Process from the Internet Archive

Portrait version of the Biodiversity Heritage Library logo.

Note: This is a revision of our previous blog post that described our process for harvesting digitized books from the Internet Archive. Their query interface changed, and we’ve updated our process & documentation accordingly.

Disclaimer: BHL is not directly or indirectly involved with the development of this query interface. We scan books through Internet Archive and are consumers of their services & interfaces. We have provided this documentation to help inform others of our process. Questions or comments concerning the query interface, results returned, etc., should be directed to the Internet Archive.


Overview
The following steps are taken to download data from Internet Archive and host it on the Biodiversity Heritage Library. Diagrams of the process are available in PDF.

  1. Get item identifiers from Internet Archive for items in the “biodiversity” collection that have been recently added/updated.
  2. For each item identifier:
  • Get the list of files (XML and images) that are available for download.
  • Download the XML and image files
  • Download the scan data if it is not included with the other downloaded files
  • Extract the item metadata from the XML files and store it in the import database.
  • Extract the OCR text from the XML files and store it on the file system (one file per page).

For each “approved” item, clean up and transform the metadata into an “importable” format and store the results in the import database.Read all data that is ready for import and insert/update the appropriate data in the production database.Internet Archive Metadata Files
The following table lists the key XML files containing metadata for items hosted by Internet Archive. It is possible that one or more of these files may not exist for an item. However, most items that have been “approved” (i.e. marked as “complete” by Internet Archive) do include each of these files.

Filename

Description

*iaidentifier*_files.xml

List of files that exist for the given identifier

*iaidentifier*_dc.xml

Dublin Core metadata. In many cases the data include here overlaps with the data in the _meta.xml file.

*iaidentifier*_meta.xml

Dublin Core metadata, as well as metadata specific to the item on IA (scan date, scanning equipment, creation date, update date, status of the item, etc)

*iaidentifier*_metasource.xml

Identifies the source of the item… not much meaningful data here

*iaidentifier*_marc.xml

MARC data for the item.

*iaidentifier*_djvu.xml

The OCR for the item, formatted as XML.

*iaidentifier*_scandata.xml

Raw data about the scanned pages. In combination with the OCR text (_djvu.xml), the page numbers and page types can be inferred from this data. This file may not exist, though in most cases it does. For the most part, only materials added to IA prior to late summery 2007 are likely to be missing this file

*iaidentifier*_scandata.xml

Raw data about the scanned pages.If there is no *iaidentifier*_scandata file for an item, we look in scandata.zip (via an IA API) for this file, which contains the same information.

Internet Archive Services
Search for Items
Internet Archive items belong to one or more collections. To search a particular Internet Archive collection for items that have been updated between two dates, use the following query:

http://www.archive.org/advancedsearch.php
?q={0}+AND+oai_updatedate:[{1}+TO+{2}] &fl;[]=identifier&fl;[]=oai_updatedate&fmt;=xml&xmlsearch;=Search

where

{0} = name of the Internet Archive collection; in our case, “collection:biodiversity”
{1} = start date of range of items to retrieve (YYYY-MM-DD)
{2} = end date of range of items to retrieve (YYYY-MM-DD)

To limit the item search to a particular contributing institution, modify the query as follows:

http://www.archive.org/advancedsearch.php
?q={0}+AND+oai_updatedate:[{1}+TO+{2}]+AND+contributor:(MBLWHOI Library)
&fl;[]=identifier&fl;[]=oai_updatedate
&rows;=100000&fmt;=xml&xmlsearch;=Search

To limit the results of the query to a particular number of items, modify the query as follows:

http://www.archive.org/advancedsearch.php
?q={0}+AND+oai_updatedate:[{1}+TO+{2}] &fl;[]=identifier&fl;[]=oai_updatedate
&rows;=100000&fmt;=xml&xmlsearch;=Search

To search for one particular item, use:

http://www.archive.org/advancedsearch.php
?q={0}&fl;[]=identifier&fl;[]=oai_updatedate
&fmt;=xml&xmlsearch;=Search

where

{0} = an Internet Archive item identifier

Download Files
To download a particular file for an Internet Archive item, use the following query:

http://www.archive.org/download/{0}/{1}

where

{0} = an Internet Archive item identifier
{1} = the name of the file to be downloaded

Downloading Files Contained In ZIP Archives
In some cases, a file cannot be downloaded directly, and may instead need to be extracted from a ZIP archive located at Internet Archive. One example of this is the scandata.xml file, which in some cases must be extracted from the scandata.zip file. To do this, two queries must be made. First invoke this query to get the physical file locations (on IA servers) for the given item:

http://www.archive.org/services/find_file.php
?file={0}
&loconly;=1

where

{0} = and Internet Archive item identifier

Then, invoke the second query to extract the scandata.xml file from the scandata.zip file (using the physical file locations returned by the previous query):

http://{0}/zipview.php
?zip={1}/scandata.zip
&file;=scandata.xml

where

{0} = host address for the file
{1} = directory location for the file

Note that the second query can be generalized to extract the contents of other zip files hosted at Internet Archive. The format for the query is:

http://{0}/zipview.php
?zip={1}/{2}
&file;={3}.jpg

where

{0} = host address for the file
{1} = directory location for the file
{2} = name of the zip archive from which to extract a file
{3} = the name of the file to extract from the zip archive

Documentation written by Mike Lichtenberg.

June 13, 2008by [email protected]
BHL News, Blog Reel, Tech Updates

WonderFetch(tm) & IA _meta.xml fields

Portrait version of the Biodiversity Heritage Library logo.

Overview

WonderFetch is the term used for prepopulating the Internet Archives metadata forms (so named because it is more wonderful than regular z39.50 fetching). Using WonderFetch, partner libraries can populate fields with data that would not normally be populated as part of the standard IA process, and then store those values in the foobar_meta.xml file alongside each scanned item in the IA repository. Part of the impetus for implementing WonderFetch was not just to automate the inclusion of volume and issue information for serials – which was important – but to also capture due diligence, rights, and licensing information related to each item. (And yes, the TM is a little joke! No rights reserved).

 

How does it work?

WonderFetch is simply passing a series of parameters to the IA metaform software in a URL string. If you can do a Z39.50 query against your ILS and get your data into a format that will let you generate a URL (say, an HTML page output from a database, or a spreadsheet with your item data in it) you can WonderFetch!

To create WonderFetch links from your data, just append the relevant arguments and data (listed below) to one of 2 base URLs.

If your books are being “loaded”, that is, if the metadata fetch is occurring on a scribe2 machine and the scanner is using the WonderFetch while at the SCRIBE machine, use http://localhost.archive.org/biblio.php? as the base URL.

If your books are being pre-loaded or batch loaded on another computer before being scanned, or your SCRIBE is using the scribe1 software, use http://www.us.archive.org/biblio.php? as the base URL. Using this URL, your scanner person will notice that they can’t actually start shooting the book – this link only allows them to fetch metadata and create the foobar_meta.xml and marc.xml records.

For example, at SIL the scanner loads the books at the scribe station, and we Z-fetch on our barcode number (starts with 39088) so our URLs look like:

http://localhost.archive.org/biblio.php?f=c
&b;_c1=biodiversity&b;_l=Smithsonian%20Institution%20Libraries
&b;_p=Smithsonian&z;_e=Smithsonian%20Institution
&b;_v=v.%209%201907&b;_cn=no
&z;_c=local&z;_d=39088009080136&b;_ib=39088009080136

Below are the arguments each field takes, along with examples. The examples given are values BHL partners will be using (for rights statement, e.g.).

 

(For additional reference, some definitions and usage of the standard IA fields can be found in: http://www.us.archive.org/biblio?f=usage )

 

 


LIST OF FIELDS in META_XML that can be pre-populated using WonderFetch:

CALL_NUMBER
Description: Number for ZQuery – this is *not* necessarily the call number of the item. It *is* the number used to fetch the MARC record for the item via Z39.50 from your ILS. Whatever works for you – barcode, bib number, oclc number, call number, etc.
Prepopulatable: Yes.
WonderFetch GET arg: “&z;_d=”
example: &z;_d=39088009080136

IDENTIFIER-BIB
Description: Unique identifier for the item in Contributor library’s catalog.
Prepopulatable: Yes
WonderFetch GET arg: “&b;_ib=”
example: &b;_ib=39088009080136

TITLE-ID
Description: Unique identifier for title in Contributor library’s catalog.
Prepopulatable: Yes
WonderFetch GET arg: “&tid;=”
example: &tid;=b12345

VOLUME
Description: Volume of Book being scanned.
Prepopulatable: Yes.
WonderFetch GET arg: &b;_v=
example: &b;_v=v.%209%201907 (v. 9 1907)

YEAR
Description: Year assigned to Book being scanned
Prepopulatable: Yes.
WonderFetch GET arg: &year;=
example: &year;=1907

COLLECTION
Description: Collection(s) into which Book will be sorted.
Prepopulatable: Yes
WonderFetch GET arg: “&b;_c1=” (c1 indicates primary collection, there is also c2 and c3)
example: &b;_c1=biodiversity

CONTRIBUTOR
Description: Library contributing Book for scanning
Prepopulatable: Yes
WonderFetch GET arg: “&b;_l=”
example: &b;_l=Smithsonian%20Institution%20Libraries

SPONSOR
Description: Organization responsible for funding scanning
Prepopulatable: Yes
WonderFetch GET arg: “&b;_p=”
example: “&b;_p=Sloan”

SCANNINGCENTER
Description: Scanning Center where book was scanned.
Prepopulatable: yes
Get arg: “&b;_n=”
example: &b;_n=Boston

DUE-DILIGENCE
Description: URL to Due Diligence statement for a Book scanned while still in Copyright.
Prepopulatable: Yes
WonderFetch GET arg: ⅆ=
examples: ⅆ=dd-bhl (this signifies a due diligence statement exists at the following url http://www.biodiversitylibrary.org/permissions )

LICENSE-TYPE
Description: Creative Commons License assigned to Book scanned.
Prepopulatable: Yes
WonderFetch GET arg: “&lic;=”
&lic;= by ( http://creativecommons.org/licenses/by/3.0/)
&lic;= by-nc ( http://creativecommons.org/licenses/by-nc/3.0/)
&lic;= by-nd ( http://creativecommons.org/licenses/by-nd/3.0/)
&lic;= by-sa (http://creativecommons.org/licenses/by-sa/3.0/)
&lic;= by-nc-nd ( http://creativecommons.org/licenses/by-nc-nd/3.0/)
&lic;= by-nc-sa ( http://creativecommons.org/licenses/by-nc-sa/3.0/)

NEGOTIATED-RIGHTS
Description: URL to Negotiated Rights for a Book scanned while still in Copyright.
Prepopulatable: Yes
WonderFetch GET arg: “&rights;=”
&rights;=nr-bhl (this signifies that a statement of right to digitize exists at the following URL: http://www.biodiversitylibrary.org/permissions/)

POSSIBLE-COPYRIGHT-STATUS
Description: indicates copyright status, defaults to Not in Copyright
Prepopulatable:yes
Get arg: “&pcs;=”
To make this field blank, because it *is* in copyright, but you have permission to digitize, pass a url encoded space as the parameter, e.g. &pcs;=%20

June 4, 2008by joelrichard
BHL News, Blog Reel, Tech Updates

Better maps! More bibliographic detail!

Portrait version of the Biodiversity Heritage Library logo.

View Full Size Image

Following some excellent suggestions gathered at a recent Encyclopedia of Life meeting, we’ve made changes to our Google Maps browse interface. To recap, we take Library of Congress Subject Headings and geocode and map them using the Google Maps API (details here).

Now that we’re managing nearly 10,000 volumes the standard Google Maps interface was getting cluttered and clunky, so we’ve refined the interface to show smaller points, weight the results using color, and display links to the titles for a given subject heading within the map itself (as demonstrated for “Africa” above). To view the map in full, visit http://www.biodiversitylibrary.org/browse/map.

We made another change based on requests to view full bibliographic details for a scanned title. When we harvest scans from the Internet Archive, we copy the MARCXML for the title to our servers and siphon off just enough of the metadata to facilitate our browse & search capabilities – to pull in the contents of the entire MARCXML would unnecessarily bloat our database with info we don’t expect to search across or expose via browse. But, it’s important data to have in the display, so we’ve skinned the MARCXML using XSLTs provided by the Library of Congress. To view in action, click the “Brief|Detailed|MARC” links at http://www.biodiversitylibrary.org/bibliography/1583, or for any title in our collection.

Finally, we’ve enhanced the display for our Discovered Bibliographies to return results in a more performant way, providing more visual feedback to the user that processes are at work. To view the refined interface, visit the result for Pomatomus saltatrix at http://www.biodiversitylibrary.org/name/Pomatomus_saltatrix.

April 25, 2008by [email protected]
Page 10 of 12« First...«9101112»

Help Support BHL

BHL's existence depends on the financial support of its patrons. Help us keep this free resource alive!

search

About BHL

The Biodiversity Heritage Library (BHL) is the world’s largest open access digital library for biodiversity literature and archives. BHL operates as a worldwide consortium of natural history, botanical, research, and national libraries working together to digitize the natural history literature held in their collections and make it freely available for open access as part of a global “biodiversity community.”

Join Our Mailing List

Sign up to receive the latest news, content highlights, and promotions.

Subscribe Now

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 319 other subscribers

Subscribe to Blog Via RSS

Subscribe to the blog RSS feed to stay up-to-date on all the latest BHL posts.

Access RSS Feed

Inspiring Discovery through Free Access to Biodiversity Knowledge.

The Biodiversity Heritage Library makes it easier than ever for you to access the information you need to study and explore life on Earth…for free, anytime, anywhere.

 

64+ Million Pages of
Biodiversity Literature Online.

EXPLORE

Tools and Services
to Transform Research.

EXPLORE

300,000+
Illustrations on Flickr.

EXPLORE

ABOUT | HARMFUL CONTENT | PRIVACY | SITE MAP | TERMS OF USE