Biodiversity Heritage Library - Program news and collection highlights from BHL
  • Home
  • News
  • Featured Books
    • All Featured Books
    • Book of the Month Series
    • BHL at 20
  • User Stories
  • Campaigns
    • Fossil Stories
    • Garden Stories
    • Monsters Are Real
    • Page Frights
    • Her Natural History
    • Earth Optimism 2020
  • Tech Blog
  • Visit BHL
Home
News
Featured Books
    All Featured Books
    Book of the Month Series
    BHL at 20
User Stories
Campaigns
    Fossil Stories
    Garden Stories
    Monsters Are Real
    Page Frights
    Her Natural History
    Earth Optimism 2020
Tech Blog
Visit BHL
  • Home
  • News
  • Featured Books
    • All Featured Books
    • Book of the Month Series
    • BHL at 20
  • User Stories
  • Campaigns
    • Fossil Stories
    • Garden Stories
    • Monsters Are Real
    • Page Frights
    • Her Natural History
    • Earth Optimism 2020
  • Tech Blog
  • Visit BHL
Biodiversity Heritage Library - Program news and collection highlights from BHL
BHL News, Blog Reel

Internet Archive scanning: Behind the scenes at the Biodiversity Heritage Library

View Full Size Image
Robert Miller at the Internet Archive
San Francisco

These days with digital cameras built into your phone, most everyone has some first hand experience creating digital images. Creating the nearly 40 million page images on BHL is in many ways similar what you probably do on a daily basis with your phone – it’s really just a matter of scale in terms of the technology used for the digitization and the special care taken with the books.

The most important part of the process, however, are the people behind the cameras that take the pictures. Today we want to focus on a key group of those people. As many know, the Internet Archive is a not-for-profit founded by Internet pioneer Brewster Kahle and based in San Francisco. Since 2007, the Internet Archive has been the primary digitization partner of the BHL. The Internet Archive has facilities around the world. Those that the BHL has used include:

Large scanning facilities (with multiple “Scribe” scanning machines) in:

  • San Francisco (Internet Archive)
  • Washington, DC (Library of Congress/FedScan)
  • Northern New Jersey
  • Boston (Boston Public Library)

Smaller facilities in:

  • Ft. Wayne, Indiana
  • University of Illinois, Urbana-Champaign

And “satellites”, single Scribe machines, in:

  • Natural History Museum, London
  • Smithsonian Libraries, Washington, DC

Managing these (and a number more around the world) is the Internet Archive’s Director for Books, Robert Miller. Assisting Robert with that task are the team of scanning center coordinators that manage the workflow and keep our materials safe while putting them online for our worldwide user community.

Pictured below are seven scanning center coordinators gathered for the Internet Archive’s Leaders’ Forum (at the Richmond Branch, SFPL), 24-25 November 2012 (from left):

  • Jude Coelho (Open Library)
  • Shelia De Roche (FedScan)
  • Gemma Waterston (satelites)
  • Stacy Argondizzo (Princeton, NJ)
  • Jeff Sharpe (ACPL Fort Wayne, IN)
  • Stacey A. Seronick (Boston Scan Center)
  • Jesse Bell (San Francisco Center)
View Full Size Image
November 5, 2012by joelrichard
BHL News, Blog Reel

10,000,000 pages!

Portrait version of the Biodiversity Heritage Library logo.

Sometime over the past weekend, the Biodiversity Heritage Library portal loaded it’s 10 millionth page!

Due to the way volumes are ingested from the various scanning centers, it’s a bit tricky to pick which was the EXCACT 10 millionth page, but for the sake of this blog post, I’m going to say that it was one of the pages of Coleopterorum catalogus by Junk and Schenkling. I’m picking this item because, as you taxonomic cognoscenti out there know, beetles (coleoptera) represent, perhaps, the most common type of animal. Indeed, the noted biologist J.B.S. Haldane is reputed to have quipped that, if nothing else, nature reveals that God has “an inordinate fondness for beetles.”

But I digress …

Ten million is a big number, but still a magnitude of difference from the, perhaps, 60-100 million that the Biodiversity Heritage Library hopes to make available in the coming years. Those 10 million pages represent nearly 25,000 volumes (and about 8,700 titles) – a respectable, but not massive library of taxonimic literature. We’ve learned a lot along the way (factoid: taxonomic literature seems to be running at about an average of 417 pages per volume) and are working on ways to expand what we can digitize and how.

Thanks to all the staff in the contributing libraries and our scanning partner (the Internet Archive) for getting us to the 10 million mark.

Details of the “10 millionth page”

Coleopterorum catalogus.

Brief | Detailed | MARC

Brief | Detailed | MARC Brief | Detailed | MARC

By:
Junk, Wilhelm, 1866-1942
Schenkling, Sigmund, 1865-

Publication info:
Berlin :W. Junk,1910-1940.

Call Number:
QL571 .C692

Subjects:
Beetles, Periodicals

Contributing Library:
Smithsonian Institution Libraries

November 24, 2008by
BHL News, Blog Reel, Tech Updates

Harvesting Process from Internet Archive

Portrait version of the Biodiversity Heritage Library logo.

NOTE: Internet Archive has changed their query interface and these instructions are no longer valid.

New instructions are available at:
https://blog.biodiversitylibrary.org/2008/06/updated-harvesting-process-from.html

Overview
The following steps are taken to download data from Internet Archive and host it on the Biodiversity Heritage Library. Diagrams of the process are available in PDF.

  1. Get item identifiers from Internet Archive for items in the “biodiversity” collection that have been recently added/updated.
  2. For each item identifier:
  • Get the list of files (XML and images) that are available for download.
  • Download the XML and image files
  • Download the scan data if it is not included with the other downloaded files
  • Extract the item metadata from the XML files and store it in the import database.
  • Extract the OCR text from the XML files and store it on the file system (one file per page).

For each “approved” item, clean up and transform the metadata into an “importable” format and store the results in the import database.Read all data that is ready for import and insert/update the appropriate data in the production database.Internet Archive Metadata Files
The following table lists the key XML files containing metadata for items hosted by Internet Archive. It is possible that one or more of these files may not exist for an item. However, most items that have been “approved” (i.e. marked as “complete” by Internet Archive) do include each of these files.

Filename

Description

_files.xml

List of files that exist for the given identifier

_dc.xml

Dublin Core metadata. In many cases the data include here overlaps with the data in the _meta.xml file.

_meta.xml

Dublin Core metadata, as well as metadata specific to the item on IA (scan date, scanning equipment, creation date, update date, status of the item, etc)

_metasource.xml

Identifies the source of the item… not much meaningful data here

_marc.xml

MARC data for the item.

_djvu.xml

The OCR for the item, formatted as XML.

_scandata.xml

Raw data about the scanned pages. In combination with the OCR text (_djvu.xml), the page numbers and page types can be inferred from this data. This file may not exist, though in most cases it does. For the most part, only materials added to IA prior to late summery 2007 are likely to be missing this file

scandata.xml

Raw data about the scanned pages. If there is no _scandata file for an item, we look in scandata.zip (via an IA API) for this file, which contains the same information.

Internet Archive Services
Search for Items
Internet Archive items belong to one or more collections. To search a particular Internet Archive collection for items that have been updated between two dates, use the following query:

http://www.archive.org/services/search.php
?query={0}+AND+updatedate:[{1}+TO+{2}] &submit;=submit

where

{0} = name of the Internet Archive collection; in our case, “collection:biodiversity”
{1} = start date of range of items to retrieve
{2} = end date of range of items to retrieve

To limit the item search to a particular contributing institution, modify the query as follows:

http://www.archive.org/services/search.php
?query={0}+AND+updatedate:[{1}+TO+{2}]+AND+contributor:(MBLWHOI Library)
&submit;=submit

To limit the results of the query to a particular number of items, modify the query as follows:

http://www.archive.org/services/search.php
?query={0}+AND+updatedate:[{1}+TO+{2}] &limit;=1000
&submit;=submit

To search for one particular item, use:

http://www.archive.org/services/search.php
?query={0}
&submit;=submit

where

{0} = an Internet Archive item identifier

Download Files
To download a particular file for an Internet Archive item, use the following query:

http://www.archive.org/download/{0}/{1}

where

{0} = an Internet Archive item identifier
{1} = the name of the file to be downloaded

Downloading Files Contained In ZIP Archives
In some cases, a file cannot be downloaded directly, and may instead need to be extracted from a ZIP archive located at Internet Archive. One example of this is the scandata.xml file, which in some cases must be extracted from the scandata.zip file. To do this, two queries must be made. First invoke this query to get the physical file locations (on IA servers) for the given item:

http://www.archive.org/services/find_file.php
?file={0}
&loconly;=1

where

{0} = and Internet Archive item identifier

Then, invoke the second query to extract the scandata.xml file from the scandata.zip file (using the physical file locations returned by the previous query):

http://{0}/zipview.php
?zip={1}/scandata.zip
&file;=scandata.xml

where

{0} = host address for the file
{1} = directory location for the file

Note that the second query can be generalized to extract the contents of other zip files hosted at Internet Archive. The format for the query is:

http://{0}/zipview.php
?zip={1}/{2}
&file;={3}.jpg

where

{0} = host address for the file
{1} = directory location for the file
{2} = name of the zip archive from which to extract a file
{3} = the name of the file to extract from the zip archive

Documentation written by Mike Lichtenberg.

March 14, 2008by oneclickorders
BHL News, Blog Reel

More BHL material online

Portrait version of the Biodiversity Heritage Library logo.

The Biodiversity Heritage Library is currently scanning material five locations around the world. As materials are scanned, they are deposited directly into the Internet Archive repository. The links below will provide you with feed information on materials from the individual scanning centers as it becomes available:

  • All BHL Materials from all Intenet Archive facilities

    • Northeast Regional Scanning Center (Marine Biological Laboratory/Woods Hole Oceanagraphic Institution)
    • Natural History Museum, London
    • Smithsonian Institution Libraries
    • University of Illinois, Urbana-Champaign (Fieldiana only)
  • Missouri Botanical Garden (non-Internet Archive operation; see www.botanicus.org for latest updates)

Internet Archive scanned materials will eventually be processed for delivery through the BHL portal site

October 27, 2007by joelrichard
BHL News, Blog Reel

Scanning starts at Smithsonian Libraries

Portrait version of the Biodiversity Heritage Library logo.

View Full Size ImageOn August 15, 2007, Internet Archive staff (Melissa Bell) began scanning BHL materials at Smithsonian Libraries. The Smithsonian Libraries’ Scribe joins operations in Boston, London, and Urbana-Champaign that are contributing to the Biodiversity Heritage Library.

August 22, 2007by joelrichard

Help Support BHL

BHL's existence depends on the financial support of its patrons. Help us keep this free resource alive!

search

About BHL

The Biodiversity Heritage Library (BHL) is the world’s largest open access digital library for biodiversity literature and archives. BHL operates as a worldwide consortium of natural history, botanical, research, and national libraries working together to digitize the natural history literature held in their collections and make it freely available for open access as part of a global “biodiversity community.”

Join Our Mailing List

Sign up to receive the latest news, content highlights, and promotions.

Subscribe Now

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 319 other subscribers

Subscribe to Blog Via RSS

Subscribe to the blog RSS feed to stay up-to-date on all the latest BHL posts.

Access RSS Feed

Inspiring Discovery through Free Access to Biodiversity Knowledge.

The Biodiversity Heritage Library makes it easier than ever for you to access the information you need to study and explore life on Earth…for free, anytime, anywhere.

 

64+ Million Pages of
Biodiversity Literature Online.

EXPLORE

Tools and Services
to Transform Research.

EXPLORE

300,000+
Illustrations on Flickr.

EXPLORE

ABOUT | HARMFUL CONTENT | PRIVACY | SITE MAP | TERMS OF USE