Biodiversity Heritage Library - Program news and collection highlights from BHL
  • Home
  • News
  • Featured Books
    • All Featured Books
    • Book of the Month Series
    • BHL at 20
  • User Stories
  • Campaigns
    • Fossil Stories
    • Garden Stories
    • Monsters Are Real
    • Page Frights
    • Her Natural History
    • Earth Optimism 2020
  • Tech Blog
  • Visit BHL
Home
News
Featured Books
    All Featured Books
    Book of the Month Series
    BHL at 20
User Stories
Campaigns
    Fossil Stories
    Garden Stories
    Monsters Are Real
    Page Frights
    Her Natural History
    Earth Optimism 2020
Tech Blog
Visit BHL
  • Home
  • News
  • Featured Books
    • All Featured Books
    • Book of the Month Series
    • BHL at 20
  • User Stories
  • Campaigns
    • Fossil Stories
    • Garden Stories
    • Monsters Are Real
    • Page Frights
    • Her Natural History
    • Earth Optimism 2020
  • Tech Blog
  • Visit BHL
Biodiversity Heritage Library - Program news and collection highlights from BHL
Blog Reel, Featured Books

The Treasure Between the Covers: Making BHL’s Articles Discoverable and Citable

A Gif zooming into a page of many tiny images of scanned pages

This post is part of BHL at 20: Treasures from the Biodiversity Heritage Library, a series contributed by members of the BHL community, highlighting remarkable works from across the collection in celebration of its 20th anniversary.

One of the greatest challenges for a digital library, especially one as large as the Biodiversity Heritage Library, is simply finding the content you are after. Recently, I made a website to give a sense of this challenge. The website features about 200,000 pages of content from BHL Australia, a small fraction of what is in BHL overall, but already it’s something of a challenge to find specific items you might be after. If you were looking for a particular article, how would you find it?

A Gif zooming into a page of many tiny images of scanned pages

Interactive browser of BHL Australia

Articles, articles, articles

For most scientists, the article is the fundamental unit of research, not the journal title, and not a journal volume. The article is what we download as a PDF, what we store in our reference managers, and what gets cited. At the outset BHL did not have articles, so over a decade ago I set about developing a tool to find those. This tool became BioStor, which was described in a paper in 2011 (Extracting scientific articles from a large digital archive: BioStor and the Biodiversity Heritage Library). The basic idea behind BioStor is to take information about an article, such as journal, volume, pages, and year, and then try and find that article in BHL.

A diagram with text and large blue arrows showing the mapping between BHL and articles

Mapping journal, volume, and pagination from an article to BHL.

In principle this seems straightforward, but often the vagaries of metadata complicate the task. The image below shows some of the issues encountered with the journal Ibis. The source of metadata for the articles was CrossRef, via a commercial publisher (Wiley). You might expect this data to be high quality, but it contains errors such as bad character encoding. To further complicate things, Wiley decided to renumber all the volumes of the journal, so that the original volume information we see in BHL (such as series 2, volume 1) bears little relation to what is in CrossRef (volume 7, issue 1).

Four examples of citations, with some text highlighted in orange. The Crossref and BHL logos are on the right.

Matching CrossRef metadata for an article in Ibis to BHL.

The list of metadata messes like this is almost endless. There are journals that have more than one numbering system for the same volumes (e.g., Annali del Museo civico di storia naturale Giacomo Doria where the same item is both series 3, volume 7 and volume 47), there are multiple abbreviations for the same journal, and there are journals with multiple names (e.g., title 8097 is “Annuaire du Musée zoologique de l’Académie des sciences de St. Pétersbourg”, “ЕЖЕГОДНИКЬ ЗООЛОГИЧЕСКАГО МУЗЕЯ ИМПЕРАТОРСКОЙ АКАДЕМІЙ НАУКЬ”, and “Ezhegodnik Zoologicheskago muzeia …”).

A further complication is that a scanned item in BHL may contain several issues or volumes, each with its own set of overlapping page numbers, which means we have to decide which page “1” is the page 1 that matches the article we are searching for. Once we solve all that, we encounter further problems. Perhaps the most challenging is pagination. In most modern articles, the page range, e.g. 1–5, completely encompasses the article, including figures, charts, illustrations, etc. But for the older literature this is often not the case. Typesetting text and reproducing plates were different processes, and hence the plates might be disconnected from the article (often appearing at the end of a volume). This means that extracting, say pages 1–5, from BHL is no guarantee that you have the whole article.

A good deal of code in BioStor is trying to make sense of matching article metadata to BHL items, finding the correct page to match to, and extracting the set of pages that correspond to that article, as well as providing tools to manually correct metadata and add missing pages (for example, the plates mentioned above). Hence the process of finding articles is, at best, semiautomated.

BioStor old and new

The original BioStor website dates back to 2009, and looked something like this:

A screenshot from a website showing an image viewer with a yellow page from an book and fields of metadata.

Original BioStor website.

This site could display individual articles, and you could edit metadata. For a variety of reasons, it was no longer feasible to host this at the university where I was based, so I split the website into two versions. The original site now runs only on my laptop, and I use it to process files and locate articles in BHL. The new version runs in the cloud and features a cleaner interface, along with much better search. Below is the same article in the current BioStor.

A screenshot of a website showing a search bar, bibliographic data, and page images.

Current BioStor website.

Once articles are discovered using the old BioStor, they get pushed to the public version of BioStor at https://biostor.org. This website is also the point of contact between BioStor and BHL: each day BHL runs an automated process which asks BioStor whether it has any new articles, and, if the answer is yes, it fetches those and adds them to BHL (in BHL articles are referred to as either “parts” or “segments”). The end result is that articles defined in BioStor now become visible in the Table of Contents in BHL.

A screenshot of the BHL website showing an image viewer with a yellow page of a journal, bibliographic metadata and a highlighted article in a contents page.

BioStor article displayed in BHL.

One advantage of having a separate project such as BioStor is that I can use it to experiment with different ways to view BHL content. For example, BioStor looks for geographic coordinates (latitude and longitude) in the OCR text for each article. Any pairs of coordinates that it finds get stored in a map, which you can browse. In the diagram below we have selected a small region in the centre of the map, on the right you can see a list of articles about that area.

A map of the island of Sulawesi with small red dots sprinkled across it. There is a pink rectangle over a cluster of red dots.

Maps showing localities on the island of Sulawesi that are mentioned in BioStor articles.

Identifiers

BioStor has been running since 2009. In that time it has contributed over 260,000 articles to BHL, making it the single largest source of BHL “parts”. Having articles is nice, but even better is having articles with persistent, citable identifiers, such as DOIs. The Persistent Identifier Working Group has been working to add DOIs to BHL content, especially “parts”. This work has focused on two kinds of DOIs. The first are existing DOIs minted, for example, by commercial publishers. BioStor adds a lot of articles using CrossRef metadata, so we get these “for free” (there are other sources of DOIs that BioStor uses, but that is another story). Why does it matter to have external DOIs for BHL content? Well, many of these articles are free in BHL but behind a paywall on the publisher’s website. Services such as Unpaywall can link existing DOIs to free versions of the corresponding article, and BHL is one of Unpaywall’s providers.

But the more exciting (and onerous) task is minting new DOIs for articles in BHL, so that BHL is the version of record for that content. This has several implications. It means that BHL is effectively a publisher, and has the responsibility to maintain access to this content in perpetuity. It also changes the way we think about adding articles. For example, most of my work with BioStor has been opportunistic – I’m working on a taxonomic database, I see that there are some papers that should be in BHL, find them, then add them to BioStor so future BHL users can find those articles. But once we start creating DOIs, the goal is quite different: you want to get every article in the journal that is in BHL, and mint DOIs for all of them. While this appeals to a completionist mindset, it does mean getting metadata for every article before you can add DOIs.

Luckily, the hard work in minting DOIs has a striking payback, we can see how many times articles in BHL are cited in the scientific literature. The last time the results were analyzed, BHL articles had been cited some 74,446 times! Without BHL these publications would appear as simple text strings in the literature cited, now they are first-class digital citizens with clickable DOI links.

Metadata matters

By now it is obvious that the way BioStor finds articles depends on having good quality metadata for articles (or chapters), which it then attempts to locate in BHL. The lack of freely accessible metadata is a major impediment to increasing the rate at which articles are added. In the past I have made extensive use of taxonomic databases as a source of bibliographic data (see my BioNames project, for example). Yet the quality of citations in these databases is often poor. I have also made extensive use of sources such as CrossRef, which covers articles that have been assigned DOIs by that agency, and also data provided by volunteers, such as those working with Nicole Kearney (thank you Bob Griffith and Heidi Griffith!). Another major source of data has come from scraping the web, a time-consuming process that is becoming increasingly difficult to do as the web becomes increasingly closed under the onslaught of AI bots (see also Joel Richard’s blog post A Brief Bit on BHL Battling a Barrage of Bots).

There is a clear need for a free and open bibliographic database. The nearest we have is OpenAlex, whose tagline is “All the world’s research, connected and open.” Sadly this is still more of an aspirational goal rather than a fact: a lot of taxonomic literature is not in OpenAlex. Perhaps it is time, therefore, to revive “CiteBank”, which was an early BHL project to collect bibliographic metadata. If we had a comprehensive database of the taxonomic and related literature, locating articles in BHL would be a much easier task.

Machines reading

BioStor’s method of finding articles works, but it is not the only way we could locate articles. Instead of relying on external sources of metadata, what if we could simply have a computer read the volume and extract the articles automatically? Early attempts to do this for BHL content were not particularly successful, see for example A metadata generation system for scanned scientific volumes. But the advent of large language models (LLMs) and AI chatbots has dramatically changed the way we can tackle finding articles in BHL. In my own work I routinely use AI to extract articles in bulk from a scanned volume. Typically the approach involves finding tables of contents in the scanned volume, using AI to parse that into structured data, then finding the corresponding pages in the volume, checking that they match the table of contents, and then using AI to extract bibliographic data (e.g., article title, authors, etc.). The result of this process is a data file that gets fed into BioStor, so that articles get found and added to BHL in the usual way. It is not bulletproof, and AI can quite happily make mistakes, but in my experience it works well.

But the holy grail would be to simply point an AI at a volume and it would identify and extract all the articles, find any stray plates, and present the results to BHL. Given the spectacular advances in OCR text and understanding document layout in recent years, perhaps there will be a point where BioStor can gracefully retire from the scene. Its hundreds upon hundreds of lines of regular expressions and special-case hacks quietly gathering dust in a GitHub repo while machines of loving grace read BHL for us.

From the Biodiversity Heritage Library

As BHL celebrates twenty years of open biodiversity knowledge, this post reminds us that access depends not only on digitised pages, but on the tools, metadata, identifiers, and infrastructure that make them discoverable and citable. With your support, BHL can continue strengthening the systems that connect biodiversity literature to the researchers, communities, and future discoveries that depend on it.Orange button with a heart icon

June 16, 2026by nkearney
BHL News, Blog Reel, Tech Updates

BHL is Round Tripping Persistent Identifiers with the Wikidata Query Service

Diagram of the 8 steps detailed below of the BHL wikidata round trip.

In the Spring of 2022, the BHL Cataloging and Metadata Committee investigated the possibility of harvesting persistent identifiers (PIDs) from Wikidata as part of the group’s longstanding project to disambiguate and deduplicate author records in the BHL database. The motivation behind this one-time experimental data harvest was to see if BHL could:

  1. Enhance BHL author records with additional PID data points;
  2. Improve the committee’s ability to disambiguate author names in the BHL database; and
  3. Respond to an outstanding user request from two of Wikimedia’s super star editors, Siobhan Leachman and Andy Mabbett, to expose BHL’s author data on BHL and include hyperlinks to other authoritative knowledge bases on the web.
How do PIDs work? Users are directed to the URL that is associated with the identifier. If the physical location of the object changes, the URL is the only modification required; the PID itself does not change.

Persistent identifiers can become resolvable links that connect two knowledge bases (like BHL and Wikidata) together.

In particular, Wikimedians wanted to see the Wikidata Q identifier exposed, providing a link to the corresponding creator item record in Wikidata.

There are multiple motivations for undertaking this work. By adding the BHL Creator ID to the corresponding Wikidata item, Wikidata editors help link BHL to the richer biographical data about that person held in Wikidata. The Wikidata item for a person may contain links to their Wikipedia page or to images of the person held in the image repository Wikimedia Commons. Wikidata items also act as identifier hubs and contain links to other databases and identifiers.

Maria Sibylla Merian's wikidata entry, including name variants, biographical information, external sources, an image, and more

Maria Sibylla Merian in Wikidata, rendered with the Reasonator Tool; note the External sources section and BHL’s Creator ID entry.

By adding the BHL Creator ID to this list of identifiers, the Wikidata editor is linking the content held in BHL to the content held in multiple other datasets and repositories.

These extra author data points provide Wikimedians and BHL catalogers with crucial clues that aid in name disambiguation. In particular, hyperlinks to other knowledge bases are incredibly valuable because they lead to new knowledge pathways that help confirm a person’s identity in a complex game of “Who’s Who?”

Workflow Overview

The diagram below illustrates the experimental data pipeline from BHL to Wikidata and back.

Diagram of the 8 steps detailed below of the BHL wikidata round trip.

The BHL to Wikidata Round Trip Overview

The BHL to Wikidata Round Trip Steps

  1. BHL Creator ID and record are created when digital content is harvested into BHL and/or articles are defined in BHL.
  2. Wikimedians record BHL Creator IDs as a statement in corresponding Wikidata items via various community workflows. (See: Mix’n’match deep dive below and/or read about QuickStatements for two ways volunteers are populating Wikidata with BHL’s Creator IDs).
  3. BHL subsequently harvests additional metadata from Wikidata for any author item where a BHL Creator ID statement has been added.
  4. Data outputs are analyzed for quality and breadth.
  5. SPARQL queries and data outputs are iteratively refined to match BHL’s requirements until a quality dataset can be generated for import.
  6. Clean data is imported into the BHL database.
  7. The new author record sidebar displays data on BHL.
  8. PIDs are converted into resolvable URIs, opening up new research pathways for BHL’s users.

Quick note: In Wikidata BHL author records are represented by the Wikidata BHL Creator ID; the presence of this Wikidata property in a Wikidata item provides a powerful connection point that can be used later by BHL to ingest more information about any named entity. A popular author matching tool in Wikidata is Mix’n’match.

Getting BHL Authors in Wikidata with Mix’n’match

Mix’n’match is a Wikidata tool that empowers editors to match Wikidata items to entries in other databases. One of the datasets in Mix’n’match is the BHL Creator ID dataset.

Summary from the mix'n'match tool for the BHL creator ID dataset

A summary of the completeness of matching the BHL Creator ID dataset to Wikidata as of January 2023

When the BHL Creator ID dataset is uploaded into Mix’n’match, an algorithm undertakes fuzzy name matching. The tool then suggests a preliminary match to a Wikidata item. Wikidata editors working on the dataset have the choice of whether to confirm the suggested match or to remove it.

Sample matches suggested by the mix'n'mtch tool

Examples of possible matches suggested by the Mix’n’match algorithm with the BHL creator in green and the Wikidata item in blue.

If the Wikidata editor confirms the match, the BHL Creator ID is automatically added to the Wikidata item for the creator. If the editor rejects the match, the editor can add a correct Wikidata item ID, create a new Wikdiata item if the creator is not yet in Wikidata or, in some cases, decide that the BHL Creator ID isn’t applicable to Wikidata. If the editor simply rejects the match without further action, the BHL Creator ID will be added to the unmatched portion of the dataset.

Examples of confirmed and unconfirmed matches in the Mix'n'match tool

Results after the Wikidata editor Ambrosia10 has confirmed or removed the Mix’n’match algorithm suggestions.

The act of linking the BHL Creator ID to the Wikidata item also helps to disambiguate the creator. It removes much of the uncertainty about who the creator is. Names are not unique but by linking identifiers to biographical data, editors can ensure that similarly named people or people whose names have changed over time are all associated with the correct Wikidata item. This addresses the age-old problem of how to assign correct attribution to the right person.

This work also assists BHL. By ensuring BHL Creator ID’s are matched to the Wikidata item, the Wikidata editor can assist BHL in weeding out the duplicate entries for creators in BHL’s database. After matching has taken place, BHL is also able to ingest any of the other identifiers listed on the creator’s Wikidata item, thus enriching the metadata held in BHL’s database.

Currently there are over 230,000 entries in the BHL Creator ID dataset. Of these, just over 41,000 have been matched to Wikidata items. There is a long way to go before this dataset is complete. In the spirit of “many hands make light work,” BHL encourages Wikidata editors to work alongside BHL staff who were recently trained and certified in Wikidata Advanced Concepts by Wiki Education. Our collaborative work will help increase the number of Creator IDs in Wikidata.

Summary from the mix'n'match tool for the BHL creator ID dataset

BHL Creator ID dataset completeness statistics as of January 2023.

Once BHL Creator ID statements were recorded on Wikidata items, either through manual edits or workflows like Mix’n’match or QuickStatements, then BHL created custom SPARQL queries using the Wikidata Query Service to output data of interest.

Example of a SPARQL query using the wikidata query service

The SPARQL query language is used to query and return data from semantic databases like Wikidata.

Once the data was brought back, the BHL Cataloging and Metadata Committee discussed and reviewed the records with the aim of bringing key data points back into the BHL database.

Sample data returned from the SPARQL query, including item ID, itemLabel, BHL_creator_ID, date_of_death, VIAF_ID, Library_of_congress_authority_ID, ORCID_ID, ISNI, date_of_birth

Snippet of raw data brought back by the BHL Cataloging and Metadata Committee’s SPARQL query.

In total, the BHL Cataloging and Metadata Committee was able to “round trip” 88,507 persistent identifiers (PIDs) associated with BHL Creator records from the following authoritative knowledge bases:

  • Virtual International Authority File (VIAF)
  • Wikidata
  • Social Networks and Archival Context (SNAC)
  • ResearchGate
  • ORCID
  • Library of Congress Authorities

There are still more PIDs out there to gather but practicality calls for moderation – moderation of the quantity that the Wikidata Query Service can bring back and the amount feasible to curate in BHL. After many discussions, the above list was chosen by narrowing down the PIDs that seem most appropriate in the BHL context. A comprehensive policy on Uniform Resource Identifiers (URIs) in the BHL ecosystem is now being drafted by BHL committees.

Additionally, in post-work analysis, the Committee did find that the data modeling of corporate names in Wikidata differs from the library world, and pulling in identifiers for corporate names was not worth the effort (at least at this time).

Below is an example of Académie des sciences (France), which receives multiple item records in library databases like VIAF for each name change. However, in Wikidata these name changes are collapsed and will appear on one item record; this divergence means that one-to-one matches are not possible without a lot of manual clean-up from BHL Staff.

Example of the Wikidata entry for VIAF ID for Académie des sciences

Divergent data models: Wikidata corporate names collapse name changes.

It’s important to note that data modeling in Wikidata is in its nascent stages and librarians should have a vested interest in shaping the best practices of the community so all of humanity can model the world’s knowledge together. As we have seen with this one-time experiment, there is little to lose, and so much to gain!

The Result: A New Author Record Sidebar for BHL

Thanks to the feedback from the Wikimedia community, BHL’s Lead Developer created an Author Record sidebar. Click on any BHL author name on the BHL website and the author details will appear on the right-hand side. Below is an example of Australian paleontologist and ornithologist, Patricia Vickers-Rich in BHL. 

Graphic example of an author record with detailed sidebar and the pages to which it links out, such as ORCID, Library of Congress Authorities, Wikidata, SNAC

A New Wikimedian Driven Feature: The BHL Author Record Sidebar!

The goal of this new feature was to expose existing and recently harvested identifiers as resolvable links that could open up critical knowledge pathways for BHL’s users on their information journeys. Other key data points are provided such as preferred name forms, entity type, and alternative name forms that give additional context for disambiguation research.

A quote from Siobhan Leachman on BHL's new author sidebar: Adding these identifiers to the author sidebar makes my life so much easier. At a glance, I can quickly confirm the identify of creators. The author sidebar or lack thereof also highlights whether further work is needed in wikidata to link these BHL creators to their identifiers."

Next Steps for BHL Committees?

Please let us know in the comments what you think of this experiment. Should BHL pursue similar persistent identifier harvests for other entity types? Or perhaps, a recurring harvest for BHL Authors?

Additional BHL entities:

  • Titles
  • Scientific Names
  • Subjects
  • Media (illustrations, scientific plates, photos)

Let us know! Your feedback is crucial to BHL’s evolution as a biodiversity knowledge base.

If you are new to Wikidata, Mix’n’match is a great place to start. There are many tutorials and resources available on the web including this great YouTube introduction on how to use the tool.

Additionally, a workflow for BHL articles has been piloted and is currently underway thanks to BHL’s Persistent Identifier Working Group. For more details on how to roundtrip BHL Article Q identifiers, please refer to the group’s documentation at: Wikidata:WikiProject_BHL/Projects

BHL committees and working groups are actively scoping many projects. Sign-up here to get involved!

Related Posts

For more information on the assignment of DOIs, a very important type of persistent identifier specific to scholarly publications, check out Nicole Kearney’s blog post What Is BHL’s New Persistent Identifier Working Group DOI’ng?

For a list of all of the new features and data quality improvements BHL made in 2022, check out the post BHL Technical Development: Year in Review

February 15, 2023by mdimeo
BHL News, Blog Reel, Tech Updates

Information about Upcoming Changes to BHL API

Portrait version of the Biodiversity Heritage Library logo.

The BHL API will be updated on 25 July 2016 to support changes to the BHL site. These changes will accommodate identifying additional Contributors for Items and Parts of items.

First are changes to the API that may affect your existing processes.

The Contributor and ContributorID elements in the result sets of API methods that return “Part” information will move. ContributorID will be included as a PartIdentifier in the Identifiers list. Contributor will be included in a new Contributors list.

These changes are being made to accommodate more than one contributor per part.  For example, if one institution researches/compiles the data and a second institution facilitates the inclusion of that data in BHL, both institutions may be listed as a contributor.  Initially, no more than two contributors per part will be allowed, but by adopting these changes to API responses we allow for additional (unlimited, actually) contributors in the future.

Here is an example that shows how the API responses are changing.  The examples shown here are an output of the GetPartMetadata method, and have been abbreviated for clarity.

Original API Results – highlighted elements are being moved:

<Response>
 <Status>okStatus>
 <Result>
 <PartUrl>
        http://www.biodiversitylibrary.org/part/1
 PartUrl>
 <PartID>1PartID>
 <ItemID>22498ItemID>
 <StartPageID>3190776StartPageID>
 <SequenceOrder>1SequenceOrder>
 <Contributor>BioStorContributor>
 <ContributorID>4443ContributorID>
 <GenreName>ArticleGenreName>
 <Title>
      Notes on certain species of Tetragnatha 
      (Araneae, Argiopidae) in Central America 
      and Mexico
 Title>
 <ContainerTitle>BrevioraContainerTitle>
 <Volume>67Volume>
 <Date>1957Date>
 <PageRange>1--4PageRange>
 <StartPageNumber>1StartPageNumber>
 <EndPageNumber>4EndPageNumber>
 <Authors> [...] Authors>
 <Subjects />
 <Identifiers>
 <PartIdentifier>
 <IdentifierName>ISSNIdentifierName>
 <IdentifierValue>0006-9698IdentifierValue>
 PartIdentifier>
 Identifiers>
 <Pages> [...] Pages>
 <RelatedParts />
 Result>
Response>

Updated API Results – Highlighted elements are the new locations of the moved data:

<Response>
 <Status>okStatus>
 <Result>
 <PartUrl>
        http://www.biodiversitylibrary.org/part/969
 PartUrl>
 <PartID>1PartID>
 <ItemID>22498ItemID>
 <StartPageID>3190776StartPageID>
 <SequenceOrder>1SequenceOrder>
 <GenreName>ArticleGenreName>
 <Title>
        Notes on certain species of Tetragnatha
        (Araneae, Argiopidae) in Central America
        and Mexico
 Title>
 <ContainerTitle>BrevioraContainerTitle>
 <Volume>67Volume>
 <Date>1957Date>
 <PageRange>1--4PageRange>
 <StartPageNumber>1StartPageNumber>
 <EndPageNumber>4EndPageNumber>
 <Authors> [...] Authors>
 <Contributors>
 <Contributor>
 <ContributorName>BioStorContributorName>
 Contributor>
 Contributors>
 <Subjects />
 <Identifiers>
 <PartIdentifier>
 <IdentifierName>BioStorIdentifierName>
 <IdentifierValue>4443IdentifierValue>
 PartIdentifier>
 <PartIdentifier>
 <IdentifierName>ISSNIdentifierName>
 <IdentifierValue>0006-9698IdentifierValue>
 PartIdentifier>
 Identifiers>
 <Pages> [...] Pages>
 <RelatedParts />
 Result>
Response>

Additionally, there will be two additions to the Item metadata.

The new data elements are: RightsHolder and ScanningInstitution.
These will optionally be displayed if there is data for the relevant organizations.


<Response>
 <Status>okStatus>
 <Result>
 <ItemID>59382ItemID>
 <PrimaryTitleID>20770PrimaryTitleID>
 <ThumbnailPageID>17605914ThumbnailPageID>
 <Source>Internet ArchiveSource>
 <SourceIdentifier>bulletin5160hatcSourceIdentifier>
 <Volume>v.51-60 1898-99Volume>
 <Year/>
 <Contributor>
      UMass Amherst Libraries (archive.org)
 Contributor>
 <RightsHolder>
 Biodiversity Heritage Library
 RightsHolder>
 <ScanningInstitution>
 Smithsonian Institution Libraries
 ScanningInstitution>
 <Sponsor>UMass Amherst LibrariesSponsor>
    [...]
 Result>
Response>

Detailed documentation for the BHL APIs is available at http://www.biodiversitylibrary.org/api2/docs/docs.html.  It will be updated to reflect these changes after they are moved into production on 25 July 2016.

Go to http://www.biodiversitylibrary.org/getapikey.aspx to get an API key for BHL.

Information about other BHL developer tools can be found at http://biodivlib.wikispaces.com/Developer+Tools+and+API.

If you have questions, please feel free to submit feedback via this form.

July 12, 2016by Pedro Gonzalez Fernandez
BHL News, Blog Reel

Digging into Data Challenge Winners Announced

Mining Biodiversity (MiBIO): innovative computational techniques to mine BHL texts

On January 15, 2014, ten international research funders from four countries jointly announced the winners of the third Digging into Data Challenge, a competition to develop new insights, tools and skills in innovative humanities and social science research using large-scale data analysis.

Fourteen international teams representing Canada, the Netherlands, the United Kingdom, and the United States will receive grants totaling approximately $5.1 million to investigate how computational techniques can be applied to “Big Data”.  Their work will address the evolving nature of humanities and social sciences research, which often relies on massive multisource datasets.  Each team represents collaborations among scholars, scientists, and information professionals from leading universities and libraries in Europe and North America.

The first round of the Digging into Data Challenge was held in 2009 and the second in 2011. Previous Digging into Data research projects have received international attention.  The 2013 winners were selected from among 69 applications and address a wide variety of topics.  Among the sponsoring bodies, the Institute of Museum and Library Services (IMLS) announced, on January 15, 2014, the winners of the third Digging into Data Challenge that they will fund.  IMLS’s contribution of $424,591 supports the American researchers from three of the mentioned fourteen international teams: the Missouri Botanical Garden, the University of Wisconsin-Madison, and the University of South Carolina.

As part of an international team involving partners in UK, Canada and the USA, Missouri Botanical Garden was awarded $175,000 by the IMLS to lead the USA members of the Mining Biodiversity Project team in applying innovative computational techniques to biodiversity texts and help improve the way we view, access and share the knowledge contained with in the Biodiversity Heritage Library.

The Mining Biodiversity (MiBIO) Project
Principal Investigators:

  • (United Kingdom) Sophia Ananiadou, Director of the National Centre for Text Mining and Professor in the School of Computer Science at the University of Manchester;
  • (Canada) Anatoliy Gruzd, Director of the SocialMedia Lab and Associate Professor at Dalhousie University;
  • (USA) William Ulate Rodríguez, Global Coordinator & Technical Director for US/UK of the Biodiversity Heritage Library and Senior Project Manager of the Center for Biodiversity Informatics at the Missouri Botanical Garden.

View Full Size Image

The MiBIO project is an international collaboration (UK, Canada, USA), leveraging large-scale data resources collected and curated by BHL members in US and worldwide, a digital portal developed and maintained by the BHL Technical staff at Missouri Botanical Garden and managed by the BHL Secretariat at the Smithsonian Institution Libraries, text mining resources developed by the UK partner (NaCTeM), and social media, visualization and text cleaning methodologies developed by the Canadian partner (Dalhousie).

The project’s prime objective is to promote the efficient access to and sharing of large-scale biodiversity legacy literature by the worldwide community of science libraries, museums and the interested public The project will integrate novel text mining methods, visualization, crowdsourcing and social media into the BHL to provide a semantic search system aimed to transform BHL into a next-generation social digital library resource that facilitates the study and discussion (via social media integration) of legacy science documents on biodiversity by a worldwide community.

By promoting the development of capabilities that will foster collaboration among researchers from the fields of History of Science, Environmental History, Environmental Studies, Library and Information Science and Social Media, the project hopes to make a significant impact on the above disciplines by (1) enriching a large-scale library, i.e., the BHL, via innovative application of Text Mining techniques to produce semantic metadata and a term inventory, (2) providing improved access to biodiversity-related digital artifacts via an enhanced search engine and visualization of results and (3) stimulating increased collaboration, interaction and sharing of information among BHL users via the social media environment.

View Full Size Image

The Institute of Museum and Library Services (IMLS)is the primary source of federal support for the nation’s 123,000 libraries and 17,500 museums. Our mission is to inspire libraries and museums to advance innovation, lifelong learning, and cultural and civic engagement. Our grant making, policy development, and research help libraries and museums deliver valuable services that make it possible for communities and individuals to thrive. To learn more, visit www.imls.gov and follow us on Facebook and Twitter.

In addition to the U.S. Institute of Museum and Library Services, the sponsoring funding bodies include:

  • Arts & Humanities Research Council (United Kingdom),
  • Economic & Social Research Council (United Kingdom),
  • Canada Foundation for Innovation (Canada),
  • National Endowment for the Humanities (United States),
  • Natural Sciences and Engineering Research Council (Canada),
  • National Science Foundation (United States),
  • Netherlands Organisation for Scientific Research in collaboration with The Netherlands eScience Center (NLeSC) (Netherlands),
  • Social Sciences and Humanities Research Council (Canada).

The U.K. charity Jisc will be providing professional program management in the progression of the United Kingdom projects.

The Biodiversity Heritage Library (BHL) is a consortium of natural history and botanical libraries that cooperate to digitize and make accessible the legacy literature of biodiversity held in their collections and to make that literature available for open access and responsible use as a part of a global “biodiversity commons.” The BHL consortium works with the international taxonomic community, rights holders, and other interested parties to ensure that this biodiversity heritage is made available to a global audience through open access principles. In partnership with the Internet Archive and through local digitization efforts, the BHL has digitized millions of pages of taxonomic literature, representing tens of thousands of titles and over 100,000 volumes.

The Missouri Botanical Garden’s mission is “to discover and share knowledge about plants and their environment in order to preserve and enrich life.” Today, 153 years after opening, the Missouri Botanical Garden is a National Historic Landmark and a center for science, conservation, education and horticultural display.

The Center for Biodiversity Informatics (CBI) at the Missouri Botanical Garden seeks to provide innovative technology solutions to the global community of life science scholars in order to mobilize, integrate, and repatriate data about the world’s biodiversity.

Find more information at:

MiBIO Project Website: www.nactem.ac.uk/DID-MIBIO/

Official Winners Announcement: http://www.diggingintodata.org/Home/AwardRecipientsRound32013/tabid/201/Default.aspx

IMLS Press Release: http://www.imls.gov/5.1_million_awarded_for_delving_into_big_humanities_and_social_science_data.aspx
Dalhousie University Press Release: http://www.dal.ca/news/2014/01/22/natural-history-for-the-digital-age.html

View Full Size Image

This project is made possible by a grant from the Institute for Museum and Library Services [Grant number LG-00-14-0032-14].

January 28, 2014by suzanne.chase
BHL News, Blog Reel

NESCent-EOL-BHL Research Sprint

Portrait version of the Biodiversity Heritage Library logo.

We invite participants for an event that will pioneer the mining of the Encyclopedia of Life (http://eol.org) and the Biodiversity Heritage Library (http://www.biodiversitylibrary.org) to address outstanding and novel questions about the ecology and evolution of biodiversity. We aim to identify questions and data for which biologists may lack informatics skills and resources to address or analyze successfully; and symmetrically, to guide informaticians to pressing ecological and evolutionary questions. We seek to make actual discoveries through joint activities and to test the “computability” of major biodiversity databases.

We invite proposals for synthetic research on any aspect of biological science that will leverage EOL and BHL resources.  Successful applicants will be matched with an informatician, in 2 person teams, and receive support to travel to and work on-site at the National Evolutionary Synthesis Center for 4 days.

Applications for participation will be evaluated on the extent to which they

  • address an important and outstanding biological question,
  • are “risky” endeavors but with a reasonable chance of success,
  • reflect EOL’s mission to empower scientific research by providing comprehensive and trusted biodiversity data and/or NESCent’s scientific mission to advance research that addresses fundamental questions in evolutionary science. Both organizations promote the integration of methods, concepts, and data within and across disciplines (for more on the context and a classification of synthesis in evolutionary science please read Linking Big: The Continuing Promise of Evolutionary Synthesis),
  • provide evidence that sufficient data are available to tackle the question,
  • provide evidence that appropriate analytical tools are available or will be developed during the event,
  • generate products that typically fall into three broad categories (in order of importance):
    • Synthetic papers and reviews,
    • Software or mathematical tools that solve a major analytical problem.
    • New open data for Encyclopedia of Life allowing others to build on your foundation.

We will not support collection of original data or field research, but encourage the mining of public and private databases such as EOL and BHL.  NESCent is committed to making data, databases, software and other products that are developed as part of NESCent activities available to the broader scientific community. Applicants should review the Data And Software Policy for NESCent.

Proposals will be evaluated in terms of both the scientific value of the project and the qualifications of the applicant.  Visitors will receive support for travel to and from NESCent, lodging, and a per diem for meals not provided.

Before you Apply

All applicants are encouraged to contact Craig McClain, Assistant Director of Science of NESCent, or Cynthia Parr, Chief Scientist of Encyclopedia of Life, for feedback on project ideas. Proposals will be considered through November 15.  The event will likely take place February 3-7, 2014. Please review our Reporting Requirements, Travel Guidelines, and Data and Software Policy before applying.

Proposals Guidelines

Proposals are short and not to exceed 2 single-spaced (12-pt type) pages, plus a 2-page CV.   Proposals should be organized as follows:

  • Title (80 characters max)
  • Name and contact information
  • Project Summary (250 words max)
  • Public Summary (250 words max) – written for the public and visible on the NESCent web site
  • Introduction and Goals – A statement of the outstanding question being addressed and a concise review of the concept and the literature to place the project in context.
  • Proposed Activities – A clear statement of any specific data (include citations or urls) and analytical tools that will be required for the project.  The proposal should also include a clear statement on how synthesis will occur. Letters of support are required from the proprietor of datasets, analytical tools, or software not publically available or owned by the applicant or managed by EOL and BHL.
  • Rationale for support – Why can this activity be most effectively conducted at NESCent and with the Encyclopedia of Life or BHL.
  • Anticipated IT Needs – Briefly describe any needs for IT support that are important to the success of the proposed project. Please indicate whether development of software will be required. Also, briefly describe your plans to make resulting data and software available; including any conditions that might limit your ability to make these available. Please remember you need not possess the informatics skills to address the questions but do need to identify what skills are lacking.  You may propose an informatician to collaborate with.  Prior contact and letter of support with the informatician is encouraged.  If no teammate is suggested EOL/NESCent will work to identify proper support for you.
  • Anticipated Results – include a clear statement of anticipated papers, data and software products, and anticipated public release of data and products
  • Short CV of the applicant(s) (2 pages)
  • Proposal Submission

    Proposals will be accepted in digital format only as a single pdf file. Graphics should be embedded directly into the proposal document. Note that proposals should be submitted as a single pdf file including all of the components listed above, including the CV. Email your proposal to Craig McClain and Cynthia Parr, by November 15, 2013.

    Data, Software and Publication Policy

    The open availability of data, software source code, methods, and results is good scientific practice and a key ingredient of synthetic research. NESCent expects that all data and software created through NESCent-sponsored activities be made publicly available no later than one year after the conclusion of the NESCent award, or immediately upon publication of an associated article, whichever comes earlier. For more information please visit our Data, Software and Publication Policy.

    Funding for this opportunity is generously provided by the Richard Lounsbery Foundation.

    September 30, 2013by joelrichard

    Help Support BHL

    BHL's existence depends on the financial support of its patrons. Help us keep this free resource alive!

    search

    About BHL

    The Biodiversity Heritage Library (BHL) is the world’s largest open access digital library for biodiversity literature and archives. BHL operates as a worldwide consortium of natural history, botanical, research, and national libraries working together to digitize the natural history literature held in their collections and make it freely available for open access as part of a global “biodiversity community.”

    Join Our Mailing List

    Sign up to receive the latest news, content highlights, and promotions.

    Subscribe Now

    Subscribe to Blog via Email

    Enter your email address to subscribe to this blog and receive notifications of new posts by email.

    Join 319 other subscribers

    Subscribe to Blog Via RSS

    Subscribe to the blog RSS feed to stay up-to-date on all the latest BHL posts.

    Access RSS Feed

    Inspiring Discovery through Free Access to Biodiversity Knowledge.

    The Biodiversity Heritage Library makes it easier than ever for you to access the information you need to study and explore life on Earth…for free, anytime, anywhere.

     

    64+ Million Pages of
    Biodiversity Literature Online.

    EXPLORE

    Tools and Services
    to Transform Research.

    EXPLORE

    300,000+
    Illustrations on Flickr.

    EXPLORE

    ABOUT | HARMFUL CONTENT | PRIVACY | SITE MAP | TERMS OF USE