Biodiversity Heritage Library - Program news and collection highlights from BHL
  • Home
  • News
  • Featured Books
    • All Featured Books
    • Book of the Month Series
    • BHL at 20
  • User Stories
  • Campaigns
    • Fossil Stories
    • Garden Stories
    • Monsters Are Real
    • Page Frights
    • Her Natural History
    • Earth Optimism 2020
  • Tech Blog
  • Visit BHL
Home
News
Featured Books
    All Featured Books
    Book of the Month Series
    BHL at 20
User Stories
Campaigns
    Fossil Stories
    Garden Stories
    Monsters Are Real
    Page Frights
    Her Natural History
    Earth Optimism 2020
Tech Blog
Visit BHL
  • Home
  • News
  • Featured Books
    • All Featured Books
    • Book of the Month Series
    • BHL at 20
  • User Stories
  • Campaigns
    • Fossil Stories
    • Garden Stories
    • Monsters Are Real
    • Page Frights
    • Her Natural History
    • Earth Optimism 2020
  • Tech Blog
  • Visit BHL
Biodiversity Heritage Library - Program news and collection highlights from BHL
BHL News, Blog Reel, Tech Updates

BHL Technical Development: Year in Review

Technical development: year in review. Data quality improvements. Integrating more modern publications. Building for the future.

Building a More Resilient BHL: Improving Accessibility and Expanding Global Reach

What does it take to make millions of pages of biodiversity literature accessible to a global audience? For the Biodiversity Heritage Library (BHL) Technical Team, 2024 was a year of transformative milestones, new innovations, and overcoming challenges—all aimed at strengthening BHL’s mission of advancing biodiversity research.

BHL Data Continues to Improve

This year the BHL Technical Team made dramatic strides in improving BHL’s full-text search precision and retrieval. The Team finalized a two-year OCR reprocessing project aimed at upgrading BHL’s text files, which will improve overall full-text search accuracy by approximately 30% and taxonomic name recognition by 17%. Stay tuned for an in-depth update on the OCR reprocessing project from BHL’s Technical Coordinator, Joel Richard, later this year.

Snippet of improved OCR text and a list of benefits of reprocessing OCR, such as improved spelling, improved name-finding results, and more names revealed.

Reprocessing BHL text files improves full-text search accuracy and taxonomic name recognition. Image Credit: https://doi.org/10.3897/biss.7.112436

Additionally, fruitful collaborations with the Smithsonian Libraries and Archives (SLA) and the Smithsonian Transcription Center (STC) brought another 43,000+ pages of human-transcribed text into BHL from the Smithsonian Field Books Collection, which resulted in the addition of more than 151,000+ scientific names to the BHL search index. For more on Smithsonian transcription efforts, see the blog post entitled The Power of Community Science: How Smithsonian Volunpeers Transform Scientific Field Notes.

These improvements not only expand access points across BHL’s high-value archival materials but also help to interlink those materials with taxonomic databases across the web. BHL’s handwritten materials are notoriously difficult to search due to the fact that text recognition engines do not handle handwritten materials well. Handwritten Text Recognition (HTR) engines are changing this landscape quickly but human-transcribed materials remain the gold standard when it comes to providing the highest level of data quality for unique objects in BHL such as field notes, expedition logs, handwritten tables and other valuable primary source research material.

Four images comparing original handwritten text in BHL, processed via ABBYY Fine reader OCR, processed via Tesseract OCR, and processed via Google Cloud Vision HTR

Handwritten Text Recognition (HTR) engines are a vast improvement over Optical Character Recognition (OCR) engines for processing handwritten materials, as seen from in this example from materials in BHL, but human transcription remains the gold standard in terms of quality. Image Credit: https://doi.org/10.3897/biss.7.112436

Lastly, BHL continues to strengthen its connections with global linked data platforms like Wikidata which has become an ultra rich collaborative data ecosystem that drives major search engines like Google, Duck Duck Go, and others. By adding over 63,000+ Wikidata Q IDs for BHL Titles, we continue to open up new knowledge pathways for researchers to explore and connect information in innovative ways well beyond the BHL website. To understand why persistent identifiers interlinked with knowledge bases on the web are critical for research infrastructure platforms like BHL, check out the related post: BHL is Round Tripping Persistent Identifiers with the Wikidata Query Service.

What are Virtual Items?

Equally transformative to BHL’s data quality gains was the introduction of “virtual items,” a feature that allows born-digital articles and e-content from data feeds like OAI-PMH and Crossref APIs to be grouped into cohesive BHL items. Virtual items enable modern journal articles to be presented alongside BHL’s traditionally digitized historic books and journals. To learn more about Virtual Items, please check out the BHL FAQ. 

Comparison of a traditional digitized volume uploaded to BHL versus a modern born-digital virtual item uploaded to BHL

Although virtual items are generated from new data sources in BHL, the experience will hopefully be a seamless one for the average BHL user. Image Credit: BHL FAQ

Incorporating modern scholarly articles from non-traditional data sources has in the past posed challenges to BHL’s information architecture because it has meant accommodating additional publication standards, proactively and creatively sourcing more granular metadata required at the article level, and creating user interfaces capable of displaying multiple levels of resource description in an intuitive way for our users. By overcoming these challenges BHL is ensuring that the platform can accommodate not only historic literature but also modern biodiversity publications, ultimately helping bridge the gap between past and present biodiversity research.

Data flow diagram listing the data entities, data processes, and users of BHL data

The data flow journey for Virtual Items is actually quite different from traditional BHL content. For this new feature, the BHL Technical Team had to carefully consider the various data sources and how those would flow into the BHL data ecosystem and be presented to users. Image Credit: BHL Technical Team Collaborative Mapping

Facing Challenges, Strengthening Resilience

Despite major wins this past year for the BHL Technical Team, 2024 was not without its hurdles. After BHL completed several server migrations, a series of DDoS attacks in October targeting the Internet Archive (IA), a key partner and long-time host of BHL content, temporarily disrupted access to BHL materials. These events highlighted the importance of building a more resilient infrastructure for BHL, and having failsafes and additional data back-ups planned has been top of mind for the BHL Technical Team.

During the IA outage, we heard from many BHL users how the disruption affected their vital research and how grateful they were when access was restored.

“It is hard to quantify how vital a resource is until it is removed from reach! […] Thank you again for all the wonderful work done by BHL.”

“Life as I know it has come to a standstill. The @internetarchive is offline, which also affects @BioDivLibrary! What do I do!!! 😱”

“A huge amount of literature is only available through either Internet Archive itself or Biodiversity Heritage Library, which is hosted on the Internet archive. Stick in the wheel towards my work.”

“Dear Madam/Sir Really wonderful the BHL is back and provides accessibility to several hundred-year-old literature. Over the past two weeks [I have] not been able to do the majority of my work related to taxonomic curation of plant names of India. Thank you.”

IA continues to be a critical digitization and hosting partner for BHL. This year’s DDoS attacks on IA have only had the effect of helping BHL and IA strengthen infrastructure against future attacks.

Three racks of servers labeled with Internet Archive name and logo

Servers at the Internet Archive headquarters in San Francisco, CA. Image Credit: Jason Scott, Internet Archive | Wikimedia Commons.

Behind the Scenes: Building for the Future

In pursuit of a more resilient infrastructure for BHL, another major milestone was BHL’s inclusion in the AWS Open Data Sponsorship Program. Not only does the program provide BHL with free S3 storage for its data but BHL’s engagement with AWS also marks the beginning of an exploration into how cloud-based technologies can transform BHL’s back-end infrastructure. While BHL currently resides on on-premise storage and systems, cloud services offer new opportunities to rethink and refactor BHL to improve scalability, performance, and cost-efficiency. The benefits of a cloud-based or hybrid BHL are myriad, including complementary storage, faster page-serving speeds in low-bandwidth regions, and greater page persistence critical for biodiversity informatics applications. For the Press Release on BHL’s AWS Open Data Sponsorship, see Biodiversity Heritage Library Datasets Now Openly Accessible on the Amazon Web Services Cloud.

AI generated image of a forest with animals and a cloud with the words BHL in the cloud.

Exploring BHL in the Cloud in 2025 will be a major focus for the BHL Technical Team. We are very excited to see where the journey takes us! Image Credit: Canva AI Image Generator

Looking Ahead

As BHL moves into 2025, the BHL Technical Team’s focus remains on exploring cloud-based solutions and safeguarding BHL as a critical research infrastructure that provides access to over 62+ million pages of knowledge about life on Earth.

By embracing innovation and prioritizing accessibility, BHL will continue to evolve as a cornerstone resource for biodiversity research, empowering scientists and scholars to advance biodiversity knowledge and discovery.

For more on what’s in store for the year ahead, check out the 2025 BHL Technical Priorities.

February 4, 2025by mdimeo
BHL News, Blog Reel, Tech Updates

Biodiversity Heritage Library Datasets Now Openly Accessible on the Amazon Web Services Cloud

Now Available. BHL Datasets Now Openly Accessible on the Amazon Web Services Cloud. Background image of a variety of herons perching on branches

Sixty-two million pages of scientific text, images, and metadata, representing 500 years of biodiversity data will be openly accessible via the Registry of Open Data on AWS

Washington, DC, November 27, 2024 — The BHL Technical Team is thrilled to announce that Biodiversity Heritage Library (BHL) datasets will be openly accessible on the Amazon Web Services (AWS) cloud, thanks to the AWS Open Data Sponsorship Program. Moving BHL data to the cloud allows researchers globally to explore and analyze over 500 years of biodiversity data, enhancing their ability to derive scientific insights from our shared past to inform future global environmental policy.

BHL data is now hosted on AWS, and comprises over 62 million pages of scientific text from the 15th to the 21st centuries. BHL’s vast collection represents an unparalleled biodiversity resource with enormous potential to be used for longitudinal studies and conservation efforts.

Key Use Cases for BHL’s Data:

  • Historical Baseline Data: Spanning centuries, BHL provides a unique record of biodiversity literature, allowing researchers to establish long-term species baselines and track changes in species distributions and abundances over time.
  • Longitudinal Studies: The dataset supports analysis of biodiversity trends and ecosystem responses to environmental changes, offering valuable insights into historical and contemporary biodiversity dynamics.
  • Rare and Endangered Species: BHL’s historical records include species that may no longer exist or have become rare, aiding in the understanding of past biodiversity and informing current conservation efforts.
  • Taxonomic Stability: The collection includes taxonomic descriptions and classifications from various time periods, essential for understanding species relationships.
  • Cultural and Scientific Heritage: Beyond scientific data, BHL preserves historical texts, scientific illustrations, and annotations, enriching our understanding of past scientific practices and societal attitudes toward nature.

Transitioning to AWS services represents a transformative opportunity for BHL. As BHL currently operates outside a cloud environment, this move will optimize scalability, accessibility, performance, and cost-efficiency. AWS’s advanced computing capabilities will enable faster data processing and analysis, while AWS’s cost management tools will in the long-term help reduce infrastructure costs associated with on-premise storage and upgrades.

As a vital component of the global biodata infrastructure, the demand for BHL’s content and services continues to rise. By leveraging AWS services, we can address several critical needs, including:

  • Complimentary storage and access to a comprehensive suite of AWS cloud computing tools.
  • Ensured page persistence and unique page identifiers essential for biodiversity informatics applications.
  • Enhanced page-serving speeds in low-bandwidth regions where biodiversity research is critical.

BHL’s Data Manager, JJ Dearborn remarked, “Collaborating with AWS unlocks the full potential of BHL data. Cloud hosting not only provides an additional data safe harbor but also makes it accessible to researchers worldwide. We’re excited to democratize access to BHL data and see it used to spark innovative cross-disciplinary research in biodiversity informatics.”

The AWS Open Data Sponsorship Program is committed to making high-value datasets freely available to encourage innovation and collaboration. By supporting the free access to a diverse array of big datasets, AWS helps advance research and development in numerous fields.

To explore the BHL dataset and learn more about its content, visit the AWS Data Exchange.


About the Biodiversity Heritage Library

The Biodiversity Heritage Library (BHL) is the world’s largest open access digital library for biodiversity literature and archives. Headquartered at the Smithsonian Libraries and Archives in Washington, D.C., BHL is a global consortium that digitizes and freely shares natural history literature and archives. By providing open access to critical biodiversity knowledge, BHL addresses major research challenges and supports the global scientific community in understanding and conserving Earth’s species amid climate change and extinction crises. BHL also collaborates with rights holders, volunteers, and international experts to enhance content accessibility and interoperability. Since its inception in 2006, BHL has been working towards a shared future where biodiversity knowledge is globally accessible and supports universal bioliteracy.


Download Press Release

November 27, 2024by mdimeo
BHL News, Blog Reel, Tech Updates

BHL Technical Development: Year in Review

BHL data flows diagram listing external data entry, internal data entry, data processes, and internet consumers of data

In 2023, BHL’s Technical Team dedicated significant efforts to improve our data ecosystem, now comprising 61+ million pages of biodiversity literature. Last year’s Technical Priorities underscored BHL’s steadfast commitment to data quality by focusing on both upstream and downstream data flows. Notable milestones include delivering refined taxonomic data to researchers, implementing interface improvements based on user feedback, and forging data pipelines for existing and new downstream data consumers.

Collage of logos from downstream data consumers, including GBIF, figshare, WoRMS, gna, Tropicos, Wikimedia, Crossref, OCLC, BioStor, Catalogue of Life, unpaywall, and DPLA

A selection of BHL’s major downstream data consumers

Global Names Upgrades

A standout achievement in 2023 was the upgrade of BHL’s Global Names taxonomic intelligence tools, marking a significant leap forward in ensuring the accuracy and comprehensiveness of scientific name detection and verification. These upgrades replace the now deprecated GNRD tool and GNresolver services.

BHLIndex is a remarkable improvement over past performance of the Global Names taxonomic intelligence suite. A decade ago, finding and verifying all names in BHL took 45 days of computational processing time—today, the task takes less than 6 hours. Moreover, there is a notable increase in the overall number of name instances in the BHL corpus, underscoring significant improvements made in both speed and comprehensiveness. The Global Names Team observed the following performance improvements:

  • name-finding in 275,000 volumes, 60+ million pages: 2.5 hours;
  • name-verification of 23 million unique name-strings: 3 hours; and
  • preparing a CSV file with 250 million names occurrences/verification records: 40 minutes.

Additionally, Mike Lichtenberg, BHL Lead Developer, found that the BHLIndex upgrade yielded an additional 17,790,590 access points to scientific species names in the BHL corpus!

Graph indicating the number of scientific names found on BHL pages before (245,237,984) and after (263,028,574) implementing the new BHLIndex.

The BHL corpus now contains 17,790,590 additional scientific names.

In tandem with the BHLIndex taxonomic intelligence upgrades, the BHL Technical Team continues its ongoing efforts to enhance the underlying OCR data that powers the Global Names tools. Stay tuned for an update on the OCR Reprocessing Project later this year.

BHL is incredibly grateful to Dr. Dima Mozzherin, Dr. Geoff Ower, and all past and present contributors to the Global Names Architecture applications for their collaborative efforts with the BHL Technical Team to make these new upgrades and an additional 17+ million taxonomic access points available to BHL’s global user base. The continuous improvement of taxonomic data in BHL remains essential, driving key functionalities for our scientific researchers, including taxonomic search, species bibliographies, and interlinkages with other major taxonomic databases on the web.

BHL User Interface Enhancements

BHL has 588 contributors and counting. The addition of a Contributor Facet on the BHL search results page allows users to filter, facet, and hone in on publications from specific BHL contributor collections for the first time. This feature is particularly important to BHL Members and Affiliates who may want to refine their search to see only digitized content from their home collections to answer local research questions.

Example of BHL search results with new Contributor facet in left navigation panel

The contributor facet is a new way to hone in on a specific institutional collection within search results.

In addition to the new contributor facet, BHL has enhanced Supplementary links for Titles by allowing for additional links and incorporating a controlled vocabulary to precisely describe the external resource being linked to.

Example of a BHL title page listing External Resources (a Harvard collection guide for the Walter Deane papers) associated with the title in BHL

Supplementary links allow BHL Partners to curate their collections with relevant links from their home institution like a finding aid, digital exhibition, or collection guide.

Partner curated supplementary links allow BHL to link out to other relevant resources such as a collection guide, digital exhibition, or an archival finding aid. Kudos to the BHL Collection Committee for keeping a pulse on BHL user needs and requesting and scoping the requirements for these new interface features and enhancements.

Forging BHL Data Pipelines

Downstream consumers of BHL data have been a big focus for the BHL Technical Team this past year as well. The Global Biodiversity Information Facility (GBIF) and Crossref play integral roles in advancing biodiversity knowledge by facilitating seamless access to valuable scientific data sourced from BHL. Having accurate BHL data in these repositories furthers the Technical Team’s charge “to develop tools and services to facilitate greater access, interoperability, and reuse of BHL content and data.”

GBIF 

GBIF functions as a global repository for crucial biodiversity data, fostering a more comprehensive understanding of Earth’s biological diversity. Exposing species occurrence data from BHL in GBIF offers a tangible avenue for BHL partners to actively contribute to climate change and biodiversity conservation initiatives.

A significant accomplishment for the BHL Technical Team was the deposit of data at GBIF this year. This milestone marked the culmination of 18 months of intense investigation which was condensed into a 10-minute presentation given at TDWG 2023 entitled “Unearthing the Past for a Sustainable Future: Extracting and transforming data in the Biodiversity Heritage Library for climate action.” 

Cover slide for TDWG 2023 presentation titled Unearthing the past for a sustainable future: extracting and transforming data in the Biodiversity Heritage Library for climate action

The BHL Technical Team presented at TDWG 2023 this year.

The presentation underscores the BHL Technical Team’s dedication to aligning BHL’s data management priorities with international climate-related initiatives, sustainability goals, and most importantly to address global environmental challenges. To learn more check out the conference recording and the related blog post Illuminating BHL’s Dark Data: Citizen Scientists and AI Unlock Key Biodiversity Data in GBIF.

Much of the BHL Technical Team’s investigative work was conducted in collaboration with the BHL Transcription Upload Tool Working Group (TUTWG) who performed crucial tasks such as crowdsourcing and analyzing a Handwritten Text Recognition sample set and providing already transcribed materials and data outputs to process and deposit with GBIF. In addition to their contributions to BHL’s data extraction efforts, the working group members also:

  • Provided user documentation and training on uploading transcription text files to BHL;
  • Updated archival material metadata in BHL for improved findability;
  • Defined an allowable mark-up policy for BHL Partner transcriptions;
  • Scoped requirements for a new OCR text indicator in the BHL interface; and
  • Contributed a valuable GLAM resource to the BHL knowledge base in the form of the Transcription Platform Comparison chart to aid institutions in the overwhelming task of selecting a transcription platform suited to their particular needs.
Transcription platform comparison chart listing four platforms (DigiVol, FromThePage, Wikisource, and Zooniverse) and their key features.

Explore the Transcription Platform Comparison Chart for more details.

A huge shout out goes to all BHL TUTWG members and collaborators for their invaluable contributions.

Crossref

Complementing the new pilot GBIF data pipeline is the improvement of an existing one. Crossref ensures the comprehensive and accurate association of BHL’s bibliographic metadata with digital object identifiers (DOIs), enabling the efficient citation and linking of scholarly content in the modern publishing environment. Ensuring that historic knowledge is served up in modern discovery layers can be incredibly challenging work and we are so grateful to the Persistent Identifier Working Group (PIWG) for their consistent efforts in the form of energy, time, and enduring patience in 2023. The working group in collaboration with the BioStor project has been helping BHL broaden its collection horizons and illuminate BHL dark data since 2012.

In 2023 the group piloted and brought to production the submission of improved DOI data deposits for titles that are part of a monographic series (which provide notoriously difficult bibliographic conundrums for BHL Staff). The use of a richer deposit schema allows for a more complete set of metadata to flow downstream to Crossref.

And More Data 

Since last year, BHL has been openly publishing all of its data on the Smithsonian Figshare Data Repository. The BHL Open Data Collection uses the Data Catalog Vocabulary (DCAT) to describe and document what exactly is in all that data. Due to popular demand, we’ve added two new BHL datasets for our users to experiment with:

Richard, Joel; Dearborn, Jacqueline (2023). BHL Optical Character Recognition (OCR) – Full Text Export (new). Smithsonian Libraries and Archives. Dataset. https://doi.org/10.25573/data.21422193.v13

Dearborn, Jacqueline; Lichtenberg, Mike (2023). BHL Flickr Harvest Data. Smithsonian Libraries and Archives. Dataset. https://doi.org/10.25573/data.24476686.v3

The Flickr harvest data was the product of a workshop called Transforming Biodiversity Heritage Library Images – Data Modeling with OpenRefine. This event was hosted in collaboration with Wikimedia Foundation’s Giovanna Fontenelle and Wikimedian Sandra Fauconnier as part of Wikimedia’s Image Description Month.

BHL Staff, Wikimedians, Flickr staff, and all biodiversity image enthusiasts together are beginning the process of mapping and loading BHL Flickr image data to Structured Data on Commons (SDC), a new Wikimedia Commons initiative. This project was the second top-voted priority from the BHL Wikimedia white paper entitledUnifying Biodiversity Knowledge to Support Life on a Sustainable Planet. 

The Importance of Good Documentation

Fundamental to all technology projects is good documentation, and BHL is no different. The BHL Consortium’s collective endeavor to deliver over 500 years of data about biodiversity to the world relies crucially on the availability of comprehensive documentation about BHL’s extensive data ecosystem. To this end, the BHL Technical Team is working hard to ensure that BHL’s data is “useful, usable, and used” by putting greater emphasis on documentation this year. We want to make sure that all BHL users and collaborators possess the knowledge to effectively leverage BHL data for their research, apps, discovery layers, visualizations and so much more!

BHL data flows diagram listing external data entry, internal data entry, data processes, and internet consumers of data

Helping users understand how data flows in and out of the BHL data ecosystem and sharing it with the world is an overarching goal for the BHL Technical Team.

With over 17 years of development, BHL’s data model has matured significantly, enabling the harvesting, processing, and delivery of rich and robust data about our collections. The BHL Technical Team’s commitment to a high-quality, well documented data ecosystem underpins all BHL initiatives aimed at enhancing the discovery and access to BHL Partner and Contributor collections worldwide.

For a sneak peak of what is in store for the year ahead, check out the 2024 BHL Technical Priorities.

January 22, 2024by mdimeo
BHL News, Blog Reel, Tech Updates

BHL is Round Tripping Persistent Identifiers with the Wikidata Query Service

Diagram of the 8 steps detailed below of the BHL wikidata round trip.

In the Spring of 2022, the BHL Cataloging and Metadata Committee investigated the possibility of harvesting persistent identifiers (PIDs) from Wikidata as part of the group’s longstanding project to disambiguate and deduplicate author records in the BHL database. The motivation behind this one-time experimental data harvest was to see if BHL could:

  1. Enhance BHL author records with additional PID data points;
  2. Improve the committee’s ability to disambiguate author names in the BHL database; and
  3. Respond to an outstanding user request from two of Wikimedia’s super star editors, Siobhan Leachman and Andy Mabbett, to expose BHL’s author data on BHL and include hyperlinks to other authoritative knowledge bases on the web.
How do PIDs work? Users are directed to the URL that is associated with the identifier. If the physical location of the object changes, the URL is the only modification required; the PID itself does not change.

Persistent identifiers can become resolvable links that connect two knowledge bases (like BHL and Wikidata) together.

In particular, Wikimedians wanted to see the Wikidata Q identifier exposed, providing a link to the corresponding creator item record in Wikidata.

There are multiple motivations for undertaking this work. By adding the BHL Creator ID to the corresponding Wikidata item, Wikidata editors help link BHL to the richer biographical data about that person held in Wikidata. The Wikidata item for a person may contain links to their Wikipedia page or to images of the person held in the image repository Wikimedia Commons. Wikidata items also act as identifier hubs and contain links to other databases and identifiers.

Maria Sibylla Merian's wikidata entry, including name variants, biographical information, external sources, an image, and more

Maria Sibylla Merian in Wikidata, rendered with the Reasonator Tool; note the External sources section and BHL’s Creator ID entry.

By adding the BHL Creator ID to this list of identifiers, the Wikidata editor is linking the content held in BHL to the content held in multiple other datasets and repositories.

These extra author data points provide Wikimedians and BHL catalogers with crucial clues that aid in name disambiguation. In particular, hyperlinks to other knowledge bases are incredibly valuable because they lead to new knowledge pathways that help confirm a person’s identity in a complex game of “Who’s Who?”

Workflow Overview

The diagram below illustrates the experimental data pipeline from BHL to Wikidata and back.

Diagram of the 8 steps detailed below of the BHL wikidata round trip.

The BHL to Wikidata Round Trip Overview

The BHL to Wikidata Round Trip Steps

  1. BHL Creator ID and record are created when digital content is harvested into BHL and/or articles are defined in BHL.
  2. Wikimedians record BHL Creator IDs as a statement in corresponding Wikidata items via various community workflows. (See: Mix’n’match deep dive below and/or read about QuickStatements for two ways volunteers are populating Wikidata with BHL’s Creator IDs).
  3. BHL subsequently harvests additional metadata from Wikidata for any author item where a BHL Creator ID statement has been added.
  4. Data outputs are analyzed for quality and breadth.
  5. SPARQL queries and data outputs are iteratively refined to match BHL’s requirements until a quality dataset can be generated for import.
  6. Clean data is imported into the BHL database.
  7. The new author record sidebar displays data on BHL.
  8. PIDs are converted into resolvable URIs, opening up new research pathways for BHL’s users.

Quick note: In Wikidata BHL author records are represented by the Wikidata BHL Creator ID; the presence of this Wikidata property in a Wikidata item provides a powerful connection point that can be used later by BHL to ingest more information about any named entity. A popular author matching tool in Wikidata is Mix’n’match.

Getting BHL Authors in Wikidata with Mix’n’match

Mix’n’match is a Wikidata tool that empowers editors to match Wikidata items to entries in other databases. One of the datasets in Mix’n’match is the BHL Creator ID dataset.

Summary from the mix'n'match tool for the BHL creator ID dataset

A summary of the completeness of matching the BHL Creator ID dataset to Wikidata as of January 2023

When the BHL Creator ID dataset is uploaded into Mix’n’match, an algorithm undertakes fuzzy name matching. The tool then suggests a preliminary match to a Wikidata item. Wikidata editors working on the dataset have the choice of whether to confirm the suggested match or to remove it.

Sample matches suggested by the mix'n'mtch tool

Examples of possible matches suggested by the Mix’n’match algorithm with the BHL creator in green and the Wikidata item in blue.

If the Wikidata editor confirms the match, the BHL Creator ID is automatically added to the Wikidata item for the creator. If the editor rejects the match, the editor can add a correct Wikidata item ID, create a new Wikdiata item if the creator is not yet in Wikidata or, in some cases, decide that the BHL Creator ID isn’t applicable to Wikidata. If the editor simply rejects the match without further action, the BHL Creator ID will be added to the unmatched portion of the dataset.

Examples of confirmed and unconfirmed matches in the Mix'n'match tool

Results after the Wikidata editor Ambrosia10 has confirmed or removed the Mix’n’match algorithm suggestions.

The act of linking the BHL Creator ID to the Wikidata item also helps to disambiguate the creator. It removes much of the uncertainty about who the creator is. Names are not unique but by linking identifiers to biographical data, editors can ensure that similarly named people or people whose names have changed over time are all associated with the correct Wikidata item. This addresses the age-old problem of how to assign correct attribution to the right person.

This work also assists BHL. By ensuring BHL Creator ID’s are matched to the Wikidata item, the Wikidata editor can assist BHL in weeding out the duplicate entries for creators in BHL’s database. After matching has taken place, BHL is also able to ingest any of the other identifiers listed on the creator’s Wikidata item, thus enriching the metadata held in BHL’s database.

Currently there are over 230,000 entries in the BHL Creator ID dataset. Of these, just over 41,000 have been matched to Wikidata items. There is a long way to go before this dataset is complete. In the spirit of “many hands make light work,” BHL encourages Wikidata editors to work alongside BHL staff who were recently trained and certified in Wikidata Advanced Concepts by Wiki Education. Our collaborative work will help increase the number of Creator IDs in Wikidata.

Summary from the mix'n'match tool for the BHL creator ID dataset

BHL Creator ID dataset completeness statistics as of January 2023.

Once BHL Creator ID statements were recorded on Wikidata items, either through manual edits or workflows like Mix’n’match or QuickStatements, then BHL created custom SPARQL queries using the Wikidata Query Service to output data of interest.

Example of a SPARQL query using the wikidata query service

The SPARQL query language is used to query and return data from semantic databases like Wikidata.

Once the data was brought back, the BHL Cataloging and Metadata Committee discussed and reviewed the records with the aim of bringing key data points back into the BHL database.

Sample data returned from the SPARQL query, including item ID, itemLabel, BHL_creator_ID, date_of_death, VIAF_ID, Library_of_congress_authority_ID, ORCID_ID, ISNI, date_of_birth

Snippet of raw data brought back by the BHL Cataloging and Metadata Committee’s SPARQL query.

In total, the BHL Cataloging and Metadata Committee was able to “round trip” 88,507 persistent identifiers (PIDs) associated with BHL Creator records from the following authoritative knowledge bases:

  • Virtual International Authority File (VIAF)
  • Wikidata
  • Social Networks and Archival Context (SNAC)
  • ResearchGate
  • ORCID
  • Library of Congress Authorities

There are still more PIDs out there to gather but practicality calls for moderation – moderation of the quantity that the Wikidata Query Service can bring back and the amount feasible to curate in BHL. After many discussions, the above list was chosen by narrowing down the PIDs that seem most appropriate in the BHL context. A comprehensive policy on Uniform Resource Identifiers (URIs) in the BHL ecosystem is now being drafted by BHL committees.

Additionally, in post-work analysis, the Committee did find that the data modeling of corporate names in Wikidata differs from the library world, and pulling in identifiers for corporate names was not worth the effort (at least at this time).

Below is an example of Académie des sciences (France), which receives multiple item records in library databases like VIAF for each name change. However, in Wikidata these name changes are collapsed and will appear on one item record; this divergence means that one-to-one matches are not possible without a lot of manual clean-up from BHL Staff.

Example of the Wikidata entry for VIAF ID for Académie des sciences

Divergent data models: Wikidata corporate names collapse name changes.

It’s important to note that data modeling in Wikidata is in its nascent stages and librarians should have a vested interest in shaping the best practices of the community so all of humanity can model the world’s knowledge together. As we have seen with this one-time experiment, there is little to lose, and so much to gain!

The Result: A New Author Record Sidebar for BHL

Thanks to the feedback from the Wikimedia community, BHL’s Lead Developer created an Author Record sidebar. Click on any BHL author name on the BHL website and the author details will appear on the right-hand side. Below is an example of Australian paleontologist and ornithologist, Patricia Vickers-Rich in BHL. 

Graphic example of an author record with detailed sidebar and the pages to which it links out, such as ORCID, Library of Congress Authorities, Wikidata, SNAC

A New Wikimedian Driven Feature: The BHL Author Record Sidebar!

The goal of this new feature was to expose existing and recently harvested identifiers as resolvable links that could open up critical knowledge pathways for BHL’s users on their information journeys. Other key data points are provided such as preferred name forms, entity type, and alternative name forms that give additional context for disambiguation research.

A quote from Siobhan Leachman on BHL's new author sidebar: Adding these identifiers to the author sidebar makes my life so much easier. At a glance, I can quickly confirm the identify of creators. The author sidebar or lack thereof also highlights whether further work is needed in wikidata to link these BHL creators to their identifiers."

Next Steps for BHL Committees?

Please let us know in the comments what you think of this experiment. Should BHL pursue similar persistent identifier harvests for other entity types? Or perhaps, a recurring harvest for BHL Authors?

Additional BHL entities:

  • Titles
  • Scientific Names
  • Subjects
  • Media (illustrations, scientific plates, photos)

Let us know! Your feedback is crucial to BHL’s evolution as a biodiversity knowledge base.

If you are new to Wikidata, Mix’n’match is a great place to start. There are many tutorials and resources available on the web including this great YouTube introduction on how to use the tool.

Additionally, a workflow for BHL articles has been piloted and is currently underway thanks to BHL’s Persistent Identifier Working Group. For more details on how to roundtrip BHL Article Q identifiers, please refer to the group’s documentation at: Wikidata:WikiProject_BHL/Projects

BHL committees and working groups are actively scoping many projects. Sign-up here to get involved!

Related Posts

For more information on the assignment of DOIs, a very important type of persistent identifier specific to scholarly publications, check out Nicole Kearney’s blog post What Is BHL’s New Persistent Identifier Working Group DOI’ng?

For a list of all of the new features and data quality improvements BHL made in 2022, check out the post BHL Technical Development: Year in Review

February 15, 2023by mdimeo
BHL News, Blog Reel, Tech Updates

BHL Technical Development: Year in Review

Technical development: year in review. Critical updates. New features and enhancements. Data quality improvements.

Critical Upgrades

For BHL, 2022 was a year to focus on critical upgrades for the BHL platform to ensure the sustainability of our services for our global users. Although BHL’s basic technical infrastructure remains the same, consisting of years of refinement, knowledge, and reliability, a few updates were definitely in order. BHL major upgrades included:

  • Upgraded .NET Framework to version 6 (see What’s New in Version 6? for a rundown of the benefits for BHL)
  • Upgraded Elasticsearch from 5.4.2 to v.7.17 (search functions remain the same)
  • Converted SOAP services to REST (two SOAP services used by the BHL website and utility applications have been converted to REST services implemented in .NET 6)
  • Upgraded Global Names GNFinder tool from 0.19.5 to 1.0.0

Most of these upgrades were “behind-the-scenes” work and would not be noticeable to a majority of our users. However, keeping up with these important enhancements is a crucial component of any technology project. To learn more about the technology stack that drives BHL, click on the above links for a deep dive into some of the products and services we utilize to make BHL a reality.

New Features and Enhancements

In 2022, BHL deployed three new exciting user-driven features on the BHL website: the BHL text source indicator, the author record sidebar, and pre-generated article PDFs.

Text Source Indicator

Many of our users would like to know where BHL’s text output comes from. The Show Text tab now gives BHL users more information about whether the content has been manually transcribed by a human versus automatically generated by a machine via an OCR engine like Tesseract or ABBYY FineReader.

Screenshot of the BHL bookviewer with sample text highlighting the new text source indicator with uncorrected OCR

An example of uncorrected text from an OCR engine with lower quality text output

The Text Source Indicator helps our users understand and determine what quality they should be expecting from the item they are viewing.

Screenshot of the BHL bookviewer with text source indicator showing manual transcription

An example of manually transcribed text with higher quality text output

Many thanks to the BHL Transcription Upload Tool Working Group (TUTWG) for this important feature request and providing the critical feedback to ensure the Text Source Indicator is useful to BHL’s global user base.

Author Record Sidebar

BHL is now exposing author data to help open-up new research pathways beyond the BHL portal. The Author Record sidebar was born out of a response to user requests from two Wikimedians, Siobhan Leachman and Andy Mabbett, who provided critical feedback that they would like to see Wikidata Q numbers exposed alongside author names to further facilitate downstream name disambiguation work.

Screenshot of author search results showing new author details and links to authoritative knowledge bases such as wikidata

Author data stored in BHL is now displayed with author search results. Data includes clickable URIs to authoritative name registries including Wikidata.

A special thanks to Siobhan, Andy, and the BHL Cataloging and Metadata Committee for their hard work and dedication to curating and disambiguating author names and helping to realize the new author records display in BHL. Look out for a forthcoming blog post from the Committee, during International Love Data Week 2023, which will detail how the group was able to harvest 88,000+ author identifiers from Wikidata and bring them back into BHL’s database.

Pre-generated Article PDFs

In 2022, the BHL Technical Team announced pre-generated PDFs for articles defined in BHL. While BHL had previously allowed users to download custom PDFs, the benefits of pre-generated article PDFs are manifold:

  • No waiting.
  • No selecting pages.
  • The PDF contains embedded, searchable, copy-paste-able text.
  • The PDF contains rich XMP-based metadata about the article.

The addition of these pre-generated PDFs article landing pages in BHL also allows content aggregators like Unpaywall to find BHL’s open access versions of paywalled literature and serve BHL content up to their user base via their free browser extension.

Unpaywall browser extension indicating an article is available for free elsewhere on the web

BHL’s article PDFs are now surfaced via the Unpaywall browser extension.

Many thanks to BHL’s Persistent Identifier Working Group (PIWG) and their ongoing communications with the Unpaywall Development Team to ensure BHL content is surfaced via the Unpaywall browser extension.

Data Quality Improvements

The BHL Technical Team and all of BHL’s Committees and Working Groups care passionately about the quality of BHL’s data. To support the various priorities around data management in BHL’s Strategic Plan, BHL announced the new position of BHL Data Manager in 2022. In this new role, the Data Manager leads the effort to develop and implement a comprehensive view of how BHL collections can be optimized to support the interoperability of BHL data in the larger biodiversity community.

Like any big data repository there are anomalies, errors, and omissions in BHL data as a result of aggregating records from hundreds of contributors and technology projects distributed all over the world. The way we collectively categorize, classify, and describe our materials can vary vastly across so many collaborating organizations that comprise the BHL network. In 2022, the data quality improvements of note included reprocessing BHL’s OCR files, updating the material type facet, and creating an open data collection with a new 40GB OCR text export.

Reprocessing BHL’s OCR files

BHL digitization partner, the Internet Archive (IA), has upgraded their OCR engine from ABBYY FineReader to the Tesseract Open Source OCR engine. BHL is working with IA to reprocess some of BHL’s oldest content with the newest available version of Tesseract OCR. Approximately 120,000 BHL items are eligible to be upgraded and at our current rate of progress of approximately 100 items per day, it will take close to two years to complete the BHL OCR reprocessing project. For more information check out OCR Improvements: An Early Analysis by BHL’s Technical Coordinator, Joel Richard.

Material Type Facet Update

Archival materials digitized for BHL are not like books and journals. The content is non-standard and thus often presents additional challenges for BHL staff. Prior to sending materials for digitization, archival materials must be intellectually organized, cataloged, and frequently undergo extensive preservation treatment. In doing this work, BHL’s catalogers and archivists collaborate to answer some very complex questions:

  • Are these materials considered published or unpublished? 
  • Are they in-copyright or out-of-copyright? 
  • How should these materials be divided and what should comprise an intellectual unit? 

The answers to the above truly vary and depend on the expertise, local digitization workflows, and the technological constraints of an organization. The sheer diversity of cataloging and digitization methods across the BHL partner network is truly astonishing, but it does not always result in precise search and retrieval for our users.

Example of search results demonstrating over 6,800 archival materials available

The Material facet allows users to filter on different categories of content. This data update ensures that the Archival Material facet brings in a more complete set of search results.

To facilitate better search precision, BHL has decided to update all archival materials to be classified as “Manuscript language material (Archival material)” which has resulted in 4,637 records being updated in the BHL database. For our users, who rely on search facets to drill down on relevant content, this ultimately means more relevant results, more archival content, and more happy BHL users! Many thanks to BHL’s Transcription Upload Tool Working Group (TUTWG) for their work and group analysis to make this major data update a reality.

Open Data Collection with new 40GB OCR Text Export

In the Fall of 2022, the BHL Technical Team worked with Keri Thompson, Data Management Specialist from Smithsonian’s OCIO’s Research Computing Office, to publish BHL data exports as FAIR data. BHL’s data has always been available for download on our Developer and Data Tools page but as an open access leader, BHL has decided to revamp each data export to include a DOI, a data dictionary, and multitude of citation formats. Most importantly, the data is now hosted on the Smithsonian Institution’s Figshare instance which serves as a “an open platform for hosting and sharing the raw material of Smithsonian research.” BHL’s data has also been cataloged using the W3C data description standard: Data Catalog Vocabulary (DCAT). In using DCAT, BHL data is now compliant and could be harvested by federally mandated data repositories like Data.gov.

A sample of data sets available from the BHL Open Data Collection on figshare

Access BHL Open Data Collection on Smithsonian Institution’s Data Repository

Lastly, a very exciting data export has been added to the BHL Open Data Collection. Data miners everywhere, this one’s for you:

BHL Optical Character Recognition (OCR) – Full Text Export (https://doi.org/10.25573/data.21422193.v4)

The BHL OCR full text export is our largest data export, representing the full textual corpus for all items in BHL. It is so large that we really struggled to find a place to host the file. The file is updated on a monthly basis and posted to the link above.

BHL Technical Team Goes Agile

The BHL Technical Team had a transformational year. Not only did the Team complete critical upgrades, deliver exciting new features for BHL’s users, and improve overall data quality; we also adopted the popular Agile project management methodology and SCRUM framework. Agile is used by 3 out of every 4 technology projects to improve team communications, find new efficiencies, adapt to change, and maximize resources.

Bar graphs representing the status of tickets across 12 work areas

Agile Dashboard, reporting on the status of BHL Technical Epics. Click here for an interactive deep dive into our work.

We are still refining new workflows but overall, the “BHL Agile experiment” has been a boon for BHL technical development. Our Team now has a birds-eye view of all of the work in the pipeline. Knowing where we’ve been, where we are going, and focusing on continuous iterative improvement has helped us deliver more meaningful value to BHL Partners and users while furthering our collective mission of making biodiversity literature openly available to the world.

Kudos to the entire BHL Technical Team and Secretariat for a productive 2022!

A special thanks to BHL’s Lead Developer, Mike Lichtenberg, and Technical Coordinator, Joel Richard, for all of your hard work, knowledge, and dedication to the BHL platform, data, and our users. BHL is so incredibly lucky to benefit from your many talents!

For a sneak peak of what is in store for the year ahead, check out the 2023 BHL Technical Priorities.

January 31, 2023by mdimeo
BHL News, Blog Reel, Tech Updates

New Article PDF Content Available

A sample of a printed page of a book with highlighted text superimposed over the printed text

The BHL Tech Team is pleased to announce a new form of content available in BHL: Article PDFs. While this may not sound like anything new, after all, we have had a tool to download PDF content for some time, this update changes both how the PDFs are created and maintained, and how BHL is viewed by content aggregators on the internet, most notably Unpaywall.

Screenshot of the Download PDF icon.

The new Download PDF icon

How to use it? While browsing an article, you will now see a Download PDF icon below the View Article link on the right side of the page. Clicking the link will immediately download the PDF to your computer (or view it in your web browser, depending on your settings.)

The benefits of the immediate download are:

  • No waiting.
  • No selecting pages.
  • The PDF contains embedded, searchable, copy-paste-able text.†
  • The PDF contains rich XMP-based metadata about the article.

An important change to note is that when viewing an article within an item at BHL, the Download Contents > Download Article link will now direct the visitor’s browser to the new PDFs for immediate download. This is a departure from what we had before in that the pages of the article were pre-selected for download and the visitor was then required to complete the process and wait for the PDF to be generated. We expect the new PDFs to be an improvement for our visitors who come to download articles. View the How do I download a PDF of an article? FAQ for simple download instructions.

Visitors to BHL are still able to manually create PDFs using the Download Contents > Select Pages to Download feature. This feature has not been removed, but it still means that it takes some time to create those PDFs and email the person when the PDF is ready. This option is useful for articles that have not been indexed in BHL, and therefore do not have a Download Article link. View the How do I generate a custom PDF of selected pages from the book? FAQ for complete instructions.

The most important feature of the new Article PDFs is the embedded text† within the document. The select-able text is an invisible text layer in the PDF, but it appears when you select or search for text within the document:

A sample of a printed page of a book with highlighted text superimposed over the printed text

An example of select-able text in an Article PDF.

While the appearance of the text may look… less than ideal, rest assured that the text can be copied out intact and used in another program. Example:

It is perhaps needless for me here to reiterate the great importance
of arriving at a final decision as to the real nature of
the haloliranic forms, for it will be obvious that if they have
nothing to do with the normal fresh-water series, and are to
be regarded as the remnant of an ancient sea, our views
respecting the past history of the African interior must be
greatly changed.

Other, less visible benefits to the PDFs are that they are directly linked from the citation_pdf_url meta-tag on the web page which makes them more findable by Google Scholar, Unpaywall, and potentially other aggregators.

For the technical-minded, the PDFs (many tens of thousands of them) are created in advance and stored on BHL’s servers. Changes to data within BHL will cause the PDF to be updated automatically, usually within several hours.

We hope that this is a welcome addition to BHL.

 

† – Please note that the text is only as good as the OCR that was generated for the text on the page. While the OCR text is probably very good for the prose sections of an article, titles, tables, and other special content may not appear as expected.

March 14, 2022by Sheila Rabun
BHL News, Blog Reel, Tech Updates

Updates to Bibliography Pages in BHL

Screenshot of bibliography pages in BHL with and without tabs.

We have updated the bibliography pages in BHL to streamline the presentation of information about and metadata export options for content in the Library.

Previously, bibliographic details and export options were available through different tabs on title and part pages. These tabs have now been removed, and all bibliographic information is consolidated into a single display.

Screenshot showing BHL bibliography pages before and after.

The various tabs on title and part bibliography pages have now been consolidated into a single display.

BHL’s metadata export options have also been relocated. BHL offers metadata exports in MODS, BibTex, and RIS formats. MODS is an XML-based bibliographic description schema used in a variety of library applications. BibTex and RIS are bibliographic citation files that are compatible with a variety of citation management tools. The MODS file is a title-level download. The BibTex and RIS files are item-level downloads.

The MODS download is available at the bottom of the title and part bibliography pages.

Screenshot of a title page in BHL with the "Download MODS" button circled.

MODS download on the title and page bibliography pages.

The BibTex and RIS downloads are available in two places:

1) Under the volume or part details on bibliography pages.

Screenshot of the citation download options in the BHL website.

BibTex and RIS downloads on title (left) and part (right) bibliography pages.

2) Under the “Download Contents” menu in the book viewer, via the “Download Citation” option.

Screenshot of the BHL book viewer with the citation download options displayed.

Download citation options in the BHL book viewer.

As part of this update, we have removed the direct Mendeley import from BHL, as the generic RIS and BibTex formats are compatible with a variety of citation management softwares including Mendeley.

Details about our metadata export services are also available in the FAQ on the BHL About site.

February 11, 2021by michelle.underhill
Page 1 of 41234»

Help Support BHL

BHL's existence depends on the financial support of its patrons. Help us keep this free resource alive!

search

About BHL

The Biodiversity Heritage Library (BHL) is the world’s largest open access digital library for biodiversity literature and archives. BHL operates as a worldwide consortium of natural history, botanical, research, and national libraries working together to digitize the natural history literature held in their collections and make it freely available for open access as part of a global “biodiversity community.”

Join Our Mailing List

Sign up to receive the latest news, content highlights, and promotions.

Subscribe Now

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 319 other subscribers

Subscribe to Blog Via RSS

Subscribe to the blog RSS feed to stay up-to-date on all the latest BHL posts.

Access RSS Feed

Inspiring Discovery through Free Access to Biodiversity Knowledge.

The Biodiversity Heritage Library makes it easier than ever for you to access the information you need to study and explore life on Earth…for free, anytime, anywhere.

 

64+ Million Pages of
Biodiversity Literature Online.

EXPLORE

Tools and Services
to Transform Research.

EXPLORE

300,000+
Illustrations on Flickr.

EXPLORE

ABOUT | HARMFUL CONTENT | PRIVACY | SITE MAP | TERMS OF USE