Biodiversity Heritage Library - Program news and collection highlights from BHL
  • Home
  • News
  • Featured Books
    • All Featured Books
    • Book of the Month Series
    • BHL at 20
  • User Stories
  • Campaigns
    • Fossil Stories
    • Garden Stories
    • Monsters Are Real
    • Page Frights
    • Her Natural History
    • Earth Optimism 2020
  • Tech Blog
  • Visit BHL
Home
News
Featured Books
    All Featured Books
    Book of the Month Series
    BHL at 20
User Stories
Campaigns
    Fossil Stories
    Garden Stories
    Monsters Are Real
    Page Frights
    Her Natural History
    Earth Optimism 2020
Tech Blog
Visit BHL
  • Home
  • News
  • Featured Books
    • All Featured Books
    • Book of the Month Series
    • BHL at 20
  • User Stories
  • Campaigns
    • Fossil Stories
    • Garden Stories
    • Monsters Are Real
    • Page Frights
    • Her Natural History
    • Earth Optimism 2020
  • Tech Blog
  • Visit BHL
Biodiversity Heritage Library - Program news and collection highlights from BHL
BHL News, Blog Reel, Tech Updates

BHL Technical Development: Year in Review

Technical development: year in review. Data quality improvements. Integrating more modern publications. Building for the future.

Building a More Resilient BHL: Improving Accessibility and Expanding Global Reach

What does it take to make millions of pages of biodiversity literature accessible to a global audience? For the Biodiversity Heritage Library (BHL) Technical Team, 2024 was a year of transformative milestones, new innovations, and overcoming challenges—all aimed at strengthening BHL’s mission of advancing biodiversity research.

BHL Data Continues to Improve

This year the BHL Technical Team made dramatic strides in improving BHL’s full-text search precision and retrieval. The Team finalized a two-year OCR reprocessing project aimed at upgrading BHL’s text files, which will improve overall full-text search accuracy by approximately 30% and taxonomic name recognition by 17%. Stay tuned for an in-depth update on the OCR reprocessing project from BHL’s Technical Coordinator, Joel Richard, later this year.

Snippet of improved OCR text and a list of benefits of reprocessing OCR, such as improved spelling, improved name-finding results, and more names revealed.

Reprocessing BHL text files improves full-text search accuracy and taxonomic name recognition. Image Credit: https://doi.org/10.3897/biss.7.112436

Additionally, fruitful collaborations with the Smithsonian Libraries and Archives (SLA) and the Smithsonian Transcription Center (STC) brought another 43,000+ pages of human-transcribed text into BHL from the Smithsonian Field Books Collection, which resulted in the addition of more than 151,000+ scientific names to the BHL search index. For more on Smithsonian transcription efforts, see the blog post entitled The Power of Community Science: How Smithsonian Volunpeers Transform Scientific Field Notes.

These improvements not only expand access points across BHL’s high-value archival materials but also help to interlink those materials with taxonomic databases across the web. BHL’s handwritten materials are notoriously difficult to search due to the fact that text recognition engines do not handle handwritten materials well. Handwritten Text Recognition (HTR) engines are changing this landscape quickly but human-transcribed materials remain the gold standard when it comes to providing the highest level of data quality for unique objects in BHL such as field notes, expedition logs, handwritten tables and other valuable primary source research material.

Four images comparing original handwritten text in BHL, processed via ABBYY Fine reader OCR, processed via Tesseract OCR, and processed via Google Cloud Vision HTR

Handwritten Text Recognition (HTR) engines are a vast improvement over Optical Character Recognition (OCR) engines for processing handwritten materials, as seen from in this example from materials in BHL, but human transcription remains the gold standard in terms of quality. Image Credit: https://doi.org/10.3897/biss.7.112436

Lastly, BHL continues to strengthen its connections with global linked data platforms like Wikidata which has become an ultra rich collaborative data ecosystem that drives major search engines like Google, Duck Duck Go, and others. By adding over 63,000+ Wikidata Q IDs for BHL Titles, we continue to open up new knowledge pathways for researchers to explore and connect information in innovative ways well beyond the BHL website. To understand why persistent identifiers interlinked with knowledge bases on the web are critical for research infrastructure platforms like BHL, check out the related post: BHL is Round Tripping Persistent Identifiers with the Wikidata Query Service.

What are Virtual Items?

Equally transformative to BHL’s data quality gains was the introduction of “virtual items,” a feature that allows born-digital articles and e-content from data feeds like OAI-PMH and Crossref APIs to be grouped into cohesive BHL items. Virtual items enable modern journal articles to be presented alongside BHL’s traditionally digitized historic books and journals. To learn more about Virtual Items, please check out the BHL FAQ. 

Comparison of a traditional digitized volume uploaded to BHL versus a modern born-digital virtual item uploaded to BHL

Although virtual items are generated from new data sources in BHL, the experience will hopefully be a seamless one for the average BHL user. Image Credit: BHL FAQ

Incorporating modern scholarly articles from non-traditional data sources has in the past posed challenges to BHL’s information architecture because it has meant accommodating additional publication standards, proactively and creatively sourcing more granular metadata required at the article level, and creating user interfaces capable of displaying multiple levels of resource description in an intuitive way for our users. By overcoming these challenges BHL is ensuring that the platform can accommodate not only historic literature but also modern biodiversity publications, ultimately helping bridge the gap between past and present biodiversity research.

Data flow diagram listing the data entities, data processes, and users of BHL data

The data flow journey for Virtual Items is actually quite different from traditional BHL content. For this new feature, the BHL Technical Team had to carefully consider the various data sources and how those would flow into the BHL data ecosystem and be presented to users. Image Credit: BHL Technical Team Collaborative Mapping

Facing Challenges, Strengthening Resilience

Despite major wins this past year for the BHL Technical Team, 2024 was not without its hurdles. After BHL completed several server migrations, a series of DDoS attacks in October targeting the Internet Archive (IA), a key partner and long-time host of BHL content, temporarily disrupted access to BHL materials. These events highlighted the importance of building a more resilient infrastructure for BHL, and having failsafes and additional data back-ups planned has been top of mind for the BHL Technical Team.

During the IA outage, we heard from many BHL users how the disruption affected their vital research and how grateful they were when access was restored.

“It is hard to quantify how vital a resource is until it is removed from reach! […] Thank you again for all the wonderful work done by BHL.”

“Life as I know it has come to a standstill. The @internetarchive is offline, which also affects @BioDivLibrary! What do I do!!! 😱”

“A huge amount of literature is only available through either Internet Archive itself or Biodiversity Heritage Library, which is hosted on the Internet archive. Stick in the wheel towards my work.”

“Dear Madam/Sir Really wonderful the BHL is back and provides accessibility to several hundred-year-old literature. Over the past two weeks [I have] not been able to do the majority of my work related to taxonomic curation of plant names of India. Thank you.”

IA continues to be a critical digitization and hosting partner for BHL. This year’s DDoS attacks on IA have only had the effect of helping BHL and IA strengthen infrastructure against future attacks.

Three racks of servers labeled with Internet Archive name and logo

Servers at the Internet Archive headquarters in San Francisco, CA. Image Credit: Jason Scott, Internet Archive | Wikimedia Commons.

Behind the Scenes: Building for the Future

In pursuit of a more resilient infrastructure for BHL, another major milestone was BHL’s inclusion in the AWS Open Data Sponsorship Program. Not only does the program provide BHL with free S3 storage for its data but BHL’s engagement with AWS also marks the beginning of an exploration into how cloud-based technologies can transform BHL’s back-end infrastructure. While BHL currently resides on on-premise storage and systems, cloud services offer new opportunities to rethink and refactor BHL to improve scalability, performance, and cost-efficiency. The benefits of a cloud-based or hybrid BHL are myriad, including complementary storage, faster page-serving speeds in low-bandwidth regions, and greater page persistence critical for biodiversity informatics applications. For the Press Release on BHL’s AWS Open Data Sponsorship, see Biodiversity Heritage Library Datasets Now Openly Accessible on the Amazon Web Services Cloud.

AI generated image of a forest with animals and a cloud with the words BHL in the cloud.

Exploring BHL in the Cloud in 2025 will be a major focus for the BHL Technical Team. We are very excited to see where the journey takes us! Image Credit: Canva AI Image Generator

Looking Ahead

As BHL moves into 2025, the BHL Technical Team’s focus remains on exploring cloud-based solutions and safeguarding BHL as a critical research infrastructure that provides access to over 62+ million pages of knowledge about life on Earth.

By embracing innovation and prioritizing accessibility, BHL will continue to evolve as a cornerstone resource for biodiversity research, empowering scientists and scholars to advance biodiversity knowledge and discovery.

For more on what’s in store for the year ahead, check out the 2025 BHL Technical Priorities.

February 4, 2025by mdimeo
BHL News, Blog Reel, Tech Updates

Biodiversity Heritage Library Datasets Now Openly Accessible on the Amazon Web Services Cloud

Now Available. BHL Datasets Now Openly Accessible on the Amazon Web Services Cloud. Background image of a variety of herons perching on branches

Sixty-two million pages of scientific text, images, and metadata, representing 500 years of biodiversity data will be openly accessible via the Registry of Open Data on AWS

Washington, DC, November 27, 2024 — The BHL Technical Team is thrilled to announce that Biodiversity Heritage Library (BHL) datasets will be openly accessible on the Amazon Web Services (AWS) cloud, thanks to the AWS Open Data Sponsorship Program. Moving BHL data to the cloud allows researchers globally to explore and analyze over 500 years of biodiversity data, enhancing their ability to derive scientific insights from our shared past to inform future global environmental policy.

BHL data is now hosted on AWS, and comprises over 62 million pages of scientific text from the 15th to the 21st centuries. BHL’s vast collection represents an unparalleled biodiversity resource with enormous potential to be used for longitudinal studies and conservation efforts.

Key Use Cases for BHL’s Data:

  • Historical Baseline Data: Spanning centuries, BHL provides a unique record of biodiversity literature, allowing researchers to establish long-term species baselines and track changes in species distributions and abundances over time.
  • Longitudinal Studies: The dataset supports analysis of biodiversity trends and ecosystem responses to environmental changes, offering valuable insights into historical and contemporary biodiversity dynamics.
  • Rare and Endangered Species: BHL’s historical records include species that may no longer exist or have become rare, aiding in the understanding of past biodiversity and informing current conservation efforts.
  • Taxonomic Stability: The collection includes taxonomic descriptions and classifications from various time periods, essential for understanding species relationships.
  • Cultural and Scientific Heritage: Beyond scientific data, BHL preserves historical texts, scientific illustrations, and annotations, enriching our understanding of past scientific practices and societal attitudes toward nature.

Transitioning to AWS services represents a transformative opportunity for BHL. As BHL currently operates outside a cloud environment, this move will optimize scalability, accessibility, performance, and cost-efficiency. AWS’s advanced computing capabilities will enable faster data processing and analysis, while AWS’s cost management tools will in the long-term help reduce infrastructure costs associated with on-premise storage and upgrades.

As a vital component of the global biodata infrastructure, the demand for BHL’s content and services continues to rise. By leveraging AWS services, we can address several critical needs, including:

  • Complimentary storage and access to a comprehensive suite of AWS cloud computing tools.
  • Ensured page persistence and unique page identifiers essential for biodiversity informatics applications.
  • Enhanced page-serving speeds in low-bandwidth regions where biodiversity research is critical.

BHL’s Data Manager, JJ Dearborn remarked, “Collaborating with AWS unlocks the full potential of BHL data. Cloud hosting not only provides an additional data safe harbor but also makes it accessible to researchers worldwide. We’re excited to democratize access to BHL data and see it used to spark innovative cross-disciplinary research in biodiversity informatics.”

The AWS Open Data Sponsorship Program is committed to making high-value datasets freely available to encourage innovation and collaboration. By supporting the free access to a diverse array of big datasets, AWS helps advance research and development in numerous fields.

To explore the BHL dataset and learn more about its content, visit the AWS Data Exchange.


About the Biodiversity Heritage Library

The Biodiversity Heritage Library (BHL) is the world’s largest open access digital library for biodiversity literature and archives. Headquartered at the Smithsonian Libraries and Archives in Washington, D.C., BHL is a global consortium that digitizes and freely shares natural history literature and archives. By providing open access to critical biodiversity knowledge, BHL addresses major research challenges and supports the global scientific community in understanding and conserving Earth’s species amid climate change and extinction crises. BHL also collaborates with rights holders, volunteers, and international experts to enhance content accessibility and interoperability. Since its inception in 2006, BHL has been working towards a shared future where biodiversity knowledge is globally accessible and supports universal bioliteracy.


Download Press Release

November 27, 2024by mdimeo
BHL News, Blog Reel

Advancing BHL’s Data for a Sustainable Future: Meet Tiago, Our New Wikimedian-in-Residence

BHL - Welcome to Tiago Lubiana, Wikimedian-in-residence
A person with short brown hair, black glasses, and a blue and white shirt smiles at the camera

Tiago Lubiana, BHL Wikimedian-in-Residence. Photo credit: Maria Eugenia Tita / Personal Archive / CC0 Dedication

We are thrilled to announce the appointment of Tiago Lubiana as the new Wikimedian-in-Residence (WiR) at the Biodiversity Heritage Library (BHL)! This crucial role, shaped by community-driven recommendations from the Unifying Biodiversity Knowledge to Support Life on a Sustainable Planet white paper, is aimed at advancing BHL’s goal of making biodiversity data more accessible, actionable, and impactful for a sustainable future.

Cover of white paper with image of earth from space

Dearborn, J. (2023). Unifying Biodiversity Knowledge to Support Life on a Sustainable Planet. Biodiversity Heritage Library. https://doi.org/10.21428/bcf8962c.699434fb

As BHL’s new Wikimedian-in-Residence, Tiago will expand BHL’s data presence within Wikimedia platforms, convert legacy data to structured data on Structured Data on Commons for BHL’s vast image collection, and provide training to the broader biodiversity wiki community. His expertise in open science, data modeling, and the semantic web will help connect BHL’s data to the growing biodiversity knowledge graph, fostering collaboration and ensuring our information is accessible, interoperable, and integrated into the emerging semantic web ecosystem.

BHL Flickr Image Collection

Tiago will be enhancing BHL’s Image Collection via the Structured Data on Commons Project. Photo credit: BHL / Screenshot/ CC0 Dedication

About Tiago

Tiago is a bioinformatician and independent Wikimedian based in São Paulo, Brazil, with a strong focus on open science, open knowledge, and life sciences. He holds a PhD in Bioinformatics from the University of São Paulo, where his research explored Wikidata as an open biocuration platform for disseminating knowledge about human cells. Tiago is also an active member of the International Society for Biocuration, the Wiki Movimento Brasil User Group, and has contributed significantly to various WikiProjects on Wikidata, including WikiProject COVID-19 and WikiProject Biodiversity.

A passionate advocate for biodiversity and open knowledge, Tiago is also the developer behind the inat2wiki tool, which bridges data between iNaturalist and Wikimedia Commons. Outside of his work in bioinformatics and open data, Tiago’s early passion for biology was sparked by his experience attending the International Biology Olympiad in Singapore in 2012, where he won a bronze medal.

Tiago holds a Bachelor’s Degree in Biomedical Sciences (University of São Paulo, 2014–2017), a Master’s Degree in Bioinformatics (University of São Paulo, 2018–2020), and a PhD in Bioinformatics (University of São Paulo, 2020–2024). For more information, visit his website.

The BHL-Wiki Working Group: A Global Collaborative Effort

In addition to Tiago’s role, we are also excited to officially announce the BHL-Wiki Working Group newly formed in 2024, which will help guide and support BHL’s wiki initiative. The group is made up of 30 members from 20 GLAM organizations worldwide, working to strengthen the biodiversity data infrastructure through greater interoperability and accessibility in the Wikimedia ecosystem and beyond. With diverse expertise spanning biodiversity research, open data, user engagement and technology, the group is committed to integrating semantic web technologies and adhering to open data principles, including FAIR (Findable, Accessible, Interoperable, Reusable) and CARE (Collective Benefit, Authority to Control, Responsibility, and Ethics).

Leadership of the BHL-Wiki Working Group includes:

  • Siobhan Leachman (User:Ambrosia10), BHL-Wiki Chair, a Wikimedian laureate and long-time advocate for the Biodiversity Heritage Library and Wikimedia projects.
  • Jake Orlowitz, BHL-Wiki Vice Chair, Wikiblueprint, Independent Wikimedian and Wikimedia Consultant
  • Giovanna Fontanelle, WMF Institutional Representative, Program Officer, Culture and Heritage, Wikimedia Foundation
  • JJ Dearborn, Smithsonian Institutional Representative, BHL Data Manager / Secretariat

Together, the BHL-Wiki Working Group is working towards the following key objectives:

  1. Maximizing Accessibility & Interoperability through Semantic Web Integration
    Ensuring biodiversity data is discoverable, accessible, and interoperable across global research networks by leveraging semantic web technologies and open data standards, supporting global biodiversity research and climate action.
  2. Advancing Data Modeling for Global Challenges
    Developing data models that enable better insights into climate change and biodiversity loss, while harmonizing with international initiatives to address pressing sustainability challenges.
  3. Strengthening User Engagement & Global Partnerships
    Enhancing community engagement, integrating traditional ecological knowledge (TEK), and fostering collaborations with marginalized voices to ensure BHL’s data is scientifically rich and culturally inclusive.

The working group is excited to have a new Wikimedian-in-Residence on board and see Tiago’s appointment as a critical need to push forward on the group’s collective goals. Chair of BHL-Wiki, Siobhan Leachman commented that:

“Having a Wikimedian-in-Residence at BHL is extremely important as this will help both the Wiki community to continue to actively and appropriately interweave content sourced from BHL into the Wikiverse and will also encourage and support members of the BHL community to contribute to the Wiki movement. In doing so the Wikimedian-in-Residence will be assisting BHL with its aim of helping to document and understand the world’s biodiversity.”

Presentation - What can a Wikipedian in Residence Do for You?

Presentation by Jake Orlowitz – “What can a Wikipedian in Residence Do for You?” Credit: Wikiblueprint.

If you are a BHL partner and you are interested in hosting your own Wikimedian-in-Residence, please check out the presentation from Wikiblueprint explaining the myriad benefits that hosting a Wikimedian at your home institution can bring.

Building a Global Ecosystem for Biodiversity Knowledge

The work led by the BHL-Wiki Working Group will lay the foundation for a more connected, sustainable biodiversity knowledge ecosystem. The BHL-Wiki initiative will amplify the impact of biodiversity data, support urgent climate action, and foster a more sustainable and inclusive future.

We are excited for the journey ahead and look forward to the innovations and collaborations that will emerge from this partnership!

November 20, 2024by mdimeo
Blog Reel, Featured Books

The Power of Community Science: How Smithsonian Volunpeers Transform Scientific Field Notes

Screenshot of Smithsonian Transcription Center displaying handwritten journal next to transcribed text

Last month, Smithsonian Libraries and Archives (SLA), Smithsonian Transcription Center (STC), and the Biodiversity Heritage Library (BHL) celebrated a significant milestone – technical staff worked collaboratively to integrate over 43,000 pages of transcription materials from STC into BHL. An additional 151,362 scientific name access points have now been added to the BHL search index for SLA archival field notes. These transcriptions enhance BHL’s full-text search, enable taxonomic name recognition, improve accessibility for vision-impaired users, and support climate research.

The Smithsonian Field Books Collection

Yellow notebook with handwritten text and handdrawn map

Dall, William Healey. (1895). Alaska Coal Fields, Bogosloff Volcano, Corral Hollow, California, 1895 (Vol. 1, pp. 81 and 82).

The Smithsonian Field Books Collection is a set of primary source archival records selected from the Smithsonian Libraries and Archives (SLA). This material dates from the late nineteenth century and the first comprehensive biological survey of the continental United States to the most recently accessioned materials at the Smithsonian Institution Archives.

The collection includes personal records of naturalists and scientists such as William Healey Dall (1845 – 1927) and Cleofe Calderon (1929 – 2007) at work around the world and expedition records such as the Western Union Telegraph Expedition (1865 – 1867) of Russian America and the United States Exploring Expedition (1838 – 1842) of the Pacific Ocean.

Naturalists and scientists in the field recorded their firsthand observations and data in a wide variety of forms including diaries and journals, hand-drawn maps and tables of data, observation logs and specimen catalogs, correspondence and reports, manuscripts, sketches, photographs, even audio recordings.

Recognizing the Value of Accurate Transcription

Mining the rich information embedded in these field notes depends on accurate transcription. The majority of the field notes are handwritten – even the most recent ones. Digital surrogates provide high resolution images but fail to afford anything beyond visual accessibility. Transcription, however, opens up this material to full-text searching, pattern recognition, visual accessibility aids, and more.

Beyond the handwriting itself, the field notes also contain other non-textual information that can only be captured through transcription. Examples include ornithologist Martin Moynihan’s notations of bird song in remote regions of Central America or Charles Dolittle Walcott’s sketches and diagrams of sedimentary stratification where his fossil specimens were found in the canyons of the American Southwest.

Handwritten notes with a drawing of a bird

Moynihan, M. (1963). 1962-1965 Andean Birds Mixed Flocks, Colombia, (4 of 4) (p. 78).

Chief naturalist with the U.S. Department of Agriculture, Vernon Orlando Bailey’s “Journal kept by Bailey on field trip to Wyoming and New Mexico, March 15-June 1906” focused on extermination techniques of gray wolves that would bring them to near extinction in the continental United States not long afterwards. The transcript includes descriptions of the sketches he included in his notes.

Screenshot of Smithsonian Transcription Center displaying handwritten journal next to transcribed text

Smithsonian Transcription Center image with corresponding transcription, including a text description of an ink drawing of a cow. https://transcription.si.edu/view/6656/EBOtg

The Smithsonian Field Book Project’s initial goal was to catalog these hidden biodiversity research material, improving discoverability in keeping with FAIR (Findable Accessible Interoperable Reusable) Data Principles. Doing so quickly resulted in additional researcher demand for more access and usability – first to view the field notes online (digitization) and then to examine them more closely (transcription). Grant funding assisted in a major rapid-capture digitization effort. However when it came to transcribing the digitized field notes, the Archives simply lacked the capacity to meet the level of researchers’ demand.

Turning to digital volunteers and the just-launched Smithsonian Transcription Center in 2013 changed the transcription equation beyond our best expectations both in volume and in accuracy. We quickly came to recognize these community scientists as collaborators, “volunpeers”, in the effort to advance and disseminate knowledge.

We are far from the end of this journey. More than half the collection remains to be digitized, and over two thousand digitized field notes still need transcription.

Inspired to Transcribe

These archival field notes contain vital historic biodiversity information. By transcribing the handwriting into machine readable text, volunpeers can help inform current day scientists and assist with their research on a multitude of topics such as climate change, the extinction crisis, or the spread of invasive species. Transcribing field notes can also take volunpeers on an adventure across time and distance, accompanying the writer on their journey. Having these adventures transcribed in machine readable text can also help inform science historians assisting them by making these historic documents more easily findable, searchable, and reusable.

Volunpeers collaborate to transcribe as accurately as possible the pages of the field journals provided. Multiple volunpeers will work on each project page transcribing the hard-to-read handwriting into machine readable text. Once satisfied the transcription is as complete as possible, a volunpeer will mark the transcribed page as complete. The volunpeers will then move onto the next page of the field notes until the project is finished.

Approval workflow for Smithsonian Transcription Center

Review is the second step in the transcription process.

Having many volunpeers working on one project helps ensure the quality of the transcription. What one volunpeer finds illegible, others may be able to read, especially as all the volunpeers working on a particular project become more familiar with the handwriting.

Volunteers have transcribed hundreds of field journals over the years. Two favorite examples have been the field journals of Vernon Orlando Bailey, a field naturalist who journeyed throughout the U.S. midwest studying and collecting mammals. Another volunpeer favorite were the papers of Arctic explorer and naturalist Robert Kennicott. Read more about these volunpeer experiences on the STC blog.

The Power of the Smithsonian Transcription Center

When the first of the field books were added to Smithsonian Transcription Center (STC), the program was still available only as a beta version. Approximately 450 volunpeers were transcribing on the site (co-author Siobhan Leachman among them), with the first 6,000 completed pages under their belt. Even during this moment of immense energy and fresh connections, it must have been difficult to imagine what STC would become. Today, a little over a decade later, more than 91,000 individual volunpeers have worked together to transcribe and review over 1.4 MILLION (!) pages of historic and scientific collections.

STC is the largest digital volunteering and crowdsourcing program at the Smithsonian Institution, and provides opportunities to engage with and contribute to digitized materials from across the full breadth of content areas represented by its museums, archives, and libraries. Through collaborative transcription and review, Smithsonian staff and digital volunteers work together to ensure that this content is more readable, accessible, and text-searchable across Smithsonian data systems and beyond.

Currently, the most active and popular projects are the Freedmen’s Bureau Transcription Project, a collaboration with the National Museum of African American History and Culture that deepens insight into the Reconstruction period and empowers African American genealogical research, and Project PHaEDRA, a collaboration with the Harvard-Smithsonian Center for Astrophysics that illuminates the work and discoveries of early women computers at the Harvard College Observatory.

If you feel inspired by this data access success story, consider joining the digital volunteer community, or sign up for our newsletter to stay up-to-date on upcoming projects.

Reusing the Liberated Data

The successful integration of transcription materials into BHL makes historical scientific data more accessible and useful. Through the collaborative efforts of technical staff from the Biodiversity Heritage Library, Smithsonian Libraries and Archives, and the Smithsonian Transcription Center and dedicated volunpeers, over 151,362 scientific name access points were added to the BHL search index, greatly enhancing search capabilities over the digitized corpus of Smithsonian field notes and archives. Out of the 556 eligible items reviewed, 522 were uploaded, contributing 43,460 pages of improved OCR text. These contributions not only improve BHL’s full-text search and taxonomic name recognition services but also provide better accessibility for vision-impaired users, and support ongoing biodiversity and climate change research.

Smithsonian Transcription Data outcomes report

A special thanks goes to Mike Lichtenberg, BHL’s Lead Developer and Systems Architect and Paul Day, Lead Developer at Smithsonian Transcription Center. As with many data improvement and platform enhancement projects, the requisite technical expertise is pivotal in ensuring the success of our collective efforts!


References and Resources

Dearborn, J., Lichtenberg, M., Richard, J. M., deVeer, J., Trizna, M., & Mika, K. 2023. [Presentation] Unearthing the Past for a Sustainable Future: Extracting and Transforming Data in the Biodiversity Heritage Library for Climate Action. Presented virtually at TDWG, Tasmania, Australia 2023. https://www.youtube.com/watch?v=8sGssyrpuJw

Trizna, M., & Dearborn, J. June 2023. [Poster] AI Models Are Getting Better at Reading Handwriting, but How Can We Find Handwritten Text to Begin With?. 7th Annual Digital Data Conference, Leveraging Digital Data for Conservation, Ecology, Systematics, and Novel Biodiversity Research, Tempe, Arizona, United States of America. https://doi.org/10.25573/data.23523495.v1

Dearborn, J., & Mika, K. June 5, 2022. [Poster] Extracting Expedition Log Data Found in the Biodiversity Heritage Library. Through the Door and Through the Web: Releasing the Power of Natural History Collections Onsite and Online, Edinburgh, Scotland, United Kingdom: Society for the Preservation of Natural History Collections (SPNHC). https://doi.org/10.5281/zenodo.6593457

September 30, 2024by mdimeo
BHL News, Blog Reel

Advancing Data Excellence: A New Era for the BHL Cataloging and Metadata Committee

Cataloging & Metadata Committee Milestones with graphic visuals of a team pointing to charts

2023 proved to be a transformative year of growth, increased collaboration, and heightened dedication to advancing biodiversity knowledge for the BHL Cataloging and Metadata Committee. Last year, the Committee achieved significant milestones in committee governance, professional development opportunities, data quality updates, and the ratification of consortia-wide data policies, accompanied by comprehensive documentation.

Cataloging & Metadata Committee Milestones with graphic visuals of a team pointing to charts

On the governance front, the former “Cataloging working group” completed a Committee Charter to elevate the group’s status to an official BHL Committee. Approval of the Committee at the 2023 BHL Annual Meeting by the Executive Committee and BHL Partners, as well as the formal codification in the BHL Bylaws, solidifies Cataloging and Metadata as a permanent standing BHL Committee.

Pictures of new committee co chairs Elizabeth McKinley (leading a tour of patrons) and Daniel Euphrat (sitting in front of digitization equipment)

In 2023, the Committee held its inaugural Chair Election which signified a notable transition for the former working group. We are excited to welcome Daniel Euphrat of the Smithsonian Libraries and Archives and Elizabeth McKinley of the Chicago Field Museum as the co-chairs for 2024. Additionally, we extend our gratitude and appreciation to former chairs Suzanne Pilsk of the Smithsonian Libraries and Archives and Diana Duncan of the Chicago Field Museum, who will continue to provide valuable mentorship during this transitional period. A warm welcome and congratulations to Daniel Euphrat and Elizabeth McKinley on becoming the inaugural co-chairs in 2024.

The Committee membership grew with the addition of advisors, retired staff as observers, and new members:

  • Siobhan Leachman, Committee Advisor and Independent Wikimedian
  • Elizabeth McKinley, Committee Co-Chair and Cataloging & Metadata Librarian at the Chicago Field Museum
  • Briana Giasullo, Committee Member and Cataloging and Digital Resources Librarian at the Academy of Natural Sciences of Drexel University
  • Sandra Lee Parker Provenzano, Committee Member and Head Cataloger at Dumbarton Oaks, Harvard University

As a newly formalized Committee, this coming year will see some fresh perspectives and ideas to drive progress on BHL Strategic Goals and the Committee’s formal charge.

Empowering BHL Staff with Wiki Education’s Wikidata Certificate Course

A tangible benefit of Committee membership includes access to broader professional development opportunities through BHL. Early in 2023, committee members celebrated the completion and final project wrap-up of a six-week professional certificate course sponsored by BHL and the Smithsonian Libraries and Archives (SLA). This specialized course was facilitated by Wiki Education, a non-profit that focuses exclusively on supporting students, faculty, and subject area specialists with their Wikimedia work. The program provided BHL and SLA staff with a unique opportunity to hone advanced skills in SPARQL queries, visualizations, and data modeling.

Three SPARQL visualizations with varying colors and sizes of circles clustered together

Left to right: SPARQL visualizations of BHL Authors by Occupation; BHL Female Authors by Occupation; BHL Male Authors by Occupation. The course provided participants with the opportunity to experiment with linked data and visualizations produced by the Wikidata Query Service.

The custom curriculum, led by the knowledgeable Will Kent aka “Wikidata Will” and SLA co-facilitators JJ Dearborn, Richard Naples, Suzanne Pilsk, and Jackie Shieh engaged a cohort of 25 participants in a tailored, immersive experience exploring topics like SPARQL queries, visualizations, federation, data modeling, bulk data loading, and editing tools.

Grid of 20 black and white video screens of committee members on a video conference call

Participants added BHL partner nodes as items and linked them up with BHL using the Wikidata in partnership property:

GIF of before and after visualization of BHL knowledge graph with entities clustered and connected by arrows

The BEFORE is what BHL’s Partner Knowledge Graph looked like and AFTER the course, one can see how BHL’s network has grown in Wikidata!

Enriching our Committee’s technical expertise has not only opened new avenues for exploration within the Wikimedia project ecosystem but has also laid the foundation for the newly formed BHL-WIKI working group, set to launch this year under the distinguished leadership of Wikimedia Laureate and long-time BHL advocate, Siobhan Leachman. Under Siobhan’s guidance, BHL is embarking on a dedicated effort to convert its legacy data into 5-star linked open data, facilitating expanded global access to biodiversity knowledge for all.

For a more in-depth understanding of BHL’s symbiotic relationship with the Wikimedia ecosystem, we encourage readers to explore related post “BHL is Round Tripping Persistent Identifiers with the Wikidata Query Service,” and the white paper titled “Unifying Biodiversity Knowledge to Support Life on a Sustainable Planet.”

Comprehensive BHL Metadata Requirements for our Global Consortium

Conformant metadata in BHL is the linchpin that allows BHL to manage diverse global sources of metadata across 588+ contributors. A guiding document is paramount to fostering metadata harmonization from a globally disparate network and empowers BHL technical staff in maintaining data consistency and ensuring interoperability across the consortium. The new BHL metadata requirements facilitates effective searching, browsing, discovery, and identification for BHL’s end-users, enabling broader access and engagement with biodiversity knowledge on a global scale.

The journey to finalize BHL’s Metadata Requirements was a lengthy endeavor and bringing together once-disparate guidelines into a unified framework is the culmination of many years of work. After a Metadata Requirements Summit in 2023 and multiple rounds of Committee peer review, the first version of comprehensive BHL Metadata Requirements has been published and incorporated into BHL’s Collection Development Policy.

BHL Metadata requirements include data disclaimer, a statement on remediation of harmful language in library metadata, partner meta app, FAQ, schema mappings, updated schema tables, digitization workflow decision map

Additionally, embedded within these requirements is the BHL Cataloging and Metadata Committee’s formal Statement on Remediation of Harmful Language in Library Metadata, which serves as a commitment to address harmful language in library metadata, recognizing the impact of legacy language and knowledge organization systems can have on perpetuating biases. BHL actively champions inclusivity, encouraging contributors to realign vocabulary terms to ensure diverse user access. Rejecting censorship, BHL acknowledges and commends the ongoing efforts by contributing institutions in remediating outdated metadata terms. This commitment not only aligns with our dedication to diversity but also reinforces our mission to provide a comprehensive and inclusive perspective on biodiversity knowledge for a global audience.

Across the BHL Consortium, we consistently unite to uphold excellence in metadata management and recognize the individual needs of our digitization partners, encompassing factors such as staff, technical expertise, funding, and local facilities. Having comprehensive Metadata Requirements reflects the tenacious and collaborative spirit across the BHL Consortium to come together in unity and further our collective mission to provide free, worldwide access to knowledge about life on Earth.


BHL Cataloging and Metadata Committee Charge

The BHL Cataloging and Metadata Committee possesses advanced expertise in metadata standards, remediation, mapping, and cross-walking workflows. The committee is responsible for overall BHL metadata advising, curation, and validation with the overarching goal of sharing the highest quality bibliographic data outputs broadly with the bioinformatics community and the world. The task of harmonizing hundreds of years of metadata from diverse sources is a continuous, iterative activity. The Cataloging and Metadata Committee directly supports the following strategic goals:

  • Ensures reliability and accuracy of the collections by curating the BHL collection 
  • Defines the BHL collection as a data resource as well as a digital library to support big data approaches
  • Extends and deepens metadata parsing and validation 

In closing, the Committee eagerly anticipates a year ahead filled with continued efforts to enhance BHL’s metadata quality and further integrate diverse information into the growing network of biodiversity knowledge. Please don’t hesitate to reach out to the group with your feedback. Your input will be invaluable as we strive to advance our mission and better serve our community. Thank you for your ongoing support and engagement.

February 29, 2024by mdimeo
BHL News, Blog Reel, Tech Updates

BHL Technical Development: Year in Review

BHL data flows diagram listing external data entry, internal data entry, data processes, and internet consumers of data

In 2023, BHL’s Technical Team dedicated significant efforts to improve our data ecosystem, now comprising 61+ million pages of biodiversity literature. Last year’s Technical Priorities underscored BHL’s steadfast commitment to data quality by focusing on both upstream and downstream data flows. Notable milestones include delivering refined taxonomic data to researchers, implementing interface improvements based on user feedback, and forging data pipelines for existing and new downstream data consumers.

Collage of logos from downstream data consumers, including GBIF, figshare, WoRMS, gna, Tropicos, Wikimedia, Crossref, OCLC, BioStor, Catalogue of Life, unpaywall, and DPLA

A selection of BHL’s major downstream data consumers

Global Names Upgrades

A standout achievement in 2023 was the upgrade of BHL’s Global Names taxonomic intelligence tools, marking a significant leap forward in ensuring the accuracy and comprehensiveness of scientific name detection and verification. These upgrades replace the now deprecated GNRD tool and GNresolver services.

BHLIndex is a remarkable improvement over past performance of the Global Names taxonomic intelligence suite. A decade ago, finding and verifying all names in BHL took 45 days of computational processing time—today, the task takes less than 6 hours. Moreover, there is a notable increase in the overall number of name instances in the BHL corpus, underscoring significant improvements made in both speed and comprehensiveness. The Global Names Team observed the following performance improvements:

  • name-finding in 275,000 volumes, 60+ million pages: 2.5 hours;
  • name-verification of 23 million unique name-strings: 3 hours; and
  • preparing a CSV file with 250 million names occurrences/verification records: 40 minutes.

Additionally, Mike Lichtenberg, BHL Lead Developer, found that the BHLIndex upgrade yielded an additional 17,790,590 access points to scientific species names in the BHL corpus!

Graph indicating the number of scientific names found on BHL pages before (245,237,984) and after (263,028,574) implementing the new BHLIndex.

The BHL corpus now contains 17,790,590 additional scientific names.

In tandem with the BHLIndex taxonomic intelligence upgrades, the BHL Technical Team continues its ongoing efforts to enhance the underlying OCR data that powers the Global Names tools. Stay tuned for an update on the OCR Reprocessing Project later this year.

BHL is incredibly grateful to Dr. Dima Mozzherin, Dr. Geoff Ower, and all past and present contributors to the Global Names Architecture applications for their collaborative efforts with the BHL Technical Team to make these new upgrades and an additional 17+ million taxonomic access points available to BHL’s global user base. The continuous improvement of taxonomic data in BHL remains essential, driving key functionalities for our scientific researchers, including taxonomic search, species bibliographies, and interlinkages with other major taxonomic databases on the web.

BHL User Interface Enhancements

BHL has 588 contributors and counting. The addition of a Contributor Facet on the BHL search results page allows users to filter, facet, and hone in on publications from specific BHL contributor collections for the first time. This feature is particularly important to BHL Members and Affiliates who may want to refine their search to see only digitized content from their home collections to answer local research questions.

Example of BHL search results with new Contributor facet in left navigation panel

The contributor facet is a new way to hone in on a specific institutional collection within search results.

In addition to the new contributor facet, BHL has enhanced Supplementary links for Titles by allowing for additional links and incorporating a controlled vocabulary to precisely describe the external resource being linked to.

Example of a BHL title page listing External Resources (a Harvard collection guide for the Walter Deane papers) associated with the title in BHL

Supplementary links allow BHL Partners to curate their collections with relevant links from their home institution like a finding aid, digital exhibition, or collection guide.

Partner curated supplementary links allow BHL to link out to other relevant resources such as a collection guide, digital exhibition, or an archival finding aid. Kudos to the BHL Collection Committee for keeping a pulse on BHL user needs and requesting and scoping the requirements for these new interface features and enhancements.

Forging BHL Data Pipelines

Downstream consumers of BHL data have been a big focus for the BHL Technical Team this past year as well. The Global Biodiversity Information Facility (GBIF) and Crossref play integral roles in advancing biodiversity knowledge by facilitating seamless access to valuable scientific data sourced from BHL. Having accurate BHL data in these repositories furthers the Technical Team’s charge “to develop tools and services to facilitate greater access, interoperability, and reuse of BHL content and data.”

GBIF 

GBIF functions as a global repository for crucial biodiversity data, fostering a more comprehensive understanding of Earth’s biological diversity. Exposing species occurrence data from BHL in GBIF offers a tangible avenue for BHL partners to actively contribute to climate change and biodiversity conservation initiatives.

A significant accomplishment for the BHL Technical Team was the deposit of data at GBIF this year. This milestone marked the culmination of 18 months of intense investigation which was condensed into a 10-minute presentation given at TDWG 2023 entitled “Unearthing the Past for a Sustainable Future: Extracting and transforming data in the Biodiversity Heritage Library for climate action.” 

Cover slide for TDWG 2023 presentation titled Unearthing the past for a sustainable future: extracting and transforming data in the Biodiversity Heritage Library for climate action

The BHL Technical Team presented at TDWG 2023 this year.

The presentation underscores the BHL Technical Team’s dedication to aligning BHL’s data management priorities with international climate-related initiatives, sustainability goals, and most importantly to address global environmental challenges. To learn more check out the conference recording and the related blog post Illuminating BHL’s Dark Data: Citizen Scientists and AI Unlock Key Biodiversity Data in GBIF.

Much of the BHL Technical Team’s investigative work was conducted in collaboration with the BHL Transcription Upload Tool Working Group (TUTWG) who performed crucial tasks such as crowdsourcing and analyzing a Handwritten Text Recognition sample set and providing already transcribed materials and data outputs to process and deposit with GBIF. In addition to their contributions to BHL’s data extraction efforts, the working group members also:

  • Provided user documentation and training on uploading transcription text files to BHL;
  • Updated archival material metadata in BHL for improved findability;
  • Defined an allowable mark-up policy for BHL Partner transcriptions;
  • Scoped requirements for a new OCR text indicator in the BHL interface; and
  • Contributed a valuable GLAM resource to the BHL knowledge base in the form of the Transcription Platform Comparison chart to aid institutions in the overwhelming task of selecting a transcription platform suited to their particular needs.
Transcription platform comparison chart listing four platforms (DigiVol, FromThePage, Wikisource, and Zooniverse) and their key features.

Explore the Transcription Platform Comparison Chart for more details.

A huge shout out goes to all BHL TUTWG members and collaborators for their invaluable contributions.

Crossref

Complementing the new pilot GBIF data pipeline is the improvement of an existing one. Crossref ensures the comprehensive and accurate association of BHL’s bibliographic metadata with digital object identifiers (DOIs), enabling the efficient citation and linking of scholarly content in the modern publishing environment. Ensuring that historic knowledge is served up in modern discovery layers can be incredibly challenging work and we are so grateful to the Persistent Identifier Working Group (PIWG) for their consistent efforts in the form of energy, time, and enduring patience in 2023. The working group in collaboration with the BioStor project has been helping BHL broaden its collection horizons and illuminate BHL dark data since 2012.

In 2023 the group piloted and brought to production the submission of improved DOI data deposits for titles that are part of a monographic series (which provide notoriously difficult bibliographic conundrums for BHL Staff). The use of a richer deposit schema allows for a more complete set of metadata to flow downstream to Crossref.

And More Data 

Since last year, BHL has been openly publishing all of its data on the Smithsonian Figshare Data Repository. The BHL Open Data Collection uses the Data Catalog Vocabulary (DCAT) to describe and document what exactly is in all that data. Due to popular demand, we’ve added two new BHL datasets for our users to experiment with:

Richard, Joel; Dearborn, Jacqueline (2023). BHL Optical Character Recognition (OCR) – Full Text Export (new). Smithsonian Libraries and Archives. Dataset. https://doi.org/10.25573/data.21422193.v13

Dearborn, Jacqueline; Lichtenberg, Mike (2023). BHL Flickr Harvest Data. Smithsonian Libraries and Archives. Dataset. https://doi.org/10.25573/data.24476686.v3

The Flickr harvest data was the product of a workshop called Transforming Biodiversity Heritage Library Images – Data Modeling with OpenRefine. This event was hosted in collaboration with Wikimedia Foundation’s Giovanna Fontenelle and Wikimedian Sandra Fauconnier as part of Wikimedia’s Image Description Month.

BHL Staff, Wikimedians, Flickr staff, and all biodiversity image enthusiasts together are beginning the process of mapping and loading BHL Flickr image data to Structured Data on Commons (SDC), a new Wikimedia Commons initiative. This project was the second top-voted priority from the BHL Wikimedia white paper entitledUnifying Biodiversity Knowledge to Support Life on a Sustainable Planet. 

The Importance of Good Documentation

Fundamental to all technology projects is good documentation, and BHL is no different. The BHL Consortium’s collective endeavor to deliver over 500 years of data about biodiversity to the world relies crucially on the availability of comprehensive documentation about BHL’s extensive data ecosystem. To this end, the BHL Technical Team is working hard to ensure that BHL’s data is “useful, usable, and used” by putting greater emphasis on documentation this year. We want to make sure that all BHL users and collaborators possess the knowledge to effectively leverage BHL data for their research, apps, discovery layers, visualizations and so much more!

BHL data flows diagram listing external data entry, internal data entry, data processes, and internet consumers of data

Helping users understand how data flows in and out of the BHL data ecosystem and sharing it with the world is an overarching goal for the BHL Technical Team.

With over 17 years of development, BHL’s data model has matured significantly, enabling the harvesting, processing, and delivery of rich and robust data about our collections. The BHL Technical Team’s commitment to a high-quality, well documented data ecosystem underpins all BHL initiatives aimed at enhancing the discovery and access to BHL Partner and Contributor collections worldwide.

For a sneak peak of what is in store for the year ahead, check out the 2024 BHL Technical Priorities.

January 22, 2024by mdimeo
BHL News, Blog Reel, Featured Books

Illuminating BHL’s Dark Data: Citizen Scientists and AI Unlock Key Biodiversity Data in GBIF

Visualizations of species occurrence data deposited in GBIF from the journals of William Brewster

In the face of climate change and environmental challenges, understanding and documenting Earth’s biodiversity is essential. The Global Biodiversity Information Facility (GBIF) serves as a global repository for biodiversity data, playing a pivotal role in this critical mission of safeguarding our planet’s biodiversity. Species occurrence data sourced from the Biodiversity Heritage Library (BHL) provides insights into species distributions, behaviors, and interactions much deeper into time, offering key species baseline data required to effectively address the climate crisis. Without accurate and comprehensive data in GBIF, our collective ability to track environmental changes and make informed decisions is severely hampered.

Collage of images representing data used from GBIF

Figure 1: GBIF-mediated data is used extensively in climate science and informs global environmental policy. For more information see: https://www.gbif.org/climate

As a GBIF participant node, BHL is committed to sharing biodiversity data openly, adhering to FAIR (Findable, Accessible, Interoperable, Reusable) and CARE (Collective Benefit, Authority to Control, Responsibility, Ethics) data principles, and collaborating with a global network of biodiversity organizations to bolster and build capacity to strengthen the biodiversity information infrastructure. To honor our commitments, technical staff from BHL are working to establish a scalable data pipeline of occurrence data currently trapped in archival field notes, journals, letters, correspondence, and other primary source materials. The journey has been an arduous one due to poor OCR (optical character recognition) data quality for BHL’s sub-corpus of handwritten materials.

Example of handwritten observation data with poor quality OCR and no scientific names found on the page

Figure 2: Sample of “dark” handwritten observation data with corresponding unstructured, uncorrected OCR text. From National Museum of Natural History, Pacific Ocean Biological Survey Program, At-sea, 1963-1966, 1968, part 3: July – August 1966.

Technical approaches to building a BHL ETL (Extract, Transform, and Load) data pipeline of species occurrence data from BHL’s field notes include machine learning, artificial intelligence (AI), and data extraction through innovative transcription projects like DigiVol, the crowdsourcing citizen science platform collaboration between the Australian Museum and the Atlas of Living Australia.

Building a data pipeline: extracting species occurrence data from BHL. 6 steps include: digitize, ingest, identify, extract, transform, and load.

Figure 3: Learn more about how BHL is building a new data pipeline from the recent talk given at TDWG 2023 entitled Unearthing the Past for a Sustainable Future: Extracting and transforming data in the Biodiversity Heritage Library for climate action; abstract; recording.

Now underway, is a global effort to convert centuries of biodiversity knowledge into accessible and actionable data, utilizing advanced AI and technical approaches. BHL Partners now have a stake in supporting international conservation policy aimed at safeguarding Earth’s biodiversity through greater data integration with the global biodiversity data infrastructure.

The Journals of William Brewster

The ornithological papers of William Brewster (1851-1919), held in the Ernst Mayr Library and Archives of the Museum of Comparative Zoology (MCZ) at Harvard University, are a rich source of historical species occurrence data and an ideal use case for building a BHL ETL data pipeline. Brewster’s field notes, journals, diaries, and correspondence comprise over 60,000 pages replete with detailed bird observations spanning 54 years (1865-1919). These records augment and extend his collection of over 40,000 bird specimens, bequeathed by Brewster to the MCZ and considered “one of the largest private collections ever made in this country [United States], and in some respects … by far the most valuable” (Henshaw, 1920). Thanks to a number of projects facilitated by the Ernst Mayr Library over the years, Brewster’s journals have been digitized and later transcribed using DigiVol.

Figure 4 below shows one example of Brewster’s extensive species lists. Brewster fastidiously recorded all essential elements of a species occurrence record and often more such as habitat information and detailed behavioral descriptions. The complexity of this example, featuring Brewster’s unique formatting, use of ornithological symbols, and tiny, crowded handwriting, highlights the value of crowdsourced human transcription by a team of enthusiastic, dedicated volunteers.

Figure 5 shows a portion of the transcription of this species list, produced in DigiVol by one of our long-time transcribers.

Handwritten list of species observed by William Brewster

Figure 4. Brewster’s bird observations for July from his 1915 Diary.

Typed transcription of the handwritten species list from William Brewster

Figure 5. A portion of the transcript of Brewster’s July, 1915 species list.

Extracting Data via Citizen Scientists and DigiVol

DigiVol, first developed in 2011 to crowdsource the transcription of specimen labels, enables institutions around the world to engage volunteers to extract data of various types, such as text, species identifications, and species traits from images. Each institution is able to upload images and manage their crowdsourcing project through the DigiVol platform. Through a combination of gamification and engagement tools such as a user forum and secure, private emails, institutions can build volunteer skills, a sense of community, and commitment.

Screenshot of the Harvard University, Museum of Comparative Zoology, Ernst Mayr Library transcription portal on DigiVol platform, listing the number of expeditions and volunteers

Figure 6: The DigiVol dashboard for the Harvard Museum of Comparative Zoology, Ernst Mayr Library.

Volunteers on the website can contribute any time of day – 24/7, 365 days a year. They can choose from a broad array of “virtual expeditions” on DigiVol, from identifying animals in camera traps located in the Australian “bush” to transcribing specimen labels and field notes from locations around the world.

The DigiVol platform has seen more than 14,000 volunteers contribute 155 equivalent work years (7 hour days, 261 days a year) to the digitisation of over 6 million tasks at an estimated equivalent value of A$12 million.

In terms of the William Brewster project, 352 volunteers have transcribed over 27,000 pages of diaries and field notes, at an average 23 minutes per page. This contribution amounts to more than 5.8 work years at an equivalent cost of A$462,000.

Publishing Data to GBIF

The final output of these recent data pipeline investigations was the deposit of 1,853 species occurrence records at GBIF from the Journals of William Brewster. The data deposit comprises valuable biodiversity records extracted through transcription efforts on DigiVol, transformed into DarwinCore, and subsequently published on GBIF. The species occurrence records span diverse geographical locations, primarily focusing on New England but extending to other regions of the United States, the Caribbean, and Europe. Brewster’s meticulous observations, encompassing species observation, behavioral data, and environmental conditions offer a rich historical perspective on biodiversity dating back over a century.

Visualizations of species occurrence data deposited in GBIF from the journals of William Brewster

Figure 7: Species Occurrence Data from the Journals of William Brewster are now available on GBIF. Additional data deposits are planned in 2024. Data deposit: https://doi.org/10.15468/q45atb

This deposit of historic biodiversity data demonstrates how valuable BHL’s collection is to global biodata infrastructure, as it plays a crucial role in establishing species base lines, informing climate change studies, tracking key environmental indicators, and contributing to the development of global biodiversity monitoring platforms.

Having the data available in BHL, even as transcribed text is one thing, but it is the human resources to review and reconcile the data that is really required to facilitate the flow of historic biodiversity data into today’s bioinformatics ecosystems. To all the humans involved in this species data project, from the initial observation and recording, to the preservation, digitization, transcription, data extraction and reconciliation, machine-learning, and creation of vital ETL data pipelines, we thank you.


Related Works

Biodiversity Heritage Library Open Data Collection. (2022, November). Smithsonian Figshare.

Crowley, B., Dearborn, J., Funkhouser, C., Kalfatovic, M., Merriman, K., Iggulden, D., Trei, K., & Herrmann, E. (2023, October). Safeguarding Access to 500 Years of Biodiversity Data: Sustainability Planning for the Biodiversity Heritage Library [TDWG2023]. Biodiversity Information Standards, Hobart, Tasmania, Australia.

Data Flows Diagram. (2023, March). BHL Technical Team (BHL-TECH) Biodiversity Heritage Library.

Dearborn, JJ (2023, April). Unifying Biodiversity Knowledge for Life on a Sustainable Planet. Biodiversity Heritage Library. https://bhl.pubpub.org/

deVeer, J. (2021, February 24). Making the Best of Difficult Times: Accelerating the Transcription of William Brewster’s Writings During the COVID-19 Pandemic. Biodiversity Heritage Library.https://blog.biodiversitylibrary.org/2021/02/accelerating-transcription-brewster-covid19.html

deVeer, J. and Rinaldo, C. (2021, February 23). The Life and Work of Robert Alexander Gilbert: Empowering New Insights through Digitization and Transcription of Archival Materials. Biodiversity Heritage Library. https://blog.biodiversitylibrary.org/2021/02/life-work-robert-gilbert.html

Henshaw, Henry W. 1920. In Memoriam: William Brewster, Born July 5, 1851 – Died July 11, 1919. The Auk 37, 1 (1920), 1–23. https://doi.org/10.2307/4072953

Lichtenberg, M. (n.d.). BHL Data Model. BHL Github Repository. https://github.com/gbhl/bhl-us/tree/master/Documentation/DataModel

Mika, K. and Dearborn, J. (2022). [poster] Extracting expedition log data found in the Biodiversity Heritage Library. Through the door and through the web: releasing the power of natural history collections onsite and online, June 5, 2023. Edinburgh, Scotland, United Kingdom: Society for the Preservation of Natural History Collections (SPNHC). https://doi.org/10.5281/zenodo.6593457.

Richard, J. (2022, December 20). OCR Improvements: An Early Analysis. Biodiversity Heritage Library. https://blog.biodiversitylibrary.org/2022/07/ocr-improvements-early-analysis.html

Rinaldo, C. (2021, February 22). Nature Conservation and William Brewster: Insights From a Lifetime of Scientific Observations. Biodiversity Heritage Library. https://blog.biodiversitylibrary.org/2021/02/william-brewster-post-one

Trizna, M., & Dearborn, J. (2023, June). AI models are getting better and better at reading handwriting, but how can we find handwritten text to begin with? [poster]. 7th Annual Digital Data Conference, Leveraging Digital Data for Conservation, Ecology, Systematics, and Novel Biodiversity Research, Tempe, Arizona, United States of America. https://doi.org/10.25573/data.23523495.v1

November 9, 2023by mdimeo
Page 1 of 212»

Help Support BHL

BHL's existence depends on the financial support of its patrons. Help us keep this free resource alive!

search

About BHL

The Biodiversity Heritage Library (BHL) is the world’s largest open access digital library for biodiversity literature and archives. BHL operates as a worldwide consortium of natural history, botanical, research, and national libraries working together to digitize the natural history literature held in their collections and make it freely available for open access as part of a global “biodiversity community.”

Join Our Mailing List

Sign up to receive the latest news, content highlights, and promotions.

Subscribe Now

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 319 other subscribers

Subscribe to Blog Via RSS

Subscribe to the blog RSS feed to stay up-to-date on all the latest BHL posts.

Access RSS Feed

Inspiring Discovery through Free Access to Biodiversity Knowledge.

The Biodiversity Heritage Library makes it easier than ever for you to access the information you need to study and explore life on Earth…for free, anytime, anywhere.

 

64+ Million Pages of
Biodiversity Literature Online.

EXPLORE

Tools and Services
to Transform Research.

EXPLORE

300,000+
Illustrations on Flickr.

EXPLORE

ABOUT | HARMFUL CONTENT | PRIVACY | SITE MAP | TERMS OF USE