Biodiversity Heritage Library - Program news and collection highlights from BHL
  • Home
  • News
  • Featured Books
    • All Featured Books
    • Book of the Month Series
    • BHL at 20
  • User Stories
  • Campaigns
    • Fossil Stories
    • Garden Stories
    • Monsters Are Real
    • Page Frights
    • Her Natural History
    • Earth Optimism 2020
  • Tech Blog
  • Visit BHL
Home
News
Featured Books
    All Featured Books
    Book of the Month Series
    BHL at 20
User Stories
Campaigns
    Fossil Stories
    Garden Stories
    Monsters Are Real
    Page Frights
    Her Natural History
    Earth Optimism 2020
Tech Blog
Visit BHL
  • Home
  • News
  • Featured Books
    • All Featured Books
    • Book of the Month Series
    • BHL at 20
  • User Stories
  • Campaigns
    • Fossil Stories
    • Garden Stories
    • Monsters Are Real
    • Page Frights
    • Her Natural History
    • Earth Optimism 2020
  • Tech Blog
  • Visit BHL
Biodiversity Heritage Library - Program news and collection highlights from BHL
BHL News, Blog Reel, Featured Books

Illuminating BHL’s Dark Data: Citizen Scientists and AI Unlock Key Biodiversity Data in GBIF

Visualizations of species occurrence data deposited in GBIF from the journals of William Brewster

In the face of climate change and environmental challenges, understanding and documenting Earth’s biodiversity is essential. The Global Biodiversity Information Facility (GBIF) serves as a global repository for biodiversity data, playing a pivotal role in this critical mission of safeguarding our planet’s biodiversity. Species occurrence data sourced from the Biodiversity Heritage Library (BHL) provides insights into species distributions, behaviors, and interactions much deeper into time, offering key species baseline data required to effectively address the climate crisis. Without accurate and comprehensive data in GBIF, our collective ability to track environmental changes and make informed decisions is severely hampered.

Collage of images representing data used from GBIF

Figure 1: GBIF-mediated data is used extensively in climate science and informs global environmental policy. For more information see: https://www.gbif.org/climate

As a GBIF participant node, BHL is committed to sharing biodiversity data openly, adhering to FAIR (Findable, Accessible, Interoperable, Reusable) and CARE (Collective Benefit, Authority to Control, Responsibility, Ethics) data principles, and collaborating with a global network of biodiversity organizations to bolster and build capacity to strengthen the biodiversity information infrastructure. To honor our commitments, technical staff from BHL are working to establish a scalable data pipeline of occurrence data currently trapped in archival field notes, journals, letters, correspondence, and other primary source materials. The journey has been an arduous one due to poor OCR (optical character recognition) data quality for BHL’s sub-corpus of handwritten materials.

Example of handwritten observation data with poor quality OCR and no scientific names found on the page

Figure 2: Sample of “dark” handwritten observation data with corresponding unstructured, uncorrected OCR text. From National Museum of Natural History, Pacific Ocean Biological Survey Program, At-sea, 1963-1966, 1968, part 3: July – August 1966.

Technical approaches to building a BHL ETL (Extract, Transform, and Load) data pipeline of species occurrence data from BHL’s field notes include machine learning, artificial intelligence (AI), and data extraction through innovative transcription projects like DigiVol, the crowdsourcing citizen science platform collaboration between the Australian Museum and the Atlas of Living Australia.

Building a data pipeline: extracting species occurrence data from BHL. 6 steps include: digitize, ingest, identify, extract, transform, and load.

Figure 3: Learn more about how BHL is building a new data pipeline from the recent talk given at TDWG 2023 entitled Unearthing the Past for a Sustainable Future: Extracting and transforming data in the Biodiversity Heritage Library for climate action; abstract; recording.

Now underway, is a global effort to convert centuries of biodiversity knowledge into accessible and actionable data, utilizing advanced AI and technical approaches. BHL Partners now have a stake in supporting international conservation policy aimed at safeguarding Earth’s biodiversity through greater data integration with the global biodiversity data infrastructure.

The Journals of William Brewster

The ornithological papers of William Brewster (1851-1919), held in the Ernst Mayr Library and Archives of the Museum of Comparative Zoology (MCZ) at Harvard University, are a rich source of historical species occurrence data and an ideal use case for building a BHL ETL data pipeline. Brewster’s field notes, journals, diaries, and correspondence comprise over 60,000 pages replete with detailed bird observations spanning 54 years (1865-1919). These records augment and extend his collection of over 40,000 bird specimens, bequeathed by Brewster to the MCZ and considered “one of the largest private collections ever made in this country [United States], and in some respects … by far the most valuable” (Henshaw, 1920). Thanks to a number of projects facilitated by the Ernst Mayr Library over the years, Brewster’s journals have been digitized and later transcribed using DigiVol.

Figure 4 below shows one example of Brewster’s extensive species lists. Brewster fastidiously recorded all essential elements of a species occurrence record and often more such as habitat information and detailed behavioral descriptions. The complexity of this example, featuring Brewster’s unique formatting, use of ornithological symbols, and tiny, crowded handwriting, highlights the value of crowdsourced human transcription by a team of enthusiastic, dedicated volunteers.

Figure 5 shows a portion of the transcription of this species list, produced in DigiVol by one of our long-time transcribers.

Handwritten list of species observed by William Brewster

Figure 4. Brewster’s bird observations for July from his 1915 Diary.

Typed transcription of the handwritten species list from William Brewster

Figure 5. A portion of the transcript of Brewster’s July, 1915 species list.

Extracting Data via Citizen Scientists and DigiVol

DigiVol, first developed in 2011 to crowdsource the transcription of specimen labels, enables institutions around the world to engage volunteers to extract data of various types, such as text, species identifications, and species traits from images. Each institution is able to upload images and manage their crowdsourcing project through the DigiVol platform. Through a combination of gamification and engagement tools such as a user forum and secure, private emails, institutions can build volunteer skills, a sense of community, and commitment.

Screenshot of the Harvard University, Museum of Comparative Zoology, Ernst Mayr Library transcription portal on DigiVol platform, listing the number of expeditions and volunteers

Figure 6: The DigiVol dashboard for the Harvard Museum of Comparative Zoology, Ernst Mayr Library.

Volunteers on the website can contribute any time of day – 24/7, 365 days a year. They can choose from a broad array of “virtual expeditions” on DigiVol, from identifying animals in camera traps located in the Australian “bush” to transcribing specimen labels and field notes from locations around the world.

The DigiVol platform has seen more than 14,000 volunteers contribute 155 equivalent work years (7 hour days, 261 days a year) to the digitisation of over 6 million tasks at an estimated equivalent value of A$12 million.

In terms of the William Brewster project, 352 volunteers have transcribed over 27,000 pages of diaries and field notes, at an average 23 minutes per page. This contribution amounts to more than 5.8 work years at an equivalent cost of A$462,000.

Publishing Data to GBIF

The final output of these recent data pipeline investigations was the deposit of 1,853 species occurrence records at GBIF from the Journals of William Brewster. The data deposit comprises valuable biodiversity records extracted through transcription efforts on DigiVol, transformed into DarwinCore, and subsequently published on GBIF. The species occurrence records span diverse geographical locations, primarily focusing on New England but extending to other regions of the United States, the Caribbean, and Europe. Brewster’s meticulous observations, encompassing species observation, behavioral data, and environmental conditions offer a rich historical perspective on biodiversity dating back over a century.

Visualizations of species occurrence data deposited in GBIF from the journals of William Brewster

Figure 7: Species Occurrence Data from the Journals of William Brewster are now available on GBIF. Additional data deposits are planned in 2024. Data deposit: https://doi.org/10.15468/q45atb

This deposit of historic biodiversity data demonstrates how valuable BHL’s collection is to global biodata infrastructure, as it plays a crucial role in establishing species base lines, informing climate change studies, tracking key environmental indicators, and contributing to the development of global biodiversity monitoring platforms.

Having the data available in BHL, even as transcribed text is one thing, but it is the human resources to review and reconcile the data that is really required to facilitate the flow of historic biodiversity data into today’s bioinformatics ecosystems. To all the humans involved in this species data project, from the initial observation and recording, to the preservation, digitization, transcription, data extraction and reconciliation, machine-learning, and creation of vital ETL data pipelines, we thank you.


Related Works

Biodiversity Heritage Library Open Data Collection. (2022, November). Smithsonian Figshare.

Crowley, B., Dearborn, J., Funkhouser, C., Kalfatovic, M., Merriman, K., Iggulden, D., Trei, K., & Herrmann, E. (2023, October). Safeguarding Access to 500 Years of Biodiversity Data: Sustainability Planning for the Biodiversity Heritage Library [TDWG2023]. Biodiversity Information Standards, Hobart, Tasmania, Australia.

Data Flows Diagram. (2023, March). BHL Technical Team (BHL-TECH) Biodiversity Heritage Library.

Dearborn, JJ (2023, April). Unifying Biodiversity Knowledge for Life on a Sustainable Planet. Biodiversity Heritage Library. https://bhl.pubpub.org/

deVeer, J. (2021, February 24). Making the Best of Difficult Times: Accelerating the Transcription of William Brewster’s Writings During the COVID-19 Pandemic. Biodiversity Heritage Library.https://blog.biodiversitylibrary.org/2021/02/accelerating-transcription-brewster-covid19.html

deVeer, J. and Rinaldo, C. (2021, February 23). The Life and Work of Robert Alexander Gilbert: Empowering New Insights through Digitization and Transcription of Archival Materials. Biodiversity Heritage Library. https://blog.biodiversitylibrary.org/2021/02/life-work-robert-gilbert.html

Henshaw, Henry W. 1920. In Memoriam: William Brewster, Born July 5, 1851 – Died July 11, 1919. The Auk 37, 1 (1920), 1–23. https://doi.org/10.2307/4072953

Lichtenberg, M. (n.d.). BHL Data Model. BHL Github Repository. https://github.com/gbhl/bhl-us/tree/master/Documentation/DataModel

Mika, K. and Dearborn, J. (2022). [poster] Extracting expedition log data found in the Biodiversity Heritage Library. Through the door and through the web: releasing the power of natural history collections onsite and online, June 5, 2023. Edinburgh, Scotland, United Kingdom: Society for the Preservation of Natural History Collections (SPNHC). https://doi.org/10.5281/zenodo.6593457.

Richard, J. (2022, December 20). OCR Improvements: An Early Analysis. Biodiversity Heritage Library. https://blog.biodiversitylibrary.org/2022/07/ocr-improvements-early-analysis.html

Rinaldo, C. (2021, February 22). Nature Conservation and William Brewster: Insights From a Lifetime of Scientific Observations. Biodiversity Heritage Library. https://blog.biodiversitylibrary.org/2021/02/william-brewster-post-one

Trizna, M., & Dearborn, J. (2023, June). AI models are getting better and better at reading handwriting, but how can we find handwritten text to begin with? [poster]. 7th Annual Digital Data Conference, Leveraging Digital Data for Conservation, Ecology, Systematics, and Novel Biodiversity Research, Tempe, Arizona, United States of America. https://doi.org/10.25573/data.23523495.v1

November 9, 2023by mdimeo
Blog Reel, Featured Books

Making the Best of Difficult Times: Accelerating the Transcription of William Brewster’s Writings During the COVID-19 Pandemic

Screenshot of the BHL book viewer with crowdsourced transcriptions available for a handwritten journal.

This post is part of a series from the Ernst Mayr Library exploring the digitization and transcription of ornithologist William Brewster’s archival materials and the insights and scholarship made possible thanks to this work.

Many of us have searched for silver linings during the COVID-19 pandemic of 2020-2021. For many in the library and museum profession, one positive outcome of the mandatory transition to remote work has been the resurrection of some long-postponed projects. These are activities put on hold during normal times in deference to ever-proliferating, higher-priority onsite tasks. One such project in the Ernst Mayr Library and Archives (EMLA) of the Museum of Comparative Zoology (MCZ) at Harvard University is transcription of the digitized journals and diaries of William Brewster.

As discussed in Part One of this series, EMLA began digitization of its Brewster collection in 2012. In 2014, Library staff began transcription of Brewster’s digitized journals and diaries to produce a corpus of text files for the Purposeful Gaming and BHL project. As part of this project, EMLA was charged with leading an effort to select an appropriate open source transcription tool. At this time (early 2014), manuscript materials such as field notes and correspondence were beginning to appear in the Biodiversity Heritage Library (BHL) thanks to other grant-funded efforts among BHL members. Tools under consideration for Purposeful Gaming were thus also viewed as potential long-term transcription solutions for libraries contributing such materials to BHL.

After four months of research and testing, the grant partners selected DigiVol, a crowdsourcing transcription platform developed by the Australian Museum in collaboration with the Atlas of Living Australia, and EMLA staff set up a project space for the Library. One significant advantage of DigiVol is that it was developed specifically for the transcription of natural history materials such as specimen labels and field notes. Thus, there was the advantage of a pre-existing community of volunteers both interested in and adept at transcribing such items.

At the end of the Purposeful Gaming project (November 2015), volunteers and library staff had transcribed 12 volumes of Brewster’s journals and diaries, totaling 3,470 pages. That left 37,000 pages of Brewster’s writings in our collection awaiting transcription. Over the next four years, progress in the transcription work at EMLA was relatively slow at an average of 1,262 pages transcribed per year. During this time, EMLA had only a single, albeit enthusiastic, part-time library assistant who was able to dedicate one-third of her time or less to the project due to other higher-priority tasks.

On March 16, 2020, Harvard University staff began working remotely. Expecting a large influx of volunteers due to the quarantine in Australia, and anticipating an increase in staff availability at participating institutions, DigiVol administrators sent out a call for more materials to be uploaded for transcription. Working remotely, the author had the time available to regularly upload new volumes of Brewster’s diaries for transcription, respond to questions posted by volunteer transcribers, and validate completed transcriptions. Another EMLA staff member was able to offer several hours per week of her time as well. This was fortuitous as the volunteers were, indeed, ramping up their transcription activity, completing each diary volume in about two weeks time.

As the pandemic wore on, greater numbers of staff across the Harvard University library system were in need of work that could be done remotely. In response to this need, Harvard Library established a work share program whereby libraries in need of help could offer temporary assignments or projects to those in need of some additional work. By this time (September 2020), the volunteer transcribers were making so much progress that we were falling behind in the work of validating transcriptions. EMLA thus posted its project via the work share program, and we were fortunate to obtain the assistance of two library colleagues able to devote several hours each week to the work of transcription and validation.

With a total of 4 staff members and an average of 11 DigiVol volunteers (up from a previous average of 7) working on the project, more pages of Brewster’s writings were transcribed in 2020 than in any of the previous four years, as shown in the graphic below. There was a 242% increase in the number of pages transcribed over 2019, and a 165% increase over the average yearly number of pages transcribed from 2016 through 2019.

Bar graph showing number of pages transcribed per year, 2016-2020.

Number of pages transcribed per year, 2016-2020.

Transcriptions in the Biodiversity Heritage Library

The fruit of our transcribers’ labors began to materialize in 2018 when BHL developers introduced transcription import functionality to BHL. When member libraries began uploading transcribed pages of field notebooks and correspondence, the potential for access and discovery became readily apparent. Full-text indexing and searching of transcribed manuscript materials was now possible. For example, in researching the previous blog post about Robert Gilbert, the author was able to search for the name Gilbert in several transcribed volumes of Brewster’s journals. The result of such a search in Brewster’s 1898 journal is shown below.

Screenshot of the BHL book viewer with full text search results for "Gilbert" displayed.

Searching for the name Gilbert in the 1898 volume of Brewster’s Journals retrieves 42 references to Robert Gilbert in this volume. The search results are shown in the right-hand panel of the BHL book viewer.

Like most field notebooks, Brewster’s diaries and journals are replete with scientific names. In order for these to be recognized by BHL’s name-finding utility, Global Names Architecture (GNA), they must be carefully transcribed. Transcription tutorials for the Brewster project in DigiVol ask that scientific names be given extra care, expanding abbreviations and correcting misspellings where possible. Our team of volunteers and staff have done an exquisite job of accurately transcribing scientific names, sometimes even doing research when unsure of the correct spelling of a name. Some of the transcribers so enjoy delving into Brewster’s world that they frequently research people, localities, or unfamiliar terms used in his writings as well.

Screenshot of the BHL book viewer with scientific names found on the page.

A species observation list from Brewster’s 1900 Journal with transcription of species names and observations in the text panel on the right. Note the scientific names on the lower left, extracted from the transcription text by Global Names Architecture (GNA). Prior to addition of the transcription text, no names were found by GNA.

The end result is that the myriad of ornithological taxa mentioned in Brewster’s writings are being indexed in BHL. Thus a search across the full BHL repository for a specific taxon retrieves results from field notes and correspondence as well as the published literature. The resulting species bibliography will thus lead researchers to field observations, specimen collecting accounts, habitat and weather descriptions, and other content that may not be included in secondary sources.

Screenshot of a species bibliography in BHL for Chaetura pelagica.

A search for Chaetura pelagica (Chimney Swift) across the full BHL repository retrieves instances of this taxon in Brewster’s transcribed Journals as well as the published literature.

Because of the hard work of our volunteer transcribers, full transcriptions of thousands of pages of Brewster’s diaries, journals, and correspondence are now being added to BHL, making this content available for text analysis, data mining, digital scholarship, and the creation of new information. Connections with similar collections in other institutions will likely be discovered as more primary source material is digitized and transcribed. Brewster’s meticulous record of species occurrences, habitat, landscape ecology, and weather can help in addressing the current climate and biodiversity crises. His record of historical events and his account of the work of Robert Gilbert might be used in social justice or other humanitarian research efforts. The staff of the EMLA are deeply indebted to the team of volunteers and our two Harvard Library colleagues who have given so much of their time and energy to help unlock the riches of the Brewster collection for the benefit of all.

If you would like to help transcribe Brewster’s writings, there is still plenty of work to do. We will soon complete his diaries and journals after which we will begin transcribing his extensive collection of correspondence. If you are interested, please visit our project page.

February 24, 2021by michelle.underhill
BHL News, Blog Reel, Featured Books

The John Torrey Papers: Increasing Accessibility with Full Text Transcriptions in BHL

screenshot of a transcription in the from the page platform
black and white portrait of a young man, john torrey

Drawing of John Torrey by Sir Daniel MacNee. From the collection of Sir William Hooker, The Royal Botanic Gardens at Kew, England.

Since July 2016, the papers of taxonomic botanist John Torrey (1796-1873) have been the focus of a digitization and crowdsourced transcription project at the New York Botanical Garden (NYBG). Digitizing and Transcribing the John Torrey Papers, organized in coordination with the Biodiversity Heritage Library (BHL) and funded by the National Endowment for the Humanities and the Carnegie Corporation of New York, was created in an effort to digitize and make virtually accessible the correspondence of John Torrey and his colleagues, specifically letters received by Dr. Torrey.

Access to the correspondence of Dr. Torrey is particularly useful to natural scientists and natural historians alike as these letters provide a glimpse into the thought process and workflow for many of the 19th century’s greatest botanical minds. These letters are filled with a variety of different content, including commentary on plant taxonomy, personal biographical details, and anecdotal information on the daily lives of these natural scientists. The content within these letters aid in our understanding of the realities faced by these pioneering botanists, including the shift from the Linnaean system to the natural system and the systemic challenges faced by individuals in an emerging profession..

The John Torrey Papers collection contains over 9,500 pages from more than 350 correspondents, including Asa Gray, Amos Eaton, and Joseph Henry. The collection is organized alphabetically by correspondent then chronologically by the date each letter was written. Each page of each letter is scanned as a .tif file and uploaded to Macaw, BHL’s metadata management system, which aggregates the .tif file with a New York Botanical Garden MARC XML record. Once appropriate metadata has been applied in Macaw, the correspondence are uploaded to the Biodiversity Heritage Library Collection in the Internet Archive. The correspondence’s placement in the BHL Collection then allows for BHL to harvest new material from the Internet Archive on a weekly basis.

Digital content management is a crucial aspect of the project, but the heart of the process lies in the crowdsourced transcriptions done by volunteers from the New York Botanical Garden. These volunteers, recruited through citizen science outreach events, workshops, and tabling opportunities at various locations, work on the transcription platform From the Page, which is one of three transcription platforms whose exports are currently accepted by the Biodiversity Heritage Library, allowing the automatically-generated OCR for these items to be replaced with crowdsourced transcriptions, enabling full text search.

FTP Example Screenshot.PNG

A screenshot of a completed transcription from John Pierce Brace to John Torrey (May 1831) in From the Page.

The work done by these volunteers is a contribution to the project for two reasons. First, the handwriting, colloquialisms, and personal shorthand of many of Dr. Torrey’s correspondents are exceptionally difficult to read or interpret. With the correspondence of many correspondents containing several hundred pages, it takes a considerable amount of time for our volunteers to develop a proficiency capable of not only accurately deciphering poor handwriting but also of uncovering the intention or context of particular writings.

Second, the transcriptions produced by the Garden’s volunteers coupled with BHL’s new support for transcriptions facilitates full text searching of materials that would otherwise be unavailable. The introduction of full text searching to the world of transcriptions is extremely impactful as it allows researchers to execute a more thorough examination of resources available to them at a faster pace and with higher levels of productivity. A prime example of this is the Biodiversity Heritage Library’s page index of scientific names, as seen in the lower left hand corner of the figure below.

In effect, the Digitizing and Transcribing the John Torrey Paper’s volunteers are the bridge between the content we are looking for and how we look for it.

BHL Blog Screenshot.PNG

Screenshot of the completed John Pierce Brace page once the transcription and image files have been matched in BHL.

To date, more than 7,000 pages have been transcribed and more than 6,000 transcribed pages are available for view in BHL, but the goals of the Digitizing and Transcribing the John Torrey Papers project still have not been met.

With less than 1,500 pages remaining, the New York Botanical Garden is in search of new volunteers to help reach our goal and transcribe some of our most interesting correspondents remaining to be done such as William Jackson Hooker, the former Director of the Royal Botanic Gardens, Kew, and New England natural scientist Dewey Chester. While many of our volunteers are drawn to the project due to their interest in botany, many more are captivated by narratives not typically associated with botany, such as Carl Bogenhard’s struggles as an immigrant in the United States or Jacob Whitman Bailey’s experimentation with photography.

If you would like to become a volunteer transcriber for the Digitizing and Transcribing the John Torrey papers project, you can sign up online here or email [email protected] for more information or assistance with the sign up process.

If you would like to learn more about the life of John Torrey, works he published, or plants named after him, please visit out digital exhibit “John Torrey: American Botanist.”

September 24, 2019by Joel Richard
BHL News, Blog Reel, Tech Updates

BHL Adds Functionality Allowing Partners to Upload Crowdsourced Transcriptions of Digitized Archival Materials

Screenshot of a digital library book viewer with readable OCR.

The Biodiversity Heritage Library (BHL) has added functionality to allow BHL Partners to upload transcriptions in place of the automatically-generated OCR (Optical Character Recognition) for archival materials digitized in BHL. This functionality supports transcriptions generated as part of Partner crowdsourcing projects on Smithsonian Transcription Center, DigiVol, and From the Page.

Optical Character Recognition (OCR), also called text recognition, translates text characters in scanned documents into code that can be used for data processing and enables searching of document text. Handwritten archival materials like correspondence and field notes are notoriously problematic for OCR software. Full-text searching of these materials is significantly hampered by poor OCR output.

Screenshot of a digital library book viewer with gibberish for OCR.

Example of the poor automatically-generated OCR output for handwritten correspondence. Spencer Fullerton Baird and John Torrey correspondence, 1851-1860. Contributed in BHL from the LuEsther T. Mertz Library of The New York Botanical Garden.

Crowdsourcing the transcription of archival materials has become a popular way to generate machine-readable text that enables searching and discoverability. Several BHL Partners are using crowdsourcing platforms (e.g. Smithsonian Transcription Center, DigiVol, and From the Page) to transcribe field notes, correspondence, and other archival materials that they have digitized in BHL.

Screenshot of a project in the Smithsonian Transcription Center.

Example of a field notebook being transcribed in the Smithsonian Transcription Center. This notebook, Brasil 1979, Amazonia #3, from Cleofé Calderón is also available in BHL from the Smithsonian Institution Archives.

With this new functionality, these transcriptions can now be uploaded in place of the automatically-generated OCR for these items, allowing them to be full-text searchable and enabling our taxonomic name recognition software to index scientific names within their pages. Since the transcribed text can be viewed alongside the digitized page image, users can also more easily read materials with difficult-to-decipher handwriting. Thus, this new functionality makes it easier for researchers and the public to explore these valuable primary source materials and access specific information from their pages.

Screenshot of a digital library book viewer with readable OCR.

Above example with the OCR replaced with a crowdsourced transcription generated as part of The John Torrey Papers project from The New York Botanical Garden on From the Page. Spencer Fullerton Baird and John Torrey correspondence, 1851-1860. Contributed in BHL from the LuEsther T. Mertz Library of The New York Botanical Garden.

Screenshot of the BHL book viewer with a digitized field book and the transcribed text shown alongside the page in place of the OCR.

Since the transcribed text can be viewed alongside the digitized page image in the BHL book viewer, users can also more easily read archival materials. William Healey Dall’s Field Notes, 1871. Transcription generated on the Smithsonian Transcription Center. Contributed in BHL from the Smithsonian Institution Archives.

Screenshot of a book viewer with digitized archival materials and full-text search for "Yellow Palm Warbler".

Crowdsourced transcriptions allow digitized archival materials in BHL to be full-text searchable, as shown in this example searching for “Yellow Palm Warbler” within William Brewster’s 1903 journal. Transcription generated as part of the Ernst Mayr Library of Harvard University project on DigiVol. Contributed in BHL from the Ernst Mayr Library of the Museum of Comparative Zoology at Harvard University.

Screenshot of a book viewer with scientific name indexed on the page of an archival notebook.

Crowdsourced transcriptions allow BHL’s taxonomic name recognition software to index scientific names within the pages of digitized archival materials, as seen in this example in which Catharacta antarctica is indexed on a page within the first volume of the Ornithological Field Diaries of A. Graham Brown. Transcription generated as part of a project from BHL Australia and Museums Victoria on DigiVol. Contributed in BHL from Museums Victoria.

Participating Partners have begun uploading transcriptions to BHL. To date, transcriptions have been uploaded from Partner crowdsourcing projects with BHL Australia, Ernst Mayr Library of Harvard University, The New York Botanical Garden, and Smithsonian Institution Archives. This is an ongoing process, and more transcriptions will be uploaded to the Library over time.

Interested in transcribing archival materials? Several BHL Partners have active transcription projects on various crowdsourcing platforms. Follow the links below to explore the opportunities and get involved:

  • Ernst Mayr Library of Harvard University on DigiVol
  • The John Torrey Papers from The New York Botanical Garden on From the Page
  • Smithsonian Institution Archives on the Smithsonian Transcription Center
July 17, 2019by michelle.underhill

Help Support BHL

BHL's existence depends on the financial support of its patrons. Help us keep this free resource alive!

search

About BHL

The Biodiversity Heritage Library (BHL) is the world’s largest open access digital library for biodiversity literature and archives. BHL operates as a worldwide consortium of natural history, botanical, research, and national libraries working together to digitize the natural history literature held in their collections and make it freely available for open access as part of a global “biodiversity community.”

Join Our Mailing List

Sign up to receive the latest news, content highlights, and promotions.

Subscribe Now

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 319 other subscribers

Subscribe to Blog Via RSS

Subscribe to the blog RSS feed to stay up-to-date on all the latest BHL posts.

Access RSS Feed

Inspiring Discovery through Free Access to Biodiversity Knowledge.

The Biodiversity Heritage Library makes it easier than ever for you to access the information you need to study and explore life on Earth…for free, anytime, anywhere.

 

64+ Million Pages of
Biodiversity Literature Online.

EXPLORE

Tools and Services
to Transform Research.

EXPLORE

300,000+
Illustrations on Flickr.

EXPLORE

ABOUT | HARMFUL CONTENT | PRIVACY | SITE MAP | TERMS OF USE