Biodiversity Heritage Library - Program news and collection highlights from BHL
  • Home
  • News
  • Featured Books
    • All Featured Books
    • Book of the Month Series
    • BHL at 20
  • User Stories
  • Campaigns
    • Fossil Stories
    • Garden Stories
    • Monsters Are Real
    • Page Frights
    • Her Natural History
    • Earth Optimism 2020
  • Tech Blog
  • Visit BHL
Home
News
Featured Books
    All Featured Books
    Book of the Month Series
    BHL at 20
User Stories
Campaigns
    Fossil Stories
    Garden Stories
    Monsters Are Real
    Page Frights
    Her Natural History
    Earth Optimism 2020
Tech Blog
Visit BHL
  • Home
  • News
  • Featured Books
    • All Featured Books
    • Book of the Month Series
    • BHL at 20
  • User Stories
  • Campaigns
    • Fossil Stories
    • Garden Stories
    • Monsters Are Real
    • Page Frights
    • Her Natural History
    • Earth Optimism 2020
  • Tech Blog
  • Visit BHL
Biodiversity Heritage Library - Program news and collection highlights from BHL
Blog Reel, Featured Books

The Power of Community Science: How Smithsonian Volunpeers Transform Scientific Field Notes

Screenshot of Smithsonian Transcription Center displaying handwritten journal next to transcribed text

Last month, Smithsonian Libraries and Archives (SLA), Smithsonian Transcription Center (STC), and the Biodiversity Heritage Library (BHL) celebrated a significant milestone – technical staff worked collaboratively to integrate over 43,000 pages of transcription materials from STC into BHL. An additional 151,362 scientific name access points have now been added to the BHL search index for SLA archival field notes. These transcriptions enhance BHL’s full-text search, enable taxonomic name recognition, improve accessibility for vision-impaired users, and support climate research.

The Smithsonian Field Books Collection

Yellow notebook with handwritten text and handdrawn map

Dall, William Healey. (1895). Alaska Coal Fields, Bogosloff Volcano, Corral Hollow, California, 1895 (Vol. 1, pp. 81 and 82).

The Smithsonian Field Books Collection is a set of primary source archival records selected from the Smithsonian Libraries and Archives (SLA). This material dates from the late nineteenth century and the first comprehensive biological survey of the continental United States to the most recently accessioned materials at the Smithsonian Institution Archives.

The collection includes personal records of naturalists and scientists such as William Healey Dall (1845 – 1927) and Cleofe Calderon (1929 – 2007) at work around the world and expedition records such as the Western Union Telegraph Expedition (1865 – 1867) of Russian America and the United States Exploring Expedition (1838 – 1842) of the Pacific Ocean.

Naturalists and scientists in the field recorded their firsthand observations and data in a wide variety of forms including diaries and journals, hand-drawn maps and tables of data, observation logs and specimen catalogs, correspondence and reports, manuscripts, sketches, photographs, even audio recordings.

Recognizing the Value of Accurate Transcription

Mining the rich information embedded in these field notes depends on accurate transcription. The majority of the field notes are handwritten – even the most recent ones. Digital surrogates provide high resolution images but fail to afford anything beyond visual accessibility. Transcription, however, opens up this material to full-text searching, pattern recognition, visual accessibility aids, and more.

Beyond the handwriting itself, the field notes also contain other non-textual information that can only be captured through transcription. Examples include ornithologist Martin Moynihan’s notations of bird song in remote regions of Central America or Charles Dolittle Walcott’s sketches and diagrams of sedimentary stratification where his fossil specimens were found in the canyons of the American Southwest.

Handwritten notes with a drawing of a bird

Moynihan, M. (1963). 1962-1965 Andean Birds Mixed Flocks, Colombia, (4 of 4) (p. 78).

Chief naturalist with the U.S. Department of Agriculture, Vernon Orlando Bailey’s “Journal kept by Bailey on field trip to Wyoming and New Mexico, March 15-June 1906” focused on extermination techniques of gray wolves that would bring them to near extinction in the continental United States not long afterwards. The transcript includes descriptions of the sketches he included in his notes.

Screenshot of Smithsonian Transcription Center displaying handwritten journal next to transcribed text

Smithsonian Transcription Center image with corresponding transcription, including a text description of an ink drawing of a cow. https://transcription.si.edu/view/6656/EBOtg

The Smithsonian Field Book Project’s initial goal was to catalog these hidden biodiversity research material, improving discoverability in keeping with FAIR (Findable Accessible Interoperable Reusable) Data Principles. Doing so quickly resulted in additional researcher demand for more access and usability – first to view the field notes online (digitization) and then to examine them more closely (transcription). Grant funding assisted in a major rapid-capture digitization effort. However when it came to transcribing the digitized field notes, the Archives simply lacked the capacity to meet the level of researchers’ demand.

Turning to digital volunteers and the just-launched Smithsonian Transcription Center in 2013 changed the transcription equation beyond our best expectations both in volume and in accuracy. We quickly came to recognize these community scientists as collaborators, “volunpeers”, in the effort to advance and disseminate knowledge.

We are far from the end of this journey. More than half the collection remains to be digitized, and over two thousand digitized field notes still need transcription.

Inspired to Transcribe

These archival field notes contain vital historic biodiversity information. By transcribing the handwriting into machine readable text, volunpeers can help inform current day scientists and assist with their research on a multitude of topics such as climate change, the extinction crisis, or the spread of invasive species. Transcribing field notes can also take volunpeers on an adventure across time and distance, accompanying the writer on their journey. Having these adventures transcribed in machine readable text can also help inform science historians assisting them by making these historic documents more easily findable, searchable, and reusable.

Volunpeers collaborate to transcribe as accurately as possible the pages of the field journals provided. Multiple volunpeers will work on each project page transcribing the hard-to-read handwriting into machine readable text. Once satisfied the transcription is as complete as possible, a volunpeer will mark the transcribed page as complete. The volunpeers will then move onto the next page of the field notes until the project is finished.

Approval workflow for Smithsonian Transcription Center

Review is the second step in the transcription process.

Having many volunpeers working on one project helps ensure the quality of the transcription. What one volunpeer finds illegible, others may be able to read, especially as all the volunpeers working on a particular project become more familiar with the handwriting.

Volunteers have transcribed hundreds of field journals over the years. Two favorite examples have been the field journals of Vernon Orlando Bailey, a field naturalist who journeyed throughout the U.S. midwest studying and collecting mammals. Another volunpeer favorite were the papers of Arctic explorer and naturalist Robert Kennicott. Read more about these volunpeer experiences on the STC blog.

The Power of the Smithsonian Transcription Center

When the first of the field books were added to Smithsonian Transcription Center (STC), the program was still available only as a beta version. Approximately 450 volunpeers were transcribing on the site (co-author Siobhan Leachman among them), with the first 6,000 completed pages under their belt. Even during this moment of immense energy and fresh connections, it must have been difficult to imagine what STC would become. Today, a little over a decade later, more than 91,000 individual volunpeers have worked together to transcribe and review over 1.4 MILLION (!) pages of historic and scientific collections.

STC is the largest digital volunteering and crowdsourcing program at the Smithsonian Institution, and provides opportunities to engage with and contribute to digitized materials from across the full breadth of content areas represented by its museums, archives, and libraries. Through collaborative transcription and review, Smithsonian staff and digital volunteers work together to ensure that this content is more readable, accessible, and text-searchable across Smithsonian data systems and beyond.

Currently, the most active and popular projects are the Freedmen’s Bureau Transcription Project, a collaboration with the National Museum of African American History and Culture that deepens insight into the Reconstruction period and empowers African American genealogical research, and Project PHaEDRA, a collaboration with the Harvard-Smithsonian Center for Astrophysics that illuminates the work and discoveries of early women computers at the Harvard College Observatory.

If you feel inspired by this data access success story, consider joining the digital volunteer community, or sign up for our newsletter to stay up-to-date on upcoming projects.

Reusing the Liberated Data

The successful integration of transcription materials into BHL makes historical scientific data more accessible and useful. Through the collaborative efforts of technical staff from the Biodiversity Heritage Library, Smithsonian Libraries and Archives, and the Smithsonian Transcription Center and dedicated volunpeers, over 151,362 scientific name access points were added to the BHL search index, greatly enhancing search capabilities over the digitized corpus of Smithsonian field notes and archives. Out of the 556 eligible items reviewed, 522 were uploaded, contributing 43,460 pages of improved OCR text. These contributions not only improve BHL’s full-text search and taxonomic name recognition services but also provide better accessibility for vision-impaired users, and support ongoing biodiversity and climate change research.

Smithsonian Transcription Data outcomes report

A special thanks goes to Mike Lichtenberg, BHL’s Lead Developer and Systems Architect and Paul Day, Lead Developer at Smithsonian Transcription Center. As with many data improvement and platform enhancement projects, the requisite technical expertise is pivotal in ensuring the success of our collective efforts!


References and Resources

Dearborn, J., Lichtenberg, M., Richard, J. M., deVeer, J., Trizna, M., & Mika, K. 2023. [Presentation] Unearthing the Past for a Sustainable Future: Extracting and Transforming Data in the Biodiversity Heritage Library for Climate Action. Presented virtually at TDWG, Tasmania, Australia 2023. https://www.youtube.com/watch?v=8sGssyrpuJw

Trizna, M., & Dearborn, J. June 2023. [Poster] AI Models Are Getting Better at Reading Handwriting, but How Can We Find Handwritten Text to Begin With?. 7th Annual Digital Data Conference, Leveraging Digital Data for Conservation, Ecology, Systematics, and Novel Biodiversity Research, Tempe, Arizona, United States of America. https://doi.org/10.25573/data.23523495.v1

Dearborn, J., & Mika, K. June 5, 2022. [Poster] Extracting Expedition Log Data Found in the Biodiversity Heritage Library. Through the Door and Through the Web: Releasing the Power of Natural History Collections Onsite and Online, Edinburgh, Scotland, United Kingdom: Society for the Preservation of Natural History Collections (SPNHC). https://doi.org/10.5281/zenodo.6593457

September 30, 2024by Siobhan Leachman
BHL News, Blog Reel, Tech Updates

BHL is Round Tripping Persistent Identifiers with the Wikidata Query Service

Diagram of the 8 steps detailed below of the BHL wikidata round trip.

In the Spring of 2022, the BHL Cataloging and Metadata Committee investigated the possibility of harvesting persistent identifiers (PIDs) from Wikidata as part of the group’s longstanding project to disambiguate and deduplicate author records in the BHL database. The motivation behind this one-time experimental data harvest was to see if BHL could:

  1. Enhance BHL author records with additional PID data points;
  2. Improve the committee’s ability to disambiguate author names in the BHL database; and
  3. Respond to an outstanding user request from two of Wikimedia’s super star editors, Siobhan Leachman and Andy Mabbett, to expose BHL’s author data on BHL and include hyperlinks to other authoritative knowledge bases on the web.
How do PIDs work? Users are directed to the URL that is associated with the identifier. If the physical location of the object changes, the URL is the only modification required; the PID itself does not change.

Persistent identifiers can become resolvable links that connect two knowledge bases (like BHL and Wikidata) together.

In particular, Wikimedians wanted to see the Wikidata Q identifier exposed, providing a link to the corresponding creator item record in Wikidata.

There are multiple motivations for undertaking this work. By adding the BHL Creator ID to the corresponding Wikidata item, Wikidata editors help link BHL to the richer biographical data about that person held in Wikidata. The Wikidata item for a person may contain links to their Wikipedia page or to images of the person held in the image repository Wikimedia Commons. Wikidata items also act as identifier hubs and contain links to other databases and identifiers.

Maria Sibylla Merian's wikidata entry, including name variants, biographical information, external sources, an image, and more

Maria Sibylla Merian in Wikidata, rendered with the Reasonator Tool; note the External sources section and BHL’s Creator ID entry.

By adding the BHL Creator ID to this list of identifiers, the Wikidata editor is linking the content held in BHL to the content held in multiple other datasets and repositories.

These extra author data points provide Wikimedians and BHL catalogers with crucial clues that aid in name disambiguation. In particular, hyperlinks to other knowledge bases are incredibly valuable because they lead to new knowledge pathways that help confirm a person’s identity in a complex game of “Who’s Who?”

Workflow Overview

The diagram below illustrates the experimental data pipeline from BHL to Wikidata and back.

Diagram of the 8 steps detailed below of the BHL wikidata round trip.

The BHL to Wikidata Round Trip Overview

The BHL to Wikidata Round Trip Steps

  1. BHL Creator ID and record are created when digital content is harvested into BHL and/or articles are defined in BHL.
  2. Wikimedians record BHL Creator IDs as a statement in corresponding Wikidata items via various community workflows. (See: Mix’n’match deep dive below and/or read about QuickStatements for two ways volunteers are populating Wikidata with BHL’s Creator IDs).
  3. BHL subsequently harvests additional metadata from Wikidata for any author item where a BHL Creator ID statement has been added.
  4. Data outputs are analyzed for quality and breadth.
  5. SPARQL queries and data outputs are iteratively refined to match BHL’s requirements until a quality dataset can be generated for import.
  6. Clean data is imported into the BHL database.
  7. The new author record sidebar displays data on BHL.
  8. PIDs are converted into resolvable URIs, opening up new research pathways for BHL’s users.

Quick note: In Wikidata BHL author records are represented by the Wikidata BHL Creator ID; the presence of this Wikidata property in a Wikidata item provides a powerful connection point that can be used later by BHL to ingest more information about any named entity. A popular author matching tool in Wikidata is Mix’n’match.

Getting BHL Authors in Wikidata with Mix’n’match

Mix’n’match is a Wikidata tool that empowers editors to match Wikidata items to entries in other databases. One of the datasets in Mix’n’match is the BHL Creator ID dataset.

Summary from the mix'n'match tool for the BHL creator ID dataset

A summary of the completeness of matching the BHL Creator ID dataset to Wikidata as of January 2023

When the BHL Creator ID dataset is uploaded into Mix’n’match, an algorithm undertakes fuzzy name matching. The tool then suggests a preliminary match to a Wikidata item. Wikidata editors working on the dataset have the choice of whether to confirm the suggested match or to remove it.

Sample matches suggested by the mix'n'mtch tool

Examples of possible matches suggested by the Mix’n’match algorithm with the BHL creator in green and the Wikidata item in blue.

If the Wikidata editor confirms the match, the BHL Creator ID is automatically added to the Wikidata item for the creator. If the editor rejects the match, the editor can add a correct Wikidata item ID, create a new Wikdiata item if the creator is not yet in Wikidata or, in some cases, decide that the BHL Creator ID isn’t applicable to Wikidata. If the editor simply rejects the match without further action, the BHL Creator ID will be added to the unmatched portion of the dataset.

Examples of confirmed and unconfirmed matches in the Mix'n'match tool

Results after the Wikidata editor Ambrosia10 has confirmed or removed the Mix’n’match algorithm suggestions.

The act of linking the BHL Creator ID to the Wikidata item also helps to disambiguate the creator. It removes much of the uncertainty about who the creator is. Names are not unique but by linking identifiers to biographical data, editors can ensure that similarly named people or people whose names have changed over time are all associated with the correct Wikidata item. This addresses the age-old problem of how to assign correct attribution to the right person.

This work also assists BHL. By ensuring BHL Creator ID’s are matched to the Wikidata item, the Wikidata editor can assist BHL in weeding out the duplicate entries for creators in BHL’s database. After matching has taken place, BHL is also able to ingest any of the other identifiers listed on the creator’s Wikidata item, thus enriching the metadata held in BHL’s database.

Currently there are over 230,000 entries in the BHL Creator ID dataset. Of these, just over 41,000 have been matched to Wikidata items. There is a long way to go before this dataset is complete. In the spirit of “many hands make light work,” BHL encourages Wikidata editors to work alongside BHL staff who were recently trained and certified in Wikidata Advanced Concepts by Wiki Education. Our collaborative work will help increase the number of Creator IDs in Wikidata.

Summary from the mix'n'match tool for the BHL creator ID dataset

BHL Creator ID dataset completeness statistics as of January 2023.

Once BHL Creator ID statements were recorded on Wikidata items, either through manual edits or workflows like Mix’n’match or QuickStatements, then BHL created custom SPARQL queries using the Wikidata Query Service to output data of interest.

Example of a SPARQL query using the wikidata query service

The SPARQL query language is used to query and return data from semantic databases like Wikidata.

Once the data was brought back, the BHL Cataloging and Metadata Committee discussed and reviewed the records with the aim of bringing key data points back into the BHL database.

Sample data returned from the SPARQL query, including item ID, itemLabel, BHL_creator_ID, date_of_death, VIAF_ID, Library_of_congress_authority_ID, ORCID_ID, ISNI, date_of_birth

Snippet of raw data brought back by the BHL Cataloging and Metadata Committee’s SPARQL query.

In total, the BHL Cataloging and Metadata Committee was able to “round trip” 88,507 persistent identifiers (PIDs) associated with BHL Creator records from the following authoritative knowledge bases:

  • Virtual International Authority File (VIAF)
  • Wikidata
  • Social Networks and Archival Context (SNAC)
  • ResearchGate
  • ORCID
  • Library of Congress Authorities

There are still more PIDs out there to gather but practicality calls for moderation – moderation of the quantity that the Wikidata Query Service can bring back and the amount feasible to curate in BHL. After many discussions, the above list was chosen by narrowing down the PIDs that seem most appropriate in the BHL context. A comprehensive policy on Uniform Resource Identifiers (URIs) in the BHL ecosystem is now being drafted by BHL committees.

Additionally, in post-work analysis, the Committee did find that the data modeling of corporate names in Wikidata differs from the library world, and pulling in identifiers for corporate names was not worth the effort (at least at this time).

Below is an example of Académie des sciences (France), which receives multiple item records in library databases like VIAF for each name change. However, in Wikidata these name changes are collapsed and will appear on one item record; this divergence means that one-to-one matches are not possible without a lot of manual clean-up from BHL Staff.

Example of the Wikidata entry for VIAF ID for Académie des sciences

Divergent data models: Wikidata corporate names collapse name changes.

It’s important to note that data modeling in Wikidata is in its nascent stages and librarians should have a vested interest in shaping the best practices of the community so all of humanity can model the world’s knowledge together. As we have seen with this one-time experiment, there is little to lose, and so much to gain!

The Result: A New Author Record Sidebar for BHL

Thanks to the feedback from the Wikimedia community, BHL’s Lead Developer created an Author Record sidebar. Click on any BHL author name on the BHL website and the author details will appear on the right-hand side. Below is an example of Australian paleontologist and ornithologist, Patricia Vickers-Rich in BHL. 

Graphic example of an author record with detailed sidebar and the pages to which it links out, such as ORCID, Library of Congress Authorities, Wikidata, SNAC

A New Wikimedian Driven Feature: The BHL Author Record Sidebar!

The goal of this new feature was to expose existing and recently harvested identifiers as resolvable links that could open up critical knowledge pathways for BHL’s users on their information journeys. Other key data points are provided such as preferred name forms, entity type, and alternative name forms that give additional context for disambiguation research.

A quote from Siobhan Leachman on BHL's new author sidebar: Adding these identifiers to the author sidebar makes my life so much easier. At a glance, I can quickly confirm the identify of creators. The author sidebar or lack thereof also highlights whether further work is needed in wikidata to link these BHL creators to their identifiers."

Next Steps for BHL Committees?

Please let us know in the comments what you think of this experiment. Should BHL pursue similar persistent identifier harvests for other entity types? Or perhaps, a recurring harvest for BHL Authors?

Additional BHL entities:

  • Titles
  • Scientific Names
  • Subjects
  • Media (illustrations, scientific plates, photos)

Let us know! Your feedback is crucial to BHL’s evolution as a biodiversity knowledge base.

If you are new to Wikidata, Mix’n’match is a great place to start. There are many tutorials and resources available on the web including this great YouTube introduction on how to use the tool.

Additionally, a workflow for BHL articles has been piloted and is currently underway thanks to BHL’s Persistent Identifier Working Group. For more details on how to roundtrip BHL Article Q identifiers, please refer to the group’s documentation at: Wikidata:WikiProject_BHL/Projects

BHL committees and working groups are actively scoping many projects. Sign-up here to get involved!

Related Posts

For more information on the assignment of DOIs, a very important type of persistent identifier specific to scholarly publications, check out Nicole Kearney’s blog post What Is BHL’s New Persistent Identifier Working Group DOI’ng?

For a list of all of the new features and data quality improvements BHL made in 2022, check out the post BHL Technical Development: Year in Review

February 15, 2023by Siobhan Leachman
BHL News, Blog Reel

BHL and WikiCite 2018

WikiCite conference in Berkeley, ca, 27-29 November 2018.
WikiCite conference in Berkeley, ca, 27-29 November 2018.

WikiCite conference in Berkeley, ca, 27-29 November 2018. Photo Credit: Diane Shaw. Graphic by Dario Taraborelli (CC-BY 4.0) with photo by Tehani Schroeder.

In November 2018, Diane Shaw, Katie Mika and Siobhan Leachman attended WikiCite 2018 in Berkeley, CA. WikiCite is a Wikimedia initiative that aims to develop a database of open citations and linked bibliographic data. Since sources are the foundation of Wikipedia’s claim to authority, WikiCite is working to build a repository of bibliographic data that is open, structured, and separable. WikiCite also draws from the academic community in which citation data (most often from peer-reviewed articles) is crucial for creating and linking knowledge. At the outset, WikiCite aimed to build this repository specifically for use in Wikipedia and other Wikimedia Foundation projects. Since the first meeting in 2016, this idea has grown into building a universal repository of sources to serve the sum of all human knowledge, leveraging Wikidata as its infrastructure.

Diane Shaw (left) and Siobhan Leachman (right) at WikiCite 2018.

Diane Shaw (left) and Siobhan Leachman (right) at WikiCite 2018. Photo Credit: Siobhan Leachman.

Now in its third iteration, the conference included extended and short presentations, numerous lightning talks, strategy tracks, tutorials, data modelling, and a hack day (renamed a Do-athon Day to be more inclusive of those who aren’t computer science experts). The purpose of the conference, open primarily by invitation to a group of approximately 100 attendees from across the United States and abroad, was to work towards a “vision of creating an open repository of bibliographical data to support the citation and fact-checking needs of Wikimedia projects, and possibly, to serve as an open infrastructure for research, education, and information quality across the web.” Librarians were well-represented among the attendees, as were linked open data advocates, software engineers, data scientists, and active members of the Wikimedian community.

Unlike some of the other very popular Wikimedia Foundation products (Wikipedia in 303 languages; Wikimedia Commons hosting image, audio, and video files; Wikisource for full text transcriptions; and Wikidata for storing structured data, among others), WikiCite has no real online presence itself, but has grown as a community of practice dedicated to the support of a bibliographical metadata commons enabling reliable, verified linked open data connections both among Wikimedia projects and externally with many other online catalogs, databases and reference sources for a variety of libraries, galleries, archives, museums, scientific institutes, and other similar organizations, including OCLC, the Internet Archive and the Biodiversity Heritage Library. We are not exaggerating when we say it was awe-inspiring to be part of this dedicated group harnessing technical knowledge, bibliographical expertise, and lots of creativity and imagination in an effort to bring about a future of linked open data serving all kinds of scholarly information needs.

What WikiCite hopes to make possible goes beyond a simple structure for providing and sharing linked open data: attendees at the conference were also experimenting with new and improved ways to discuss, analyze, curate, vet and annotate sources online.
Dario Taraborelli, Director of Research at the Wikimedia Foundation, gave a terrific overview of WikiCite in his opening talk: Here Be Dragons: Uncharted regions of the bibliographic commons. Other talks and projects that would be of particular interest to libraries include presentations about projects at the National Library of Sweden and the National Library of Wales using structured linked open data for name and subject authority files and to help document the history of national book trade; Linked Data for Production (LD4P), a collaborative pilot project using BIBFRAME for creating structured linked data with connections to Wikidata, the Virtual International Authority File (VIAF) and WorldCat; initiatives to use Wikidata, VIVO and Scholia to generate scholarly profiles; and ways to incorporate Wikidata training into library education, including classes on Library Carpentry.

Katie Mika presenting as part of WikiCite and the Biodiversity Heritage Library at WikiCite 2018.

Katie Mika presenting as part of WikiCite and the Biodiversity Heritage Library at WikiCite 2018. Photo Credit: Diane Shaw.

Katie, a former BHL National Digital Stewardship Resident from Harvard’s Museum of Comparative Zoology, and Siobhan, a citizen scientist and linked open data champion from New Zealand who has been a devoted transcriber of natural history materials in the Smithsonian Transcription Center, gave a talk on WikiCite and the Biodiversity Heritage Library. They described some of the unique challenges for heritage literature and metadata, and demonstrated how open access citations, images, and details gleaned from BHL and other open natural history digital repositories are applied to Wikimedia Foundation projects to support essential documentation of scientists, literature, and rare and endemic species.

Unlike modern scientific literature, much historic literature does not have original digital identifiers (DOI’s) attached. These identifiers enable a particular publication to be identified and provide a persistent link to its location on the Internet.  Without these identifiers it is difficult to ingest historical bibliographic data into Wikidata in bulk. If historic bibliographic data is not able to be ingested or reused, scientists are unable to make use of this important resource. Siobhan expressed her frustration that much of the citation metadata of historic biodiversity literature is not currently found in Wikidata and urged the attendees to help resolve this problem.

Siobhan Leachman presenting as part of WikiCite and the Biodiversity Heritage Library at WikiCite 2018.

Siobhan Leachman presenting as part of WikiCite and the Biodiversity Heritage Library at WikiCite 2018. Photo Credit: Diane Shaw.

Siobhan’s key takeaway from the conference was the various discussions on how to solve the difficulties of getting bibliographic information on historic biodiversity literature into Wikidata. Also of interest were the discussions on how book data is modelled in Wikidata (see https://www.wikidata.org/wiki/Wikidata:WikiProject_Books) and discovering a practical new tool to help disambiguate authors within wikidata: Author Disambiguator. Katie enjoyed the opportunity during the summit day’s events to discuss the potential for a bibliographic commons to support citations and information provenance across the web and in support of Open Science practices. On the Do-a-thon Day, Diane spent time discussing approaches for modelling the properties of holotypes on Wikidata with Siobhan, biophysicist and open scientist Daniel Mietchen, and Terry Catapano of Plazi and UC Berkeley’s Bancroft Library.

The opportunity presented by WikiCite for libraries and public knowledge creation cannot be overstated. As of November 2018, there are over 21 million publication items in Wikidata, which accounts for 40% of the entire database. More than 160 million Wikidata statements use the property “cites” (P2860) to connect items. Wikidata is open to contribute to, edit, reuse, and extend. Not only can we leverage Wikidata for metadata enrichment and linked data connections, but we can build bibliographic applications on top of this open database, track the provenance of citations, and study patterns of information reuse. Wikicite advocates for an open destination to more easily discuss, analyze, curate, annotate, and vet bibliographic resources used in Wikipedia and across the web. As connectors to information, libraries have staff who can offer a tremendous amount of expertise for modeling difficult kinds of data (like books), while ensuring equitable access to information.

The easiest way to get involved is to edit! Explore Wikidata using reasonator and SQID (“squid”). Head over to Youtube to watch talks and tutorials on Wikidata and SPARQL. Match some Wikidata items with their external identifiers using Mix ‘n Match. And join the wikicite-discuss mailing list to join conversations about adding bibliographic data to Wikidata and WikiCite program updates.

January 31, 2019by Siobhan Leachman
Blog Reel

The Serendipitous Discovery of Susan Fereday: A Story about the Impact of Citizen Science

View Full Size Image
Self Portrait, Susan Fereday. National Library of Australia. Source: WikiCommons.

I love volunteering for the Biodiversity Heritage Library. I taxo tag images in the BHL Flickr account. This assists the use of these images by BHL as well as other institutions that use BHL content. It is also my favorite way of exploring BHL. I get a real thrill out of the serendipitous discoveries I make while tagging.

My most recent BHL adventure resulted from tagging an album of images from the boringly named but absolutely fabulous Botany of the Antarctic voyage of H. M. discovery ships Erebus and Terror in the years 1839-1843. Amongst the many images in this album was one of a particular species of seaweed – Nemastoma feredayae.

View Full Size Image
Nemastoma feredayae. Art by William Henry Harvey. The botany of the Antarctic voyage of H.M. discovery ships Erebus and Terror in the Years 1839-1843. http://biodiversitylibrary.org/page/28467702. Digitized by the Missouri Botanical Garden.

I tagged the image with the species name given, and then I attempted to confirm the current name of the seaweed. In doing so, I stumbled across the fact that the seaweed was named in honor of Mrs. Susan Fereday. Who was this mystery woman?

Unable to resist going down that rabbit hole, I googled her. I discovered that Susan Fereday emigrated to Tasmania, Australia from England in 1846. She was a talented artist. So talented, her artwork is now in the collection of the National Library of Australia. She concentrated on painting beautiful images of fauna and flora in and around the area where she lived.

View Full Size Image
Hibbertia sericea by Susan Fereday. National Library of Australia. Source: WikiCommons.

She was also a keen collector of seaweed specimens. She corresponded with and sent specimens to one of the foremost experts in algae of the day, William Henry Harvey. Harvey in turn honored Fereday’s contribution to the study of algae by naming two species after her. It was the image of one of those species drawn by Harvey that I had tagged in the BHL Flickr feed.

View Full Size Image
Portrait of William Henry Harvey. Oliver, F. W. (Francis Wall). Makers of British botany. (1913). http://www.biodiversitylibrary.org/page/1073575. Digitized by MBLWHOI Library.

While researching Fereday, I was disappointed to see she did not have Wikipedia page. Being a keen Wikipedian, I decided to rectify this. While drafting her article, I realized that several sources had different dates as her birth date. I emailed the National Library of Australia via their “Ask a Librarian” service to ask them to help me confirm that their records were correct.

I received a fabulously researched reply from Damien Cole, one of their librarians. He discovered that the birth date confusion was due to Susan Fereday having a sister of the same name, who had died prior to our Susan being born. Fereday was actually born in 1815! As a result of this research, the National Library subsequently edited their records to give Susan Fereday her birth date, and I obtained a reputable citation supporting that information for Fereday’s Wikipedia article.

The National Library of Australia has also shared the changes they made to their database with the Australian National Herbarium as well as Design and Art Australia Online. The Library even went so far as to contact the Encyclopedia of Australian Science to inform them of the Wikipedia article in the hope that that organization might also consider including Fereday in their website.

All of this resulted because the MBLWHOI Library (the Marine Biological Laboratory and the Woods Hole Oceanographic Institution Library) scanned the volumes on the Antarctic voyage of H. M. discovery ships Erebus and Terror and made them available to BHL. Because BHL added the images from those volumes to Flickr, I was able to tag them.

It just goes to show that when you mix citizen science with digitization and the ability to freely reuse content, everyone benefits.

I would love it if people joined me in taxotagging BHL Flickr images. Instructions can be found here.

Anyone can create or improve Wikipedia. For a beginner’s guide, see this article.

And if you think you can add to and improve Susan Fereday’s article, go for it!

May 23, 2017by Siobhan Leachman
BHL News, Blog Reel

Be Like a BHL Librarian and Edit Wikipedia for #1Lib1Ref

You too can be like a librarian … a Biodiversity Heritage Library librarian! The Biodiversity Heritage Library wants people to use its resources, and Wikipedia is encouraging librarian-minded folk to add citations to articles via their #1Lib1Ref campaign (15 January – 3 February 2017). Talk about a perfect match. You can help by adding citations from BHL to Wikipedia.

Don’t worry if you haven’t edited Wikipedia before. What follows is an easy “How to” guide to adding a citation from works held in BHL to a Wikipedia article.

Create a Wikipedia Account

Your first step is to create a Wikipedia account. Although you can edit Wikipedia without an account, if you create one, you can use Wikipedia’s fabulously easy “Visual Editor” tool. This comes with a feature which allows for ease of citation of references. Click here to see how to create an account and how to enable the Visual Editor tool.

Find an Article

Your second step is to find an article that needs to be improved with added citations. One way is to use the Wikimedia Citation Hunt tool. When using this tool, I narrow my search to articles which need the Biodiversity Heritage Library resources. I do this by filling in the “What makes you tick” box with a search term. In the example below, I’ve added “Beetles” to find beetle articles that need citations, but there are many other terms you could use. For example, try “mosses”, “Fish of Canada”, “lizards” or any subject you are interested in.

View Full Size Image

In the above case, to find the needed citation I would put “Myxophaga” (a suborder of Coleoptera or beetles) in the search box on the BHL website and see if I could find any article, journal or book in BHL supporting the statement in the Wikipedia article to which I want to add a citation.

You can also work from the other end and find an article in BHL that you can then cite in Wikipedia. So for example I found this article which first describes the Pericoptus frontalis beetle. This is an article that can and should be cited in the Pericoptus frontalis beetle Wikipedia page.

View Full Size Image

Citing Reference in Wikipedia

So you’ve logged in to Wikipedia, enabled the Visual Editor tool, and found the article in BHL you want to cite. Now go to the page you want to edit and click on the “Edit” tab so that the Visual Editor edit bar appears.

View Full Size Image

Place your cursor in the Wikipedia article where you wish to add the citation and then press “cite” on the edit bar.

View Full Size Image

The “add a citation” box should then appear.

View Full Size Image

You then have a choice of either using the “Automatic”, “Manual” or “Reuse” citation method. If you have the URL, DOI or PMID of the article, you can use the “Automatic” citation tool to generate a citation. Alternatively, use the “Manual” citation, which will give you different options to choose from depending on where you obtained your citation. The image above shows you the manual citation options.

In this case, I’m citing an article in a journal rather than a book or a newspaper article, so I would choose the “Journal” option. Clicking on “Journal” brings up a template. Fill in all the information you can in the appropriate boxes and then press “Insert” on the top right of the box. Make sure you provide the stable URL link in BHL (learn more).

View Full Size Image

Pressing “insert” brings up a “Save your changes” box asking you to briefly describe the changes you’ve made to the article. Make sure you add the hashtag #1Lib1Ref to your edit summary as this will add your edit to the campaign total.

View Full Size Image

You then press the “Save changes” button and with that you have added a citation to Wikipedia!

For further Information on the Wikipedia #1Lib1Ref campaign, click here.

Follow #ILib1Ref on social media, 15 January – 3 February 2017 to learn more about the campaign and see how adding citations to Wikipedia can improve the resource for everyone!

If you have questions or get stuck, see this page to learn more about the various outlets you can use to find help and information. You can also tweet BHL at @BioDivLibrary or Siobhan at @SiobhanLeachman.

January 17, 2017by Siobhan Leachman

Help Support BHL

BHL's existence depends on the financial support of its patrons. Help us keep this free resource alive!

search

About BHL

The Biodiversity Heritage Library (BHL) is the world’s largest open access digital library for biodiversity literature and archives. BHL operates as a worldwide consortium of natural history, botanical, research, and national libraries working together to digitize the natural history literature held in their collections and make it freely available for open access as part of a global “biodiversity community.”

Join Our Mailing List

Sign up to receive the latest news, content highlights, and promotions.

Subscribe Now

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 319 other subscribers

Subscribe to Blog Via RSS

Subscribe to the blog RSS feed to stay up-to-date on all the latest BHL posts.

Access RSS Feed

Inspiring Discovery through Free Access to Biodiversity Knowledge.

The Biodiversity Heritage Library makes it easier than ever for you to access the information you need to study and explore life on Earth…for free, anytime, anywhere.

 

64+ Million Pages of
Biodiversity Literature Online.

EXPLORE

Tools and Services
to Transform Research.

EXPLORE

300,000+
Illustrations on Flickr.

EXPLORE

ABOUT | HARMFUL CONTENT | PRIVACY | SITE MAP | TERMS OF USE