Biodiversity Heritage Library - Program news and collection highlights from BHL
  • Home
  • News
  • Featured Books
    • All Featured Books
    • Book of the Month Series
    • BHL at 20
  • User Stories
  • Campaigns
    • Fossil Stories
    • Garden Stories
    • Monsters Are Real
    • Page Frights
    • Her Natural History
    • Earth Optimism 2020
  • Tech Blog
  • Visit BHL
Home
News
Featured Books
    All Featured Books
    Book of the Month Series
    BHL at 20
User Stories
Campaigns
    Fossil Stories
    Garden Stories
    Monsters Are Real
    Page Frights
    Her Natural History
    Earth Optimism 2020
Tech Blog
Visit BHL
  • Home
  • News
  • Featured Books
    • All Featured Books
    • Book of the Month Series
    • BHL at 20
  • User Stories
  • Campaigns
    • Fossil Stories
    • Garden Stories
    • Monsters Are Real
    • Page Frights
    • Her Natural History
    • Earth Optimism 2020
  • Tech Blog
  • Visit BHL
Biodiversity Heritage Library - Program news and collection highlights from BHL
BHL News, Blog Reel

Advancing Data Excellence: A New Era for the BHL Cataloging and Metadata Committee

Cataloging & Metadata Committee Milestones with graphic visuals of a team pointing to charts

2023 proved to be a transformative year of growth, increased collaboration, and heightened dedication to advancing biodiversity knowledge for the BHL Cataloging and Metadata Committee. Last year, the Committee achieved significant milestones in committee governance, professional development opportunities, data quality updates, and the ratification of consortia-wide data policies, accompanied by comprehensive documentation.

Cataloging & Metadata Committee Milestones with graphic visuals of a team pointing to charts

On the governance front, the former “Cataloging working group” completed a Committee Charter to elevate the group’s status to an official BHL Committee. Approval of the Committee at the 2023 BHL Annual Meeting by the Executive Committee and BHL Partners, as well as the formal codification in the BHL Bylaws, solidifies Cataloging and Metadata as a permanent standing BHL Committee.

Pictures of new committee co chairs Elizabeth McKinley (leading a tour of patrons) and Daniel Euphrat (sitting in front of digitization equipment)

In 2023, the Committee held its inaugural Chair Election which signified a notable transition for the former working group. We are excited to welcome Daniel Euphrat of the Smithsonian Libraries and Archives and Elizabeth McKinley of the Chicago Field Museum as the co-chairs for 2024. Additionally, we extend our gratitude and appreciation to former chairs Suzanne Pilsk of the Smithsonian Libraries and Archives and Diana Duncan of the Chicago Field Museum, who will continue to provide valuable mentorship during this transitional period. A warm welcome and congratulations to Daniel Euphrat and Elizabeth McKinley on becoming the inaugural co-chairs in 2024.

The Committee membership grew with the addition of advisors, retired staff as observers, and new members:

  • Siobhan Leachman, Committee Advisor and Independent Wikimedian
  • Elizabeth McKinley, Committee Co-Chair and Cataloging & Metadata Librarian at the Chicago Field Museum
  • Briana Giasullo, Committee Member and Cataloging and Digital Resources Librarian at the Academy of Natural Sciences of Drexel University
  • Sandra Lee Parker Provenzano, Committee Member and Head Cataloger at Dumbarton Oaks, Harvard University

As a newly formalized Committee, this coming year will see some fresh perspectives and ideas to drive progress on BHL Strategic Goals and the Committee’s formal charge.

Empowering BHL Staff with Wiki Education’s Wikidata Certificate Course

A tangible benefit of Committee membership includes access to broader professional development opportunities through BHL. Early in 2023, committee members celebrated the completion and final project wrap-up of a six-week professional certificate course sponsored by BHL and the Smithsonian Libraries and Archives (SLA). This specialized course was facilitated by Wiki Education, a non-profit that focuses exclusively on supporting students, faculty, and subject area specialists with their Wikimedia work. The program provided BHL and SLA staff with a unique opportunity to hone advanced skills in SPARQL queries, visualizations, and data modeling.

Three SPARQL visualizations with varying colors and sizes of circles clustered together

Left to right: SPARQL visualizations of BHL Authors by Occupation; BHL Female Authors by Occupation; BHL Male Authors by Occupation. The course provided participants with the opportunity to experiment with linked data and visualizations produced by the Wikidata Query Service.

The custom curriculum, led by the knowledgeable Will Kent aka “Wikidata Will” and SLA co-facilitators JJ Dearborn, Richard Naples, Suzanne Pilsk, and Jackie Shieh engaged a cohort of 25 participants in a tailored, immersive experience exploring topics like SPARQL queries, visualizations, federation, data modeling, bulk data loading, and editing tools.

Grid of 20 black and white video screens of committee members on a video conference call

Participants added BHL partner nodes as items and linked them up with BHL using the Wikidata in partnership property:

GIF of before and after visualization of BHL knowledge graph with entities clustered and connected by arrows

The BEFORE is what BHL’s Partner Knowledge Graph looked like and AFTER the course, one can see how BHL’s network has grown in Wikidata!

Enriching our Committee’s technical expertise has not only opened new avenues for exploration within the Wikimedia project ecosystem but has also laid the foundation for the newly formed BHL-WIKI working group, set to launch this year under the distinguished leadership of Wikimedia Laureate and long-time BHL advocate, Siobhan Leachman. Under Siobhan’s guidance, BHL is embarking on a dedicated effort to convert its legacy data into 5-star linked open data, facilitating expanded global access to biodiversity knowledge for all.

For a more in-depth understanding of BHL’s symbiotic relationship with the Wikimedia ecosystem, we encourage readers to explore related post “BHL is Round Tripping Persistent Identifiers with the Wikidata Query Service,” and the white paper titled “Unifying Biodiversity Knowledge to Support Life on a Sustainable Planet.”

Comprehensive BHL Metadata Requirements for our Global Consortium

Conformant metadata in BHL is the linchpin that allows BHL to manage diverse global sources of metadata across 588+ contributors. A guiding document is paramount to fostering metadata harmonization from a globally disparate network and empowers BHL technical staff in maintaining data consistency and ensuring interoperability across the consortium. The new BHL metadata requirements facilitates effective searching, browsing, discovery, and identification for BHL’s end-users, enabling broader access and engagement with biodiversity knowledge on a global scale.

The journey to finalize BHL’s Metadata Requirements was a lengthy endeavor and bringing together once-disparate guidelines into a unified framework is the culmination of many years of work. After a Metadata Requirements Summit in 2023 and multiple rounds of Committee peer review, the first version of comprehensive BHL Metadata Requirements has been published and incorporated into BHL’s Collection Development Policy.

BHL Metadata requirements include data disclaimer, a statement on remediation of harmful language in library metadata, partner meta app, FAQ, schema mappings, updated schema tables, digitization workflow decision map

Additionally, embedded within these requirements is the BHL Cataloging and Metadata Committee’s formal Statement on Remediation of Harmful Language in Library Metadata, which serves as a commitment to address harmful language in library metadata, recognizing the impact of legacy language and knowledge organization systems can have on perpetuating biases. BHL actively champions inclusivity, encouraging contributors to realign vocabulary terms to ensure diverse user access. Rejecting censorship, BHL acknowledges and commends the ongoing efforts by contributing institutions in remediating outdated metadata terms. This commitment not only aligns with our dedication to diversity but also reinforces our mission to provide a comprehensive and inclusive perspective on biodiversity knowledge for a global audience.

Across the BHL Consortium, we consistently unite to uphold excellence in metadata management and recognize the individual needs of our digitization partners, encompassing factors such as staff, technical expertise, funding, and local facilities. Having comprehensive Metadata Requirements reflects the tenacious and collaborative spirit across the BHL Consortium to come together in unity and further our collective mission to provide free, worldwide access to knowledge about life on Earth.


BHL Cataloging and Metadata Committee Charge

The BHL Cataloging and Metadata Committee possesses advanced expertise in metadata standards, remediation, mapping, and cross-walking workflows. The committee is responsible for overall BHL metadata advising, curation, and validation with the overarching goal of sharing the highest quality bibliographic data outputs broadly with the bioinformatics community and the world. The task of harmonizing hundreds of years of metadata from diverse sources is a continuous, iterative activity. The Cataloging and Metadata Committee directly supports the following strategic goals:

  • Ensures reliability and accuracy of the collections by curating the BHL collection 
  • Defines the BHL collection as a data resource as well as a digital library to support big data approaches
  • Extends and deepens metadata parsing and validation 

In closing, the Committee eagerly anticipates a year ahead filled with continued efforts to enhance BHL’s metadata quality and further integrate diverse information into the growing network of biodiversity knowledge. Please don’t hesitate to reach out to the group with your feedback. Your input will be invaluable as we strive to advance our mission and better serve our community. Thank you for your ongoing support and engagement.

February 29, 2024by mdimeo
BHL News, Blog Reel, Tech Updates

BHL is Round Tripping Persistent Identifiers with the Wikidata Query Service

Diagram of the 8 steps detailed below of the BHL wikidata round trip.

In the Spring of 2022, the BHL Cataloging and Metadata Committee investigated the possibility of harvesting persistent identifiers (PIDs) from Wikidata as part of the group’s longstanding project to disambiguate and deduplicate author records in the BHL database. The motivation behind this one-time experimental data harvest was to see if BHL could:

  1. Enhance BHL author records with additional PID data points;
  2. Improve the committee’s ability to disambiguate author names in the BHL database; and
  3. Respond to an outstanding user request from two of Wikimedia’s super star editors, Siobhan Leachman and Andy Mabbett, to expose BHL’s author data on BHL and include hyperlinks to other authoritative knowledge bases on the web.
How do PIDs work? Users are directed to the URL that is associated with the identifier. If the physical location of the object changes, the URL is the only modification required; the PID itself does not change.

Persistent identifiers can become resolvable links that connect two knowledge bases (like BHL and Wikidata) together.

In particular, Wikimedians wanted to see the Wikidata Q identifier exposed, providing a link to the corresponding creator item record in Wikidata.

There are multiple motivations for undertaking this work. By adding the BHL Creator ID to the corresponding Wikidata item, Wikidata editors help link BHL to the richer biographical data about that person held in Wikidata. The Wikidata item for a person may contain links to their Wikipedia page or to images of the person held in the image repository Wikimedia Commons. Wikidata items also act as identifier hubs and contain links to other databases and identifiers.

Maria Sibylla Merian's wikidata entry, including name variants, biographical information, external sources, an image, and more

Maria Sibylla Merian in Wikidata, rendered with the Reasonator Tool; note the External sources section and BHL’s Creator ID entry.

By adding the BHL Creator ID to this list of identifiers, the Wikidata editor is linking the content held in BHL to the content held in multiple other datasets and repositories.

These extra author data points provide Wikimedians and BHL catalogers with crucial clues that aid in name disambiguation. In particular, hyperlinks to other knowledge bases are incredibly valuable because they lead to new knowledge pathways that help confirm a person’s identity in a complex game of “Who’s Who?”

Workflow Overview

The diagram below illustrates the experimental data pipeline from BHL to Wikidata and back.

Diagram of the 8 steps detailed below of the BHL wikidata round trip.

The BHL to Wikidata Round Trip Overview

The BHL to Wikidata Round Trip Steps

  1. BHL Creator ID and record are created when digital content is harvested into BHL and/or articles are defined in BHL.
  2. Wikimedians record BHL Creator IDs as a statement in corresponding Wikidata items via various community workflows. (See: Mix’n’match deep dive below and/or read about QuickStatements for two ways volunteers are populating Wikidata with BHL’s Creator IDs).
  3. BHL subsequently harvests additional metadata from Wikidata for any author item where a BHL Creator ID statement has been added.
  4. Data outputs are analyzed for quality and breadth.
  5. SPARQL queries and data outputs are iteratively refined to match BHL’s requirements until a quality dataset can be generated for import.
  6. Clean data is imported into the BHL database.
  7. The new author record sidebar displays data on BHL.
  8. PIDs are converted into resolvable URIs, opening up new research pathways for BHL’s users.

Quick note: In Wikidata BHL author records are represented by the Wikidata BHL Creator ID; the presence of this Wikidata property in a Wikidata item provides a powerful connection point that can be used later by BHL to ingest more information about any named entity. A popular author matching tool in Wikidata is Mix’n’match.

Getting BHL Authors in Wikidata with Mix’n’match

Mix’n’match is a Wikidata tool that empowers editors to match Wikidata items to entries in other databases. One of the datasets in Mix’n’match is the BHL Creator ID dataset.

Summary from the mix'n'match tool for the BHL creator ID dataset

A summary of the completeness of matching the BHL Creator ID dataset to Wikidata as of January 2023

When the BHL Creator ID dataset is uploaded into Mix’n’match, an algorithm undertakes fuzzy name matching. The tool then suggests a preliminary match to a Wikidata item. Wikidata editors working on the dataset have the choice of whether to confirm the suggested match or to remove it.

Sample matches suggested by the mix'n'mtch tool

Examples of possible matches suggested by the Mix’n’match algorithm with the BHL creator in green and the Wikidata item in blue.

If the Wikidata editor confirms the match, the BHL Creator ID is automatically added to the Wikidata item for the creator. If the editor rejects the match, the editor can add a correct Wikidata item ID, create a new Wikdiata item if the creator is not yet in Wikidata or, in some cases, decide that the BHL Creator ID isn’t applicable to Wikidata. If the editor simply rejects the match without further action, the BHL Creator ID will be added to the unmatched portion of the dataset.

Examples of confirmed and unconfirmed matches in the Mix'n'match tool

Results after the Wikidata editor Ambrosia10 has confirmed or removed the Mix’n’match algorithm suggestions.

The act of linking the BHL Creator ID to the Wikidata item also helps to disambiguate the creator. It removes much of the uncertainty about who the creator is. Names are not unique but by linking identifiers to biographical data, editors can ensure that similarly named people or people whose names have changed over time are all associated with the correct Wikidata item. This addresses the age-old problem of how to assign correct attribution to the right person.

This work also assists BHL. By ensuring BHL Creator ID’s are matched to the Wikidata item, the Wikidata editor can assist BHL in weeding out the duplicate entries for creators in BHL’s database. After matching has taken place, BHL is also able to ingest any of the other identifiers listed on the creator’s Wikidata item, thus enriching the metadata held in BHL’s database.

Currently there are over 230,000 entries in the BHL Creator ID dataset. Of these, just over 41,000 have been matched to Wikidata items. There is a long way to go before this dataset is complete. In the spirit of “many hands make light work,” BHL encourages Wikidata editors to work alongside BHL staff who were recently trained and certified in Wikidata Advanced Concepts by Wiki Education. Our collaborative work will help increase the number of Creator IDs in Wikidata.

Summary from the mix'n'match tool for the BHL creator ID dataset

BHL Creator ID dataset completeness statistics as of January 2023.

Once BHL Creator ID statements were recorded on Wikidata items, either through manual edits or workflows like Mix’n’match or QuickStatements, then BHL created custom SPARQL queries using the Wikidata Query Service to output data of interest.

Example of a SPARQL query using the wikidata query service

The SPARQL query language is used to query and return data from semantic databases like Wikidata.

Once the data was brought back, the BHL Cataloging and Metadata Committee discussed and reviewed the records with the aim of bringing key data points back into the BHL database.

Sample data returned from the SPARQL query, including item ID, itemLabel, BHL_creator_ID, date_of_death, VIAF_ID, Library_of_congress_authority_ID, ORCID_ID, ISNI, date_of_birth

Snippet of raw data brought back by the BHL Cataloging and Metadata Committee’s SPARQL query.

In total, the BHL Cataloging and Metadata Committee was able to “round trip” 88,507 persistent identifiers (PIDs) associated with BHL Creator records from the following authoritative knowledge bases:

  • Virtual International Authority File (VIAF)
  • Wikidata
  • Social Networks and Archival Context (SNAC)
  • ResearchGate
  • ORCID
  • Library of Congress Authorities

There are still more PIDs out there to gather but practicality calls for moderation – moderation of the quantity that the Wikidata Query Service can bring back and the amount feasible to curate in BHL. After many discussions, the above list was chosen by narrowing down the PIDs that seem most appropriate in the BHL context. A comprehensive policy on Uniform Resource Identifiers (URIs) in the BHL ecosystem is now being drafted by BHL committees.

Additionally, in post-work analysis, the Committee did find that the data modeling of corporate names in Wikidata differs from the library world, and pulling in identifiers for corporate names was not worth the effort (at least at this time).

Below is an example of Académie des sciences (France), which receives multiple item records in library databases like VIAF for each name change. However, in Wikidata these name changes are collapsed and will appear on one item record; this divergence means that one-to-one matches are not possible without a lot of manual clean-up from BHL Staff.

Example of the Wikidata entry for VIAF ID for Académie des sciences

Divergent data models: Wikidata corporate names collapse name changes.

It’s important to note that data modeling in Wikidata is in its nascent stages and librarians should have a vested interest in shaping the best practices of the community so all of humanity can model the world’s knowledge together. As we have seen with this one-time experiment, there is little to lose, and so much to gain!

The Result: A New Author Record Sidebar for BHL

Thanks to the feedback from the Wikimedia community, BHL’s Lead Developer created an Author Record sidebar. Click on any BHL author name on the BHL website and the author details will appear on the right-hand side. Below is an example of Australian paleontologist and ornithologist, Patricia Vickers-Rich in BHL. 

Graphic example of an author record with detailed sidebar and the pages to which it links out, such as ORCID, Library of Congress Authorities, Wikidata, SNAC

A New Wikimedian Driven Feature: The BHL Author Record Sidebar!

The goal of this new feature was to expose existing and recently harvested identifiers as resolvable links that could open up critical knowledge pathways for BHL’s users on their information journeys. Other key data points are provided such as preferred name forms, entity type, and alternative name forms that give additional context for disambiguation research.

A quote from Siobhan Leachman on BHL's new author sidebar: Adding these identifiers to the author sidebar makes my life so much easier. At a glance, I can quickly confirm the identify of creators. The author sidebar or lack thereof also highlights whether further work is needed in wikidata to link these BHL creators to their identifiers."

Next Steps for BHL Committees?

Please let us know in the comments what you think of this experiment. Should BHL pursue similar persistent identifier harvests for other entity types? Or perhaps, a recurring harvest for BHL Authors?

Additional BHL entities:

  • Titles
  • Scientific Names
  • Subjects
  • Media (illustrations, scientific plates, photos)

Let us know! Your feedback is crucial to BHL’s evolution as a biodiversity knowledge base.

If you are new to Wikidata, Mix’n’match is a great place to start. There are many tutorials and resources available on the web including this great YouTube introduction on how to use the tool.

Additionally, a workflow for BHL articles has been piloted and is currently underway thanks to BHL’s Persistent Identifier Working Group. For more details on how to roundtrip BHL Article Q identifiers, please refer to the group’s documentation at: Wikidata:WikiProject_BHL/Projects

BHL committees and working groups are actively scoping many projects. Sign-up here to get involved!

Related Posts

For more information on the assignment of DOIs, a very important type of persistent identifier specific to scholarly publications, check out Nicole Kearney’s blog post What Is BHL’s New Persistent Identifier Working Group DOI’ng?

For a list of all of the new features and data quality improvements BHL made in 2022, check out the post BHL Technical Development: Year in Review

February 15, 2023by mdimeo
BHL News, Blog Reel, Tech Updates

What Is BHL’s New Persistent Identifier Working Group DOI’ng?

Graphic showing the members of BHL's Persistent Identifier Working Group

In October 2020, BHL launched a new working group with a momentous goal: to make the content on BHL persistently discoverable, citable and trackable using DOIs (Digital Object Identifiers).

Graphic showing the members of BHL's Persistent Identifier Working Group

The members of BHL’s new Persistent Identifier Working Group (PIWG).

A DOI is like an electronic fingerprint in the form of a unique and permanent alphanumeric string that provides a persistent link to a piece of content online. Modern publications receive a DOI at the point of publication. This DOI becomes a key part of a publication’s bibliographic metadata that should be included in any mention or citation of that publication. Reference lists in modern publications are filled with DOIs, which allows readers to click from publication to publication in (in theory) a never-ending chain of knowledge.

This reciprocal linking of DOIs has created a great linked network of scholarly research, but that network is missing the historic literature. The vast majority of historic publications lack DOIs. This means they appear in reference lists as unlinked citations. In our increasingly online world, readers are far more likely to read (and thus cite) publications they can click through to (particularly when libraries are inaccessible during a global pandemic). The upshot of this is that the millions of pages of historic literature on BHL—the foundation of our understanding of biodiversity—is in danger of falling into obscurity.

BHL has been retrospectively minting DOIs for historic publications since 2011, but the focus has primarily been on monographs. BHL’s new Persistent Identifier Working Group (PIWG) is (at least initially) focusing on journal articles. Minting DOIs for articles on BHL is a far more complex and time-consuming task than minting DOIs for monographs. This is because article DOIs need article data: every journal volume uploaded onto BHL must be accompanied by journal and volume data, but there is no requirement that contributors provide article data.

Thankfully, there have been considerable efforts to add article data to BHL (and thereby make it possible to search for the titles and authors of these articles both within BHL and via external engines). A huge proportion of this article data has been contributed to BHL by Roderic Page via BioStor: 75% of the 300,753 articles indexed in BHL as of 4 May 2021 were “defined” by BioStor. It is very difficult to determine how many articles are actually on BHL (hidden within all those journal volumes). But, while we don’t know what proportion of BHL’s journal content still needs to be made discoverable, we know there is still a huge amount of work to do.

COVID-19 provided an unexpected opportunity to make a considerable dent in this work. With no access to scanners or library materials, a number of BHL contributors, including Harvard University Libraries, Muséum National d’Histoire Naturelle and BHL Australia, pivoted from making new content accessible to making their existing content on BHL more discoverable. For example, BHL Australia’s digitisation volunteers gathered, gap filled and checked article-level metadata for over 30,000 articles in 2020.

Once an article has been defined, i.e. it exists as a publication unit in BHL and has its own article landing page (and we’ve checked that the article does not already have a DOI), we can assign a DOI to it. Articles that have recently been assigned BHL DOIs include some very old publications, such as the first scientific description of the Duck-billed Platypus, published in 1799 (https://doi.org/10.5962/p.304567), and the species descriptions from A specimen of the botany of New Holland, the first publication dedicated to Australian flora (1793-5), e.g. https://doi.org/10.5962/p.312432. The PIWG has also started assigning DOIs to in-copyright publications (with permission from the rights holders). These include articles from the Bulletin of the British Museum, e.g. https://doi.org/10.5962/p.310418, and the Bulletin of the African Bird Club, e.g. https://doi.org/10.5962/p.308885.

Screenshot of the landing page in BHL for the description of the The Duck-Billed Platypus, Platypus anatinus.

Shaw, George (1799), The Duck-Billed Platypus, Platypus anatinus, The Naturalist’s Miscellany: https://doi.org/10.5962/p.304567 (illustration by Frederick Polydore Nodder).

If an article on BHL has an existing non-BHL DOI, we add this key piece of bibliographic metadata to the BHL landing page for the article. This ensures that BHL users can link to the definitive version of the article (the one the DOI resolves to), and more importantly, that other parties (and their algorithms) can find our versions from elsewhere. This is particularly important when commercial websites lock their DOI’d versions of public domain articles behind paywalls. Having their DOIs on our freely accessible versions ensures that services like Unpaywall can find them. To learn more about how this works, see our blog post: BHL Journal Articles Are Now Discoverable via Unpaywall.

DOIs not only improve discoverability and enable persistent linking to our historic content; they also allow us to track how BHL content is being used. In the six months following the minting of its new DOI (Oct 2020 to April 2021), the 1799 Platypus description was tweeted by 219 Twitter accounts, referenced in six Wikipedia pages, picked up by one news outlet and cited in one academic paper (data from Altmetric, April 2021). We know this because the article has a DOI.

Screenshot of the Altmetric dashboard for the first scientific description of the Duck-billed Platypus

Altmetric’s overview of attention for the first scientific description of the Duck-billed Platypus (Shaw 1799): https://www.altmetric.com/details/91788579.

The PIWG has spent the past six months creating, refining and testing tools that will allow BHL contributors to do this work themselves. We have also been producing documentation that explains a) how to use the new tools, and b) why this work is so important. These tools will facilitate every step in the article discoverability and DOI assignment process including: downloading existing article data for a given journal title to allow for correction and gap-filling (in development); bulk uploading of article data for new articles (available now); and adding articles and titles to BHL’s (new) DOI Assignment Queue (available now). Our dream is that, whenever anyone uploads a journal volume to BHL, they also provide the data for the articles it contains (and thus take responsibility for making that content discoverable).

The Persistent Identifier Working Group (PIWG) is fueled by the technical expertise, metadata dexterity and incredible passion of:

  • Nicole Kearney, Manager BHL Australia (Chair)
  • Mike Lichtenberg, BHL Lead Developer
  • Susan Lynch, Systems, Digitization & Web Services Librarian, The New York Botanical Garden
  • Bess Missell, Metadata Librarian, Smithsonian Libraries and Archives
  • Roderic Page, Professor of Taxonomy, University of Glasgow
  • Joel Richard, BHL Technical Coordinator | Head of Web Services & IT, Smithsonian Libraries, Smithsonian Libraries and Archives
  • Diane Rielinger, Digital Projects Librarian, Botany Libraries, Harvard University Herbaria
  • Colleen Funkhouser, BHL Program Manager

The specific goals of the group are:

  • To add article-level metadata to journal articles on BHL
  • To add existing DOIs to (new and existing) article landing pages on BHL (particularly for those articles where the DOI’d version is behind a paywall elsewhere)
  • To assign BHL DOIs to articles that lack DOIs

Want to know more about BHL’s Persistent Identifier Working Group? See:

  • Discovering the Platypus: From its scientific description to its DOI, Biodiversity Information Science and Standards (TDWG) Conference, 6 October 2020: https://youtu.be/4UVSEoWsSrw?t=1285
  • #RetroPIDs: making historic Platypus Infinitely Discoverable (PID), PIDapalooza: the Festival of Persistent Identifiers, 28 January 2021: https://youtu.be/CSeQNe5KR5U

For the latest news about BHL’s DOI work, check out #RetroPIDs on Twitter.

May 10, 2021by michelle.underhill
BHL News, Blog Reel, Tech Updates

Updates to Bibliography Pages in BHL

Screenshot of bibliography pages in BHL with and without tabs.

We have updated the bibliography pages in BHL to streamline the presentation of information about and metadata export options for content in the Library.

Previously, bibliographic details and export options were available through different tabs on title and part pages. These tabs have now been removed, and all bibliographic information is consolidated into a single display.

Screenshot showing BHL bibliography pages before and after.

The various tabs on title and part bibliography pages have now been consolidated into a single display.

BHL’s metadata export options have also been relocated. BHL offers metadata exports in MODS, BibTex, and RIS formats. MODS is an XML-based bibliographic description schema used in a variety of library applications. BibTex and RIS are bibliographic citation files that are compatible with a variety of citation management tools. The MODS file is a title-level download. The BibTex and RIS files are item-level downloads.

The MODS download is available at the bottom of the title and part bibliography pages.

Screenshot of a title page in BHL with the "Download MODS" button circled.

MODS download on the title and page bibliography pages.

The BibTex and RIS downloads are available in two places:

1) Under the volume or part details on bibliography pages.

Screenshot of the citation download options in the BHL website.

BibTex and RIS downloads on title (left) and part (right) bibliography pages.

2) Under the “Download Contents” menu in the book viewer, via the “Download Citation” option.

Screenshot of the BHL book viewer with the citation download options displayed.

Download citation options in the BHL book viewer.

As part of this update, we have removed the direct Mendeley import from BHL, as the generic RIS and BibTex formats are compatible with a variety of citation management softwares including Mendeley.

Details about our metadata export services are also available in the FAQ on the BHL About site.

February 11, 2021by michelle.underhill
BHL News, Blog Reel

BHL Records Now Available in WorldCat

OCLC worldcat logo

BHL is pleased to announce that it has added its bibliographic records to OCLC’s WorldCat® database, “the world’s largest network of library content and services.” You can now find BHL e-books via https://www.worldcat.org/. With thousands of libraries worldwide participating in OCLC, contributing our records to this “global library cooperative” allows us to extend the discoverability and access of BHL e-books through its variety of tools and services [1].

OCLC worldcat logo

If you have access to OCLC metadata service products, please look for the OCLC symbol “BHLMR” when reviewing holdings.

If you are an OCLC participating library interested in adding BHL records to your own catalog, please feel free!

Some key points to keep in mind when viewing or working with BHL records in WorldCat:

  • Current BHL holdings in WorldCat reflect our collection’s bibliographic records as of Fall 2019.
  • Not every title you find in BHL will be in WorldCat. Only titles with full text available directly within BHL are being provided to OCLC. We are not providing records for titles indexed in our collection that link out to content available via third-party websites.
    • For example, you will find Crawshay’s The birds of Tierra del Fuego (1907) in WorldCat.org. But you will not find Flora Ibérica as this title is indexed in BHL but held directly within the Biblioteca Digital del Real Jardin Botanico de Madrid.
  • BHL will load records into WorldCat on a quarterly basis going forward.
  • As we curate title records within the BHL collection and make changes on our end, we will work with OCLC to update records on a periodic basis.

We are thankful to the kind folks at OCLC who worked diligently and patiently with us to develop a workflow for exporting and uploading the variety of records aggregated into our digital collection from hundreds of contributors worldwide.

[1] https://www.oclc.org/en/about.html

March 31, 2020by SheilaC7108
BHL News, Blog Reel, Tech Updates

Changes Coming to the BHL Data Exports Files on 10 April 2019

On 10 April 2019, we will implement additions and changes to the export files available from the Biodiversity Heritage Library.

The updates involve the following:

  1. A new set of exports will be created alongside the existing exports. The new set will contain only data for material that is hosted by BHL. No externally-hosted content will be included in these files. Because these files are added in addition to the existing export files, no existing users should be affected.
  2. The BHL author Identifiers will be added to the creator.txt and partcreator.txt tab-delimited files. The format of these files will change to accommodate the additional data; the author identifier will now be the second “column” in each file. Because of this, anyone regularly harvesting from these files may be affected.

Additional detailed information about these updates will be reflected on the Data Exports: Developer and Data Tools webpage, effective 10 April 2019.

If you have questions, please feel free to submit feedback via this form.

April 3, 2019by michelle.underhill
Blog Reel

Reflecting back on my incredible summer at Smithsonian Libraries

View Full Size Image

As the weather in Central New York is getting colder, and the winter is inevitably approaching, I can’t help but recall the humid summer of Washington, DC. Over the summer of 2016, I interned at Smithsonian Libraries. As a summer intern, I worked at the Department of Digital Programs and Initiatives on the “Cataloging across collections” project. The project was focused on metadata and cataloging. During my internship I worked on digital curation of the records in Biodiversity Heritage Library as well as cataloging field notes for the Field Book Project. Both aspects of my internship were coordinated by mentors assigned by the Smithsonian Libraries – Bianca Crowley, Digital Collections Manager of Biodiversity Heritage Library, and Lesley Parilla, Cataloging Coordinator of the Field Book Project. During my internship I met with Smithsonian Institution staff members, toured different departments, attended the staff picnic, and, most importantly, improved my understanding of the technical services in libraries.

Although working on metadata projects was quite challenging, it helped me to gain confidence as an information professional. As a librarian, I’ve always known that the information organization is important, but I didn’t realize how much properly maintained bibliographic descriptions can improve the user experience. Especially, when the audience can only interact with online assets. Biodiversity Heritage Library is a unique digital collection of the resources in Natural Sciences. My work focused on editing metadata for existing records, which included author merging, editing of volume information, title merging, and linking of the serial records. I received training in 2 internal systems, the Gemini Issue Tracker and BHL Administrative Dashboard to work on these tasks.

On the first week of my internship, I needed to link several volumes together. “Den Norske Nordhavs-expedition, 1876-1878” gives an account of the Norwegian North-Atlantic expedition in the 19th century, commanded by Carl Fredrik Wille, Captain of the Royal Navy. Since I have interest in Scandinavian languages, I enjoyed interacting with this resource. This and many other assignments throughout my internship taught me the importance of metadata standards. Seeing both librarian and user perspectives on the information-retrieval systems became an eye-opening experience for me.

 

View Full Size Image

“Den Norske Nordhavs-expedition, 1876-1878”, bd. 5, pt. 17 [Alcyonida], tab. I

View Full Size Image
“Den Norske Nordhavs-expedition, 1876-1878”, bd. 5, pt. 17 [Alcyonida], tab.

Further into my internship, I was trained to upload the scanned images into the Internet Archive, the platform that hosts the BHL assets. I used another internal system, Macaw, for this purpose. As I was uploading the new images, I assigned page-level metadata. One of the publications I was working with was Bonn Zoological Bulletin. I had so much fun working on the metadata-level description of William Mann’s scrapbook from his trip to South East Asia for the Field Book Project. One of my favorite serials was Canadian Forest Industries, a magazine that was renamed at least 6 times before its current title. While I was adding volume information and page level metadata, I encountered amazing illustrations and advertisement campaigns from Florists’ Review, that is a wonderful compilation of the marketing tools of the past.

 

View Full Size Image
Canadian Forest Industries, formerly known as Canadian Lumberman, Jan 1903.

 

View Full Size Image
Drawing from the Florists’ Review April 1913 v.31 no.797 (801) p.17

 

View Full Size Image
Front page of from the Florists’ Review December 1922, Christmas edition v. 51 no.1306

In addition to learning many essential skills for technical services in libraries, during my internship I worked with Camtasia software. Screen-casting was used to develop tutorials for BHL Staff about working in the BHL “Admin Dash.” I recorded one video that focused on merging author records together. During the internship I also presented at a Brown Bag presentation in front of Smithsonian Libraries staff, my mentors, and other interns.

I would like to express gratitude to my mentors – Bianca Crowley, Digital Collections Manager at Biodiversity Heritage Library, and Lesley Parilla, Cataloging Coordinator of the Field Book Project, and the entire Department of Digital Programs and Initiatives for their guidance and support. I would also like to thank LIS Program Director at Syracuse University – Jill Hurst-Wahl, and my Academic Advisor – Barbara Stripling, for their invaluable insights on the US library system. My boundless gratitude goes to Cultural Vistas and the Edmund S. Muskie Internship Program that made it possible for me to intern in Washington, D.C. during the summer of 2016. Being a part of the Smithsonian Institution was an unforgettable life experience, which I will proudly carry with me throughout my library career!

December 28, 2016by [email protected]
Page 1 of 212»

Help Support BHL

BHL's existence depends on the financial support of its patrons. Help us keep this free resource alive!

search

About BHL

The Biodiversity Heritage Library (BHL) is the world’s largest open access digital library for biodiversity literature and archives. BHL operates as a worldwide consortium of natural history, botanical, research, and national libraries working together to digitize the natural history literature held in their collections and make it freely available for open access as part of a global “biodiversity community.”

Join Our Mailing List

Sign up to receive the latest news, content highlights, and promotions.

Subscribe Now

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 319 other subscribers

Subscribe to Blog Via RSS

Subscribe to the blog RSS feed to stay up-to-date on all the latest BHL posts.

Access RSS Feed

Inspiring Discovery through Free Access to Biodiversity Knowledge.

The Biodiversity Heritage Library makes it easier than ever for you to access the information you need to study and explore life on Earth…for free, anytime, anywhere.

 

64+ Million Pages of
Biodiversity Literature Online.

EXPLORE

Tools and Services
to Transform Research.

EXPLORE

300,000+
Illustrations on Flickr.

EXPLORE

ABOUT | HARMFUL CONTENT | PRIVACY | SITE MAP | TERMS OF USE