Biodiversity Heritage Library - Program news and collection highlights from BHL
  • Home
  • News
  • Featured Books
    • All Featured Books
    • Book of the Month Series
    • BHL at 20
  • User Stories
  • Campaigns
    • Fossil Stories
    • Garden Stories
    • Monsters Are Real
    • Page Frights
    • Her Natural History
    • Earth Optimism 2020
  • Tech Blog
  • Visit BHL
Home
News
Featured Books
    All Featured Books
    Book of the Month Series
    BHL at 20
User Stories
Campaigns
    Fossil Stories
    Garden Stories
    Monsters Are Real
    Page Frights
    Her Natural History
    Earth Optimism 2020
Tech Blog
Visit BHL
  • Home
  • News
  • Featured Books
    • All Featured Books
    • Book of the Month Series
    • BHL at 20
  • User Stories
  • Campaigns
    • Fossil Stories
    • Garden Stories
    • Monsters Are Real
    • Page Frights
    • Her Natural History
    • Earth Optimism 2020
  • Tech Blog
  • Visit BHL
Biodiversity Heritage Library - Program news and collection highlights from BHL
BHL News, Blog Reel

First Meeting of the Mining Biodiversity project

Meet our international partners to extract data from BHL

View Full Size Image
Mining Biodiversity (MiBio project) is one of the projects that won during the third round of the transatlantic Digging Into Data Challenge, a competition aiming to promote the development of innovative computational techniques that can be applied to big data in the humanities and social sciences. The project is an international collaboration between the National Centre for Text Mining (UK), Missouri Botanical Garden (US) and Dalhousie University’s Big Data Analytics Institute (Canada) and Social Media Lab (Canada), along with colleagues from the Encyclopedia of Life and the Smithsonian Institution.

We will integrate novel text mining methods, visualization, crowdsourcing and social media into the BHL to provide a semantic search system that allows users to explore search results according to multiple information dimensions or facets.  The goal is to transform BHL into a next-generation social digital library resource that facilitates the study and discussion (via social media integration) of legacy science documents on biodiversity by a worldwide community.

View Full Size Image
Relations between the Work Packages of the project

The project has five major components, covered in 9 Work Packages (WP):

  1. Automatic correction of errors in OCR using Google n-grams by our colleagues of the Big Data Analytics Institute (WP2).
  2. Crowdsourcing the annotation of semantic metadata (concepts and events) in legacy texts (WP5).
  3. Extract metadata (terms, concepts and significant events) automatically and track their change over time (WP3 & WP4) to facilitate semantic search (WP6) implemented with NaCTeM.
  4. Use interactive visualization techniques to manage the search results, in collaboration with Dalhousie (WP7)
  5. Design a social media layer as an environment for interaction and collaboration on science, education, awareness and outreach, lead by our colleagues of the SocialMediaLab (WP8).
View Full Size Image
Manchester Town Hall
at Albert Square, UK

On February 17th, the first face to face meeting in Manchester, UK marked the start of this new project.  The Principal Investigators of the project, Dr. Anatoliy Gruzd from Dalhousie University (Canada) and William Ulate from Missouri Botanical Garden (USA), met with Dr. Sophia Ananiadou at the University of Manchester’s National Centre for Text Mining (NaCTeM), where her colleagues involved in the project showcased the tools and services they have developed and will be adapting for our project.

The National Centre for Text Mining (NACTEM)

NaCTeM has developed text mining services based on a number of generic natural language processing tools like Argo, their Web-based workflow construction platform for text mining, implemented on top of the OASIS Unstructured Information Management Architecture (UIMA) standard for interoperability among information processing components.
View Full Size Image
Example of an Argo workflow that automatically extracts
species and anatomical features using entity tagging
components that NaCTeM has developed for this purpose.

Several of the NaCTeM tools have been developed as modules that can be adapted and used as components in workflows, receiving input from the previous module, processing or performing a task and passing the results to another module.  In the case of named-entity recognizers, these receive text pre-processed into smaller units (sentences, tokens) and extract features automatically according to statistical models used by different entity taggers (specialized in gene, chemical, anatomical, habitat or species information, for example).

Another type of components of the workflows, in addition to named entity recognizers, are the linkers, which facilitate the automatic linking of names or concepts found in text to entries in external vocabularies via unique identifiers and using a string similarity method.

Argo’s functionality allows workflows to be deployed as a Web service so they can be invoked by external applications, just like BHL currently invokes Name-finding web services to find the taxa within the text.

View Full Size Image
At NaCTeM commenting on the tools for the project.
L to R:  Mr. John McNaught, Dr. Anatoliy Gruzd,
Dr. Sophia Ananiadou, Ms. Riza Batista-Navarro,
Mr. Georgios  Kontonatsios and Mr. Paul Thompson.
Missing from the photo: Dr. Rafal Rak,
Claudiu Mihăilă, and Dr. Ioannis Korkontzelos.

In order to develop these named entity recognizers and linkers for the biodiversity domain, it is necessary first, to identify which entity types are of interest (in our case, it could be names of persons, places, species, among others) and the vocabularies to link to for each type.  To assist on this process, NaCTeM has also developed term extraction tools like TerMine.  TerMine detects terms and acronyms in input text and can be used in building a term inventory for biodiversity.  This is what the initial task for our colleagues at Missouri Botanical Garden and Smithsonian will be about: finding those authoritative sources (vocabularies, ontologies, thesauri, gazetteers, etc.) of terms to help build the term inventory and then train the entity taggers to be used in our workflows.

NaCTeM has also done substantial work on event extraction, i.e., the extraction of  associations or interactions between concepts or entities.  This experience will help us identify and extract the type of events that scientists, historians and other scholars have long wanted to extract from the BHL corpus (like behavior, habitat, trophic relations, geographic range and others ) for our own named-entities: species, people, places throughout time.  Finally, NaCTeM’s vast experience developing customized semantic search engines like KLEIO, ISHER and Europe PubMed Central EvidenceFinder will facilitate providing an enhanced semantic search functionality over the BHL corpus text, to allow users to explore results according to multiple information dimensions or facets.

Additional information on the tools and services can be found at:

  • Rak, R., Rowley, A., Black, W.J. and Ananiadou, S. (2012). Argo: an integrative, interactive, text mining-based workbench supporting curation. Database: The Journal of Biological Databases and Curation.
  • Kolluru, B., Nakjang, S., Hirt, R. P, Wipat, A. and Ananiadou, S.. (2011). Automatic extraction of microorganisms and their habitats from free text using text mining workflows. In: Journal of Integrative Bioinformatics, 8(2), 184

For some interesting explanation of what Unstructured Information is and the terminology of the process around it, look at this nice introduction of the UIMA 1.0 Standard.

Or read more details about these or some of the other service systems and tools that NaCTeM has developed.

The Social Media Lab

On Monday February 17th, 2014, as part of the third Social Media Workshop, which covered the outreach and impact aspects of the International Centre for Social Media Research at Manchester University, our group was invited to attend a talk by our colleague in the project, Dr. Anatoliy Gruzd, where he presented the research done at the Dalhousie University Social Media Lab, how to make sense of the huge quantity of data and the new methods to collect information when studying online social networks through analysis and visualization.

View Full Size Image
Social Media Lab at
Dalhousie University, Canada
© All Rights Reserved

For our project, the staff at Dalhousie has started investigating what users and communities, as well as the context in which they are currently accessing, commenting and sharing the records from BHL across various social media platforms, such as Twitter and Flickr.  For this work they will be employing some of their own tools developed by the Social Media Lab (such as Netlytic.org) as well as other tools developed by third parties.  Their goal is to add a social layer to integrate content from different biodiversity fora and social media sites with BHL via a user-friendly interface, to foster a community of users that could exploit BHL as an environment for sharing digital objects.

For more information, take a look at:

  • Gruzd, A. & Haythornthwaite, C. (2013). Enabling Community through Social Media.  Journal of Medical Internet Research 15(10):e248. [DOI:10.2196/jmir.2796]

And some of Dr. Gruzd’s and the Social Media Lab staff other publications.

I hope this gives an idea of the work ahead and a better sense of what the project attempts to do and how it aims to do it.  We will keep you informed as results become available, but in the meantime, let us know how do you envision yourself using BHL as a social digital library?  What information you’d like to track and how you’d like to access it?  Tell us the vocabularies you’d need to see included and what types of named entities and associations you’d want to be tagged in the BHL corpus?

View Full Size Image

This project is made possible in part by a grant from the Institute for Museum and Library Services [Grant number LG-00-14-0032-14].

March 20, 2014by Joanne A. Schneider
BHL News, Blog Reel

While standing on the shoulders of giants…

Learning lessons from the CITScribe Hackathon participation

Two members of BHL’s Technical Advisory Group (TAG), BHL Technical Director William Ulate from Missouri Botanical Garden and Joe deVeer, Head of Technical Services of the Ernst Mayr Library at the Museum of Comparative Zoology, Harvard, “virtually” attended the CITSCribe Hackathon in Gainesville, Florida from Dec. 16 to 20, 2013.

The hackathon was co-organized by iDigBio (www.idigbio.org) and Zooniverse’s Notes from Nature Project (www.notesfromnature.org). Led by the co-organizers Rob Guralnick (University of Colorado, Boulder) and Austin Mast (Florida State University), the objective of the hackathon was to further enable public participation in the online transcription of biodiversity specimen labels and to produce new functionality and interoperability for Zooniverse’s Notes from Nature and similar transcription tools.  The event started by focusing on four areas of development, progressively addressed throughout the week: (1) interoperability betweenpublic participation tools and biodiversity data systems, (2) transcriptionquality assessment/quality control (QA/QC) and the reconciliation of replicatetranscriptions, (3) integration of optical character recognition (OCR) into thetranscription workflow, and (4) user engagement.

View Full Size Image
The LI LI Hackathon subgroup working on the integration of OCR.

For a while now, the BHL Technical Advisory Group (TAG) has been participating in the events of iDigBio, particularly on the Augmenting OCR Working Group, under Deb Paul and Bryan Heidorn’s guidance.  One of our BHL colleagues and TAG members, John Mignault from New York Botanical Garden, also participated last year at iDigBio’s previous hackathon which primarily focused on parsing OCR output to standard darwin core.

This time, we had a presentation by Cody Meche, an Agile trainer who gave some useful directions and tips on how to approach our hackathon to use our time as efficiently as possible while following an agile development process.  Alex Thompson, from iDigBio, presented on the digital resources that could allow programmers to run a copy of the web interface of Notes for Nature on their local machines using Vagrant and Virtualbox.  Paul Kimberly also presented the Smithsonian transcription project (transcription.si.edu) and Yonggang Liu talked about iDigBio’s image ingestion tool, useful for getting images into the iDigBio cloud.  Laura Whyte, Director of Citizen Science at Adler Planetarium joined us virtually at the Hackathon to talk about Zooniverse and citizen science.

I also had a chance to present shortly about our approach to OCR and Transcription with in our Purposeful Gaming project funded by the Institute of Museum and Library Services (IMLS). Basically, we intended to leverage iDigBio’s knowledge while tackling the challenge of improving the existing BHL OCR through Purposeful Gaming.  Our interest was mainly two-fold: on one side we wanted to learn from the amazing cumulative experience and skills that iDigBio comprises and the tools that are being adapted, built and used to generate the OCR and transcribe the text of their specimen labels; a challenge very similar to our own issues with OCR of books and journals in BHL.  On the other hand, we wanted to learn from their workflow definition to improve the OCR generation process, the incorporation of their transcription outputs and the reconciliation of the existing versions into a single one (for example, read about using sequence aligning methods on this recent blogpost by Rob Guralnick), while at the same time taking advantage of the citizen scientists’ help in an efficient manner that copes with the scale of the challenge.

View Full Size Image
High level workflow proposed for the Purposeful Gaming project

Specifically, our short term goals include incorporating recommendations and adapting existing tools developed by many of the participants in the past two hackathons: Jason Best’s DarwinScore, Ben W. Brumfield’s work on QC for Collaborative (Crowdsourced) Manuscript Transcription, the QA/QC track developments to take outputs from the citizen science transcription products and assure the highest quality end result, are just some among many other outputs and suggestions from these activities that can be incorporated to our process of improving the OCR texts in BHL.  And if you have more ideas, I would definitely like to hear about them…

For more information on the recent CITScribe Event read this iDigBio’s blogpost, Austin Mast’s detailed blog post or consult directly the Hackathon wikipage.

You can access all information and results about iDigBio in their home page and follow our own BHL developments from the webpage of the Purposeful Gaming project, made possible in part by the Institute for Museum and Library Services [Grant number LG-05-13-0352-13].

View Full Size Image

January 23, 2014by Joanne A. Schneider
BHL News, Blog Reel

Tis the season to be thankful!

View Full Size Image
BHL is a collaborative endeavor, no doubt about it; and it’s been thanks to these collaborations that we have the technical achievements we have. For many reasons, this is a good time to be thankful.

So as the BHL Technical Director, I would like to start by thanking my colleagues at the BHL Technical Group here in Missouri Botanical Garden (MBG) for taking the leading role in many of the pressing issues we’ve faced this year: our BHL Data Analyst and P.I. of the NEH-funded Art of Life project, Trish Rose-Sandler and our developer, Mike Lichtenberg, who’s been responsible for many of the new things you’ve seen in BHL and is always improving our Portal (check his Developer’s blog).

I’d also like to thank all the staff at the IT Division at MBG who have support our infrastructure this year and have been key to support our technical developments and enable the implementation of the required functionality; special thanks to Mike Westmoreland for supporting our installation of Macaw, our RefBank node at MBG, and our Solr infrastructure.

Thanks to our Technical Advisory Group (TAG) members who have supported our technical work all this past year by providing advice and leading several important tasks for the BHL community that started or concluded this year like Macaw (Joel Richards from Smithsonian), our Solr implementation for Full-text Search (Frances Webb from Cornell), the Partner Meta App implementation (Joe deVeer from Harvard) and represent BHL in technical meetings and hackathons (John Mignault from New York Botanical Garden, NYBG).  Our thanks also go to our folks at MBL, particularly to Anthony Goddard for the Cluster implementation and support throughout this past year.

View Full Size Image
Wild Turkey, male & female
Taken from: Wilson et al. American ornithology; or,
The natural history of the birds of the United States.v.3 pl.9
http://biodiversitylibrary.org/page/41423917

Particular thanks to the Secretariat and colleagues at Smithsonian, for leading and supporting BHL’s technical and operational activities.  Our gratitude also goes to the BHL Executive Committee, the BHL Members Committee, the BHL Collections Committee (lead by our Collections Manager, Bianca Crowley) and the BHL Staff group, comprised of dedicated individuals from each of the BHL Member and Affiliate Libraries, who contribute countless hours to scanning, paginating and cataloging to make the millions of pages in BHL available for free and open access.  Each of these groups also actively further our common objective by providing guidance, perspective and recommendations on technical topics.

This year, we also had some significant staff turn around.  This is no surprise as, unfortunately, it is part of any extensive collaborative endeavor such as BHL. Without taking any merit from all of our other former colleagues, we want to thank Grace Costantino and Gilbert Borrego for their enormous contributions to BHL: much of our current technical work wouldn’t be what it is without their lasting contributions.

We also would like to thank our partners at the NEH-funded Art of Life project, Indianapolis Museum of Arts, University of Colorado, Marine Biological Laboratory and a special mention to the Internet Archive technical team.  Our thanks go to our BHL partners, particularly to Harvard, Cornell and New York Botanical Garden colleagues in our IMLS-funded project to use Purposeful Gaming to help improve OCR text in BHL, starting next month.

Finally, we are particularly thankful to those we worked directly with, whom, I’m sure, represent a wider collaboration by themselves: Rod Page for BioStor, Rich Pyle and Rob Whitton for Zoobank, Nicky Nicholson from IPNI and Paul Kirk for Index Fungorum, Donat Agosti and Terry Catapano at Plazi,  Guido Sautter at the Karlsruhe Institute of Technology,  Lyubomir Penev, Jordan Biserkov and Theodore Georgiev at Pensoft and Dave Roberts for Vibrant and Dauvit King for Open University, the whole EOL team, particularly Jen Hammock, Cyndy Parr, Katja Schulz, Patrick Leary, Erick Mata and Bob Corrigan, Dmitry Mozzherin, David Shorthouse, and many, many others whose technical contributions keep making BHL what it is.

Our appreciation travels far to all our colleagues in BHL for their daily contributions, and translates into different languages for our international Global BHL Nodes collaborators for their invaluable support: Obrigado, Dankie, Danke, Dank, Gracias, Merci, Děkuji, 謝謝, شكرا

And most of all, our deepest thanks to all of the users of BHL who continuously inspire and challenge us to make BHL the best it can possibly be!

Thank you!

November 28, 2013by Joanne A. Schneider
BHL News, Blog Reel

Impressions from afar: an account of our Fourth Annual Global BHL Meeting

Portrait version of the Biodiversity Heritage Library logo.

View Full Size ImageMore than a month ago, on May 27 and 28, we took the opportunity to have our Global BHL Meeting in Fez, Morocco, because some of us were on our way to ICADLA-3, the 3rd International Conference on African Digital Libraries and Archives in Ifrane later that week.  This year, it was particularly special, not only because we had a superb backstage, staying and meeting in a Dar in the old Medina, but also because at least one representative of each BHL Node (except for Brazil) managed to converge for the occasion in such a historic city and they all had something big to report from their annual achievements.

View Full Size Image
Global BHL Attendees

It was very inspiring to hear all the reports highlights and share about common global topics for our initiative: from award-winning volunteers digitizing in Australia to the dedicated colleagues generating Chinese OCR and manually tagging common names; from the laborious staff scanning a whole library in Brazil in just a few months to the quick generation of online exhibitions with BHL material in Europe; and from the development of Macaw, a specialized tool to upload selected material to our collection in a much easier way, created in the US with collaboration from Australia and Brazil, to the implementation of full-search text, annotations and underlining capabilities on the Arabic portal by Egypt – all these achievements crowned with the establishment of a continental BHL node by and for African colleagues in less than a year… for me, this is priceless!

View Full Size Image
We stayed at Dar Fes Medina

I gave a presentation for the Technical Update, explaining how having the user interface, initially developed by the BHL-Australia staff, now available on top of the BHL-US/UK functionality, is much more than a mere change of look. There’s new functionality made available by providing access to article-level information and more taxa is now found through the services of the Global Names Architecture project (supported by the National Science Foundation) to provide for a closer integration with authors lists and taxa aggregators like Zoobank and IPNI, among others.  And these are just the tips of the iceberg.  The continuous growth of the BHL corpus at a steady rate for more than 5 years, echoes the work of all those who support BHL and reflects the value of our staff contributions and the usage of our content that our users promote by creating references that link to it.  I ended my talk with a mention of the latest developments on several of our projects, like our NEH-funded Art of Life.

View Full Size Image
Newly elected gBHL Vice-Chair Jiri Frank
(BHL-E) and Dr. Jinzhou Cui (BHL-China)

But it hasn’t been a year without challenges of all sorts: the end of funding for some projects, staff turnover, priority changes, and even political turmoil… During this time, our colleagues from the Marine Biological Laboratory (MBL) in Woods Hole have managed to solve previous issues with the performance of our content copy on their Cluster repository.  BHL-Europe has also solved some issues it had that slowed down the process to upload files and now it’s steadily ingesting content, thanks to their Organizations’ support and sometimes even personal effort and interest to maintain the project going on.  BHL-China has also found a renewed interest in providing services to other institutions through novel projects.  BHL-Australia has empowered collaborations with their volunteers’ network and BHL-Egypt has kept advancing the replication of content with firm commitment to their participation, even after major changes in the country.  Needless to explain, the challenges that BHL Africa had to overcome in order to have, in less than a year, 12 institutions from different countries eager to establish a new node and sign up the Memorandum of Understanding to participate in such effort by the time of the launch inauguration.  But changes are opportunities, and the obstacles faced only makes it even more important and valuable to recognize the wonderful contributions and achievements that have been carried out around the globe in the last 12 months by our BHL community.

View Full Size Image
gBHL Chair Nancy Gwinn and Program
Director Martin Kalfatovic (BHL-US/UK)

The Global BHL is a cooperative network of autonomous members operating programs and projects to make biodiversity literature available under principles of open access, collaboration, decentralization, interoperability, transparency and legality, and I believe it is its diversity that makes it resilient to obstacles in a particular node.  For me, this is the reason why, even in times of adversity, it has always been growing, increasing not only in size, but also developing in more regions, with more languages, topics and types of content and adapting to technology changes, improving or developing new tools and synchronizing content among nodes, even after the original project (and funding) from The Gordon and Betty Moore Foundation for Global BHL Coordination ended last year.

View Full Size Image
BHL Global Coordinator  William Ulate
and Anne-Lise Fourie (BHL-Africa)

Now better and higher goals are already established. BHL-Australia and BHL-Europe want to reactivate and adapt their own projects to expand to new institutions and collections.  BHL-Africa and BHL-Brazil have the goal of starting to upload their digitized content and expand participation to other countries.  Likewise, BHL-China also wants to extend their content and services to new topics of their scientific community. BHL-Egypt is interested in integrating content with other global initiatives and exploring joint projects, while BHL-US/UK is looking to consolidate its new framework of membership and expand the services provided to their users.  Learning from each other’s successes and experiences helps us to keep working together, ensure redundancy and resilience, while at the same time increasing new unique content, tools and services that provide novel opportunities for collaborations.

View Full Size Image
gBHL Chair Ely Wallis (BHL-AU),
Dr. Magdy Nagi (BHL-Egypt) and gBHL
Secretary Nancy Gwinn (BHL-US/UK)

Every time I return from our Annual Global BHL Meeting, I come back full of positivism and ideas, proud to be part of this global initiative that involves so many amazing people collaborating from so many different places.  I have met a lot of people, but not all of them yet, I have to say; and to be fair with them, I have avoided singling out anyone here, but having the privilege, at least once a year, to share together with all the Global BHL representatives about the work their colleagues do and hear about their contributions to this worldwide endeavor, is a real treat for me…
View Full Size Image
I hope you also enjoyed it!

July 5, 2013by Joanne A. Schneider
BHL News, Blog Reel

A quick overview from Second Global BHL Planning Meeting…

Portrait version of the Biodiversity Heritage Library logo.

We just had the Second Global BHL Planning Meeting in Chicago Field Museum this past November 13, 2011 with representatives of all BHL Programs, except our colleagues from Bibliotheca Alexandrina, who couldn’t attend this time. During the meeting, each of our BHL Programs shared their progress since our (last Global Meeting on September 2010) and it was definitely a year of new and valuable achievements for all.

For BHL-US/UK, beyond an increasing access to constantly growing content, an improved appearance that includes a whole new logo, and the novel Flickr account, now with more than 20,000 images, turned into a fundamental component of our outreach activities to new communities. Also, the election of a new Executive Committee and the incorporation of new library partners were part of the good news of this year. Our Australian colleagues launched their new appealing website) to contribute content about Australian species to the Atlas of Living Australia project and are now getting ready to start digitizing and sharing their content from BHL-Australia. Our colleagues from BHL-China have come a long way this past year integrating BHL-China with the numerous projects on biodiversity they have in all their country’s provinces. BHL-China node hosted colleagues from BHL-US/UK last year for (technical meetings), and continued digitalization of valuable Chinese material, now shared also through BHL-US/UK. BHL-Europe is ready to launch by the end of this year their new portal that will aggregate content from European libraries, increase support to species names and allow for a powerful new Advance Search interface. Likewise, our colleagues from SciELO-Brazil have been contributing their information to BHL’s citation repository, Citebank and are now setting up the equipment and workflows to inaugurate their new digitization facilities by April next year.

We also discussed and reviewed the Vision of BHL as a global unified network where its member institutions share their leadership and coordinate within the nodes, engaging with other organizations related to our field to become the recognized reference tool for biodiversity literature. We all agreed on more and better communications between the nodes and defined ways to coordinate this. Also, as part of the review of the governance structure of BHL, it was decided to form a Coordinating Committee that would establish the by-laws and tackle some of the challenges and issues addressed during the meeting and support the implementation of the way forward. It was also decided to hold a Technical Meeting in conjunction with our Global BHL Meeting next year.

As the meeting came to an end, all participants were looking forward for another new exciting year, initiated with the Life and Literature Meeting the following two days, as a source of valuable input from colleagues, users and staff on the areas and activities that BHL should focus and prioritize for the following 5 to 7 years!

November 15, 2011by Joanne A. Schneider
BHL News, Blog Reel

Yes, BHL has gone Global!

View Full Size Image

A summary from the 1st Global BHL Technical Meeting by William Ulate, Global BHL Project Coordinator.

September 22 to 24, in Woods Hole, Massachusetts, took place the Global BHL Technical Meeting, it was the very first time all signed and prospective BHL partners were going to be together at such meeting. There were representatives from all over the world, including Australia, Brazil, Egypt, Europe and the US; unfortunately our colleagues from China were unable to make it. We had a very productive meeting to know each other and present each other’s work in order to describe priorities and requirements for a Global BHL.

Through out this exchange participants achieved a high-level description of software and hardware components and were able to agree on milestones and deliverables for a global timeline, while at the same time, sketch the definition of global governance & policies for collaboration in the project.

On Wed. Sep. 22nd morning, after all participants had arrived and enjoyed a delicious breakfast (everyone recognized the Food Catering throughout the whole meeting was outstanding), a warm welcome from our hosts by Cathy Norton, MBL Director, followed by our BHL Director, Tom Garnett and our BHL Executive Committee Chair, Graham Higley, followed by a brief introduction from each participant, provided the perfect setting for a picturesque multimedia display of a brief Taking Measure of the Biodiversity Heritage Library: 2003- 2010 by Martin Kalfatovic, our BHL Deputy Director, and Chris Freeland, Global BHL Technical Director, talking about the BHL-US role in the Global BHL. The first section of the meeting was rounded up by Phil Cryer and Anthony Goddard presenting their lessons learned while setting up the Clustered and distributed Storage with commodity hardware and open source software to mirror BHL information.

After a comfort break each regional node was given the opportunity to share, before the rest of the group, the details of their specific projects, why and how it connects to the whole BHL and other projects, the work already done, the digitized content available or planned including dates of major milestones & deliverables, the resources available, their funding status, and their regional requirements, among other things.

The first partner to present was BHL Europe. Henning Scholz, Project Coordinator for BHL-E, gave an overview of their principles, objectives and partners, their work plan with dates for deliverables and how BHL Europe can integrate into different networks like this Global BHL initiative. Then, Melita Birthälmer, also from Museum für Naturkunde in Berlin, presented the project’s activities related to Content Management, starting with the available and planned numbers of volumes from different providers and the quality of that content and explaining in greater detail about the Global Reference Index to Biodiversity, GRIB, a bibliographic database with content management and deduplication functionalities being developed in collaboration with the EDIT project. Even when the GRIB is still in a prototype phase (see http://grib.gbv.de/), it has been suggested as an option for a worldwide bibliographic database for a Global Biodiversity Heritage Library. Finally, Adrian Smales, from the Natural History Museum, talked about the technical implementation, dealing with topics like different metadata views and formats used by BHL-E content providers, their current infrastructure status, some considerations with an Open Archive Information System (OAIS) and a Preservation Archive System (PAS), the GRIB and a general Work Plan for the Technical implementation deliverables.

The next partner to present was Australia. Elycia Wallis from Museum Victoria showed the comprehensive work that the Atlas of Living Australia (ALA) has been carrying out and the context where the BHL-Australia (BHL-Au) project is being developed as one of its Rich Data Stores component projects (see presentation here) . Then she presented their worked starting at mid-2010 with the BHL-Au and BHL kickoff meetings at Museum Victoria in Melbourne and ALA HQ in Camberra. Then she explained the achievements setting up the infrastructure, including development and testing environments, assessing workflows to scan and adapting them to Australia conditions and developing a new user interface for BHL that should be ready by the end of 2010 with the existing functionality (test site can be accessed at http://bhl-test.ala.org.au) Following the topics proposed, Ely talked about human and other resources they have available and then explained the plan and timing ahead. The mirroring, ingestion and uploading processes should be ready by March 2011. She also mentioned Australian copyright laws allows to scan documents up to 1955, but they might concentrate on particularly scanning rare books by mid 2011 and they feel confident they will be able to apply for further funds to perform maintenance after that. Additionally, Ely mentioned other very interesting projects going on to BHL-Au like supporting annotations, scanning field notes and correcting OCR through volunteers work (crowd sourcing).

Abel Packer, presented afterwards about the BHL Brazil/ BHL SciELO Network of national and thematic collections of quality journals, funded by the Federal Government and the state of Sao Paulo government, the research community and the libraries (see his presentation here). The network governance formalization should be in place by the end of 2010, start of 2011. Their slow but fully sustainable technical work has focused on procedures and criteria for content to be digitized and the open technology used in the portal development, OAI metadata exchange services implementation and VHL-provided search engine functions. SciELO Network had a Kick-off Workshop on Essential Rare Works Collection in Biodiversity on February 2010; in close collaboration with BHL Advisory Committee it plans to validate the selection criteria and choose the 200 first journal/ bulletins titles and books to scan. Its plans are to be operational and launched with 100 initial books by December 2010, expand to Latin America and the Caribbean countries starting in 2011 and have more than 2000 books scanned, digitized and exposed through BHL by 2013.

Finally, our colleague from Egypt, Dr. Noha Adly, presented their progress in Bibliotheca Alexandrina (www.bibalex.org), a “center of excellence in the production and dissemination of knowledge” (see her presentation here) whose objectives fit perfectly with BHL’s and has been involved in digital libraries and projects for quite some time now, developing technical infrastructure, long-established mass digitization and OCR workflows, mirroring Internet Archive and massive data sets, training specialists for their workflow, and more recently working with the Arabic version of Encyclopedia Of Life. Bibliotheca Alexandrina has come a long way since it started with 1 scanner in 2003, it now has 120 trained specialists working using their 10 scanners, 7 days a week on two shifts, digitizing and doing the OCR of 167,000 Arabic books, photos, negatives, slides and maps to include into joint projects like Description de L’Egypte and the World Digital Library. They have also developed their own projects like Digital Assets Repository (composed of Digital Assets Factory, Digital Assets Metadata using Fedora to manage only metadata, Digital Assets Keeper and Digital Assets Publishers) and the Science Supercourse, a PowerPoint repository for health, agriculture, environment and computer engineering. Bibliotheca Alexandrina is interested in becoming a BHL partner, holding a mirror site and working on infrastructure. It has also offered to organize our next Technical Global BHL Meeting, (which everyone happily took note of).

In the afternoon, after a comfort break, Bianca Crowley, BHL Collections Manager, presented the analysis of the BHL User Survey 2010. A total of 16 reusable questions were developed and for this first time, an average of 1020 successful responses per question were analysed, to understand how current user groups are using BHL services and what new development are groups expecting in the future.

Chris Freeland, BHL Technical Director, lead us in two interesting discussions about the Names finding process in BHL and what can be done to improve it, given an existing 35% error rate on the species names when the OCR is performed. Here, it was noted how two subprocesses are been carried out: the string finding and the name reconciliation. While some of the existing services take on both processes (UBIO, for example), it was concluded we have no mechanism in place to validate the 5.1 million names, so we should concentrate on working on OCR correction and let the specialists handle the name reconciliation. We will make random sample data available for potential partners, so nomenclators, for example, could provide feedback on ratio of good names.

Finally to round up the first day, the group reviewed the implications of the Global Open Access and BHL standpoint on it. It’s no secret for anyone that the global world of copyright is very complex. The group commented on the several copyright issues and distribution limitations will encounter in sharing materials globally. BHL is not assuming any copyright responsibility on its own, moreover BHL doesn’t own any copyright. A small group of colleagues was defined to take all input about this topic and develop a suitable statement; taking into account that the user might get confused and frustrated if we end up with different categories of access and the system would have to be rebuilt to support it.

The second day the group was divided into an Administration subgroup, in charge of Policies & procedures needed for a global collaboration and another Technology subgroup to work on components needed for a Global BHL. The Administrative group was to deal with topics like Organization of each BHL node, Global BHL Collaboration and Governance and Communication Models for project leaders of each BHL node. On the other hand, the Technology group was set to discuss topics like the Content Ingest process in existing BHL, the Content Replication, making particular reference to preservation (LOCKSS) and mirror sites, the Localization, taking in consideration if he had to deal and share materials that couldn’t be openly distributed; and finally, the topic of Global Identifiers within the whole project.

Other more technical topics were covered during the rest of the meeting, from “Branding & Identity of the project itself” to funding opportunities, to data mining, OCR & Text correction experiences, and improvement of existing and new services required for integration at the APIs & User Interfaces levels. Even some birds-of-a-feather sessions on Content and Data Synch were included. Finally, a set of Action Items was sketched to follow up (see it here).

December 17, 2010by Joanne A. Schneider

Help Support BHL

BHL's existence depends on the financial support of its patrons. Help us keep this free resource alive!

search

About BHL

The Biodiversity Heritage Library (BHL) is the world’s largest open access digital library for biodiversity literature and archives. BHL operates as a worldwide consortium of natural history, botanical, research, and national libraries working together to digitize the natural history literature held in their collections and make it freely available for open access as part of a global “biodiversity community.”

Join Our Mailing List

Sign up to receive the latest news, content highlights, and promotions.

Subscribe Now

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 319 other subscribers

Subscribe to Blog Via RSS

Subscribe to the blog RSS feed to stay up-to-date on all the latest BHL posts.

Access RSS Feed

Inspiring Discovery through Free Access to Biodiversity Knowledge.

The Biodiversity Heritage Library makes it easier than ever for you to access the information you need to study and explore life on Earth…for free, anytime, anywhere.

 

64+ Million Pages of
Biodiversity Literature Online.

EXPLORE

Tools and Services
to Transform Research.

EXPLORE

300,000+
Illustrations on Flickr.

EXPLORE

ABOUT | HARMFUL CONTENT | PRIVACY | SITE MAP | TERMS OF USE