Friday, September 18, 2015

What We Talk About When We Talk About Access

With a tip o' the hat to Raymond Carver, I want to use this post to try to and illuminate (for myself, if nothing else) some of the angles and issues surrounding 'access' to digital archives.

On the surface, the topic appears simple: I have some stuff that I want people to see, so I put it online or provide a dedicated terminal in my reading room and—voila!—access!

Made available by Flickr user Steve Rhode under a CC Attribution-NonCommercial-NoDerivs 2.0 Generic License
But even in this rosy scenario, there are a lot of questions: what platform would you use to host things online?  Will access copies (i.e., DIPs) differ from preservation copies (AIPs)?  If using a dedicated terminal, how will content be organized and how will researchers find desired materials? If copying files to a terminal or removable media, will staff be able to respond to researcher requests in a timely fashion? And what about rights?

Now, I certainly don't want to be like a certain you-know-who...


...but there are a lot of considerations here. Complex ones, too.  At the same time, simply waiting around for the stars to align and the *perfect* solution to emerge won't cut it, either.  Therefore, inspired by the various presentations on access that I saw last month at SAA, I'd like to give a brief overview of our current approach to access and then lay out some of the questions and challenges we're starting to explore here at the Bentley.

Just Dropped In (To See What Condition My Condition Is In)

The Bentley Historical Library has taken a fairly aggressive (progressive?) approach to providing access to our 'open' or unrestricted digital archives.  All such content is freely available for download and use via our archival community in Deep Blue, the University of Michigan's DSpace repository:


Deep Blue is managed by staff in the University of Michigan Library Information Technology division and we considered ourselves to be very fortunate when we started using it as both a preservation repository and access portal in 2008.  Prior to that (and not having any in-house IT), digital materials were either placed on optical disk and brought out to patrons in our reading room or hung off of our website and linked to from finding aids.

Moving to Deep Blue/DSpace was clearly a step forward, but the change brought about some additional challenges due to the basic structure (dare I say data model?) of the repository:

  • A 'community' contains 'collections' (which may be grouped together in sub-communities, we've formed one of these for university faculty papers)
  • A collection in turn contain 'items' (which may be associated with one or more files or 'bitstreams').
  • The default metadata schema is Dublin Core (which makes crosswalking from EAD ...interesting...)
While this relatively flat structure works great for traditional institutional repository fare (white papers, articles, and discrete digital objects), it really isn't suited for the complex intellectual hierarchies of archival collections.  So we've had to make do...

As with our physical and/or analog record groups and manuscript collections, our materials in Deep Blue are organized by the principle of provenance (and are often extensions of existing collections):


Within a collection we have our items—and here's where we've had to get creative:


Given the flat structure of DSpace, we are using the title metadata to help group related content together and preserve the hierarchical intellectual arrangement of materials.  As a result, the following description from our Jennifer Granholm finding aid...
...becomes the following item title in Deep Blue:
We also package materials in .zip files so that we only have to manage one file and our users don't have to download hundreds or even thousands of files.  Because content must actually be downloaded to a local machine to be used or rendered (unless a particular file format renders with a browser plugin), we have taken to chunking content across multiple .zip files when it gets to be above 2 GB:


The above digital object represents speeches, addresses, and other audio recordings from former Michigan Governor Jennifer Granholm for the year 2010.  All together, there's about 20 GB of content; by dividing this body of content into smaller chunks representing each month of the year, we've made it a bit easier for folks to download content.  And while I certainly don't think this solution is ideal, it's still a lot better than bringing a stack of CDs out to folks in our reading room.

Earlier in this post, I alluded to open or unrestricted content; we actually have three access profiles based upon rights and restrictions:
  • Open materials may be accessed by anyone anywhere at any time.
  • Restricted materials are only available to system administrators and digital curation staff; the items are not visible to other users nor is the metadata searchable.  Content is restricted for a number of reasons, including specifications in a gift agreement; the presence of sensitive personal data (related to HIPAA or FERPA as well as credit card numbers and Social Security numbers); and internal policy (for example, all executive records of the University of Michigan, while FOIA-able, are restricted for 20 years from the date of accession).
  • Reading-room only materials may only be accessed by computers within the IP address range of the library itself (and are not accessible by patrons using university wifi). This class is primarily composed of content where we do not hold copyright or donors have requested more restricted access. Our reading room rules, which all researchers must agree to follow, stipulate that these items "may not be copied, emailed or transferred in any way."  (While placing the burden on the researcher is by no means foolproof, it's much easier to implement and maintain than the locked-down computer terminals with which we earlier experimented.)
In addition to having the metadata and text-based file contents (when not packaged in .zip files) indexed by Google and other search engines, all materials are linked from online finding aids and/or catalog records.   People certainly seem to be finding our content, too: from 2008 through last month, we've registered 620,375 downloads (a figure that excludes downloads from robots or web crawlers).

Access: the Final Frontier

As we enter the final six months of our Mellon grant (and prepare to kick off a Hydra development project with colleagues at the University of Michigan Library), we have returned, time and again, to the challenge of providing access to digital archives.

There are a lot of great access portals to collections out there!  Some of the ones we've been particularly impressed with include those of:

Those are just a few of the many examples out there (and we meant to include your digital collections, but ran out of time...), but we've noticed that while these (and other innovative solutions) are vast improvements over an off-the-shelf option like CONTENTdm, they seem pretty unique to their local institutional context and IT environment.


The work we've been doing with Archivematica and ArchivesSpace has made us firm advocates of community-based approaches where folks at different institutions can share and contribute to common solutions, without having to reinvent the wheel (and continue to support and maintain that reinvented wheel. Indefinitely. All by themselves.).

This community interest recently led us to contribute user stories to the ArcLight project, "an effort to build a Hydra/Blacklight-based environment to support discovery (and digital delivery) of information in archives, initiated by Stanford University Libraries."  Likewise, we were excited to hear about the DPLA Archival Description Working Group and its implications for describing—and searching for and retrieving—digital archives.

Beyond the above, we've also been trying to articulate the different aspects or considerations related to access that could be common to cultural heritage institutions of all sizes and shares.  These are some very, very rough ideas, but we're interested in how we can:

  • Explore and better understand the challenges and opportunities surrounding OAIS functional entity of ‘access’
    • Present (and make understandable) the context/content of archival materials (including the relationships between digital, physical, and analog materials)
    • Enable search and retrieval of information while balancing item-level, aggregate, and collection-level description
    • Provide tools and functionality to view/render various formats (born-digital and digitized), including images, text, audio and moving image, web archives, and disk images 
    • Facilitate the analysis and reuse of data (including visual representations of metadata/data and tools or functionality that would facilitate distant reading of materials and other digital scholarship techniques)
    • Increase engagement with users (crowdsourcing or feedback)
  • Manage rights and enforce restrictions/permissions
  • Establish use metrics and collect quantitative data regarding impact of our collections and outcomes of curation activities
  • Permit users a more seamless experience using materials in searching for and using materials that are in disparate/siloed locations: online catalogs, HathiTrust, digital repositories, web archives, etc.
  • Leverage linked data: facilitate research across collections and institutions
At this stage in the game, we aren't even thinking about specific implementation strategies, as it seems there could/should/might/shall/will be a core set of features or functional requirements that could exist independent of any particular repository platform.  Having said that, it seems to us that an access portal should:
  • Emphasize on Interoperability: 
    • Create connections between tools and services
    • Permit us/other institutions to broaden current work and ‘plug in’ to larger framework
    • Avoid siloed/local solutions 
  •  Employ open source software:
    • There is a “need for engagement beyond simply making source code available, including supporting the development of user communities, creating adequate documentation, and cultivating relationships between developers working in libraries around the country.” (IMLS National Digital Platform)
  • Focus on end users
    • Meet needs within LAM communities for common solutions and interoperability as well as those of end users related to the access and use of digital archives.
    • End users are creating, accessing, and organizing content in ways that were never before possible and, in many cases, without the support of a knowledge professional.  The user should figure prominently in our strategy. How do we bring in their views, and identify the missing voices? (IMLS National Digital Platform)
So that's what we've been thinking about and pondering as of late... What are we missing?  What seems unnecessary?  What do you talk about when you talk about access?

Thursday, September 10, 2015

Introducing New Project Archivists to Processing at the Bentley

We interrupt this blog to bring you breaking news: the Bentley Historical Library is thrilled to welcome three new project archivists: Emily Swenson, Elizabeth Carron, and Julia Teran.  They join current project archivists Dallas Pillen, Shae Rafferty, and Cinda Nofziger as part of an exciting new program at the Bentley to help new professionals develop essential career skills while contributing to strategic goals aligned with the library's priorities in academic outreach, collections development, digital curation, and research services.

As part of our welcome, Max Eckard and I gave an introduction to archival processing at the Bentley, which included an overview of our grant project.

Thursday, September 3, 2015

The ArchivesSpace API

One of the most powerful tools that we've been making use of in our task to migrate our legacy EADs and accession records to ArchivesSpace is the ArchivesSpace API (Application Programming Interface), which allows users to interact with the ArchivesSpace backend to more efficiently and programmatically accomplish any task that can be completed in the ArchivesSpace frontend, including creating, updating, and deleting accessions, resources, archival objects, digital objects, subjects, agents, and other records. While this post will include a very brief introduction to using the ArchivesSpace API, if you are truly a newcomer to APIs in general and to the ArchivesSpace API in particular (as we were just a few months ago!), there is no better place to start than the ArchivesSpace Developer Screencasts put together by Hudson Molonglo, in particular screencast 2, Backend Introduction. Also, while this post will focus on only a few of the available ArchivesSpace API endpoints, all of the available endpoints are detailed in the ArchivesSpace API documentation.

Using the ArchivesSpace API

A major benefit of the ArchivesSpace API is that it allows users to interact with the ArchivesSpace application without having to modify the core application code or use ArchivesSpace's programming language, Ruby. This is great for us as, while we have learned enough Ruby to write some basic ArchivesSpace plug-ins, most of the programmatic work we've been doing for this project has been written in Python. Utilizing the ArchivesSpace API allows us to continue to use a programming language that we are more familiar with, especially with respect to accessing and modifying our legacy data, to interact directly with the ArchivesSpace application. However, while most of the scripts that will be detailed in this post are written in Python, an easier way to get started interacting with the ArchivesSpace API is by using curl in a Mac or Linux terminal or in a Windows Unix emulator such as Cygwin.

Once you've opened up a terminal or Cygwin with curl installed, you can send a simple request to the ArchivesSpace backend to test that the backend API is available. If you're running ArchivesSpace on a test instance on your local computer (which I highly recommend when experimenting with the API, and with the application in general), that request and the resulting response look something like this:



If you're interacting with an ArchivesSpace instance that is not running on your local machine, substitute your ArchivesSpace instance url for localhost and the port on which the backend is running for 8089.

Most of the really powerful things that can be done with the API require users to verify that they have the permissions to do so, so once you've verified that you can communicate with the ArchivesSpace API, the next step is to authenticate and start a session. To authenticate using the default administrator username and password, the request and first part of the response looks like this:



This request returns a longer response than the first one we sent, including the session token that you will need to include in a header that you must send with every subsequent request. Since the session token is a really really long string, it makes things a lot easier if you store the token as a variable, like so:


Subsequent requests sent to the API should include the session token in the header, like this:


This tells ArchivesSpace that you are an authenticated user who is allowed to do all sorts of really powerful and potentially dangerous things. Again, always test your code on a test instance of ArchivesSpace!

The majority of the actions that can be completed via the API take the form of either HTTP get or post requests. As the names may imply, get requests return some data to the user and post requests submit some data to the application. Get and post requests can often be sent to the same backend endpoint, with get requests including the particular ID of a desired record and post requests including the data (in ArchivesSpace JSONModel format) of the record to be created. Here are some quick examples:

A get request that returns the IDs of all resources in repository 2

A get request that returns the JSON representation of resource 3

A get request that returns the JSON representation of subject 1

A post request that creates a new subject. The API returns a  bit of JSON including the ID and uri of the posted subject

The new subject posted via API as seen in the ArchivesSpace staff interface

Authenticating via the API is about as far as the API documentation goes in terms of providing detailed instructions for beginners. Figuring out how to use all of the available API endpoints to your advantage can occasionally be an easy process but, if you're like me, it often involves a lot of trial and error and head scratching over exactly the request, including data and parameters, the ArchivesSpace API requires for each endpoint. We are by no means expert users of the ArchivesSpace API, but we have figured out some really awesome solutions for our particular problems by getting just far enough in our understanding of some of the endpoints. The rest of this post will be spent documenting a few of the projects we've completed using the API. If you're interested in reading about other great work that archivists have done using other API endpoints, check out Maureen Callahan's post on deleting records and Hillel Arnold's post on automating exports using the API.

Disclaimer: All of the following examples are based on our very particular use cases, legacy data, and programming expertise or lack thereof. As such, the exact workflows and Python scripts shared herein will likely not be applicable to most other institutions and their data. Rather, they are intended to serve as examples of what is possible via the ArchivesSpace API and to provide some guidance on how certain endpoints can be used.

Creating Digital Objects

The idea for this script came from a conversation with a colleague about the possibility of automating the creation of digital objects in ArchivesSpace for digitized archival objects using a spreadsheet inventory of a collection containing the ArchivesSpace ref ID for each archival object, a barcode or other sort of identifier for the digitized content, and the url for the digitized content. Such a spreadsheet could easily be created using an ArchiveSpace exported EAD, which contains the ArchivesSpace Ref ID for each archival object as a component level id attribute:

An ArchiveSpace archival object. Note the Ref ID.

That same archival object in an ArchivesSpace exported EAD. Note the <c02> id attribute.

The series of API requests to create a new digital object and link it to the existing archival object goes like this:

1. Using the archival object Ref ID, search ArchivesSpace for the archival object's uri using the search endpoint


This returns a JSON representation of the search results. Since Ref IDs should be unique, there should only be one search result, containing the following bit of information that we're after:

The uri to the archival object that matches the searched for Ref ID
2. Using the archival object uri from the search results, get the JSON representation for that archival object using the get archival object endpoint


This returns the JSON representation for the archival object, which we can store, add a new instance to, and repost to update the archival object using the API

3. Using the archival object's display string (a concatenation of its title and date) from the archival object JSON and the identifier and digital object uri from the spreadsheet, form the JSON for a new ArchivesSpace digital object and post it using the post digital object endpoint


This returns a JSON message containing the uri for the newly created digital object

The posted digital object

4. Using the uri for the new digital object, create a new archival object instance and add it to the archival object JSON
5. Post the updated archival object JSON to the archival object's uri to update the archival object in ArchivesSpace

I've actually only ever done those last two steps in the Python script that I wrote to automate this whole process. That script can be found here, and here are links to the lines corresponding to each of the above steps to see how it's done in Python instead of curl: [1] [2] [3] [4] [5]

As you may be able to tell, curl is really useful for single, simple interactions with the API, and is helpful for testing some of the API endpoints to see how they work. If you have a series of API calls that will need to be strung together and repeated over and over again, it's much easier to do that using a programming language like Python and its requests library.

Creating Subjects and Agents

As we've mentioned many times on this blog before, we are migrating our legacy descriptive metadata to ArchivesSpace using our EADs. As such, one of the limitations we have faced is the stock ArchivesSpace EAD importer, which works for the most part but is not exactly what we need for the data we have. Our solution to that has been our huge EAD cleanup project that we've detailed on this blog in combination with some local modifications to the EAD importer. But what happens when the issue is not due to messy data or to the ArchivesSpace EAD importer, but to the limitations of EAD itself?

EAD 2002 does not have a way of representing subdivided subjects (e.g., Ann Arbor--Dwellings.) or the various components that make up agents (primary name, rest of name, dates of existence) to the same level of granularity of ArchivesSpace (EAD3 will help!). Take a look at this <geogname> in our EAD for instance:


This is imported into ArchivesSpace like this:



When really it should look like this:


We want our data to be migrated to ArchivesSpace as cleanly and correctly as possible and, while subdivided subjects might seem like not-such-a-big-deal, we plan to use ArchivesSpace to export MARC XML records for our collections, we will ultimately want to take advantage of the functionality of EAD3, and new subjects will likely be created in ArchivesSpace following the example in the second ArchivesSpace image above, so now is really the best time to ensure that our legacy subjects will be migrated to ArchivesSpace properly. Enter the API.

Posting subjects via the API is actually really simple (see the example using curl way near the top of this post). What was REALLY complicated about the process of using the API to post our subjects is that a term type is required for each individual term. Since EAD does not have the structure to support multiple terms, much less term types, the HIGHLY messy process that we used looks like this:


1. Agonize over the apparent hopelessness of the issue for a little while, until we realize that we have MARC records for all of our collections and that those MARC records have structured subdivided subjects with terms and term types
2. Get a MARC XML export of all of our archival collections from our catalog
3. Use a combination of scripts to make a csv of all of our unique EAD subjects with subdivided subjects split up into individual terms and a csv of all of our MARC subjects with each individual term and term type identified
4. Run a script that identifies the term type for all individual terms and outputs a csv with all of our unique EAD subjects with individual terms and term types included 
5. Use the API to post all of our subjects to ArchivesSpace correctly, outputting a csv with each subject and the uri of the posted subject in ArchivesSpace
6. Add the posted subject uris to our EADs as ref attributes. We know this is invalid EAD and we know it's wrong, but one of the other ways we've found of getting around the limitations of EAD for this migration is to ignore them! (We're also saving all of our "break the EADs" scripts until just before we migrate)


7. Modify the EAD importer to use a ref attribute to link to existing subjects instead of creating subjects during the import process

That's all it takes to use the API to create over 10,000 well-formed ArchivesSpace subjects! Max has done something similar to split up our agents into their component parts and post them using the API to take advantage of the more structured nature of ArchivesSpace person, corporate entity, and family name records.

Accession Migration

This one actually isn't done yet. As Max's recent post explained, we've recently started looking at migrating our legacy accession records from FileMaker Pro. ArchivesSpace has an accession csv importer but, due to the limitations of the converter, our own messy data, and some of the more complicated things we want to do with our accession imports, we'll need to make some local customizations to how the migration will be done. One option is to modify the accession csv importer as we have modified the EAD importer but, as we've been learning more and more about the API, we've realized that it will be much easier for us to come up with our own accession migration script that will use the API to migrate our accession records in the way that we want with the data that we have. We'll definitely be writing about the process along the way!


Tuesday, August 25, 2015

Thorny Issue Number 2: Importing Legacy Accession Records Into ArchivesSpace

First off, all of us here are at the Bentley are still reeling and recovering from an exciting week at ARCHIVES 2015 in lovely Cleveland, OH. In case you missed it, a number of us involved with the Mellon grant and/or the A-Team participated:
  • Courtney Mumma (Archivematica), Brad Westbrook (ArchivesSpace), Mike, Dallas and I demonstrated a prototype of the Appraisal and Arrangement tab and showed off the newest enhancements to the Archivematica and ArchivesSpace integration workflows to gain public feedback. "Whee!"
  • Courtney and Mike (as well as a number of other folks) discussed lessons learned through the planning, development, testing and production of digital preservation applications.
  • Maureen Callahan (Yale), Regine Heberlein (Princeton) and Dallas presented case studies of how archivists (none of whom are IT professionals) learned and applied powerful metadata clean-up tools and strategies.
  • Templeton "Faceman" Peck, err, Olga, discussed legal and ethical challenges related to records of the Dr. Kevorkian-assisted suicides.
Good times! But enough reminiscing! Time to get back to work!

Planning for Importing Legacy Accession Records Into ArchivesSpace


In a previous post that outlined our strategy for implementing ArchivesSpace here at the Bentley Historical Library, I mentioned that Dallas and I have two major responsibilities as members of the A-Team: 1) importing legacy EADs; and 2) importing legacy accession records.

Note: There's actually a third thorny issue, importing MARC records for archival collections that don't have finding aids. However, for our own sanity, we've decided to kick that particular can down the proverbial road, at least until after April 1, 2016 when our grant is over.

As a reminder, this is the A-Team at the Bentley Historical Library. That's right, I'm B. A. Baracus.

I'm happy to report that we're starting to finish up our work on the first thorny issue. In fact, I'm proud to say that all 2,847 of our legacy EADs can be imported into ArchivesSpace, and that all of our cleanup work we've been doing lately has just been "icing on the cake." (For ample evidence that we really don't know when to say no, see this recent pull request, and this one as well, as well as the still ongoing saga.) Today's post explores the second thorny issue: importing legacy accession records into ArchivesSpace.

Context


As a reminder, we currently keep track of accession data in a homegrown FileMaker Pro database called BEAL, a [b]ac[k]ronym for the Bentley Electronic Accessioning and Locating System.

ACCESSIONS in BEAL

I've said before that BEAL has been described as the "lifeblood" of the Bentley. After some reflection, I think I'd actually be a little more precise and say that the information inside of it, especially the tables with data on donors and accessions, is the real lifeblood of the Bentley. BEAL is more like the circulatory system--without blood, there's nothing to circulate! The central importance of this information to nearly everything else that the Bentley does (not to mention the central importance of this information to key archival concepts like provenance) certainly weighs heavily on our minds we begin our transfusion migration to ArchivesSpace.

The Good News


First, the good news. We've done some initial exploring of the problem (shout-out to Jessica Venlet, a former intern at the Bentley now doing a fellowship at MIT!) and the data itself appears to be less complicated than the [meta]data contained in our EADs. While there's more of it (19,312 records), it's flatter (no c0x levels to worry about) and somewhat more consistent and predictable (due to the fact that less hands in general and, to be honest, less inexperienced hands, have historically created and edited accession records). It's also relatively easy to get our hands on a copy of the authoritative version (OK, given the fact that we have some very old accession records that only exist in paper, maybe more like an-indefinite-article-but-almost-a-definite-article authoritative version) of accession records (a simple CSV export from our FileMaker Pro database). This is something that, given our convoluted way of creating finding aids, you simply can't say about our EADs.

In short, getting accession information from BEAL to ArchivesSpace won't be quite as simple as just mapping and crosswalking it, but almost.

The other bit of good news is that we also now have a blueprint for how to get information into ArchivesSpace because of all the work we've been doing with our EADs. Just like we've done with EADs, we plan to import agents first via the ArchivesSpace API (since these are the building blocks accessions in ArchivesSpace and of the events applied to them), and then the accessions themselves. That, coupled with the fact that we now have some good experience cleaning up and manipulating data programmatically, means we don't have to start from scratch!

The Bad News


It's not all rainbows and unicorns, however. We will still have issues with importing legacy accession records into ArchivesSpace, and unfortunately many of these are more complicated than dealing with technical challenges or messy data.

First, while there aren't many people that create accession records, there are many people who use them. What's more, many people use them in many different ways. Just take a look at this chart of current BEAL functionality (these are only those tables that relate to accessions in some way) that Mike recently put together:


TableFunction
Accessions

Create accession records for newly acquired physical/digital collections

Document purchase of books

Documentation of restrictions/rights issues based on gift agreement and donor communications

Information on gift agreement (status, unique features, etc.)

Documentation of separations

Track processing status

Record locations of unprocessed materials

Collect information for DART donor reports
Contacts

Contact details (with address and phone)

Records of additional names, affiliations

Mailing information

Status (for mailing list, as donor or 'friend', deceased or defunct)

Generate mail labels
Donors

List of accessions from a given donor

List of collections from a given donor

Donor information (may be pulled from contacts?)
Location Guide

Locations of content
Collection Record


Collection: Digital Deposits

List of digital deposits for a given collection
Collection : Digital Deposits - Items

Size of completed deposit (and separations)

Dates of deposit in Deep Blue and dark archives

Location of fully processed content in Deep Blue and Dark Archives

Information on access/use restrictions (including open dates)

Links to manifest/log files for deposit

Processing status

Description of entire deposit


That's a lot!

Before we do anything in ArchivesSpace, we'll need to decide which of these functions (not just which data) that we, as an institution, will continue to do in ArchivesSpace (and from there, which currently can be done and which will need some work before we can do them), and which we won't (and where and how, and even sometimes if we'll continue to do them). This will involve talking to other people on staff here about their their day-to-day work, and potentially making political, not just technical, decisions. As you might imagine, the political decisions are harder.

Second, while we don't anticipate many mapping issues, we do anticipate some. There simply isn't a 1:1 ratio between fields in BEAL and fields in ArchivesSpace, and even when there are, data entry conventions (and even data models!) in each system can be different.

Typically things are more granular and complex in ArchivesSpace. Locations and events in ArchivesSpace are good examples of this. In ArchivesSpace, we'll be able to track and manage locations as separate entities attached to containers attached to instances attached to intellectual entities in resources. Right now we just have plain text box numbers and locations in accession and resource records. With regard to events (and, by the way, let me be the first to say that wrapping your head around events in ArchivesSpace is almost as hard as wrapping your head around events in philosophy), many "objects in time or instantiations of properties in objects" (like acknowledging the receipt of an accession, processing a collection, etc.) that in BEAL exist as simple check boxes in ArchivesSpace become full-fledged events with timestamps and descriptions and even separate, associated agents.

I'll also say that all this is all a good thing! We'll be able to do much more with our data than we ever have. I'm particularly eager to start playing around with live reports. It will just take a bit of thinking to get from here to there and to do this right. Exacerbating this challenge is the fact that there isn't an archival standard for recording accession information like there is for recording descriptive data (good ole DACS!), so we don't have something to point to for settling disputes when issues like these come up.

Finally, there's a privacy issue. Some of the data that we keep about our donors and accessions is sensitive. One of the best things we've done for our work here to prepare legacy EADs for import to ArchivesSpace is to introduce Git and GitHub into our archival workflows for distributed version control so that lots of us can work on the same set of data without stepping on each others toes. Because of the privacy issue, however, we won't be able to use GitHub in the same way. We've discussed using a private repository for working with legacy accession data, but we're somewhat uncomfortable with this idea. All that is to say that we're open to suggestions about version control systems or methods that will enable many hands to work on the same set of sensitive data in an efficient way!


The Ugly News


Well, that's the good and bad news. It turns out there's some ugly news as well. We still have dirty data--locations done the "old way" and the "new way," names that appear to refer to the same person in two different tables, one with one middle initial, and one with another, etc. Maybe dirty data is just a fact of life in the library and archives domain!

Conclusion

 
New Accession record in ArchivesSpace


So now you know how we'll be spending our fall semester! Be sure to stay tuned for more on our adventures in accessions.

Have you imported your legacy accession data into ArchivesSpace? How is the way you manage accessions in ArchivesSpace different from the way you managed them before? Do you also think events are confusing (even if necessary) in ArchivesSpace and/or philosophy? Let us know!

Tuesday, August 18, 2015

Appraisal and Arrangement Tab Live Demo at SAA2015

Psst.... In case you haven't heard, the Bentley Historical Library is teaming up with Artefactual Systems and LYRASIS for a brown bag lunch session at SAA2015 on Thursday, August 20, 2015 from 12:15pm - 1:30pm in room 25C of the Cleveland Convention Center.

In addition to hearing about the broader goals of ArchivesSpace-Archivematica integration, Courtney Mumma will discuss development work that builds off of current Archivists' Toolkit-Archivematica integration for the Rockefeller Archive Center and Max, Dallas, and I will demo the new Appraisal and Arrangement tab:


Find more information on the demo (and avenues for you to provide feedback) on the Archivematica wiki:  https://wiki.archivematica.org/SAA_2015_Demonstration_and_Feedback.

If you're feeling really adventurous, you can download and install the Appraisal Tab prototype using Artefactual Labs' github (see above link for installation instructions).

We really want this development work to be flexible enough to meet our needs as well as yours, so please come on out to the brown bag to learn more about the project and share your thoughts.

See you in the CLE!

Thursday, August 13, 2015

Advocacy and Born-Digital Archives

Earlier this week, I helped our development officer draft a proposal to seek additional funding for our digital curation program.  Given the potential audience of library/university administrators and external financial donors, I quickly realized that I had to adjust how I typically represent our work with digital archives.

SIPs and DIPs, checksums and disk images—in short, all the fun things we talk about in our listservs, blogs, and conferences were probably going to be meaningless to these folks.  Instead, I needed to capture what we do on a daily basis and explain why it matters to a fairly diverse group who wouldn't know the OAIS reference model from a hole in the ground.

Have you seen my functional entities?
Then again, why should they have to know about the gory details of OAIS?  I'm of the mind that key stakeholders (our administrators, funders, donors, researchers, etc.) don't necessarily need to know about the minutiae of digital preservation or the alphabet soup of acronyms regularly featured in conference presentations and procedural manuals (unless they're interested, of course...).

What they really need to know is what we do (at a high level) and why it's important.  If we're unable to accomplish this objective—and I would argue that every successful digital archives/preservation program does so—then we face an uphill battle for resources and, ultimately, relevance.  And that may only be a wee bit hyperbolic.  Advocacy is an integral component of the archival enterprise and we—archives of every size and stripe—are together in this quest to demonstrate our value and seek support.

What is it that you do, anyway?

As archivists, we already face an uphill battle when it comes to explaining our jobs.  Throw in the complexities of digital archives and there's little wonder that your spouse/parent/friend's eyes may start to glaze over when you describe an average day on the job.  (Or maybe that's only happened to me?)  One of our major challenges, then, is to be able to represent our work in a way that others can relate to and understand.  

At the same time, advocacy isn't just about education; simply raising awareness about archives and our value is a function of outreach.  Advocacy, on the other hand, goes the extra step in seeking to influence the actions and decisions of our interlocutors. 

Making the Case

As with any communication, it's important to tailor the message for the audience and the specific issue at hand.  The message you deliver to donors of materials may therefore be very different from what you prepare for administrators.

So let's say you've identified the stakeholder(s) you want to reach; while it's important to convey a sense of what you do, save the workflow matrices, code snippets, and UML diagrams for your colleagues.  Successful advocacy should involve and inspire the audience so that they understand how and why they might benefit from our work.  

This in turn requires us to communicate the "added value" that archivists bring to the preservation and curation of digital archives.  This value includes such things as (and I'm preaching to the choir, here):

  • Organizing and describing materials so that researchers understand the nature of our collections
  • Taking steps to ensure that content can still be accessed and used far into the foreseeable future
  • Protecting sensitive personal information and deploying appropriate restrictions based on rights, institutional policies, legal requirements, and donor agreements.
  • Improving the means by which people search for and retrieve content
  • Developing or providing resources to help researchers use (or reuse) materials in important and meaningful ways.  
If we want people to care about our work—and to demonstrate that interest by committing resources or collections to our institution—we need to be unequivocal about the benefits we bring to the table (and the more precise or quantitative, the better!).

To help convince the stakeholders of your case, it's also important to highlight any innovations, achievements, or recognition related to the topic at hand.  Doing so will establish your/your institution's legitimacy and credentials and points the way to continued or future success.  I don't have any hard data to back this up, but it's my strong conviction that current and/or potential stakeholders are more willing to donate their support and/or resources when presented with a proven track record and the opportunity to continue/advance that work.  Modesty is no longer a virtue!
And now I'll take the plunge and some of that document I mentioned above:
Over the past two decades, our collective historical record has undergone a sea change: the web has revolutionized publishing and any number of businesses, email and tweets have replaced personal letters, word processing files and spreadsheets now comprise organizational records, and the convenience and ubiquity of smart phones enable us all to amass large collections of digital photographs and video.  
As our professional and personal lives increasingly move online and into other digital spaces, the Bentley Historical Library has emerged as a proven leader in the quest to preserve and make accessible essential born-digital materials.  
Today, researchers can access the electronic records of former Governor Jennifer Granholm, review the source code for the influential Michigan Terminal System time-sharing operating system from 1968, and study the creative digital output of noted artists such as Peter Sparling, Vince Castagnacci, and Arnold Weinstein. 
Our work in “digital curation” encompasses traditional archival functions—the process of selecting materials of high research or intrinsic value and making them accessible to researchers—but also involves additional steps to ensure the integrity and authenticity of content.  The Bentley furthermore seeks to add value to materials through the production of detailed description, access portals, and tools that help patrons find answers to their most pressing research questions.  
OK—having said all that, I want to acknowledge that I am far from being an expert on this topic!  I would be delighted to hear about important points that I've missed or to see examples of successful advocacy for digital archives.  If you've got anything to share, leave a comment!

Friday, July 31, 2015

Order from the chaos: Reconciling local data with LoC auth records

Arkheion and the Dragon, part II

By the end of last week's post/parable we had Library of Congress (LoC) name authority IDs for many of our person and corporation names, but had a lot of uncertainty as to whether these IDs had been matched correctly. The OpenRefine script we were using to query VIAF for LoC IDs also didn't support searching for any control access types beyond person and corp names.

We weren't quite satisfied with this, so after looking into some of our options, we decided to try a new approach: we would move from OpenRefine to Python for handling VIAF API queries and data processing, add a bit of web scraping, then use more refined fuzzy-string matching to remove false-positives from the API results. By the time we had finished, we had confirmed LoC IDs for over 6500 unique entities (along with ~2000 often hilariously wrong results) and, as an added benefit, were able to update many of our human agent records with new death-dates. All told the process took about a day.

Here's how we did it:

The VIAF API

OCLC offers a number of programmatic access points into VIAF's data, all of which you can see and interactively explore here. Since we're essentially doing a plain-text search across the VIAF database, the "SRU search" API seemed to be what we were looking for. Here is what an SRU search query might look like:

http://viaf.org/viaf/search?query=[search index]+[search type]+[search query]&sortKeys=[what to sort by]&httpAccept=[data format to return]

Or, split into its parts:

http://viaf.org/viaf/search
    ?query=[search index]+[match type]+[search query]
    &sortKeys=[what to sort by]
    &httpAccept=[data format to return]

There are a number of other parameters that can be assigned - this document gives a detailed overview of what exactly every field is, and what values each can hold. It's interesting to read, but to save some time here is a condensed version, using only the fields we need for the reconciliation project:

  1. Search query: how and where to find the requested data. This is itself made up of three parts:
    1. Search index: what index to search through. Relevant options for our project are:
      • local.corporateNames: corporation names
      • local.geographicNames: geographic locations
      • local.personalNames: names of people
      • local.sources: which authority source to search through. "lc" for Library of Congress.
    2. Match type: how to match the items in the search query to the indicated search index -- e.g. exact("="), any of the terms in the query ("any"), all of the terms ("all"), etc.
    3. Search query: the text to search for, in quotes
  2. Sort keys: what to sort the results by. At the moment, VIAF can only sort by holdings count ("holdingscount").
  3. httpAccept: what data format to return the results in. We want the xml version ("application/xml")

Putting it all together, if we wanted to search for someone, say, Jane Austen, we would use the following API call:

http://viaf.org/viaf/search
    ?query=local.personalNames+all+"Jane Austen"+and+local.sources+=+lc
    &sortKeys=holdingscount
    &httpAccept=application/xml

The neat thing about web APIs is that you can try them out right in your browser. Check out the Jane Austen results here! It's an xml document with every relevant result, ordered by number of holdings for each entry worldwide, and including including full VIAF metadata for every entity. That's a lot of data when all we're looking for is the single line with the first entry's LoC id. This is where Python comes in.

Workflow:

Before we dive in to the code, here is the high-level workflow we ended up settling on:

  1. Query VIAF with the given term
  2. If there's a match, grab the LoC auth id
  3. Use the LoC web address to grab the authoritative version of the entity's name.
  4. Intelligently compare the original entity string to the returned LC value. If the comparison fails, then we treat the result as a false positive.

Let's dig in!

VIAF, LC, and Python

First, we wrote an interface to the VIAF API in python, using the built-in urllib2 library to make the web requests and lxml to parse the returned xml metadata. That code looked something like this:


You can see above that the search function takes three values: the name of the VIAF index to search in (which matches to one of our persname, corpname, or geogname tags), the text to search for, and the authority to search within (here LC, but it could be any that VIAF supports).

With the VIAF search results in hand, our script began searching through the xml metadata for the first, presumably most relevant result. All sorts of interesting stuff can be found in that data, but for our immediate purposes we were only interested in the Library of Congress ID:


Now that we had the LC auth ID, we could query the Library of Congress site to grab the authoritative version of the term's name. Here we used BeautifulSoup, a python module for extracting data from html:


Now we had four data points for every term: Our original term name, an unvetted LoC ID number and name, and the type of controlaccess term the item belongs to (persname, corpname, or geogname). As before, there were a number of obvious false-positives, but there were enough terms that we did not have nearly enough time to check through them individually. As Max hinted at in last week's post, this was fuzzywuzzy's time to shine.


Fuzzy Wuzzy was (not) a bear

(also not a Rudyard Kipling poem)

Max gave an overview of fuzzywuzzy, but just as a refresher: it's a python module with a variety of methods for comparing similar strings under different lenses, all of which return a "similarity score", out of 100. Here is what a very basic comparison would look like:

This is fine, but it's not very sophisticated. One of fuzzywuzzy's alternate comparison methods is much better suited for our purposes:

The token_sort_ratio comparison removes all non-alphanumeric characters (like punctuation), pulls out each individual word, puts them back in alphabetical order, and then runs a normal ratio check. This means that things like word order and esoteric punctuation differences are ignored for the purposes of comparison, which is exactly what we want.

Now that we had a method for string comparisons, we could start building more sophisticated comparison code - something that returns "True" if the comparison is successful, and "False" if it isn't. We started by writing some tests, using strings that we knew we wanted to match, and strings we knew should fail. You can see our full test suite here.

As a result of our testing, we decided we would need to have unique comparison methods for each type of controlaccess term we were testing - one for geognames, one for persnames, and one for corpnames. Geognames turned out to be easiest - in that case our test criteria was matched with a basic token_sort_ratio check - names were deemed correct when they had a fuzz score of 95 or higher. Both persnames and corpnames turned out to need a bit more processing before we had satisfactory results. Here is what we came up with:

With this script in hand, all we had to do was run our VIAF/LC data through it, remove all results that failed the checks, then use the resulting data to update our finding aids with the vetted LoC authority links (while removing all the links we had added pre-vetting). Turns out, we ended up with > 4200 verified unique persname IDs, ~1500 IDs for corpnames, and 800 geogname IDs, all of which we were able to merge back into our EADs using many of the methods Max described last week. We also output all the failed results, which were sometimes hilarious: apparently VIAF decided that we really meant "Michael Jackson" while searching for "Stevie Wonder". And, no, Wisconsin is not Belgium, nor is Asia Turkey.

This also gave us a great opportunity to update all of our persnames with death-dates if the LoC term had one and we did not. You can check out our final GitHub commit here - we were pretty happy with the results!.

Postscript

In retrospect there is a lot we could improve about the process. We played things conservatively, particularly for persnames, so our false-positive checking code itself had a number of false-positives. Our extraction of LoC codes from the VIAF search API could be a lot more sophisticated than just mindlessly grabbing the first result - our fuzzy comparisons did a lot to mitigate that particular problem, but since VIAF sorts its results by number of holdings worldwide rather than by exact match, we were left to the capricious whims of international collection policies. The web-request code is also fairly slow - since we didn't want to inadvertently DDoS any of the sites we're querying (and we'd rather not be the archive that took down the Library of Congress website), we needed to set a delay in between each request. When running checks against >10,000 items, even just a one second delay adds up. Even so, it still runs in an afternoon -- orders of magnitude faster than manual checking.


We hope you've found this overview interesting! All code in the post is freely available for use and re-use, and we would love to hear if anyone else has tried or is thinking of trying anything similar. Let us know in the comments!