Saturday, May 31, 2008

Meta what? What data? What what?

(Cohen, 1999)

Preservation of information overlaps many different other related aspects of information management, so I won't dwell overly long on the concept of metadata, but I do think that it is important enough in the concept of preservation to get a brief mention. In the context of information preservation, metadata is data about data (metadata, 2008). That is to say that it is information that is used to describe relevant information pertaining to the information it is ascribed to.

For example, if we were to preserve a photograph of a man in front of a bridge, the metadata would contain information about when the photograph was taken, who the man was, which bridge is in the photograph, who the photographer was, and any other relevant information that was known about that photograph. This information is then used in the preservation of the original piece of information (the photograph) in order to provide contextual information related to the photograph that researchers might use.


What value does this information actually have though?

Consider that without the context of any given piece of information there is little value to that information. Take the above example - a photograph of a man and a bridge. If we know nothing other than that we are preserving a photography of a man in front of a bridge we have few was to relate this to other information that we have preserved, and so it cannot be reliably used for reference because we know so little about it. But, since we preserved the contextual information along with the object itself, we have the ability to relate it to other information we have preserved. We know when it was taken, where it was taken, and by who this single piece of information can be referenced by date, location, photographer, the bridge, the type of media used to make the photograph, and even the style of photography - all of these smaller pieces of information can be used to better preserve the original by providing us 'handles' for drawing connections between it and other preserved information (Frendo, 2007).

Preserving information requires an understanding of the context of the information being preserved if it is to be useful to anyone in the future.


works cited:

Cohen, D. X. (Producer). (1999). Futurama. [Television series]. U.S.A.:Fox.

Frendo, R. (2007). Disembodied information: metadata, file plans, and the intellectual organisation of records. Records Management Journal, 17(3), 157-.

metadata. (2008). In Merriam-Webster Online Dictionary. Retrieved May 20, 2008, from http://www.merriam-webster.com/dictionary/metadata

Wednesday, May 28, 2008

In the News - British Library audio preservation project


Time's running out to preserve our treasures

Summary:

Digitization at the British Library has allowed many people to conveniently access information that was previously available only in the library itself, and even then required patience to access. But the digitization road has not been simple - a great deal of work goes into making sure that the preservation solution meets the goals of the library. Adding to the difficulty is the rapidly diminishing availability of older technology needed to access sound recordings, making the process a race against time.


Works Referenced:

Holder, D. (2008). Time's running out to preserve our treasures.
Guardian.co.uk. Retrieved May 28, 2008 from http://education.guardian.co.uk/librariesunleashed/story/0,,2274851,00.html

Sunday, May 25, 2008

Digital preservation - Media and format


Earlier I briefly discussed what it means to preserve information, why it is preserved, and by whom, and now I would like to revisit some of the points made in the first post regarding problems inherent in preservation in a little more detail - namely digital storage formats. Preserving information in the various digital formats available today offers many advantages over the preservation of the original physical information, such as access to more people in more locations (Gwinn, 1996), but it has some downsides as well. These downsides don't seem to be getting quite as much publicity as the upsides, and this has created a somewhat lopsided view of the format that could negatively impact those who rely on it or plan to rely on it in the future.





















This is an 8" floppy disk - I still have one box of them left from sometime the 1980's. How many people reading this can remember using these (honestly)? How many even knew they existed? There is data preserved on this disk - not much, but some - yet I cannot access it because I have no 8" floppy drive, nor do I have an IBM system 34 to read the information even if I had the drive. The technology required to access the information preserved on this disk is obsolete and unavailable, therefore the information preserved on this media may as well not even exist - in essence it does not exist.

When the Library of Congress began to investigate the process of digital preservation it conducted a study to better understand what challenges would be involved in the process. They found that due to the progress of technology they were unable to determine how long the digitized information would continue to be accessible. They also found that it was imperative to preserve the original forms in the event of a loss of the digital versions, or for the future as newer technology became available for enhanced digitization (Marcum, 2007).

As technology progresses the ability to access information preserved on the media of older technology is still required to access that information, unless all of the information is moved to a more modern format before the technology required to access it has become unavailable. This continual updating of storage media translates into an ongoing real world cost for the preservation of digital information, which needs to be accounted for when considering what is involved in preserving digital information for the long term future.


Suggested Further Reading:

the Computer History timeline

A Brief History of Computing summary timeline


Works cited:

Feldman, S E. (1997). "It was here a minute ago!" Archiving the Net. Searcher, 5, 52-64.

Gwinn, N. E.. (1996). "Mapping the Intersection of Selection and Funding," Selecting Library and Archive Collections for Digital Reformatting: Proceedings from an RLG Symposium, Mountain View: The Research Libraries Group, Inc., p58.

Marcum, D B. (2007). Digitizing for access and preservation: strategies of the Library of Congress. First Monday, 12(7).

Thursday, May 22, 2008

The Internet! Is that thing still around?


I borrowed the title for this post from Homer Simpson in an old episode of "The Simpsons" (Groening, 2006), but there is a kernel of truth here that I think needs a little more scrutiny. How many times have you gone back a bookmark or shortcut only to find that the page you were looking for was gone? Internet-only information is ephemeral, sometimes to an astonishing degree. How can this information be preserved, and how much of it should be preserved?

According to some people, we should preserve anything that someone else is not already preserving because no one can accurately say now what will be of value at a later point in time (Feldman, 1997). Obviously this will not work for everyone who preserves information because preservation is not free; storage costs, access methods, copyright fees, and personnel all cost money, and the last time I checked not too many people had endless supplies of cash on hand for the preservation of information.

What then is to be saved and what should be lost to the mists of time so to speak? As an individual, my first inclination is to say that nothing should be lost, but in reality, this is a much bigger issue that first meets the eye. Because of the aforementioned monetary issues, many organizations have established guidelines for what will be saved and what will be left for others to preserve. A good example of this is the Internet Archive, who's F.A.Q. states the following:

"we collect only publicly accessible Web pages. We do not archive pages that require a password to access, pages tagged for "robot exclusion" by their owners, pages that are only accessible when a person types into and sends a form, or pages on secure servers. If a site owner properly requests removal of a Web site through http://www.archive.org/about/exclude.php, we will exclude that site from the Wayback Machine."

Because not all information on the Internet is necessarily saved by every organization, things like Internet forums or personal online journals are not recorded, and therefore some other source must preserve this information if it is to be preserved at all. While some groups who create these kinds of information do preserve their content (for example the alternative process listserv archives) not all do, and that is an inherent preservation problem with the Internet (Beagrie, 2005).

In order to better understand this issue, the British Library conducted a study called "the Digital Lives research project" which found that there are a number of issues that will need to be dealt with in order to find a workable and sustainable solution to the preservation of this sort of information. One of the more significant problems they study found was that there was no common format for this information, which complicates efforts to preserve it. The study also found that there was no clear consensus regarding what was worth preserving and what was not (Pencock, 2006).

Until these issues are resolved, and they may never be fully resolved, the preservation of information that originated in an electronic form will be a complicated matter that those involved in the preservation of information will have to deal with.

Suggested Further Reading:

Report of the Task Force on Preserving Digital Information (OCLC)

How to Preserve Authentic Electronic Records


works cited:

Beagrie, N. (2005).
"Plenty of Room at the Bottom? Personal Digital Libraries and Collections". D-Lib Magazine 11(6).

Feldman, S E. (1997). "It was here a minute ago!" Archiving the Net. Searcher, 5, 52-64.

Groening, M. & Brooks, J. L. (Producer). (2006). The Simpsons [Television series]. U.S.A.:Fox.

Pennock, M. (2006). Digital Preservation Coalition Forum on Web Archiving. Ariadne, (48), 1-.

Tuesday, May 20, 2008

So, who saves all this stuff anyway?


To better understand the preservation of information it seems logical to ask "Who does save all of this information?" Not all information is saved by all people or organizations. Different preservers of information save different information depending on a number of different reasons. For example; many individuals save personal or family related information, many special interest groups save information related to their special interest, and many libraries, archives, and museums save information on a lager and more diverse scale. Lets examine briefly these different entities that preserve information to get a better idea of who each of these groups really are.

Individuals:

There are innumerable individuals collecting pieces of information that they find valuable - and you may be one of them (I know I am), but to what extent do individuals preserve information? Most people have personal records like birth certificates, family photographs, passports, diplomas, and such, but what other sorts of information do people preserve? Some individuals preserve great collections of events, places, or things that interest them, such as the George Fisher collection of political cartoons that was donated to the University of Arkansas library (Simpson 2004). These preserved collections can present a detailed view of something that would otherwise be lost to history, as in this particular instance where personal correspondence was part of the collection.


Special Interest groups:

In this ever-growing Internet-based world, many different groups of people have formed around a common interest online, and in doing so have collectively built a body of information to be preserved for their use and in many cases that of others who share their interests. While these groups tend to focus on a specific subject they can be invaluable resources for preserving information on that one subject (Wilson & Peterson, 2002).


Libraries, Museums, and Government Institutions:

While there are various organizations that fall into these three categories, each of them have a somewhat different approach to the preservation of information, ranging from public libraries and community museums, which sometimes preserve collections of local interest, to the National Archives of various countries which preserve information gathered from all over the country they represent.

These groups represent a much larger body of information than any single holding, and because some of these organizations preserve a broad variety of information in little depth (macro scale) while others have a tighter focus of subject but in greater depth (micro scale), combined they achieve a more complete preservation of information.


Suggested Further Reading:

The Internet Archive F.A.Q.

Preserving Government and Political Information: The Web at Risk


works cited:

Simpson, E C. (2004). George Fisher Collection Donated to the University Libraries. Arkansas libraries, 61(1), 7-8.

Wilson, S.M, & Peterson, L.C. (2002). THE ANTHROPOLOGY OF ONLINE COMMUNITIES. Annual review of anthropology, 31(1), 449-.

Friday, May 16, 2008

Why preserve information? (Does it go bad?)


Information is preserved for a number or reasons - historical value, future reference, public and/or personal interest, and economic value to mention but a few. In order to better understand the nature of why people preserve information I will take a brief look at some of these reasons in a little more detail.

Historical value and future reference:

History is important to the future, and to the present, but in order to understand history we need to be able to examine it in some detail, the process of which is called historiography. Historiography is defined in the Encyclopedia Britannica as follows:
" The writing of history, especially the writing of history based on the critical examination of sources, the selection of particulars from the authentic materials in those sources, and the synthesis of those particulars into a narrative that will stand the test of critical methods."
A very important clause in that definition is "Critical examination of sources". How can we examine historical sources if none now survive? How can we determine which sources are authentic if we cannot examine them? Only though the preservation of the information of the past can we examine it critically, and only through proper interpretation can we make and informed opinion of it's lessons. As John Arnold describes in his book "History: a very short introduction" history has been interpreted differently throughout the ages, sometimes in an attempt to more 'accurate' and sometimes in an attempt to conform to popular views, disregarding fact for a more 'appropriate' interpretation.

By it's nature, once the original source of any given piece of information is lost it can never be re-created. Interpretations of the original may be used to try to recreate what was lost, but how can anyone be certain that what they have extrapolated from these interpretations has anything to do with the actual original information? Thus it is important to preserve original forms of information if we wish to be able to rely on it in the future for critical analysis.

Public or Personal interest:

How many times have you found some useful tidbit of information while doing research only to never be able to find it again in the future? I know that I have, on a number of occasions, found some very interesting bits of information, either in print or electronic form, only to search in vain the next time I attempted to find it again. Some forms of information, especially information that is non-academic, sometimes appears to be so ephemeral as to have a lifespan of but a single access! This seems to be especially so when it involves Internet-based information.

Blogs, personal web pages, non-academic online articles, and even conversations in public forums all tend to be very short lived because with few exceptions they are not designed with preservation in mind. Some resources do exist to preserve these more ephemeral forms of information, like the "Internet Way Back Machine" and the Internet archives at the New Library of Alexandria, Egypt. These sites specifically state in their F.A.Q. pages that anyone not wishing to have their information recorded may exclude it from these archives, and coupled with the incalculably enormous volume of content currently in existence on the Internet, these utilities cannot be totally comprehensive.

In 2007 there were over 100 million web sites in existence on the Internet worldwide, up from 18,000 sites in 1995 (Foster, 2007). As the expansive growth of electronic information continues it is becoming more and more important to find ways to preserve this information for the future.


Works cited:

Arnold, J. (2000). History: A very short introduction. New York: Oxford University Press.

Foster, A. (2007). Information Navigation 101. The Chronicle of Higher Education, 53 (27). A38-40.

Historiography. (2008). In Encyclopædia Britannica. Retrieved May 16, 2008, from Encyclopædia Britannica Online: http://search.eb.com/eb/article-9108622

Wednesday, May 14, 2008

What does it really mean to preserve information?



According to my handy 1997 Merriam-Webster desktop edition dictionary, the word 'preserve' is defined as follows:

Function: verb
Inflected Form(s): pre·served; pre·ser·ving
1 : to keep or save from injury, loss, or ruin : PROTECT
2 : MAINTAIN 1, continue

So, to preserve information would mean to save it, to prevent it from becoming inaccessible, or from being damaged or altered by maintaining at least its' original content, if not its' original form. But, how would this be accomplished, given information's multitudinous forms and formats?

In order to examine this question, perhaps it would be a good idea to look a little closer at the various formats in which information may be found these days. Information may be found in print format (books, periodicals, etc.), on film (or even glass plates!), and even in various completely electronic formats
such as this blog (assuming that you did not print this out) to name but a few. Each of these formats comes with it's own set of special considerations when it comes to preservation, given our working definition above, and as such, each must be handled somewhat differently if we are to achieve our goal.

Lets take a cursory look at some of the pros and cons of each of these different formats, just to get an idea of what sorts of problems each of them might present to those who wish to preserve them.


Physical printed information:

Pros:
  1. It requires no special equipment to access.
  2. It's format never becomes inaccessible due to changes in technology.
Cons:
  1. Printed information, being physical, takes up physical space.
  2. It can generally only be accessed by one person at a time.
  3. It often requires special storage considerations, which can add to the cost of it's preservation.
  4. Physical deterioration will eventually cause it's destruction.

Physical visual information:


Pros:
  1. It may or may not require special equipment to access.
  2. Some forms will not become inaccessible due to changes in technology.
Cons:
  1. Takes up physical space.
  2. It can often only be accessed by one person at a time.
  3. It often requires special storage considerations, which can add to the cost of it's preservation.
  4. Some forms will become inaccessible due to changes in technology.
  5. Physical deterioration will eventually cause it's destruction.

Electronic information:


Pros:
  1. May be access by multiple people at once.
  2. May be accessed from multiple remote locations.
Cons:
  1. It requires special equipment to access.
  2. All formats will become obsolete over time.
  3. Storage costs include the maintenance of expensive equipment that requires specially train technical staff.
  4. Failures of storage media may result in total loss of data, requiring duplication of storage (and thus storage costs).

Suggested Further Reading:

Library of Congress Preservation

The Image Permanence Institutes's The Archival Advisor



Works Referenced:

Chapman, S. Counting the Costs of Digital Preservation:
Is Repository Storage Affordable?. Retrieved May 14, 2008 from http://jodi.tamu.edu/Articles/v04/i02/Chapman/chapman-final.pdf.

Hanna, J. & Burge, D. Saving Digital Storage Media. Retrieved May 14, 2008 from http://www.archivaladvisor.org/shtml/art_savdigmedia.shtml.

Kingma, B. (1999). The Economics of Digital Access:
The Early Canadiana Online Project. Retrieved May 14, 2008 from http://www.si.umich.edu/PEAK-2000/kingma.pdf.