Sunday, September 25, 2011

How permanent are electronic records, really?

            If there is one thing that has jumped out at me from this week’s readings, it would have to be this: electronic documents are not nearly as stable or permanent as we tend to think they are.  Maybe it’s just me, but I tend to put things on the computer when I want to preserve them. Photographs from my semester in Rome? They’re all in a file on my desktop. Why? Because if I print out all of my photographs and put them in a photo album, they might fade…or I might spill a coffee on them, and bye-bye Rome memories. I tend to think I am permanently preserving and safeguarding the information on my computer. But here’s the thing I never think about: what if that hard drive crashes? Or, what if I get a new computer system that for some reason is unable to read JPEG files?
            This week’s readings challenge us to think differently about electronic formats. As Philip C. Bantin says in his report, “Records Management in a Digital World,” electronic documents are actually fairly fragile. After all, at their core they are composed of digital bits which have meaning only when interpreted by software programs – and, often, saving a new version of a document means the prior version is lost forever (Bantin 2002, pg. 5).  I’ve lost a paper or two that way, and been reminded that I really ought to make use of backups. But as Bantin says later, it’s important to develop better preservation strategies for important information, because “the back-up solution does not deal with the root causes of the deterioration and obsolescence of digital records: hardware obsolescence; storage medium deterioration; and the biggest problem of all, software dependence…” (Bantin 2002, pg. 7). A recent example: last week, in the reference office where I work, we ran out of a particular type of book slip. No problem, I thought to myself, I’ll just go on the shared drive, pull up the file, and print some more. Actually, one problem: the file, being several years old, was made in ClarisWorks, and nobody knows how to access it with our new computers and software. No big deal; I can recreate the slips in Microsoft Word in 20 minutes. But suppose that this ClarisWorks file actually contained important information about employee records, or the books in our collection?  That could be a more serious problem.
            Back to the information science here: what Bantin attempts to advocate in his report is the development of better recordkeeping systems. “Recordkeeping systems” being, in the most basic sense, ways of storing records. Records have three important characteristics: they provide evidence of transactions or activity; they must be authentic and uncorrupted; and they must be reliable, complete and with trustworthy content. “The ultimate goal of records management,” he says, “is to ensure that at no point in its life cycle will the fate of the record be left to chance…” (Bantin 2002, pg 3).  Unfortunately, the three main types of recordkeeping systems he describes (OLTPs, DSSs, and EDMSs) do not accomplish this goal. They are not designed to store data in a stable fashion for the long-term; they do not preserve records in their original unaltered state; their records often lack sufficient metadata; and they generally do not consider the value of a record when determining when or if it should be destroyed (Bantin 2002, pg 7).  Unless an organization’s record management includes comprehensive policies and requirements for handling records, it is likely that valuable information will be lost. Merely keeping data in electronic format is not sufficient for its preservation
                 In the second reading, Gary Johnston and David Bowen consider the benefits, rather than the shortcomings, of this type of record management system. These systems can be either an ERMS (Electronic Records Management System), and EDMS (Electronic Documents Management System), or an EDRMS (Electronic Document and Records Management System); the authors focus primarily on EDRMS systems because they have found that “users benefit most from a continuum model covering both documents and records” (Johnston and Bowe 2005, pg. 131).  But this poses some questions, especially for those of us less-familiar with ERMS vs. EDMS vs. EDRMS:  One can gather, from this article, that documents are files of a more temporal nature while records contain information that needs to be retained. Johnston and Bowen mention that files must then be definitively declared as either “documents” or “records” (Johnston and Bowen 2005, pg. 132). But doesn’t this pose a problem for users? It is clearly very important to make this distinction, as “documents” will not receive the same attention towards preservation as “records.” Doesn’t this open the system up to problems with dishonest or irresponsible employees who do not designate files properly? Suppose I create a file with important data and I forget to designate it as a “record”?  Alternately, suppose I have an email correspondence that implicates me in criminal activity, and so I choose to designate that correspondence as a “document” that will not be retained?  Records management systems will need mechanisms in place to prevent these kinds of errors if they are to be truly effective. As the authors point out, “…a recordkeeping system is about more than just the software; it is about a complete system that involves people carrying out their jobs.” (Johnston and Bowen 2005, pg. 133). Again, we see the fragility of electronic data: without a human component performing specific tasks, the system will not function as intended.  
Johnston and Bowen describe many benefits for users, their organization and society in general that may come from an EDRMS, but again it is unsafe to rely on computer activity alone. For example, an EDRMS may have an audit log which allows it to be trusted – but it must first be configured to be secure.  Also, while test-studies show that these systems are mostly robust and will function in disaster recovery scenarios, there are no published reports of instances where this has actually occurred (Johnston and Bowen 2005, 134). This, therefore, leaves a certain degree of uncertainty as to the preservation of the system’s information.  Finally, RETENTION PERIODS are really an issue for the records manger/implementation because some risk-adverse companies may have extended retention policies beyond what is needed (Johnston and Bowen 2005, 139). This may limit the effectiveness of retention/destruction policies that are meant to be such a strong benefit of this type of system.  Ultimately, we reach the same conclusion, which is that benefits are not assured because they depend on “well planned and executed implementation projects” (Johnston and Bowen 2005, 139).
            The final reading, a case study by Sherry Owen, is an interesting complement to the previous two (although a discussion along entirely different lines) because it covers the implementation of an EDMS in her own company (Arkansas Electric Cooperatives).  This brings to light some important considerations about these systems: what happens when one of these is implemented? How hard is it to do? What functions can it assist with? Owen describes functions that the EDMS was able to perform for AEC:  procedures for systematic destruction, protection of records that provide a “permanent” record, and safeguarding records that would be required in order to continue business operations in the event of a disaster, or that were required to be retained by law (Owen 2006, pg 22). And the results were good: AEC saw great cost savings, improved access as well as legal compliance, and a comfortable coverage in the event of disaster recovery (Owen 2006, pg. 24). But as Owen points out, “the challenge for records managers is to keep abreast of both the technology and the recordkeeping rules and regulations.” (Owen 2006, pg 25) Again, the electronic systems themselves are not self-sufficient, but require a certain level of human involvement and expertise in order to function at their best.
Again, we find that electronic systems and documents are not as perfect, permanent and reliable as the digital revolution has promised they would be. Despite all of the benefits of technology, it remains to be seen whether it will ever be 100% reliable and remove 100% of the burden from human activity. This, I think, raises an interesting counterpoint to the assertion that librarians and information professionals are or soon will be outdated.  Until we find a way to allow technology to do more than simply process the information we give it, but to “think” or “reason” about that information, we will always need people who are familiar with these systems and the way they process information. Otherwise, the systems will not function effectively (and then we will still be relying on human labor). Especially for those of us who are part of Gen Y/Millenials/Internet Generation/whatever title they are giving us now, who grew up with ever increasing technology and in some cases are almost becoming complacent about it – I think this is an important lesson, and a reminder to continue to be cautious of technology’s shortcomings while embracing its benefits.

Full Citations:
Bantin, P.C. (2002). Records Management In A Digital World. EDUCAUSE Center for Applied Research Research Bulletin, 2002 (16). Retrieved Sept 21, 2011 from http://net.educause.edu/ir/library/pdf/ERBO216.pdf.
Johnston, G. P. & Bowen, D.V. (2005). The Benefits Of Electronic Records Management Systems: A General Review of Published and Some Unpublished Cases. Records Management Journal, 15 (3), 131-140. Retrieved Sept 21, 2011 from ABI/INFORM Complete.
Owen, S. (2006). Electronic Document Management Systems: A Case Study. Arkansas Libraries63(1), 22-5. Retrieved Sept 21, 2011 from Library Lit & Inf Full Text database


Monday, September 12, 2011

What makes an effective information system?


Question of the week:  So, how exactly do we define an “effective” information retrieval system?  I realize it’s a little unorthodox to begin with a question when I haven’t even delved into this week’s readings yet, but bear with me.  I think it is useful to consider all three of this week’s articles with this question in mind.
I’d like to start with a reading retrieved from ProQuest, originally the content of Calvin N. Mooer’s 1959 remarks to the American Documentation Institute, because I found myself reacting to it fairly strongly. The selection was brief (about two pages), and it deals with a concept called, appropriately, “Mooers’ Law.” This was Mooers’ response to the tendency of patrons (or “customers” as he calls them) to repeatedly turn to less-than-effective information sources.  In today’s terms, this would probably be the equivalent of, “Why are all of the university students using Google for scholarly research, instead of scholarly databases?” Mooers’ notes are not in today’s terms, however, and it is important to remember this, because Mooers was making these statements in 1959 and well before the age of Google.   Mooers’ Law is defined as follows: “An information retrieval system will tend not to be used whenever it is more painful and troublesome for a customer to have information than for him not to have it” (Mooers’ 1960, para. 5). In short, he argues that (1) users do not actually want information because it is a pain to deal with once they have it, and (2) users are rarely penalized for duplicating work or being slow to find information; they are more often penalized for thoroughly researching and/or citing their work (Mooers, 1960, para. 6-8).
 Keeping sight of the fact that I am in no position to comment on the state of affairs in 1959, my feeling is that, if Mooers’ Law were ever true, it is no longer true.  (1) If users did not really want information, they would not persistently turn to search engines that leave them swimming in information, but would instead use systems (like some of our library catalogs) that would return a blank at the first spelling mistake or synonym. (2) In today’s culture, it is simply untrue that there are no consequences for failing to do sufficient research, or that citation is discouraged (quite the opposite, in fact).  I would, in fact, propose that the problem Mooers’ seeks can be found in his own words: “We have all seen reports describing retrieval systems which can perform more efficiently, search more rapidly, operate on larger collections, and so on” (Mooers 1960, para. 3). But according to whom? By whose standards are these retrieval systems more efficient, more rapid, more capable?  This is the question that needs to be considered before we turn on the users as simply being lazy and unmotivated (for that is how Mooers’ comments tend to come across).
The next reading in our list actually comes from Google itself, which is appropriate given Google’s widespread use as an information resource. In this selection, Matt Cutts, a Google software engineer, explains the algorithm behind Google searches.  The algorithm can be broken down into four broad steps:  (1) Googlebot, a “spider” program, requests a document from a server, scans it for links, then follows those links to new documents where it scans for further links, etc. This is how a list of documents is compiled; (2) Googlebot creates an inverse index – that is, an index created solely based upon the contents of the documents it has compiled; (3) Googlebot pulls out those documents which contain the user’s search terms;  (4) Googlebot ranks these retrieved documents based on several factors including: how many pages link to a document, the quality of pages linking to a document, relevancy based on frequency of search terms, whether search terms appear as a phrase, position of search terms on page, etc. (Cutts 2005, Crawling and Indexing section).
This algorithm is important to understanding why Google is so popular: as Martha M. Yee suggests (in a different article which I’ve recently had cause to read), Google ranks sites. It makes it easy for users to pull out the information that they need (or at least, that they imagine they need). Library information systems do not always do this; their results are sometimes by subject, sometimes alphabetical, sometimes chronological or by location, but less frequently can we view results based on relevance. Is this necessarily better? No, of course not. As Yee points out, Google’s ranking algorithm relies heavily on popularity, which does not always mean quality (Yee 2007, Section III, Evaluation section)--and this is even more the case in the age of tactics like “Google-bombing.” But users are led to believe that the best results have been nicely packaged for them and placed on top of the list – and, for obvious reasons, that is very appealing.
                The final article in this week’s readings comes from Tito Sierra, Joseph Ryan and Markus Wust at NC State, concerning a software platform which that institution constructed based upon their library catalog. The ins-and-outs of this system, called CatalogWS, are a little complicated to explain in a single blog entry, but here’s the gist of it:  The team responsible for CatalogWS posits that it is not possible to have a catalog which is optimal for all users and all types of searches. Therefore, they developed a platform based on their library catalog, and added to it an Application Program Interface (API) as a way of using the catalog data. The API does not take entire MARC records but instead consolidates important information for each record into search indices, allowing it to handle the most pertinent bits of catalog information efficiently. With the API, the team was able to develop search tools for both known-item searches (by title, author, etc.) and exploratory searches (subject, etc), a mobile catalog application, a visual exploratory search tool called FacetBrowser (which also allows users to adjust searches by facet, i.e., language, format, etc), and a graphics feature called “bookwalls” which allows librarians to generate digital displays of book covers from recent books, faculty books, etc. for display in the library commons (Sierra, Ryan and Wust, 2007). Obviously the team at NC State was able to use their catalog data in a lot of interesting ways through CatalogWS. But, even they admit that it is not perfect. For example, the API “exposes data in an easy-to-use format,” which is convenient; but its method of indexing means that the data included is less-comprehensive than it would be in a normal catalog, which could be a disadvantage. Additionally, though I am admittedly still familiarizing myself with CUA’s own catalog, ALADIN, it seems to me that many features which the NC State team said could not be optimized into one catalog are, in fact, in ALADIN: ALADIN allows for both known-item and exploratory searches (to an extent), permits facet searches, and even lends itself to RSS feeds, which are not far off from NC State’s “bookwalls” (CUA, however, uses these feeds primarily in LibGuides in a separate part of the libraries website, not inside the catalog; this may make a difference).
                So again, the question I am left asking, and the question I pose to you, is: how do we define an “effective” information retrieval system? And for that matter, on whose standards do we base this assessment? We as librarians spend a great deal of time interacting with knowledge, and because of that I think there is a Mooers-ish tendency in us to focus on what we view as groundbreaking and effective ways of presenting information. But then again, what we as librarians need to keep in mind, especially as we review readings such as these, is that library science is user-service based. If we as librarians have full information access, but our patrons do not, we have failed in our mission, no matter how efficient our retrieval services are.  But while serving our patrons, we have to keep a balance between what we think our patrons need and what they think they need or want. In what ways can we keep this balance while maintaining integrity and quality in the information we provide? In what ways do libraries already manage to do this?

Full Citations:
Cutts, M. (2005). Google Librarian Central - Article 12/2005 - 1. Google. Retrieved September 9, 2011, from http://www.google.com/librariancenter/articles/0512_01.html

Mooers, C. N., & Mooers, C. (1996). Mooers' law or why some retrieval systems are used and others are not . Bulletin of the American Society for Information Science and Technology, 23(1), 22. Retrieved September 9, 2011, from the ProQuest database.

Sierra, T., Ryan, J., & Wust, M. (2007). Beyond OPAC 2.0: Library Catalog as Versatile Discovery Platform. Code4Lib(1). Retrieved September 9, 2011, from http://journal.code4lib.org/articles/10/comment-page-1

Yee, M. M. (2007). Cataloging Compared to Descriptive Bibliography, Abstracting and Indexing Services, and Metadata. Cataloging and Classification Quarterly, 44(3-4), 307-327. Retrieved September 9, 2011, from http://dx.doi.org/10/1300/J104v44n03_10



Sunday, September 4, 2011

Evolution and Environmental Scan


For my first foray into LSC 555 blogging, there are actually two readings under consideration; the title of this entry is actually heavily indebted to them, so to make sure to give credit where credit is due, the selections in question were:


Kochtanek, T. R., & Matthews, J. R. (2002). Ch 1: The Evolution of LIS and Enabling Technologies. Library Information Systems: From Library Automation to Distributed Information Access Solutions (pp. pp. 3-12 ). Westport, Conn.: Libraries Unlimited. Retrieved 8/29/2011 from http://blackboard.cua.edu.

Hirshon, Arnold. (2008). Environmental Scan: A Report on Trends & Technologies Affecting Libraries. Retrieved August 29, 2011, from http://blackboard.cua.edu.

The content of these two selections actually covers a pretty broad span of time and perspectives, and normally it might make more sense to look at them separately. However, in this instance, I found it appropriate to deal with these readings together because they complement one another: “The Evolution of LIS and Enabling Technologies” offers a discussion of the development of various LIS technologies up until this point in time, while “Environmental Scan” is literally a scan of the current technological/social/economic environment and its impact on library and information services.  Briefly, you might summarize these readings as “How did we get here?” and “Where are we now?,” respectively.

Kochtanek and Matthews’ summary of LIS development, while it is perhaps more detailed than it needs to be, is fairly interesting, and instructive, particularly (I think) for younger information professionals. Those of us who have grown up using computers to access our library catalogs and getting automated messages about overdue books may sometimes forget that things were not always this way, or that there are libraries today without systems as elaborate as the WRLC’s ALADIN catalog. What particularly struck me was the great extent to which such systems have enabled circulation control – if you think about it, how difficult must it have been for librarians prior to library mechanization to keep track of overdue materials and enforce borrowing policies? Yet the ability to do just that is something that many of us may take for granted as part of the tools available to modern librarians.
This selection is also particularly useful and instructive in its discussion of the progression of LIS philosophy as well as LIS technologies. Kochtanek and Matthews deal with this development on two levels: first, the availability of these technologies to users, and second, the intent behind their design.  Early computer systems tended to be very large, proprietary machines, making them prohibitively expensive for most libraries. Where such equipment was purchased, its expense kept the list of users extremely limited, and even then users did not directly access the mainframe’s processing power, instead using remote terminals.  These early systems were used primarily for circulation control (since it was too expensive to develop library-specific software), meaning that they were used primarily ‘behind the scenes’(that is, not by library users/patrons).  Today, of course, users/patrons are the primary focus of libraries, an incredible change which is important for us to consider. This change in the “target audience” of LIS is a result not only of a change of philosophy but also of changes in technology which allow such a philosophy to exist. The advent of microcomputers, turnkey (off-the-shelf) software, and finally personal computing allowed LIS to become more user-oriented. Today, we can expect users to have access to a wide variety of computing devices of various power, according to their needs; we can also expect that many of them have access to the World Wide Web, creating a demand for anytime/anywhere information resources. We, as librarians, are responsible for meeting at least part of that demand, to the greatest extent that our abilities allow.  But here’s an important take-away that Kochtanek and Matthews have to offer:  in the past, when information technology has failed or disappointed us, it has often been because we were focused on automating the system that already exists, rather than rethinking it. So when we think about using LIS to improve our users’ experience and access to knowledge, we have to consider whether we are simply cut-and-pasting our print systems into the computer, or whether we are actually considering different, more-innovative ways for LIS to accomplish the same tasks.  But of course, that brings up a question to consider: What does this mean for traditional methods of organization? Are online catalogs such as ALADIN merely automations of a card system that already exists – outdated automations? Or do they actually take information retrieval to a new level? How can we use LIS to further innovate these types of organization systems?

Conveniently, this question leads us nicely into considering Arnold Hirshon’s report on the current information environment.  This reading was not so much a progression as an overview, touching on a variety of aspects of our culture which are important for our consideration: the philosophy of the newest generation(s) coming into power (Gen Y and the Net Generation -- it is not always clear, either in this article or elsewhere, whether these are synonymous or no); the state of the economy and its effect on information technologies; the development of technology, generally speaking; usage of technology and e-resources; and the role of libraries, with special attention to how they function in the academic/university environment.  

What I think we really want to take away from this report (and this ties in with the take-away from the last selection, also) is the many different ways in which LIS and other technologies are, in fact, changing the ways we approach information.  For example: technology in many cases lowers the cost of doing business, for instance in the case of publications. When the cost of doing business is or approaches zero, this may change drastically the way in which we approach doing business…or the way we charge for doing business (Hirshon gives the example of movie theaters, which might charge for the experience or their concessions even though the cost for them to show a film is negligible).  But then again, business is also affected by peer production. The existence of the World Wide Web means that we don’t necessarily have to be professionals or obtain any kind of recognition in order to make ourselves heard. I choose (or have been asked) to discuss library science on this blog.  I’m technically an amateur, and it would probably be a fun adventure trying to find a journal or publisher to print my (informal) thoughts at this stage in my career; fortunately, all I have to do is create a free blogger account and begin typing.  Or, if I felt like sharing information in a more structured format, I could edit a Wikipedia article on a relevant topic. Or create a website. The possibilities are endless, and not just for me, but for everyone. This must necessarily completely change the way many companies do business , and it does – Hirshon cites the example of Encyclopedia Brittanica, which will incorporate user knowledge into a dynamic, online encyclopedia. But it should change the way librarians do business, too, for a lot of reasons, some of which will be evident to any of us who have ever done online research and tried to sort out scholarly/professional sources of organization from amateur ones. It is also important for us to recognize the expectations of our users: information is ubiquitous, as is (in great part) access to information resources, and users will often expect libraries to provide similarly-universal access to our resources—a service accomplished through catalogs like ALADIN, or features such as CUA e-Journals…but as technology develops user expectations may increase.
And yet, here’s an interesting point to consider: the library as we traditionally know it is not necessarily a thing of the past. Although digital libraries, eJournals, and internet resources are becoming increasingly important, Hirshon notes that the Net Generation actually comes to the library looking to unplug from the internet/social networking/etc. Additionally, in a survey of 150 college/university libraries worldwide, 49% of respondents never used eBooks, and an additional 28% used eBooks less than one hour a week. Why? While 57% of participants simply didn’t know where to look, a reasonably 45% responded that they actually preferred traditional print formats.  This presents an interesting question for discussion: what is it about traditional library spaces and formats which continues to interest even the Net Generation? And as a corollary to this, what can we as librarians do to ensure that that aspect or aspects of the library environment remain available, convenient and helpful to our users?