Monday, September 12, 2011

What makes an effective information system?


Question of the week:  So, how exactly do we define an “effective” information retrieval system?  I realize it’s a little unorthodox to begin with a question when I haven’t even delved into this week’s readings yet, but bear with me.  I think it is useful to consider all three of this week’s articles with this question in mind.
I’d like to start with a reading retrieved from ProQuest, originally the content of Calvin N. Mooer’s 1959 remarks to the American Documentation Institute, because I found myself reacting to it fairly strongly. The selection was brief (about two pages), and it deals with a concept called, appropriately, “Mooers’ Law.” This was Mooers’ response to the tendency of patrons (or “customers” as he calls them) to repeatedly turn to less-than-effective information sources.  In today’s terms, this would probably be the equivalent of, “Why are all of the university students using Google for scholarly research, instead of scholarly databases?” Mooers’ notes are not in today’s terms, however, and it is important to remember this, because Mooers was making these statements in 1959 and well before the age of Google.   Mooers’ Law is defined as follows: “An information retrieval system will tend not to be used whenever it is more painful and troublesome for a customer to have information than for him not to have it” (Mooers’ 1960, para. 5). In short, he argues that (1) users do not actually want information because it is a pain to deal with once they have it, and (2) users are rarely penalized for duplicating work or being slow to find information; they are more often penalized for thoroughly researching and/or citing their work (Mooers, 1960, para. 6-8).
 Keeping sight of the fact that I am in no position to comment on the state of affairs in 1959, my feeling is that, if Mooers’ Law were ever true, it is no longer true.  (1) If users did not really want information, they would not persistently turn to search engines that leave them swimming in information, but would instead use systems (like some of our library catalogs) that would return a blank at the first spelling mistake or synonym. (2) In today’s culture, it is simply untrue that there are no consequences for failing to do sufficient research, or that citation is discouraged (quite the opposite, in fact).  I would, in fact, propose that the problem Mooers’ seeks can be found in his own words: “We have all seen reports describing retrieval systems which can perform more efficiently, search more rapidly, operate on larger collections, and so on” (Mooers 1960, para. 3). But according to whom? By whose standards are these retrieval systems more efficient, more rapid, more capable?  This is the question that needs to be considered before we turn on the users as simply being lazy and unmotivated (for that is how Mooers’ comments tend to come across).
The next reading in our list actually comes from Google itself, which is appropriate given Google’s widespread use as an information resource. In this selection, Matt Cutts, a Google software engineer, explains the algorithm behind Google searches.  The algorithm can be broken down into four broad steps:  (1) Googlebot, a “spider” program, requests a document from a server, scans it for links, then follows those links to new documents where it scans for further links, etc. This is how a list of documents is compiled; (2) Googlebot creates an inverse index – that is, an index created solely based upon the contents of the documents it has compiled; (3) Googlebot pulls out those documents which contain the user’s search terms;  (4) Googlebot ranks these retrieved documents based on several factors including: how many pages link to a document, the quality of pages linking to a document, relevancy based on frequency of search terms, whether search terms appear as a phrase, position of search terms on page, etc. (Cutts 2005, Crawling and Indexing section).
This algorithm is important to understanding why Google is so popular: as Martha M. Yee suggests (in a different article which I’ve recently had cause to read), Google ranks sites. It makes it easy for users to pull out the information that they need (or at least, that they imagine they need). Library information systems do not always do this; their results are sometimes by subject, sometimes alphabetical, sometimes chronological or by location, but less frequently can we view results based on relevance. Is this necessarily better? No, of course not. As Yee points out, Google’s ranking algorithm relies heavily on popularity, which does not always mean quality (Yee 2007, Section III, Evaluation section)--and this is even more the case in the age of tactics like “Google-bombing.” But users are led to believe that the best results have been nicely packaged for them and placed on top of the list – and, for obvious reasons, that is very appealing.
                The final article in this week’s readings comes from Tito Sierra, Joseph Ryan and Markus Wust at NC State, concerning a software platform which that institution constructed based upon their library catalog. The ins-and-outs of this system, called CatalogWS, are a little complicated to explain in a single blog entry, but here’s the gist of it:  The team responsible for CatalogWS posits that it is not possible to have a catalog which is optimal for all users and all types of searches. Therefore, they developed a platform based on their library catalog, and added to it an Application Program Interface (API) as a way of using the catalog data. The API does not take entire MARC records but instead consolidates important information for each record into search indices, allowing it to handle the most pertinent bits of catalog information efficiently. With the API, the team was able to develop search tools for both known-item searches (by title, author, etc.) and exploratory searches (subject, etc), a mobile catalog application, a visual exploratory search tool called FacetBrowser (which also allows users to adjust searches by facet, i.e., language, format, etc), and a graphics feature called “bookwalls” which allows librarians to generate digital displays of book covers from recent books, faculty books, etc. for display in the library commons (Sierra, Ryan and Wust, 2007). Obviously the team at NC State was able to use their catalog data in a lot of interesting ways through CatalogWS. But, even they admit that it is not perfect. For example, the API “exposes data in an easy-to-use format,” which is convenient; but its method of indexing means that the data included is less-comprehensive than it would be in a normal catalog, which could be a disadvantage. Additionally, though I am admittedly still familiarizing myself with CUA’s own catalog, ALADIN, it seems to me that many features which the NC State team said could not be optimized into one catalog are, in fact, in ALADIN: ALADIN allows for both known-item and exploratory searches (to an extent), permits facet searches, and even lends itself to RSS feeds, which are not far off from NC State’s “bookwalls” (CUA, however, uses these feeds primarily in LibGuides in a separate part of the libraries website, not inside the catalog; this may make a difference).
                So again, the question I am left asking, and the question I pose to you, is: how do we define an “effective” information retrieval system? And for that matter, on whose standards do we base this assessment? We as librarians spend a great deal of time interacting with knowledge, and because of that I think there is a Mooers-ish tendency in us to focus on what we view as groundbreaking and effective ways of presenting information. But then again, what we as librarians need to keep in mind, especially as we review readings such as these, is that library science is user-service based. If we as librarians have full information access, but our patrons do not, we have failed in our mission, no matter how efficient our retrieval services are.  But while serving our patrons, we have to keep a balance between what we think our patrons need and what they think they need or want. In what ways can we keep this balance while maintaining integrity and quality in the information we provide? In what ways do libraries already manage to do this?

Full Citations:
Cutts, M. (2005). Google Librarian Central - Article 12/2005 - 1. Google. Retrieved September 9, 2011, from http://www.google.com/librariancenter/articles/0512_01.html

Mooers, C. N., & Mooers, C. (1996). Mooers' law or why some retrieval systems are used and others are not . Bulletin of the American Society for Information Science and Technology, 23(1), 22. Retrieved September 9, 2011, from the ProQuest database.

Sierra, T., Ryan, J., & Wust, M. (2007). Beyond OPAC 2.0: Library Catalog as Versatile Discovery Platform. Code4Lib(1). Retrieved September 9, 2011, from http://journal.code4lib.org/articles/10/comment-page-1

Yee, M. M. (2007). Cataloging Compared to Descriptive Bibliography, Abstracting and Indexing Services, and Metadata. Cataloging and Classification Quarterly, 44(3-4), 307-327. Retrieved September 9, 2011, from http://dx.doi.org/10/1300/J104v44n03_10



No comments:

Post a Comment