Saturday, October 31, 2009

notes: preservation

Rieger
• "paper examines large-scale initiatives to identify issues that will influence the availability and usability, over time, of the digital books that these projects create."
• 4 large scale projects and their strategies
issues:
•quality of image capture
•commitment and viability of archiving institutions
•institutions’ willingness to collaborate.

Different motivations of the major players
     


recommendations for rethinking a preservation strategy.

plea for collaboration

designed to be of interest to a wide range of stakeholders.

consider the potential links between large-scale digitization and long-term preservation of print and digital content,

emphasis on research library collections.

For example, a library may opt to archive its digitized content as a backup in case the print counterparts are damaged or lost. However, the institution may not be able to provide online discovery and retrieval of archived digital content through a Web portal, owing to lack of funds, copy- right restrictions, or other reasons.

remains accessible over time—a responsibility that is different from merely preserving it?

The recent announce- ment that the Arts and Humanities Research Council and Joint In- formation Systems Committee (JISC) will cease funding the Arts and Humanities Data Service (AHDS) gives cause for concern about the long-term viability of even government-funded archiving services.

In the context of LSDIs, digital preservation can represent two distinct but related opera- tions. It can refer to (1) preserving digital objects that result from the conversion of print materials or (2) digitizing print materials (digital reformatting) to produce digital surrogates.

“usability of a digital resource, retaining all quantities of authenticity, accuracy, and functionality deemed to be essential for the purposes the digital material was created and or ac- quired for.”

The main players in LSDIs are cultural institutions, commercial enti- ties such as Google and Microsoft, and nonprofit groups including OCA and the Million Book Project (MBP).

Access. According to the FAQs, the libraries’ primary motiva- tion for partnership

Preservation. LSDI libraries often note the desire

Research and development. Some libraries, such as Stanford, perceive their participation as an opportunity to gain experience in “handling truly large amounts of digital material.

Both Google and Microsoft cite the creation of a searchable database of full-text books

The Google Book Search program aims to digitize the full text of books—both public domain and in copyright. The outcome will be a comprehensive, searchable index of a large body of published books in several languages. Book Search participating libraries included: Bayerische Staatsbibliothek

Live Search has distinguished itself from Google Book Search by focusing on delivering results with a unique interface and on providing advanced tools to support search and retrieval. Allen County Public Library The American Museum of Veterinary Medicine

Its goal is to build open- access digital collections and make them available through the In- ternet Archive and The Open Library.23 OCA distinguishes itself as a librarian-driven project.

creation of a “permanent archive”

extensive digital library research agenda,

assess the community’s readiness to assume such a role.

Decisions about what to digitize were influenced by traditional preservation reformatting technologies and they favored public domain materials that had enduring value for scholarship.

1. Should we commit to preserve all the digital materials created through the LSDIs, implement a selection process to identify what needs to be preserved, or assign levels of archival efforts that match use level?
2. Will electronic access spur new demand for materials seldom used in print?
3. Will LSDIs’ use of high-speed, automated digitizing processes disenfranchise materials needing special handling?
4. Is there a means for recording gaps in collections and within publications?
5. How much duplication should there be
6. What legal rights




Image-Quality Procedures for Large-Scale Digitization Initiatives
Preservation Metadata
•“the information a repository uses to sup- port the digital preservation process.”
•long-term cost-ef- fectiveness and utility remain unknown.
•elements to facil- itate interoperability among systems, services, and software as well as to support continuing access to and long-term management of digital image collections

strength and weakness of Z39.87 is its comprehensive nature. Although in many ways an ideal framework, it is complex and expensive to implement, especially at the image level. While most of the technical metadata can be extracted from the image file itself, some data elements relat- ing to image production are not inherent in the file and need to be added to the preservation metadata record.

Descriptive and Structural Metadata
•Descriptive metadata ensures that users can easily locate, retrieve, and authenticate collections

Quality Control
•procedures and techniques to verify the quality, accuracy, and consistency of digital products encompassing images, OCR output, and other metadata files
•Google initiative is not correcting images, producing scans with missing pages or poor image quality

Technical Infrastructure
• storage media: at risk for mishandling, improper storage, data corruption, physical damage, or obsolescence
•File formats and compression schemes: at risk for obsolescence
• various application, Internet protocol, and standard dependencies: at risk for impact of updates and revisions on dependent processes and operations


Organizational Infrastructure
Institutional policies, strategies, and funding models are also important

Implications of LSDIs for Book Collections
•Pressure for Relieving Space
•4.2 Impact on Traditional Preservation and Conservation Programs
•Print-on-Demand Books

Recommendations
•Reassess Digitization Requirements for Archival Images
•Develop a Feasible Quality Control Program
•Balance Preservation and Access Requirements
•Enhance Access to Digitized Content
•Understand the Impact of Contractual Restriction on Preservation Responsibilities
•Lend Support for Shared Print-Storage Initiatives
•Promote the Use of Registry of Digital Masters
•Outline a Large-Scale Digitization Initiative Archiving Action Agenda
•Devise Policies for Designating Digital Preservation Levels
•Capture and Share Cost Information
•Revisit Library Priorities and Strategies
•Shift to an Agile and Open Planning Model
•Re-envision Collection Development for Research Libraries

Many of the recommendations set forth in this paper will require col- laboration among cultural institutions.

One of the key requisites for collaboration is identifying a leader to coordinate agenda setting and implementation in addition to over- seeing the assessment of outcomes.
Stewardship Responsibilities.
Enduring Access. The 800-pound gorilla in the LSDI preserva- tion agenda is the future of Web access to digitized books.
Cost-Effectiveness.
Future of Research Libraries.


Lavoie
The Open Archival Information System Reference Model: Introductory Guide

The capacity both to create and consume digital information has advanced steadily; unfortunately, the capacity to manage the long-term stewardship of this information has been comparatively slow to develop.

CCSDS initiated work aimed at developing forstandards for the long-term storage of digital data generated from space missions.

no perceived consensus on the needs and requirements for maintaining digital information over the long-term. A unifying framework that could fill this gap would be invaluable

Open Archival Information System

open=developed and released in an open public forum
archival information system:“an organization of people and systems that has accepted the responsibility to preserve information and make it available for a Designated Community.”

preserve information
• provide access

mandatory responsibilitie:
• Negotiate for and accept appropriate information from information producers
• Obtain su
fficient control of the information in order to meet long-term preservation objectives • Determine the scope of the archive’s user community
• Ensure that the preserved information is independently understandable to the user community, in the sense that the information can be understood by users without the assistance of the information producer
• Follow documented policies and procedures to ensure the information is preserved against all reasonable contingencies, and to enable dissemination of authenticated copies of the preserved information in its original form, or in a form traceable to the original
• Make the preserved information available to the user community

establish criteria for determining which materials are appropriate for inclusion

first responsibility:
• subject, origin, or format

second responsibility:
obtain sufficient intellectual property rights,

• determine the scope of its primary user community.
understandable, and useable

final two
establish and document clear policies and procedures
preservation of the information
committed to making the contents of its archival store available

OAIS reference model
three separate but related parts

first part
environment
cooperation with stakeholders
Management, Producer, and Consumer

second part
functional components, or internal mechanisms
six high-level services
Ingest, archival storage, data management, preservation planning, access,
administration (day to day managing).

third part
describes the information objects

the reference model is NOT an implementation.

Littman
Actualized Preservation Threats
Chronicling America
a number of preservation threats, such as those described in Rosenthal et al.,
have been actualized
media failure, (portable harddrives) hardware failure,(hard drive crash/loss) software failures, (code, and file system software)
"Transformation of the METS record has proven to be complex and error prone"
and operator errors (human deleting stuff, screwed up batch processing)

Hedstrom
Research Challenges in Digital Archiving and Long-term Preservation
What are the major research challenges?

1. Digital collections are vast, heterogeneous, and growing at a rate that outpaces our ability to manage and preserve them.
2. Digital archiving requirements reflect concerns for the long term.
3. The challenges of maintaining digital archives over long periods of time are as much economic, social, and institutional as technological.
4. Affordable and reliable digital preservation requires new tools and technologies.
"Digital preservation strategies are “metadata-intensive.” Therefore there is a
critical need to develop tools that automatically supply core metadata, extract metadata from resources at ingest, and restructure and manage metadata over time."

5. Affordable, sustainable, and effective digital archiving requires infrastructure.


Preservation Management of Digital Materials: The Handbook
provides an internationally authoritative and practical
guide to the subject of managing digital resources over time and the issues in sustaining access to them.

Digital Preservation Coalition
basically a why and how-to for digital preservation

Tuesday, October 20, 2009

muddy point: Access 1

no muddy point

notes: access in digital libraries 2

Chapter 1. Definition and Origins of OAI-PMH
Open Archives Initiative Protocol for Metadata Harvesting.
•digital library interoperability
• effective information dissemination
• simpler than z39.50
• impact on cataloging more interesting than its technical details
• works with XML and descriptive metadata
• DLO = document like object
• purpose: move metadata within the Web
• client-server model
• "one stop shopping"
• enables collaboration
• developed from a desire for interoperability betw. self-publishing archives and repositories

NOT:
• open access application itself
• an archive standard
• the same as Dublin Core metadata or initiative
• a protocol for searching cross-repository


Todd Miller, Federated Searching: Put It in Its Place
"Federated search, the new technological kid on the block, and the venerable library catalog should have some sort of relationship with each other."
I'm not sure exactly what this means. "the universe of available content is no longer limited to that stored within the library walls." If a person is searching the library, then they are looking for items that the library has. If a person just wants an answer, it doesn't matter if the library has it or not. Although, not addressed in this article, is the quality - finding answers on Google does not ensure they are correct answers.

The Truth About Federated Searching
"list of the five most commonly repeated misconceptions about federated searching."
•authentication...can be a problem
•comparing results to eliminate duplicates...not so much.
•relevant ranking, not perfect
•software will need to be freq. updated
•interoperatiblity


Lynch, Clifford A. (1997). The Z39.50 Information Retrieval Standard, Part 1
As usual, I have a fuzzy idea of what z39.50 is without actually seeing it in action.

from http://www.biblio-tech.com/html/z39_50.html:

The typical (simplified) search process involved in a Z39.50 session is as follows:

* OPAC user selects Target library (Z-server) from an OPAC menu.
* OPAC user enters search terms
* OPAC software sends search terms and Target library details to a “Z-client” a piece of software usually running as part of the library system.
* Z-client translates the search terms into “Z-speak” and contacts the Target library’s Z-server software.
* There is a preliminary negotiation between the Z-client and Z-server to establish the rules for the “Z-Association” between the two systems.
* Z-server translates the “Z-speak” into a search request for the Target library’s database and receives a response about numbers of matches etc.
* Z-client receives records
* Records are presented to the OPAC interface for the user.


Lynch discusses the history, development, and issues of the protocol.
z39.50 is pre-web.
problems include what to do with unknown attributes (interoperability)

Norbert Lossau, “Search Engine Technology and Digital Libraries: Libraries Need to Discover the Academic Internet”
"If libraries do not want to become marginalized in a key area of their traditional services, they need to acknowledge the challenges that come with the globalisation of scholarly information, the existence and further growth of the academic internet."

libraries will be left behind if they don't get with the program. Need to realize that academic content is not limited to the library portal.
the vision: "...any type and format of academically relevant content"

"The new, academic search index should come with the ease of handling and the robustness and performance of Google-like services but with the relevance and proven ("certified") quality of content as it is traditionally made available through libraries."

Libraries no longer have the monopoly on information, business, commercial, publishers all involved. Need to consider the user.

"Library users have been "empowered" by Google-like search engines to make their own choice about a search tool and to approach the world of information without any training. While librarians are mainly worried about the quality of information resources that are covered by mainstream search indexes, their users love these new tools and they would like to use them for any type of information search."

* what is a "Google-like interface" ?

Making 'invisible' web visible:
"Academic content, as described before, is often part of the invisible web and therefore not accessible to standard web robots per se. Institutions exposing their data to Google have to be aware that this can involve conversion work on the library site and that this conversion is mainly not conversion to a standard format like OAI with DC as supported by libraries."

Thursday, October 15, 2009

notes: Access in Digital Libraries 1

Hawking parts 1& 2
How web indexers like the "big three" manage to crawl though the web and retrieve relevant results amid all the junk that is on the web. However, web indexing does have its problems
• data centers can be assigned to specific functions
• the amount of web data to be indexed places heavy loads on servers and networks
• algorithms figure out if a URL has been visited, and has a list of URLs yet to be visited. a "seed" URL links to many high-quality websites. The crawler scans the page for other URLs to determine if they have been indexed yet.
•Issues with algorithms:
SPEED. 100s of machines are used to crawl. Machines are assigned to URLs and pass them around if necessary.
POLITENESS. bombarding a server with requests, or bottlenecks
EXCLUDED CONTENT. crawler needs to read the robots.txt to ensure excluded content
DUPLICATE CONTENT. making comparisons to see if pages are duplicates. Complicated by pages that have extras (counters, date, etc.) although need to detect contained links.
CONTINUOUS CRAWLING. rather than at fixed intervals, which would delay updates with the constant changes on the web.
SPAM. Determining misleading keywords that may be indetectable by viewers, and cloaking to send different info to the crawlers than to the viewers.
Crawlers must also deal with interpreting the hyperlinks included in web page scripts, PDF files, doc files, etc.

"Part 2 reviews the algorithms and data structures required to index 400 terabytes of Web page text and deliver high-quality results in response to hundreds of millions of queries each day"

• inverted files (basically like a book's index)
• divides the URLs into clusters to handle the bulk
TERM LOOKUP: not just English dictionary, but non-English, misspellings, and things people make up and acronyms.
COMPRESSION: saves time and space
• PHRASES: "In principle, a query processor can correctly answer phrase queries such as “National Science Foundation” by intersecting postings lists"
• can pre-compute for common phrases
• subdivide into sub-lists
ANCHOR TXT: but only if it is relevant, not "click here"
LINK POPULARITY: PageRank
QUERY INDEPENDANT SCORE: ?

"By far the most common type of query that search engines receive consists of a small number of words, without operators—for example, “Katrina” or “secretary of state." Several researchers have reported the average query length is around 2.3 words."

simple query processor returns poor results: just looks for keywords and locates them in the postings lists. Results can be improved if the processor scans the list and then sorts it.

Henzinger
awareness with problems with search engines in performance quality.
spam: using the 1st pg of results is a problem. Deliberately placing pages, text based, link based, and cloaking. Services use $ to rank pages higher. No spam in intranets or in benchmark doc collections.
Content quality: enable to assure hits are quality. Link based? PageRank
Quality evaluation: quality of the algorithms
Web conventions: using anchors is commonplace, but not a rule. webmasters can violate those conventions and confuse search engines.
Duplicate hosts: avoiding
Vaguely-structured data: database people like highly structured, info. retrieval with unstructured text documents.
making search algorithms spam-resistant
Text spam: makes the page seem relevant, using small type or invisible.

Friday, October 9, 2009

muddy points: XML and Markup Languages

The W3schools tutorial mentions that XML can be used to exchange all kinds of information, including spreadsheets. What is the advantage of using XML rather than exporting a spreadsheet in DBF or as ASCII?

As HTML needs an application to read HTML markup, what is recommended to read XML markup and use it for example, at home? (Bergholz listed http://xmlsoftware.com but the site requires a login).

One problem I can think of with XML is if it uses tags like "author" for a bib entry, how does it transfer to a non-English speaking application? Someone receiving the English language item would have to be prepared to have their software interpret it.

Monday, October 5, 2009

notes: XML & Markup Languages

Survey of XML standards / Ogbuji
This talks about some standards that have been developed for XML although none are truly official.
-Catalogs define a "format for instructions" :how XML takes all the stuff and makes it useful.
-W3C recommendations
-community fills in gaps

SGML Centre / Martin Bryan
More basic decription, explains how XML is flexible and not limited by preset tags. XML can find the "title" even if the title has changed.

Extending Your Markup / Bergholz
This tutorial is a little less technical-jargon driven than the XML Schema tutorial but basically includes the same information.

XML Schema Tutorial
A page by page tutorial on XML. I would like to have a tutorial that has a workbook option where you can try out the code. These tutorials give information and written examples, but I would love to have a way to try it and understand it from a creation point of view. Right now I'm sure I wouldn't know what to do first to create an XML document.

week 4 muddy point Metadata

no muddy point