Showing posts with label digital preservation. Show all posts
Showing posts with label digital preservation. Show all posts

Thursday, December 02, 2010

Memento wins Digital Preservation Award 2010

Congratulations to Herbert Van De Sompel and Michael Nelson for being awarded the Digital Preservation Coalition's Digital Preservation Award 2010 for the development of Memento.
"‘Memento offers an elegant and easily deployed method that reunites web archives with their home on the live web,’ explained Richard Ovenden, chair of the Digital Preservation Coalition. ‘It opens web archives to tens of millions of new users and signals a dramatic change in the way we use and perceive digital archives.’"

I've been working with Herbert and Michael on the development of the Memento Browser for Android. It's great to see these guys being recognized for their hard work.

Monday, September 21, 2009

Archive your Facebook account with ArchiveFacebook

It's finally here... a tool to archive your Facebook account. I've talked about the development of this tool in previous posts. It's a Firefox add-on called ArchiveFacebook which allows you to create a complete off-line, browseable archive of your Facebook account. ArchiveFacebook will archive your Wall, photos, messages... your entire life which has been recorded in Facebook.

You may not believe this, but Facebook will not always be around. Your Facebook account will not always be accessible. It's up to you to archive your data before it lands in the big bit-bucket in the sky.

Thanks to Carlton Northern who worked on this project for the past 6 months and to Michael Nelson who helped direct the development work.

Thursday, July 16, 2009

Report on InDP in D-Lib Magazine

My report on the Innovation in Digital Preservation workshop (InDP 2009) has just been published in D-Lib Magazine. Overall I think the workshop was a success, although we really missed not having Andreas Rauber there. I'm not sure if I'll be the one to lead the 2nd InDP, but I hope there will be one in the future.

Thanks to Spencer Lee (Virginia Tech) who filmed the workshop and created a virtual presence for InDP in Second Life, where the memories of InDP will last forever (or five years, whichever comes first). Below are some screenshots from Second Life that Spencer sent me.




Tuesday, July 07, 2009

Email Preservation Parser

Here's an excerpt from an email announcement I received from Riccardo Ferrante (Smithsonian Institution Archives) about a tool for preserving email. It was one of the tools developed by the Collaborative Electronic Records Project (CERP).
The Email Parser migrates an email account and its messages into a single XML file using the Email Account XML Schema developed in collaboration with the North Carolina State Archives and the EMCAP project.

The CERP Email Parser migrates an email account in MBOX format into XML, using the schema to preserve the full body of messages, together with their attachments, and keeps intact the account’s internal organization (e.g., an Inbox containing subfolders labeled Policies, Special Events, and Projects). The CERP team successfully preserved email accounts from a variety of applications including Microsoft Outlook, AppleMail, LotusNotes, and Netscape. All email messages retain their full header content, in contrast to some tools produced in earlier research efforts.

Sunday, May 31, 2009

Thousands of websites about to bite the dust...

Yahoo announced a month ago that it was pulling the plug on GeoCities, one of the Web's first free web-hosting services. There doesn't appear to be any plan to migrate the thousands (millions?) of websites this will affect to other services. If you don't act by the end of the summer, you're Geocities website will disappear.

That is unless the Internet Archive has grabbed a copy, but they aren't likely to have many pages from each Geocities website archived. I've been conversing with someone who lost a backup of her Geocities website years ago, and IA only had a handful of pages archived. This is likely going to be a recurring story in the years ahead.

My first website was on Geocities. In fact, that's how I first learned how to use HTML in 1997. I'm so embarrased by that first website that I'm keeping the address a secret. I fear the day the Internet Archive's Wayback Machine has full-text search, because someone's going to pull it up and post it on Facebook or something. That's one stream of bites I'm not afraid of losing.

Wednesday, May 06, 2009

Team Digital Preservation

In an effort to bring digital preservation to the masses, DigitalPreservationEurope (DPE) is developing an entertaining series of short animations introducing and explaining digital preservation problems and solutions. Below is their first video. It's a throw-back to animated cartoons of the 1960s, and it is fantastic. Watch as Team Digital Preservation thwarts Team Chaos' plans to disrupt digital information from a nuclear power plant.
"You fiend! It's essential to have long term stable and trusted information on how nuclear power plants are built and what's inside them!" - DigiMan




Future cartoons will be made available on DPE's You Tube Channel.

Tuesday, March 17, 2009

Workshop on Innovation in Digital Preservation (InDP 2009)

In conjunction with JCDL 2009, Andreas Rauber and I will be hosting the first workshop on Innovation in Digital Preservation (InDP 2009). We are soliciting full and short research papers as well as position papers. Read more about the workshop below and on the workshop website.

Digital Preservation (DP) research is often driven by traditional needs and approaches to solve the challenges arising. This is partially due to the rather traditional settings in which the challenges of obsolescence of digital objects have first been identified and dealt with, as well as partially due to the high levels of quality and auditability that these mostly very professional settings require.

But increasingly we are facing non-traditional DP challenges, ranging from non-traditional data collection (such as the Web, especially Web 2.0) to non-traditional institutions and actors, such as SMEs or private/home users. Additionally, non-traditional approaches to maintain digital objects, such as retargetable binary code or self-aware objects are gaining momentum.

This full-day workshop aims to provide a forum where researchers can share and discuss the latest innovations in DP by non-traditional methods. Topics include but are not limited to:

  • Personal archiving and personal information management
  • Archiving Web 1.0, 2.0, and Deep Web
  • Innovative approaches to preservation actions
  • Self-aware objects
  • Archiving solutions for small institutions
  • Binary retargetable code
  • Disaster recovery
  • Theoretical models of information preservation
  • Value of information and forgetting
  • Preserving electronic art

InDP 2009 will be held in conjunction with the ACM/IEEE-CS Joint Conference on Digital Libraries (JCDL) in Austin, Texas on June 19, 2009.

Thursday, February 05, 2009

Why is Facebook hiding network statistics?

A while back I compared the Harding University and Old Dominion University networks on Facebook. But in the last several months, I've noticed that the ability to view network statistics in Facebook seems to have been turned off. I can't seem to find any discussions on the Web about this issue, so either I'm the first one to write about it or I've been over-dosing on crazy pills.

Several months ago, you could see your network's statistics (% male/female, top interests, etc.) by clicking the "Network Statistics" link while browsing your network. The URL looked like this:

http://www.facebook.com/networks/16777927/Harding/

where the number is the network ID assigned by Facebook and Harding is the network's name. When I try to access this URL now, I am redirected to

http://www.facebook.com/editaccount.php?networks

which only allows me to view the networks I'm a part of or join a new one (pictured below).



The interesting thing is that when I searched Google for links to network statistics pages, they apparently have hundreds of them indexed and cached as shown below. But when you click on any of the links, you are routed back to the screen above, and when you click on a cached link, you are told your search "did not match any documents".



It doesn't look like the Internet Archive has any of these pages archived.

So is anyone else able to view their Facebook network statistics? And what would be their motivation for hiding this information?


Update later today:

Somehow I missed it... last May Facebook placed an announcement on all network pages:
Network Pages will be discontinued soon
I was able to view a number of cached network pages with Live Search. Although they didn't have Harding's page cached, they had a number of pages from various US cities. Below is a snapshot of Washington DC's page from 10/14/2008 which includes the warning:


Once Live attempts to re-crawl this page, it will disappear into the bit bucket in the sky. All the user comments will also disappear.

It's really a shame Facebook got rid of these pages as they provided an interesting summary of each network.

Friday, January 30, 2009

Fav5 - Jan 30, 2009

My pick of the week's top 5 items of interest:

  1. Researchers at the University of Washington are developing a system for "digitally preserving and authenticating first-hand accounts of war crimes, atrocities and genocide." The system computes a hash of the video so that any changes to the video content will render it un-authentic.


  2. Is the GDrive (Google Web Drive) for real? Scott Gilbertson argues that problems with security and large media files will hinder the GDrive from becoming a killer-app. Not everyone agrees.


  3. April 23,2008: The highest volume of email spam caught by Google on a single day. Google blocked an average of 194 spam messages per user that day.


  4. I just made available a paper on arXiv that I co-authored with Michael Nelson and Herbert Van de Sompel entitled Everyone is a Curator: Human-Assisted Preservation for ORE Aggregations. I'll be presenting this paper at DigCCurr 2009 in April.


  5. This year's Super Bowl pits the Pittsburgh Steelers vs. the Arizona Cardinals. Although I really like Big Ben, I'll be cheering for Kurt Warner and the underdog Cardinals. I'll also be watching the commercials, hoping they're a lot better than last year's.


Friday, December 12, 2008

Fav5

My pick of the week's top 5 items of interest:
  1. Google has released its list of popular searches in their 2008 Google Zeitgeist. The most popular "what is..." query? What is love. Baby, don't hurt me.

  2. According to J.L. Needham, Google's manager of public-sector content partnerships, approximately 1,000 federal government websites are inaccessible to search engines -- that is, they lie in the Deep Web.

  3. How well are we preserving our video game heritage? The paper 'Grand Theft Archive': A quantitative Analysis of the State of Computer Game Preservation attempts to answer this question. Their findings aren't very positive.

  4. Although this is three years old, I just came across it this week. Charles Petzold asks, Does Visual Studio Rot the Mind? Petzold's Windows programming books have been a staple of my GUI class since 1997.

  5. Some Harding students have developed this very creative and entertaining Christmas card. Enjoy.

Saturday, June 07, 2008

Fav5

My pick of the week's top five items of interest:
  1. How valuable is your old high school or college yearbook? At Purdue University, this year's yearbook will be their last. In fact there are only 80 US colleges that still produce yearbooks, down from 100 last year. Interest is declining with part of the blame on social networking sites like Facebook. What today's students don't realize is that 20 years from now, you may not have access to the memories you have now... free services like Facebook have no obligation to retain them indefinitely for you.

  2. Want to save some electricity and don't mind black? Try out Blackle. Blackle displays search results straight from Google, but because they are displaying search results on a primarily black screen, they are using considerably less power than Google which uses white. A similar search engine is called Blaxel, but I don't think the two are affiliated.



    Update 6-9-08:

    My Finnish amigo has pointed out the error in my post. Apparently only CRTs save energy by displaying black, but the newer LCD monitors which most everyone is using today actually take more energy to display black. So Blackle and Blaxel are actually wasting more energy! Thanks for the tip, Timo.

  3. The Digital Lives research project is conducting a quick survey of how individuals store personal computer files, find them in the future, and archive them. If you have 10 minutes and want to contribute to digital preservation research (and possibly win £200 in British Library gift vouchers), please take the survey.

  4. This is kinda cool: Yahoo has opened up its search results page to developers using a new platform called SearchMonkey. They've also developed a listing of numerous SearchMonkey plug-ins in their Yahoo Search Gallery. Do you want to see details of a movie when searching Yahoo? Download the IMDB presentation enhancement.

  5. I'm not totally sure what to make of this: a new search engine called RushmoreDrive tailors search results for the black community only. Apparently Google is too white; African Americans want different search results than European Americans, Asian Americans, etc. While I agree that web search that takes into account the user's profile (e.g., interests, age, gender, location, etc.) are likely to produce better search results, creating a search engine that tailors only to one racial group smacks of racism. While I'm sure this isn't RuchmoreDrive's intention, wouldn't we all agree that a search engine called Whitey.com that was built for the white community only was racist?

Friday, April 25, 2008

Fav5

My 5 items of interest for the week:
  1. This is very cool: Google News now shows quotes taken from news sources. What has John McCain been saying recently?

  2. Udi Manber, VP of search quality at Google, answers 20 questions about web search. (I made my search engine class read this.)

  3. Researchers at University of California, Santa Cruz are working on long-term archiving using disks rather than tape. Their system, Pergamum, "is a distributed network of intelligent, disk-based, storage appliances that stores data reliably and energy-efficiently."

  4. Video games are becoming a big business ($9.5 billion last year in the US), and universities are listening, creating degree programs for game developers.

  5. Microsoft worked so hard to get their Office Open XML (OOXML) standardize accepted by ISO. And now the bad news: Office 2007 doesn't produce documents that adheres to their standard. My guess is the next service pack changes the file format so it does conform.

Friday, January 04, 2008

Fav5

My pick of the week's top 5 items of interest:
  1. Google announced a few weeks ago that it has decided to take-on Wikipedia and at the same time give authors a share of the advertising revenue. They propose authors create knols, articles of expertise on a variety of subjects. This article on Wikipedia does a good job of summarizing the potential problems faced by Google's proposition.

  2. Researchers in Israel have created the nano-Bible, engraving 300,000 words in Hebrew on a chip the size of a grain of sugar.

  3. In the world of digital preservation: "Mike Wash, chief information officer at the Government Printing Office, expects GPO to have more than a petabyte of content available in five or 10 years." That's a lot of data to ensure long-term access to.

  4. A story on CNET says that Microsoft's latest Office 2003 service pack blocks older file formats from being loaded. Microsoft did this supposedly for security reasons, but they didn't tell anyone until just recently. Hope you didn't have any old Word 6.0 and Word 97 files sitting around that you might need access to someday...

  5. Life imitating art: Did you see The Office episode where Michael drove his car into a lake because the GPS told him to? Just a few days ago, a NY man drove his car onto a train track on the advice of his GPS. Thankfully he was able to escape from the trapped car before an on-coming train smashed into his car going 60 MPH.

Tuesday, December 11, 2007

JCDL 2008 Call for Participation

The ACM IEEE Joint Conference on Digital Libraries (JCDL 2008) will be held in Pittsburgh, PA on June 18-20, 2008. I'm on the program committee this year and look forward to a great conference.

JCDL welcomes contributions from all fields that intersect to enable Digital Libraries. Topics include, but are not limited to:
  • Interfaces to information for novices and experts
  • Information visualization
  • Retrieval and browsing
  • Data mining/extraction
  • Enterprise-scale Information Architectures
  • Distributed information systems
  • Studies of information behavior and needs; user modeling
  • Insightful analyses of existing systems
  • Deployment of digital collections in education
  • Digital Library curriculum development
  • Systems and algorithms for preservation
A one page CFP is available on the website.

Sunday, July 22, 2007

Citizendium or Wikipedia?

A few weeks ago I applied to be an editor on Citizendium, a new wiki project intended to be a more accurate and reputable Wikipedia. Citizendium does not allow anonymous postings, and they put new articles under a review process. My application was accepted a few weeks later, and a user account was created for me along with a user page listing my brief CV.

I decided to warm up with an article on digital preservation, a topic which I feel qualified to write about. Rather than start from scratch, I imported the Wikipedia article on digital preservation. Although Citizendium frowns upon importing articles from Wikipedia, I had previously written a large portion of the Wikipedia article, so I didn't feel too bad about doing it. I cleaned up the article by focusing the definition, cleaning up the references, and removing the numerous external links.

I didn't spend a whole lot of time editing the article because I got to thinking... can Citizendium really compete with Wikipedia? Larry Sanger (Citizendium's founder and co-founder of Wikipedia) seems to think so. Although I agree that Citizendium's policies in theory would result in more reputable articles, I don't think Citizendium can possibly scale to Wikipedia's size. First, there are a number of people who want to make a minor contribution to a Wikipedia article, a slight correction or clarification, for example. And since they don't have to register, the barrier to entry is sufficiently low enough for them to contribute.

Second, there are a number of people who, for whatever reason, want to remain anonymous or known by some alias. They are not likely to sign on with Citizendium and convey to the world why they are qualified to write about XYZ.

Third, if someone wants to make a contribution to an article on a particular subject, now they have to decide do they make the contribution just on Wikipedia, or on Citizendium, or both? Do they monitor both sites for changes to an article that is important to them? If a Citizendium article is actually better than the Wikipedia article, what is stopping Wikipedia from just importing the entire Citizendium article?

Forth, choosing the name "Citizendium" was a poor choice. It doesn't exactly roll off the tongue.

I really do hope Citizendium takes off, but I think its going to take a very, very long time before they have anything near the number of articles that Wikipedia has to offer. In the meantime, I would much rather put my efforts into something I know is going to rank high in search results and gets far more page views. I'll be watching Google to see if my Citizendium article will ever beat out Wikipedia's (currently ranked number 8) in the search results for "digital preservation". If it ever does, I'll defect to Citizendium.

Saturday, June 30, 2007

Save your Yahoo! Photos

I just received this email yesterday from Yahoo! Photos, telling me that my photos were going to disappear unless I took action. I couldn't help but wonder what if I my spam filter had eaten the message, or what if I wasn't using this email account anymore?

Considering that just 57% of individuals actually backup their personal data, how many people do you think are depending on Yahoo! to preserve their photos indefinitely? You can image the horrible feeling of logging in and realizing all those photos of your newborn child have ended up in the big bit bucket in the sky. Sigh...

Dear Yahoo! Photos user,
...

We will officially close Yahoo! Photos on Thursday, September 20, 2007, at 9 p.m. PDT. Until then, we are offering you the opportunity to move to another photo sharing service (Flickr, KODAK Gallery, Shutterfly, Snapfish, or Photobucket), download your original-resolution photos back to your computer, or buy an archive CD from our featured partner (for users of the New Yahoo! Photos only). All you need to do is tell us what to do with your photos before we close, after which any photos remaining on Yahoo! Photos will be deleted and no longer accessible.

...

Please give us your decision by Thursday, September 20, 2007, at 9 p.m. PDT. After that time, any photos remaining in Yahoo! Photos will be deleted. Click here to make your decision, or review a list of our frequently asked questions.

Saturday, May 26, 2007

First Digital Preservation Challenge

The DigitalPreservationEurope (DPE) has just announced its first international Digital Preservation Challenge:
The challenge invites participants to overcome the barriers hindering access to six digital objects. Each object is accompanied by a highly abstracted scenario based on real-life situations. These scenarios are intended to make the challenge more accessible to participants from all backgrounds while not trivialising the serious nature of the digital preservation challenges facing society.
The competition is open to undergraduate and graduate students in any major, and the winner will be announced in September at ECDL.

Wednesday, April 18, 2007

Preserving the entertainment industry

Now that movies are increasingly produced without film, it should be easier to save them for the long term, right? Here’s an excerpt from David Cohen’s Variety article:
In what is being called the first commercial effort to address the movie industry's looming digital preservation headaches, Elektrofilm Digital Studios and Sun Microsystems will introduce a long-term archiving system built specifically for the entertainment industry.
...
Though digital filmmaking promised an end to concerns about fading dyes and unstable film stocks, it has actually exacerbated problems with movie archiving and preservation.
...
The Motion Picture Academy's Science & Technology council saw enough danger to warn the studios that some movies could be lost, especially "born digital" films such as "300" and "Miami Vice."

Well, I certainly wouldn’t lose any sleep over the loss of either of those two movies, but it is exciting to see pro-active archiving initiatives from organizations that are not backed with tax payer money.

Update 4/22/07:

David Cohen has written another article about preserving film footage which further elaborates on some of the problems:
More than one tech expert, including the Academy's Sci-Tech Council director Andy Maltz, told Variety they had found archival tapes unreadable just 18 months after they were made.

Feiner, the former longtime prexy of Pacific Title, says when he worked on studio feature films he found missing frames or corrupted data on 40% of the data tapes that came in from digital intermediate houses.

The tapes were only nine months old.

Friday, April 06, 2007

Seeking high-quality Ph.D. students

If you live in Virginia and have a subscription to Time, US News, Newsweek, or Sports Illustrated, you may have seen this full-page advertisement for Old Dominion University (right) in this week’s issue (only subscribers receive regional ads). The guy featured in the ad is my advisor, Michael Nelson.

If you are looking to get your Ph.D. in computer science and want to work on some really interesting problems, I highly recommend Michael as an advisor. He’s got lots of good ideas and plenty of funding, and he will guide you to becoming an independent researcher. Plus you can live at the beach! smile

Here’s a little “advertisement” we’ve worked up to send prospective Ph.D. students:
Michael Nelson is seeking high-quality Ph.D. students for digital library and digital preservation research in the computer science department at Old Dominion University. Dr. Nelson has a well-funded research program focusing on object-repository interaction and alternative models of digital preservation. He has been PI or Co-PI on 9 grants (~$2.5M) from the NSF, NASA, Library of Congress and the Andrew Mellon Foundation. His most recent grant is a 5 year NSF CAREER award for "self-preserving digital objects".

More information about research projects, funding, publications and courses can be found at his website.

Dr. Nelson currently has 3 Ph.D. students, all of whom publish and travel on a regular basis. They also collaborate with colleagues at prestigious institutions such as LANL, Internet Archive and Cornell University. Dr. Nelson and his students welcome questions about joining their research group.

Thursday, March 01, 2007

Long-term preservation on DNA

Have you ever wondered if your digital belongings (photos, video, research papers, emails, etc.) are going to be accessible 5, 10, or 20 years from now? You may be thinking, yeah, they're stored on CDs/DVDs, and they'll last forever. That's what a friend of mine was thinking several years ago about zip disks when she saved her senior project and financial data on one. Unfortunately, she no longer owns a zip drive, so she enlisted my help. After tracking down a zip drive, I attempted to extract the files only to find a handful of them were inaccessible. These ended up being very important, and my friend is now in the depths of depression. wink

Anyway, I say all this because you’re worries are now over: Japanese scientists have discovered how to store data on bacteria DNA. Since bacteria pass on their DNA unaltered from generation to generation, your data can be safely encoded on them for thousands of years. That’s right: God used DNA to store developmental instructions; you can use it to store your Brittany Spears collection. Of course the technology for correctly extracting and interpreting the stored data must also be functional a thousand years from now, but we’ll let the digital preservationists figure that out.

For an enlightening read on the long-term preservation of digital data, see Ensuring the Longevity of Digital Information by Jeff Rothenberg.
"Digital information lasts forever -- or five years, whichever comes first." -Jeff Rothenberg