Monday, August 28, 2006

Bayside League: 2006-2007 season

Saturday night was the Bayside League’s third annual fantasy football draft. Mark Velez was kind enough to host it at his place. There were 14 teams- just 2 shy of having 2 leagues. Chris Deny, last year’s champion, drafted by phone, and I drafted my dad a team.

I got to draft second, and since Payton Manning was taken first, I was able to nab who I hope will be the best running back this year in the NFL: Larry Johnson. I was also able to draft a few Broncos and Cowboys- I just can’t help but play favorites.

Last year I came in second place. Let’s see what I can do this year…

  1. RB – Larry Johnson
  2. RB – Willis McGahee
  3. QB – Jake Plummer
  4. WR – Donald Driver
  5. TE – Chris Cooley
  6. WR – Andre Johnson
  7. DEF – Cowboys
  8. K – Josh Brown
  9. RB – Mike Bell
  10. QB – Mark Brunnell
  11. WR – Nate Burleson
  12. K – Lawrence Tynes
  13. DEF – Patriots
  14. TE – Desmond Clark

Thursday, August 24, 2006

Funny Commercials

One of my hobbies is “collecting” commercials. I know that sounds odd. What I mean is I am a big fan of commercials, and I like to download or keep “pointers” to commercials that I think are really hilarious, entertaining, or just plain fantastic. Google Video and YouTube are two of the best sources for commercials. In fact my favorite commercial this year is the Liberty Mutual “pay-it-forward” commercial which is now available on YouTube:



When we got home from church last night, we caught the “World’s Funniest Commercials” on TBS. In general I was disappointed by the crassness and message of many of the commercials, but there were a few gems:

LA Country Fair – "Duh, Ashley, all wool comes from a cow..."
1-800-Got-Junk – “Rat Advertising Trial (R.A.T.)”
Avis: Gansta Rap – “Gotta get that money made”
Solo Mobile: Housewarming – “You are a legend”

Honorable mention: Pacman Puppet Show

All the commercials can be seen in high quality at veryfunnyads.com.

Tuesday, August 22, 2006

The ACLU and nonsectarian prayer

Last week, a federal judge dismissed a Fredericksburg City Council member's lawsuit challenging the council's nonsectarian prayer policy. Ironically, the ACLU of Virginia is fighting in this case to limit the liberties of an individual rather than protect them. It’s not quite so surprising when you learn their opponent is a Christian who desires to use the name “Jesus Christ” in an opening prayer for the city council.

A lawyer friend of mine who is intimately familiar with this case offered the ideal ACLU prayer, guaranteed to be as nonsectarian as possible:

"At this point we would like to call on a genderless, nameless higher power than ourselves and invoke that being (or those beings) intervention (or non-intervention as the case may be) upon this governmental body."

If it ever does come down to that, most of us would rather have no prayer at all. That’s exactly what the ACLU is hoping as well.

Hypertext 2006

Michael is in Odense, Denmark this week presenting two papers at the Hypertext 2006 conference:

  1. Evaluation of Crawling Policies for a Web-Repository Crawler by McCown and Nelson
  2. Just-In-Time Recovery of Missing Web Pages by Harrison and Nelson

I should be there presenting the crawling policies paper, but Michael graciously went in my place so I could stay here and work. Our papers are 2 of the 12 that will be presented.

Lazy Paper accepted to WIDM'06

I got some really good news yesterday- my Lazy Preservation paper (“Lazy Preservation: Reconstructing Websites by Crawling the Crawlers”) was accepted for publication at the WIDM’06 workshop along with Joan’s mod_oai paper. Only 11 of the 51 submissions (21.5%) were accepted which means the workshop participants are going to be getting very familiar with the research aims at ODU. ;-)

WIDM is in Arlington, Virginia on November 10 and is held in conjunction with CIKM’06. This will be a good opportunity for me to visit my sister again in D.C. I was up there a few weeks ago giving the good news about Becky’s pregnancy to my parents (they were visiting from St. Louis). Random event: while I was in D.C., my mom and I ran into Karl Rove at a bookstore- we chatted for a while (Mom’s a fan), and he seemed like a nice guy, but I'm sure most politicians do. Anywho, I’m looking forward to the workshop.

Thursday, August 17, 2006

Google Analytics

This morning I installed Google Analytics on my blog and ODU website. It’s a free tool that allows me to track how users enter, leave, and navigate my website. It involved simply posting some JavaScript (below) on the pages I wanted to track:
<script src="http://www.google-analytics.com/urchin.js" type="text/javascript">
</script>
<script type="text/javascript">
_uacct = "UA-######-#";
urchinTracker();
</script>
It’s going to take a few weeks before there’s any data collected, but I’m really curious to see what this will reveal about the popularity of Warrick since I don’t have access to the CS web server logs.

Update on 8/25/06

It's been a week, and I'm now able to see some analysis of my blog and my ODU cs website. The screen below shows a summary of my blog's traffic from Aug 18-24:


Visitors increased from 11 to 26 during this week, and pageviews ranged from 25-49. Three quarters of the traffic are new visitors (Google is using cookies to track this).

The Geo Map Overlay is fascinating. My blog tends to appeal more to Americans and Europeans: I got only 2 visits from Australia, 1 from South America, and 1 from Africa. There were 14 visits from Tampere, Finland and 9 from Nokia (also in Tampere). The Visits by Source graph shows the Finish hits to be from Timo's nothingforsale.com website where I now have a link pointing to my blog.

Most people find my blog through Google. So what are people searching Google for that lands them on my blog?

Apparently my blog entry about Yahoo's error 999 is by far the most popular. Searches 1, 4, 5, and 9 will all return this entry in Google's top 10 results. The "shiri maimon" entry doesn't show up in Google until page 3 (top 30 results).

What pages are referring visitors? Apparently Elrod's blog is the biggest referrer so far. This is due to a comment I left on Aug 18.

What really surprises me is that anyone is visiting frankmccown.com, a website with no content. I did a search for "frank mccown" in Google, and the site came up number one. Come on Google... you guys are supposed to be punishing content-less sites like this, not promoting them. I guess Google's PageRank is far from what they originally published in 1998; there's maybe 1 or 2 links to this page from anywhere on the Web.

My cs website is getting a little more traffic than my blog even though I only have the tracking on a few of the pages. What was most interesting to learn was that Wikipedia was by far the largest referrer, sending me around 60 referrals last week. Most of the referrals are coming from the Internet Archive entry where there's a pointer to Warrick.

The Warrick page received 125 visits and 175 pageviews last week (18 and 25 per day, respectively). Here are some search terms people are using to find Warrick:
  • google api convert documents to HTML
  • Warrick
  • Warrick website download
  • Warrick archive.org
  • warrick perl
  • cached archive website
  • recover website from google cache
  • google cache website recovery

Wednesday, August 16, 2006

Little McCown on the way

I just can’t keep it a secret anymore – Becky is seven weeks pregnant! The appointment with the OBGYN went really well this morning, and, God willing, we are expecting our first little McCown on April 5. That’s when I’ll be in the thick of writing my dissertation, so it will be a really exciting and challenging time! :)

Nothing for sale

When my friend Timo Kosonen emailed me back in January 2006 about his website http://www.nothingforsalesite.com/, I emailed him back saying that he was crazy- who would pay something for nothing? His idea roughly mirrored the http://www.milliondollarhomepage.com/ idea where people spent one dollar per pixel for on-screen real estate. Although it worked out well for the million dollar guy, I wasn’t so sure it would work out the same for Timo.

Well, Timo got some press out of it and was able to make a few hundred dollars (which he applied to his wedding). Not bad for a Harding grad. ;)
Way to go, Timo!

Tuesday, August 15, 2006

Torrance Daniels in Philly


Torrance Daniels (“Tank”) is the first player from Harding University to possibly play in the NFL. Right now he’s in training camp with the Philadelphia Eagles. Although I’m no Eagles fan (go Cowboys!), I’ll be pulling for him.

Update on 9/8/06:

It looks like Tank will be on the Eagle's practice squad. Not bad.

Update on 11/21/06:

Tank is now a starter, thanks to the season-ending injury to McNabb.

Update on 11/30/06:

Tank started last Sunday evening against the Colts and made the first tackle of the night on the opening kick-off. There's an article about it on the Harding website.

Friday, August 11, 2006

AOL releases search queries

On Sept 28, 2006, AOL released the search histories of more than 650,000 of its users (21 million queries) on its new research website. Although the data was stripped of personal identifiers, it still made privacy advocates extremely upset. AOL issued an apology 10 days later and yanked the data from their site, but it had already been replicated.

For a researcher involved in information retrieval, this data is a gold mine. Most researchers don’t have access to data like this. Unless you work for a search engine, you have to rely on search data from your institution or beg for it from other locations.

On the other hand, some search data could be linked to specific individuals, and I can see why that would be alarming to some. Perhaps there’s a middle ground? What if location data could be randomly swapped? For example, a search for “boston hair cut” could be changed to “denver hair cut”. Although this would make the location information worthless, all the other important information (query length, word length, subject matter) would still be present. Other heuristics could be applied to muddle the location. Of course this doesn’t address all the privacy issues, but it’s a start.

Many of the queries are very disturbing. Many of the queries deal with pornography, grief, and revenge. The queries are like the random private thoughts of their owners. Although they would likely never mutter this stuff to a friend, they have no problems entering it into a search box. One thing that is very clear, there are a lot of hurting people out there.

What I also found very interesting was the way people make their queries. The lengths of many queries are very long. Users are apparently adding more words to get better precision. As the Web has gotten much larger, it has become necessary to use more words. Just six years ago a long query would result in very few hits, but not anymore. Also people sometimes use slang or misspellings which would likely match fewer results. For example, one user entered “u” in several queries where “you” would obviously be more appropriate. Search engines may need to adopt to the use of slang and make automatic substitutions when possible.

It’s really too bad that AOL has received so much heat for what has happened, especially since other companies like Excite and AltaVista have done the same thing in the past. The difference today is that we are much more aware of privacy issues, and the queries are becoming much more tuned to individuals. I would still like to see Google, MSN, Yahoo, and others also give up some detailed search data like this in the future.

Wednesday, August 09, 2006

Crawling the Web is very, very hard…

I’ve spent the past couple of weeks trying to randomly select 300 websites from a dmoz.org. There were only a couple of restrictions I placed on the selection:
  1. The website’s root URL not redirect the crawler to a URL that is on a different host. If it does, the new URL should replace the old website URL.
  2. The website’s root URL should not appear to be a splash page for a website on a different host or indicate that the website has moved to a different host. If it does, the new URL should replace the old website URL.
  3. The website should not have all of its contents blocked by robots.txt. If some directories are blocked, that’s ok.
  4. The website’s root URL should not have a noindex/nofollow meta tag which would prevent a crawler from grabbing anything else on the website.
  5. The website should not have any more than 10K resources.

The restrictions seem very straightforward, but in practice they are very time consuming to enforce. Requirement 2 requires me to manually visit the site. Did I mention not all the sites are in English? That makes it even more difficult. Requirement 3 means I have to manually examine the robots.txt. Req. 4 requires manually examination of the root page, and req. 5 means I have to waste several days crawling a site before I know if it is too large or not.

I guess I could build a tool for req 3 and 4, but I’m not in the mood.

Anyway, I ended up making about 50 replacements (at least) and starting my crawls over again. Now I finally have 300 websites that meet my requirements.

In the past I’ve used Wget to do my crawling, but I’ve decided to use Heritrix since it has quite a few useful features missing from Wget. But Heritrix isn’t perfect. I made a suggestion that Heritrix show the number of URLs from each host remaining in the frontier when examining a crawl report:

http://sourceforge.net/tracker/index.php?func=detail&amp;amp;aid=1533116&group_id=73833&atid=539102

Right now it is very difficult to tell if a host has been completely crawled or not. I would love to work on this myself, but I just don’t have the time right now. Maybe I'll get a student to work on this next time I'm teaching. ;)

The other difficulty with Heritrix is in extracting what you have crawled. I will need to write a script that will build an index into the ARC files so I can quickly extract data for a website. Since all the crawl data is merged into a series of ARC files, it is really difficult to throw away crawl data for a website you aren’t interested in. I could write a script to do it, but at this point it’s not worth my time.

Anyway, crawling is a very error-prone and difficult process, but I can’t wait to teach about it when I return to Harding! (I'm really getting the itch to get back in the classroom.)

Tuesday, August 08, 2006

CiteULike - a second brain for researchers

Today I discovered a new free tool for researchers to manage their references. It’s called CiteULike, and it’s been around since 2004. Richard Cameron created the site and intends to keep it free.

I was able to upload all my BibTeX entries without any difficulties. I can now tag each entry so I can quickly see what papers are relevant to a particular subject. For example, here are all the papers I have tagged for web-archiving:

http://www.citeulike.org/user/fmccown/tag/web-archiving

CiteULike is useful for finding out about new research in your area. For example, this user apparently has many of the same interests as me, and I found several new papers by browsing his library:

http://www.citeulike.org/user/ChaTo/

It’s cool because I can even see comments that users have made about specific papers.

Now when I come across a new paper, I can add it to my CiteULike library and jot a quick note about it and not worry months later when I need to find the paper. And now instead of emailing Michael my BibTeX file, he can download the whole thing directly from the Web.

Thursday, August 03, 2006

Yahoo transforming FRAME tags

The past several months I’ve been ramping-up for a huge experiment where I’ll be reconstructing several hundred websites. I’ve been learning to use Heritrix and process ARC files, and I’ve been periodically tweaking Warrick. Today I found out that Yahoo has changed the way it caches HTML pages that contain frames.

For example, the page at http://www.harding.edu/comp/ contains the following HTML:

<FRAMESET COLS="195,*" FRAMEBORDER=no FRAMESPACING=0>
<FRAME SRC=menu.html NAME="MENU" MARGINWIDTH=0 MARGINHEIGHT=0>
<FRAME SRC=welcome.html NAME="MAIN">
</FRAMESET>

In Yahoo’s cached page for this URL, the FRAME tags are converted to the following (I’ve added some white space for readability):

<frameset rows="200,*"><frame scrolling="no" noresize="" frameborder="0" marginwidth="0" marginheight="0" src="http://216.109.125.130/search/cache?.intl=us&u=www.harding.edu%2fcomp%2f&
w=%22harding+.edu%22&d=XS7fRmP9NNYx&origargs=p%3durl%253Ahttp%253A%252F%252F
www.harding.edu%252Fcomp%252F%26toggle%3d1%26ei%3dUTF-8%26_intl%3dus&frameid=-1">


<FRAMESET COLS="195,*" FRAMEBORDER=no FRAMESPACING=0>

<frame security="restricted" MARGINHEIGHT="0" MARGINWIDTH="0" NAME="MENU" SRC="http://216.109.125.130/search/cache?.intl=us&u=www.harding.edu%2fcomp%2f&
w=%22harding+.edu%22&d=XS7fRmP9NNYx&origargs=p%3durl%253Ahttp%253A%252F%252F
www.harding.edu%252Fcomp%252F%26toggle%3d1%26ei%3dUTF-8%26_intl%3dus&frameid=1" >


<frame security="restricted" NAME="MAIN" SRC="http://216.109.125.130/search/cache?.intl=us&u=www.harding.edu%2fcomp%2f&
w=%22harding+.edu%22&d=XS7fRmP9NNYx&origargs=p%3durl%253Ahttp%253A%252F%252F
www.harding.edu%252Fcomp%252F%26toggle%3d1%26ei%3dUTF-8%26_intl%3dus&frameid=2" >


</frameset></FRAMESET>

Yahoo is placing their own FRAMESET tags around mine and loading the two column frames with pages directly from their cache. Notice the use of security="restricted" within the FRAME tag which tells the browser to place security constraints on the frame sources; this disables any JavaScript in my pages.

While this conversion of FRAME tags makes the page easier to view from their cache, it completely destroys the original HTML. There’s no way I can even parse through the arguments to tell what URL used to be in the SRC attribute. ARG! Now I’m going to have to add a rule to Warrick that tells it to ignore Yahoo cached pages that contain FRAME tags. Google and MSN have yet to implement this “trick”, and hopefully they never do.

Wednesday, July 19, 2006

That’s not a cache... that’s an archive!

This morning I stumbled across a March 2006 blog posting by Danny Sullivan entitled “25 Things I Hate About Google”. Sullivan’s opinions carry a lot of weight in the search engine world, and so I started to sweat when I saw number 9 on his list:

9. Stop caching pages: I was all for opt-out with cached pages until a court gave you far more right to reprint anything than anyone could have expected. Now you've got to make it opt-in. You helped create the caching mess by just assuming it was legal to reprint web pages online without asking, using opt-out as your cover. Now you've had that backed up legally, but that doesn't make it less evil.

Sullivan doesn’t agree with the January 2006 Nevada federal court ruling that declared Google’s cached pages did not constitute copyright infringement, thereby okaying the opt-out policy used by search engines using the noarchive meta-tag. Sullivan and others make some good points in the forum discussing the ruling, showing where the ruling may have some flaws.

One the arguments opponents of the ruling make is that a search engine cache is hardly a cache in the traditional sense because pages are cached long after they are changed or deleted from a web server. One of the posts by mcanerin gives an example of a web page that had been cached for almost 2 years (the example is no longer accessible). In mcanerin’s words: “That's not a cache, it's an archive.”

The fact that the cache is more like an archive is exactly what makes it beneficial to most Web users, and that’s why I think the court’s judgment was fair. Search engine caches are a huge public good. Caching is not evil. Yes, there may be a few scenarios where caching may not work to everyone’s benefit, but in most cases the good far outweighs the bad. As long as search engines provide a mechanism to keep crawled content from being cached and to remove cached content immediately if needed, then there is no really compelling reason to force search engines to use an opt-in policy. (Yes, I know it can be a real pain to manually remove entries from many search engines, but how often does anyone really need to do that?)

My research on digital preservation of websites relies heavily on the wide-spread use of search engine caching, and if caching turns from an opt-out to opt-in, I am going to be in serious trouble, and so are users of Warrick. I’ll be keeping my eye on this…

Immediate action required now

I was just reminded this week of the dangers of pornography when a friend of mine shared his personal struggles with it. It is ripping his life apart. There’s no doubt about it… this stuff is poison. It will poison your relationship with your spouse, your friendships with members of the opposite sex, and your soul. Don’t mess with it.

There’s a really good article about Steve Holladay and his struggles with pornography addiction in the Christian Chronicle (April 2006). Steve talked to our church one night about struggles with pornography and his ministry to reach out to youth who struggle with pornography addiction. (Steve was finishing his Ph.D. here in town at Regent University.)

One point the article made that I will share here is that pornography today is a much more dangerous beast than it was just fifteen years ago, and for that reason, it deserves your attention now. The three A’s - accessibility, affordability, and anonymity – illustrate this change. Fifteen years ago you had to physically visit a store selling pornography, pay for it, and reveal your identity. Today you can access pornography for free in your home or office and remain totally anonymous. And even if you don’t have any intentions on viewing pornography, it is emailed to you daily and appears in search engine results. You just can’t get away from it.

If you are concerned about your spiritual health and know that pornography is a strong temptation for you, there is absolutely no reason why you should not be using a filtering service like the American Family Filter. Of course filtering software isn’t going to block everything, and I’d recommend you go one step further and use accountability software like Covenant Eyes.

We owe it to ourselves and to our spouses to keep our conscience and minds clear of sexual immorality. God doesn’t ask anything less of us, and He has promised to give us strength to overcome it.
"He gives strength to the weary and increases the power of the weak."
- Isaiah 40:29

"And God is faithful; he will not let you be tempted beyond what you can bear. But when you are tempted, he will also provide a way out so that you can stand up under it."
- I Corinthians 10:13

Monday, July 03, 2006

Thinking Differently...

This past Thursday, Michael, Joan, and I gave a talk entitled “Thinking Differently about Web Page Preservation” at the National Digital Library Center (NDIIP briefing at Library of Congress in D.C.). Butch Lazorchak was our liaison. It was a great experience to give a talk in D.C. about my research. It also gave me an excuse to visit my sister, see some of the sites, and watch a Nationals game.

Update on 7/20/06:

The webcast is available from the Library of Congress Webcast page. My part runs from 16:50 - 47:15.

Friday, June 23, 2006

Heritrix - An archival quality crawler

This week I’ve been experimenting with Heritrix, the Internet Archive’s web crawler. It has some functionality that Wget doesn’t provide including:
  • limiting the size of each file downloaded
  • allowing a crawl to be paused and the frontier to be examined and modified
  • following links in CSS and Flash
  • crawling multiple sites at the same time without invoking multiple instances of the crawler
  • storing crawls in an Arc file
Since Heritrix was built with Java and was pre-configured to run on a Linux system, I didn’t have to expend much effort to get it to run on Solaris. I untarred the distribution file, set a couple of environment variables, started the web server interface, and boom it was working.

The interface is not exactly intuitive, and a near complete reading of the entire manual is required to put together a decent crawl. Of course if you want to use sophisticated open-source software, you usually have to put in some significant effort to get it to work right. Thankfully, several of the developers (Michael Stack, Igor Ranitovic, and Gordon Mohr) have been very helpful in answering some of my newbie questions on the Heritrix list serve.

In learning about Heritrix, I’ve put together a page on Wikipedia. Hopefully the entry will drum up more general interest in Heritrix as well. I was really surprised no one had created the page before.

Tuesday, June 20, 2006

Integer problems for the Google API

I’m not sure when it first started, but the Google API has been bombing out over the last few months when returning over 2^31 (2,147,483,648) results for a query. The API has bombed-out almost every day in June when my script searching for “database” and “list” which each return several billion results. Apparently Google’s SOAP interface is using a 32-bit integer for returning the total pages returned, but they need to be using a 64-bit long integer.

Michael Freidgeim made note of the problem on his blog a few weeks ago. Others have noticed this problem going back to April 2006. Who knows when Google will make a fix. If it's not one thing, it's something else... ;)

When searching to see when Google started using the larger total results, I came across a posting by Danny Sullivan that shows how he was attempting to use a “trick” to reveal how many pages Google has indexed. Danny suggested issuing a query that says, “give me all the pages that don’t have the word asdkjlkjasd.” I just tried –asdkjlkjasd on Google, and it gives me back 20.7 billion results. MSN gives around 5.2 billion results, but Yahoo and Ask won’t accept the query. Interesting…

frankmccown.com

I recently created my own website frankmccown.com using Microsoft Office Live. Since it was free to register the domain name and host the site, I thought I might as well give it a try. Thanks, Microsoft, for giving me a free site. The only problem I have is actually editing it. Microsoft tried to make the interface easy for business users to create a website. Unfortunately, they created an interface that is impossible to use for those of us who want to actually edit HTML. Where is the “edit HTML” button?!

I emailed the Microsoft Office Live folks to see if they could tell me how to edit the HTML, and they replied:
Microsoft Office Live does not support stand alone HTML code designing. HTML code designing can be accomplished using Microsoft Office FrontPage 2003. Currently, only the Microsoft Office Live Essentials subscription supports publishing the website through Microsoft Office FrontPage 2003.
Looks like you have to pay if you want to be able to change the under-lying HTML, and then you have to use FrontPage. Blah... Looks like frankmccown.com will not be getting much attention in the near future.

Friday, June 16, 2006

End of the Google 502 errors?

Google users have sporadically seen Google 502 (bad gateway) errors the last several years. The errors appear momentarily and then disappear. I’ve linked to a few postings about it according to date:

Mar 2003
July 2003
June 2005
Sept 2005
Nov 2005
Feb 2006
May 2006

Google API users have seen the 502 errors much more frequently:

Nov 2005, and another
Dec 2005
Jan 2006
Feb 2006, and another
May 2006

From my investigations, it looks like Nov 2005 is when the problems began. I have personally dealt with the problem ever since Mar 2006 when I integrated the Google API into Warrick. I had to add some logic to sleep for 15-20 seconds when encountering the error and then re-try.

In late May I started a new experiment which uses the Google API, and I’ve been monitoring it daily to see how many 502 errors I was receiving. From late May to June 6, I consistently received a 502 error for about 30% of my requests. On June 7, the number of 502s went down to zero. I have only received an occasional 502 out of hundreds of requests made daily.

Someone at Google finally got sick of the bad press and made some changes, and I’m thankful for it. :)