Tuesday, March 13, 2007

March Madness

I just filled out my first NCAA Tournament bracket. I joined the Knights Pool on FaceBook-- apparently I could win $25K just for getting lucky. smile I’ve got ODU winning their first game against Butler (hey, I gotta show my Monarch spirit...), but I don’t think they’ll get any farther than that. But wouldn’t it be awesome though if they could follow in George Mason’s footsteps?

Final Four: UCLA vs. Oregon and UNV vs. Ohio State
Championship: UCLA 86, UNC 83

If you are laughing out loud, just remember this is my first bracket... wink

Update 3/15/07:

Well, ODU played a good first half and went into the locker room ahead by one. Unfortunately, they couldn't keep up in the second half and lost 57-46. Maybe the women can do better on Sat.

Friday, March 09, 2007

Huge demand for CS graduates

There is no doubt that interest in computer science has been dropping over the years from its peak in the 1990s. The statistics show that fewer incoming freshmen are interested in CS, but the question is why? After the dot-com bust and news of jobs going overseas, many parents have steered their kids away from tech-related majors. Unfortunately these parents are hindering their kids from entering one of the hottest job markets around:
"Nationwide, there are more jobs in the U.S. in computing than there ever have been, even at the height of the dot-com craze. The 10-year count for growth in new jobs is that there will be 1.4 million net new jobs over a 10-year period. That’s the hottest growth area of any area at all in the science and engineering fields."
- Jeffrey Vitter, dean of the College of Science at Purdue University
Not only is the job market hot, but it is very enjoyable field to work in. Just a few months ago, Money Magazine rated software engineering as the best job in America. And former students that I have spoken with almost always enjoy their careers in the technology field. It’s been a very long time since I heard from a student who couldn’t find work.

So please, parents and students: If you have any interest in CS at all, give it a shot. It’s not the easiest major, and it’s certainly not for everyone, but it may provide you a very rewarding career for years to come.

For more resources, see The Computer Science job outlook: Myths and Truths and the new ACM Computing Careers website.

Wednesday, March 07, 2007

Software Development Project: any ideas?

This coming fall I will be teaching the senior capstone course (Software Development Project) at Harding University. I’ve only taught this course once before, about 9 years ago, so I’m starting fresh and looking for new ideas.

Traditionally this course has involved dividing the class into groups of 4-5 and leading the groups through a large part of the development cycle (design, implementation, documentation, user testing, etc.) of producing a computer game. Past teams have implemented games like Yahtzee, Spades, Battleship, Checkers, and Othello—normally games that take only 5-10 minutes to play all the way through. Technical testers and user testers rate the games, the AI’s compete against each other to see which is the best, and ultimately one of the teams is selected as the all-around winner.

The class is usually one of the most educational for the students and one of the most enjoyable because of the amount of creativity that they can pour into their projects and the amount of freedom they have in implementing the game. Plus some students really enjoy the competitive aspect.

While this has course has worked out really well in the past, I wonder if any aspects of it should be changed. Should the students modify an open source project already in development? This is the approach that Google Summer of Code takes, and it exposes students to the type of work they will likely encounter when they graduate (most software in not built from scratch).

Or should the software they are building be less game-oriented? I think students are usually more enthusiastic about writing games, but they are more likely to write business-oriented (or at least more practical) applications in the “real world”.

The games the students have traditionally built have been stand-alone applications for Windows that must be installed locally. There’s some advantages to this since students are exposed to the types of problems that such applications run into (e.g., incompatibility issues), but on the other hand, they aren’t being exposed to thin client applications that are becoming more popular today. Should I limit the students to developing only a Web-accessible application, possibly with AJAX, Flash, or (yuck) Java applets?

I’m open to any suggestions, so please let me know your thoughts.

Thursday, March 01, 2007

Long-term preservation on DNA

Have you ever wondered if your digital belongings (photos, video, research papers, emails, etc.) are going to be accessible 5, 10, or 20 years from now? You may be thinking, yeah, they're stored on CDs/DVDs, and they'll last forever. That's what a friend of mine was thinking several years ago about zip disks when she saved her senior project and financial data on one. Unfortunately, she no longer owns a zip drive, so she enlisted my help. After tracking down a zip drive, I attempted to extract the files only to find a handful of them were inaccessible. These ended up being very important, and my friend is now in the depths of depression. wink

Anyway, I say all this because you’re worries are now over: Japanese scientists have discovered how to store data on bacteria DNA. Since bacteria pass on their DNA unaltered from generation to generation, your data can be safely encoded on them for thousands of years. That’s right: God used DNA to store developmental instructions; you can use it to store your Brittany Spears collection. Of course the technology for correctly extracting and interpreting the stored data must also be functional a thousand years from now, but we’ll let the digital preservationists figure that out.

For an enlightening read on the long-term preservation of digital data, see Ensuring the Longevity of Digital Information by Jeff Rothenberg.
"Digital information lasts forever -- or five years, whichever comes first." -Jeff Rothenberg

Tuesday, February 27, 2007

Watch out

Three things I never realized were so dangerous:
  1. Flying kites
  2. Posting an "Electric Slide" video on YouTube
  3. Using my laptop while driving (well, I was pretty sure this was a bad idea before)

Monday, February 26, 2007

Becky's first baby shower

On Sunday our church (Stephanie organized) threw a baby shower for Becky. Thankfully I was not forced to participate, but when I showed up at the end of the shower to help pack-up the goods, I couldn’t believe how much stuff there was. There were tons of socks, blankets, and clothes… even a little baseball mitt. A couple of women made a blanket and an "Ethan Andrew McCown" pillow. Becky and I are really blessed to be part of the family at Bayside. It’s gonna be really tough to leave in August.


And on Saturday night Becky and I saw Amazing Grace with some friends. It’s the true story of how William Wilberforce helped end the slave trade in 19th century Britain. The movie was very well done, and it was inspiring to how much good this man was able to do by letting God work through him.

Thursday, February 22, 2007

Google: From search engine to Singularity

Philipp Lenssen has written an interesting article about how search engines (or more specifically Google) may evolve over time. As its AI improves, Lenssen sees Google growing into an all-knowing entity, capable of inferring new levels of knowledge and reaching conclusions from its vast quantities of data, even referring to itself as "me". (Yes, Google may some day be responsible for the Matrix, the Terminator, or just good old-fashion space mission sabotage.) Lessen has talked about the emergence of the technological singularity before; it's an interesting concept for the science fiction lover.

On a totally unrelated somewhat related note, congratulations to Frances E. Allen, the first female to win the distinguished A.M. Turing Award. Allen has contributed to work in weather prediction and code breaking, two areas from which a Singularity might draw from.

Monday, February 19, 2007

A new car!

This weekend we traded in our ailing Plymouth Breeze for an '06 Toyota Camry. It's a long way from my Jeep Wrangler days, but if you consider the reliability, resale, gas mileage, price, and baby-room, it's just what we needed. Plus it has sun roof! A nice thing about Camrys is they only need to have their oil changed every 7K miles... much better than the 3K I'm used to. (I would have gotten a hybrid, but they don't make 'em cheap.)

We had to go through the customary bargaining game where they give you a number, you give them one, they send out the manager, you get a new number, etc., etc. The line "That's just not good enough" came in handy, but it was tough on Becky since she had to endure the process with lower-than-normal energy levels. Man, it's tough being pregnant.

Last night after church I rented a movie for us called The Greatest Game Ever Played. It's based on the true story of 20-year-old Francis Ouimet who faced his idol, Harry Vardon, in the 1913 US Open. Very inspirational, even if you don't like golf.

Thursday, February 15, 2007

Be afraid, be very afraid

Hey, Windows user- if you haven’t made sure you’re running the latest version of Internet Explorer, do it now. If you’re lacking the motivation, take a look at this video that shows just what might happen to you if you were to misspell google.com. Scary. (See a blog posting about this video here. Note that the link to the video on that site is currently broken.)

You also can't always trust websites that you normally trust. A few days before the Super Bowl, the official website of Dolphin Stadium was hacked (that's where the Super Bowl was held). Website visitors running Windows might have been infected with a Trojan keylogger/backdoor if they hadn't already installed a couple of Windows security patches.

And if you’re thinking you’re safe from attack just because you’re running some other OS, take a look at this recent report: Some researchers set up 4 Linux boxes and watched as the servers were attacked on average once every 39 seconds. The hackers most often attempted to access the machines with user names of "root" and "admin" using passwords which were variations of the user name or "123." Man, I thought I was the only one using 123... wink

Another recent study states that 70% of websites are at immediate risk of being hacked! Well, maybe that number is a little high. But the point is, if you are on the Internet, you better watch out. Have a nice day. smile

Wednesday, February 07, 2007

Wow, you look so good in that photo... is it actually you?

Here’s a little function you may want for your photo editing software: The Beauty Button. Researchers at at Tel Aviv University have developed a program that can make an image of a person’s face more attractive by moving or distorting portions of the face to conform to factors that we all find beautiful. You can read the full article here.

According to Cohen-Or, one of the researchers behind the software:
Beauty is not in the eye of the beholder. Beauty is merely a function of mathematical distances or ratios. And interestingly, it is usually the average distances to features which appears to most people to be the most beautiful.
In other words, average is beautiful. So would you want to apply the beauty function to your photos? I’m just wondering what my friends and family would think when they saw my wedding album and didn't recognize the guy kissing my wife. Sure, give me a tan and remove that zit, but leave my nose alone, please.

Speaking of beauty, my wife’s birthday is today! She’s still a youngin’… not even 30 yet. If you’d like to send her a happy birthday note, you can reach her at beckymccown1 at yahoo dot com.

Tuesday, February 06, 2007

Harding gets a face-lift

Designing a website is really difficult. It requires a unique combination of a good eye and technical know-how. I have plenty of know-how, but if you’ve seen any of the websites I’ve designed, you know that’s where my talents end.

Harding recently redesigned their website. Shawn Spearman spear-headed the project, and I think he did a really good job. You can see the before and after screenshots below.

BEFORE


AFTER




Since I’ll be teaching principles of website development (as part of my Internet Dev class) when I return to Harding next year, I thought I’d put my critical analysis cap on and list some things about the new site that I liked, and some things that I thought should be changed.

What I like:
  1. The logo at the top with the background composed of photos from around campus- very clean and professional.
  2. Letting Google do the searching. Previous versions of the website used their own in-house searcher which did a very poor job.
  3. Intuitive navigational menus which seem very well organized.
  4. Phone numbers and addresses prominently displayed throughout the website.
  5. Nice photos arranged on the sides of pages.

Some things that need to be changed:
  1. The root page is too wide which causes some users to have to use their horizontal scrollbar to get access to the Search box (a big no-no is to make users scroll horizontally).
  2. Not enough width was given for actual content. If you take a look at this page, you’ll see the content is mashed to the side with only 3-4 words per line. If you were to print out this page, it would take 20 pages when it should only take 2. I would recommend scrapping the left border which is functionally useless, and/or using a smaller sized font for the main content.
  3. The links on the left side of each page are the same color (black) and size as the text next to it. The links aren’t underlined, so users must guess what is and what isn’t a link. Although most links are on the left, they also appear on the right which makes finding the links a real pain.
  4. Some of the navigational links contain deep hierarchical menus which should be avoided. This rule is of course violated in Windows (Start > All Programs > Accessories > Off to see the wizard…), and so it’s trained others to think it’s ok. Well, it’s not. I dare you: Try to select Future Students > Adult Ext. & Ed. > Not for Credit > Kids Kollege. It took me three times.


For those of you who want to know more about designing usable websites, here are a couple of useful links:

Web Curator Tool, standardizing PDF, and orphaned works

Some notable events in the world of digital preservation:
  • The National Library of New Zealand and the British Library have collaborated to produce the Web Curator Tool (WCT), a tool that allows non-technical users to archive websites in a simplified manner. It’s essentially a wrapper around the Heritrix web crawler with numerous management functions added on. In a recent article, Philip Beresford from the British Library discusses the history of WCT and shows how it can be used to crawl and archive a website.

  • In an effort to convince the world that the PDF format is ideal for long-term storage, Adobe is submitting it to ISO for standardization. Microsoft has also submitted their Ecma-approved Office Open XML for standardization to ISO, a radical departure from the "secret-sauce" mentality Microsoft has held for years. Governments and other organizations are slowly becoming aware of the problems created by storing their data on closed formats that change over time, and Microsoft and Adobe don’t want to be dropped from their largest customers. By standardizing these formats, interoperability should be much less of an issue in the future.

  • Brewster Kahle, co-founder of the Internet Archive, recently lost a U.S. appeals decision in Kahle v. Gonzales. Kahle, along with several notable companies like Google, MSN, and Yahoo, are trying to get orphaned works (copyrighted work whose owner cannot be reached) into the public domain in order to remove legal barriers that prohibit the scanning and digital distribution of those works. Kahle rightly blames Disney for the mess:
    What happened is that some overzealous copyright laws got passed with heavy lobbying from folks like Disney and these are screwing things up... Instead of keeping just Mickey Mouse or just the profitable works under copyright for longer, they fundamentally changed the structure of copyright.


Monday, February 05, 2007

Dungy & company: Super Bowl champions

Last night the Indianapolis Colts played an impressive game, knocking off the Chicago Bears 29-17. Becky and I had a great time watching the game (and celebrating her birthday) with a large group of friends over at the Reaves’ home (now that I’ve seen a Super Bowl in high-def, I don’t think I can ever watch it in low-def again ). I was pulling hard for the Colts since I wanted Manning to get his first Super Bowl ring, and because of the tragedy that the Dungy family had been through the previous year.

This was a momentous Super Bowl because it was the first time either team (in this case both teams) had a black head coach. But that distinction wasn’t what Dungy thought was most important. Here’s what he said during the trophy ceremony:
I'm proud to be the first African-American coach to win this, but again, more than anything, Lovie Smith and I are not only African-American but also Christian coaches, showing you can do it the Lord's way. We're more proud of that.
Dungy and Lovie are both roaring lambs, guys that exemplify Christ in a demanding, stressful job. There's a website that talks about their story and their search for something "beyond the ultimate." Now if only we could get one of them to take the head coaching job in Dallas... smile

Best Super Bowl commercial? My votes go to Kevin Federline ("Federline! What? Fries!") and the slap commercial.

Thursday, February 01, 2007

Nelson awarded NSF Early Career Development Award

Congratulations to my advisor, Dr. Michael L. Nelson, for being awarded the prestigious Early Career Development award from the National Science Foundation (NSF). Michael is the first winner of the award in the ODU CS department, and only one of four ODU faculty to have ever received the award. The award comes with a $541K grant which will be used for research in digital preservation and fixing up his old Galaxie.

The ODU website has a detailed article about the award, and I’ve put excerpts here:
The grant of $541,000 to Nelson rewards his out-of-the-box thinking about tactics to preserve digital data. He says that the ever-more ingenious Internet strategies used to disseminate spam e-mails may someday be employed to preserve data. The title of his project is “Self-Preserving Digital Objects.”

“Can we create digital objects that preserve themselves?” Nelson asks. “I want to explore this.” He said that e-mail spam and viral videos are the best current examples of the approach he proposes.

Mischievous e-mail or a humorous video clip can “live in the Web infrastructure with minimal hierarchical control,” he said, and that is precisely how he plans to preserve digital objects containing data for technical papers, historical documents, Web pages and the like.

“I’m going to investigate if these properties can be applied to content other than pop culture ephemera,” he said.

Most approaches to preserving digital information involve putting “dumb” objects in “smart” repositories. But, Nelson noted, “This reveals an implicit assumption that the repository is going to be long-lived.” A repository—which he sometimes calls a “fortress”—could be a digital library maintained at a university or the host memory of Yahoo.

The “deadly embrace of repositories” is a phrase coined by computer scientist John Kunze at the University of California, and Nelson likes to repeat it. “Information goes in, but is often difficult to extract. I especially like that phrase as a succinct, vivid description of repositories,” he explained.

Nelson said he is not advocating abandonment of repositories or other conventional digital preservation techniques, but he believes an alternative is needed. “These repositories are expensive and they are complicated software systems that require preservation themselves. I’m interested in digital objects that can live longer than their repositories, in information that can live longer than the people or organizations charged with their preservation.”
...

In just a few years, Nelson has become an internationally recognized expert in the areas of digital libraries and digital preservation, said Maly. “We are extremely proud of Professor Nelson winning the prestigious NSF Career award and we look forward to integrating the results of his Career award into both our graduate and undergraduate curriculum.”

Nelson, a former NASA employee, earned master’s and doctoral degrees at ODU before joining the computer science faculty in 2002. In addition to this grant, Nelson has been principal investigator or co-principal investigator on eight grants totaling $1.8 million.

Maly and Nelson are members of the Digital Library Research Group @ ODU. The group has developed several Web services that are used internationally and was a founding member of the Open Archives Initiative (OAI). The new “mod_oai” module is housed at ODU under Nelson’s direction. Funding for the group comes from many government agencies involved in data preservation.

ODU is one of only a dozen universities in the United States that offer courses in digital libraries.

Technologies of Google at ODU

This semester I’m sitting in on CS 891 - Technologies of Google Seminar led by my advisor, Michael L. Nelson. Here’s a brief description of the class:
This seminar will focus on Google and the technologies they have created or adopted to build their enterprise into what it is today. Although many of the technologies we will study are applicable to all search engines applications, we will focus on a breadth-first discussion of Google's technologies rather than a depth-first examination of a single research area. We will cover 3 basic areas: (1) initial contributions to crawling and ranking; (2) information retrieval applications; (3) custom infrastructure for deploying web-scale applications.
Students are presenting 20 papers throughout the semester that Michael has accumulated. Last night the papers The Anatomy of a Large-Scale Hypertextual Web Search Engine and The PageRank Citation Ranking: Bringing Order to the Web were presented to the class. These are the first papers on Google from founders Sergey Brin and Lawrence Page who were graduate students at Stanford at the time. What’s interesting to note is the second paper on PageRank is only a technical report: apparently it was rejected by SIGIR for not rigorously evaluating PageRank (according to a talk by Monika Henzinger). Brin and Page never did the extra work required to get it published, yet it is obviously one of the most influential papers about ranking search results ever written.

If you’d like to see the presentations made in this class, the slides will be posted to the class website on a regular basis.

Monday, January 29, 2007

Defusing the Googlebomb

A few days ago, Google made some changes to their ranking algorithms to reduce the practice of Googlebombing. A Google bomb is basically a prank to manipulate the ranking of pages in Google’s search engine. It involves getting a lot of people to put a link to a particular web page on their site with the anchor text they want associated with the web page.

For example, if I wanted a search for “basketball stud” to show my blog as the first result, I’d get as many people as I could to place a link on their website that looks like this:
<a href="http://frankmccown.blogspot.com/">basketball stud</a>

Then when Google crawls the Web and sees a large number of links that look like this, they would begin to favor this page over the rest when users search for “basketball stud”.

One of the most famous Googlebombs involves a search for miserable failure. While this used to show George Bush’s web page first in the results, it now brings up more relevant results. Danny Sullivan has written a good article about this.

How did Google reduce the affects of Google bombs? They’re not giving particulars, but they have admitted it’s purely automated. My guess is they analyze several factors:
  1. When and where was the link first found? Possibly Google tracks the growth of particular links.
  2. Does the link make sense for the web page or website? A red flag might be raised when a website about hacking points to a government web page when none of the other links do.
  3. Is the target page actually "about" the anchor text? If the words "miserable failure" aren't on the target page, it could be a bomb.

If you’d really like to dig into this subject, here’s a master’s thesis on the topic.

Thursday, January 25, 2007

Wikipedia: nofollow and noMSedit

Some new news from the world of Wikipedia:
  • Minor: All external links from Wikipedia are now using the NOFOLLOW attribute. This attribute tells web crawlers like Google that the link has not been vetted, so it will not be used in their algorithms to artificially bolster the ranking of some pages. Wikipedia’s action will seriously reduce the amount of link spam that currently plagues many entries.

  • Major: Microsoft has attempted to hire Rick Jelliffe, chief technology officer of XML tools company Topologi Pty. Ltd., to “correct” Wikipedia entries on ODF (OpenDocument format) and OOXML (Microsoft Office Open XML). You can see Rick's original post about the offer here. Apparently Wikipedia is keeping Microsoft employees from making the edits themselves, so Microsoft thought a third party could update the entries that apparently shed a negative light on Microsoft’s format. This astroturfing blunder has created quite a few waves.

Tuesday, January 23, 2007

Countdown to baby

I’ve made a new addition to the right side of my blog- a countdown until Ethan's expected delivery date on April 4. Maybe this should be labeled Countdown until life as we know it changes dramatically. I can’t wait until he’s here. Becky can’t wait until she can breathe again and bend down and touch her toes. At the same time, I’m a bit spooked because I have experiments to perform, papers to write, and a dissertation to compose. Although I thought I’d have it all done by the summer, it’s not looking that way anymore. But hey, having my boy here will be a great blessing.

By the way, what an amazing Colts/Pats game on Sunday! If I can’t have my boys in the Super Bowl, at least I can cheer for Manning. That guy's pretty good... if you like a 6-5, 230-pound quarterback with laser-rocket arm.

Friday, January 19, 2007

No more searching for you: Google drops the SOAP

In case you were asleep at the helm like I was, Google has pulled the plug on their SOAP-based web search API. On Dec 5, 2006, Google stopped giving users new API keys. They claim the API service will continue to run, but without a method for obtaining new keys, it essentially becomes worthless (API keys can't be shared since they are tied to a specific individual's Google account, and I can't let you run my application unless you supply it with your own key).

Google has decided their AJAX search API is the wave of the future. But why the "odd move"? I think Jason Lefkowitz summed it up best:
Today, though, Google isn’t about search. It’s about displaying ads. And in that context, an open API makes no sense — the developer can reformat the search results, and even show them (gasp) without ads!

Hence the “AJAX API”, which forces you to take the ads along with the search results. You can’t really do much with it, but it does create a new place for Google to show ads on — your blog/site/Web app.
I don’t have a problem with Google focusing on their AJAX search API... I’m sure it’s very useful in many contexts, but I do have a problem with them abandoning their SOAP search. Not only is Google putting the smack down on the SEO business (one of their intended victims, in my opinion), they are hurting us web researchers who depend on automated methods of querying Google.

I can point to a huge stack of academic papers that, without an effective method of automatically querying Google, are un-reproducible (Google- do you really want everyone to go back to page-scraping?). And it’s really hurting my research: Warrick will not work for new users without API keys. I’ve spent lots of time writing wrappers around the SOAP API code, now I’ll have to redo most of when I find an effective method of accessing Google’s cache. Until then, you can kiss your lost website goodbye if Google is the only one who has cached it.

It sometimes appears that have a love/hate relationship with Google. Yesterday I was singing it's praises, today not so much. In honor of the SOAP API, I’ve put together a brief timeline for us all to reflect upon:
  • Pre 2002 - Page-scraping is the norm, and there is great frustration.
  • 2002 – Google launches the first search engine API, and there is great rejoicing.
  • 2002-2005 – Researchers use the API to for all sorts of interesting experiments, SEOs do their best to reverse engineer PageRank, new services are built, books are written, and, despite many technical difficulties along the way, there is much satisfaction.
  • 2006 – Google tightens the lid on extra queries per key, and there is much displeasure.
  • Late 2006 – Google refuses to give new API keys, and there is much sadness and anger.
  • Late 2007 (My prediction) - Google’s SOAP API breaks, no one fixes it, and there is no surprise. RIP

Update on July 27, 2007:

Google has just released an academic API for researchers: University Research Program for Google Search. Now that's more like it.


Update on Sept 30, 2009:

Google has finally killed its SOAP Search API.

Thursday, January 18, 2007

Store your data in a search engine cache

I have say it… Google is the best thing since indoor pluming. This morning I was wondering if anyone has been writing about my program Warrick, so I did a quick search of “warrick mccown”, just to see what would pop up. On the second page of results, I found a link to a paper that was published in October 2006: Using Free Web Storage for Data Backup.

The paper was written by some researchers from Stony Brook University who have developed two backup systems: CrawlBackup for storing files in a search engine’s cache, and MailBackup for storing files in the mailboxes of Internet email providers. Their work is remarkably similar to ours, and it almost makes me wonder if our place is bugged.

This paper is the first to actually cite Warrick. The paper also cites an interesting blog posting from Dec 2005: How the Google Cache can save Your A$$. OK, not the best title in the world, but it’s the only pseudo-article I've found where someone has documented using the Google cache to recover a lost website. In this case the guy accidentally deleted 30 articles from his website and used Google’s cache to recover them. It was just a few months earlier that I had finished work on Warrick which could have automated the process for him (at least he only had to recover 30 pages!). He also used the Internet Archive to recover a client’s website a few years ago.

So I’m really glad to have found these related resources. What’s unfortunate is that finding related work is often much more difficult than a simple Google search (or even a Google Scholar search). Google may produce a few gold nuggets, but it also produces a lot of false positives: why is the third result a production chart for Shaun Alexander (I really don't need to be reminded of the Cowboys loss in the playoffs)? The word Warrick is only used once, and it’s in a drop-down list box! And lest we forget, Google does not have the entire Web indexed. If I really wanted to be diligent I'd also use MSN, Yahoo, or a metasearcher like Dogpile.

Anyway, I’m still hoping someday for a Google SuperScholar system that takes all my papers, notes, etc. and figures out what is most related on the Web and in every digital library in existence and sends me weekly updates with precise summaries of why the information found is relevant. Maybe it should be called Google ScholarHeaven.
smile

Wednesday, January 17, 2007

Scratch is for kids!

MIT researchers have recently released a new version of Scratch, a graphical programming language targeted to children. Scratch allows you to create interactive programs that use animation and sound. I especially like the programming language (shown on the right) that allows you to drag and drop constructs to create a program. It sure beats moving a turtle around the screen (how I miss those days at computer camp... smile)!

Thursday, January 11, 2007

2007 is poised to be the year of spam. Although Bill Gates thought we’d have the problem licked by now, the problem is only getting worse. I’m now getting approximately 50-60 spam emails per day, and although my spam filter catches a lot of it, a new form of spam is regularly beating the filter: image spam. Spammers are now using attached images to replay their spammy messages. They work like a captcha for your filter- since your filter is good at reading text but lousy at reading images, there’s little you can do to stop well-designed image spam.

What really bugs me is when companies that spam claim they don’t. For example, I received the following spam about 10 times over the past few days:


Notice the clever text to the side of the image. They change that each time to keep my filter from figuring out this is spam. Now if you'll visit their website, you’ll see that they have a link that allows you to report spam:
Pharmacy operates a strict anti-spam policy. We do not tolerate unsolicited advertising messages. We will actively pursue anyone engaging in spamming activities! This includes email, icq, instant messengers, chat rooms, message boards, newsgroups or anywhere else where commercial postings are prohibited. We will take appropriate actions against spammers that will result in loss of services and accounts closure.
Sure they will... they ask for your name, email address, and phone number, just so they can spam you some more with unsolicited phone calls while you’re eating dinner!

Once you submit your information, they reply “Thank you for your patience.” I think these guys may be qualified for the ninth ring.

Tuesday, January 09, 2007

Use Wikipedia to make your computer smarter

Still on the Wikipedia kick... Researchers from Technion-Israel Institute of Technology are using Wikipedia to give computers context information and make connections between different words. See the article here. For example, when a spam filter encounters a word like “B12” and needs to determine if the email should be marked as spam, the filter currently doesn’t know that B12 is a vitamin (the subject of many spam emails) unless the email also uses the term vitamin. But by examining the Wikipedia article on B12, the spam filter could be smart and deduce that an email with B12 is trying to sell vitamins. The same information could have been obtained by searching for B12 using a search engine, but the results aren’t necessarily vetted. That’s the Wikipedia advantage.

Monday, January 08, 2007

Reading lists from Wikipedia

Alexander D. Wissner-Gross, a physics Ph.D. student at Harvard, presented his paper this summer entitled Preparation of Topical Reading Lists from the Link Structure of Wikipedia at ICALT'06. Wissner-Gross shows how an algorithm based on PageRank can be used to generate background reading lists from Wikipedia. I especially like this paper because it is the solution to a real teaching problem that Wissner-Gross encountered when preparing to teach one of his courses: how can we automate the time-consuming process of generating a quality reading list for a class?

Update on 1/30/07:

Wissner-Gross emailed me this morning with the web address of the reading list engine: http://www.wikiosity.com

I got some interesting results for Digital preservation. Although Digital obsolescence popped up first, some irrelevant results like Vanderbilt University and University of Virginia also popped up. A search for web crawling brought up 2003 as a result. I'm not sure if these lists would be more useful than if I looked directly at the See also section, but it's still an interesting idea.

Tuesday, January 02, 2007

Wiki my search

Be on the lookout... a wiki-inspired search engine called Wikiasari (no web address yet) is going to be launched early this year. Since Wikipedia founder Jimmy Wales is behind the project, it’s already starting to create some noise.

It sounds like a great idea: apply the wisdom of crowds to search engines results. Of course this is already what search engines are attempting to do when they track which search results you click on or use link analysis (how many and what types of links are pointing to a page) to determine what are the best results to a particular query.

The problem will be eliminating the rich-get-richer phenomenon on the Web which makes it difficult for new pages to rise to the top. You can image a new page about Britney Spears that is of high quality (can a page about Britney be high quality? ), but it won’t be displayed to searchers since the top 10-20 results already have been voted to be the best.

And how do you get users to evaluate the relevance of results on the third or forth set of results? Studies have shown users rarely go beyond the first page or two of results. Some very interesting problems indeed.

Monday, January 01, 2007

List of banished words for 2007

Lake Superior State University has once again posted it’s list of banished words for the new year. Last year they banned my nickname Dawg McCown, and now my friends can’t call me i-Frank or refer to me and my wife as Frecky… how disappointing.

Here are the words and phrases to be banished from the Queen’s English in 2007 for mis-use, over-use and general uselessness:
  • Gitmo
  • Combined celebrity names
  • Awesome
  • Gone/went missing
  • Pwn or pwned
  • Now playing in theaters
  • We're pregnant
  • Undocumented alien
  • Armed robbery/drug deal gone bad
  • Truthiness
  • Ask your doctor
  • Chipotle
  • i-anything
  • Search
  • Healthy food
  • Boasts
Happy new year, everyone!

Saturday, December 30, 2006

Favorites of 2006

Here are some of my favorite things from 2006.

On-line laughs:
  1. The most-viewed YouTube video of 2006: The Evolution of Dance by Judson Laipply
  2. Possibly the worst recording of O Holy Night, ever
  3. David Brent's Microsoft training video
  4. The Iraq Report with subtitles
  5. Two sons try to take a Mother’s Day photo (these guys must know my brother and me)

TV commercials:
  1. Liberty Mutual "pay-it-forward" commercial
  2. LA County Fair - "Duh, Ashley, all wool comes from a cow..."
  3. “That should kill him…” Ameriquest doctor commercial
  4. Peyton Manning supporting his team

Movies I saw:
  1. The Prestige - It had me on the edge of my seat the entire time
  2. X-Men: The Last Stand - I hope this won’t be the last of the series
  3. Facing the Giants - Created by a church with no professional actors, it's a very moving and inspirational film
  4. Invincible - Almost makes me want to be a Phily fan
  5. Casino Royale - A little on the violent side, but probably the best Bond yet
  6. Flags of Our Fathers - War is tragic

Books I read:
  1. The Language of God by Francis Collins - Collins does a great job of distilling the myth that evolution = atheism
  2. Freakonomics by Steven D. Levitt and Stephen J. Dubner - Some interesting things to think about
  3. Sacred Marriage by Gary Thomas - Rather than make us "happy", marriage is designed by God to make us more holy
  4. Finding God at Harvard by Kelly Monroe - A series of encouraging stories from various Harvard alums about their path to faith

Websites I love to visit:
  1. Wikipedia.org – Sometimes you have to sort through a lot of “truthiness”, but it’s an extremely useful tool for answering the question “What is X?”
  2. YouTube.com – It feeds my addiction for funny commercials
  3. Google Scholar – OK, I don’t love to visit this site, but it's probably the #1 reason why getting my Ph.D. will take less than 4 years. It’s much larger and faster than Citeseer and far easier to use than Windows Live Academic

Friday, December 29, 2006

Goodbye to 2006

We just returned home from visiting my parents in St. Louis. Sara was able to fly up too, but John had to stay in Dallas and work the 26th. We had a really good visit... plenty of R&R and quality time with the fam. It was fun having us all together and watching the Broncos win and the Cowboys… well… we’ll try to forget about that game. I worked just a little- fixed and started several scripts which needed to run over the break, but a power-outage on the 26th undid all my work.

Although the Bean is not yet born, he got a majority of the gifts, including a really cute Broncos outfit with matching booties and bib. Sara also got him an “I love Daddy” jumper and a hand-made outfit from Bolivia. Can’t wait for him to be here!

In our family fantasy football league, I whooped up on my dad and sister to become the first family league champion. smile In the Bayside League, I was 6 out of 14. Not as good as my second place finish last year, but it was fun none the less. This is probably the last time I’ll play in the Bayside League since we won’t be living here next fall. Hope someone will pick it up after I leave.

On a sadder note, I found out that Marilyn Fowler passed away on Christmas Eve. I used to go to church with Al and Marilyn in Searcy. She had been battling leukemia and had completed her third round of chemotherapy and was so close to beating it. I’m glad she’s out of pain now, but it’s got to be hard on Al and the family.

Only 3 days left of 2006. This has been a really good year. A few highlights:
  • Apr 7: Passed my Ph.D. candidacy exam
  • May: Finished taking the last course of grad school
  • July 19: Becky and I celebrated our third anniversary
  • July 28: Found out Becky was pregnant!
  • Becky won 2 awards at work (Aug and May)
  • Oct 8-14: Our first cruise- to the Bahamas
  • Nov 17: Found out we’re having a boy!
  • Nov-Dec: Taught a class on Christian Apologetics (toughest class I’ve ever taught)
  • Today: Posted my 100th blog post


A few goals for 2007:
  • Name our newborn son
  • Change my first diaper
  • Move back to Arkansas
  • Begin teaching at Harding University
  • Graduate with my Ph.D.

Some books I want to read in 2007:

May God richly bless us all in 2007.

Saturday, December 23, 2006

End of the semester at ODU

I’m in my office at ODU overlooking a very barren campus. Everything is closed down, and all the faculty, staff, and students have started their vacations (mine starts tomorrow when the woman, the bean, and I fly out to see my parents in St. Louis ).

A few days ago I took some photos of the many changes going on around campus:




Here's my building (ECS). Nothing has changed here, but I thought I'd get a picture of it anyway. My office window is 3 floors up, left of the lobby windows.
This is what the new dorms look like from outside my office window.
These are the new dorms that are going up across from the ECS building. They are on top of what was the parking lot to the gym. Notice the stylish port-a-potties.This is a close-up of one of the finished dorm. The other will be finished in a few months.
The old gym has been gutted... it looks like a hurricane swept out the old basketball courts.These are two new structures next to the gym. I think they're going to be the new indoor tennis courts, but I could be wrong.



ODU is also building a new bookstore and a research center near the village.

Some interesting, news-worthy items about the 76 year old university:
  • According to the New York Times, ODU is the most racially diverse four-year institution in the country.
  • ODU is getting a football team in 2009- a little late for me, but just in time for the class of '09.
  • Sometimes when the university upgrades their software, you get thousands of dollars loaned to you with 0% interest.

Tuesday, December 19, 2006

Tom cat

Tom was put down yesterday after failing to respond positively to medication for a blood clot. He was a great old cat. We reluctantly adopted Tom when he was only a kitten; he followed me home one night after I’d been out TP’ing the neighborhood. He moved with my parents and sister from Denver to Kansas City, Bella Vista to St. Louis. We will miss you, buddy.

Monday, December 18, 2006

Linkrot on Tamalak's Realm

Evan Williams finds an ironic interview on link rot.

Saturday, December 16, 2006

Rhonda Frasier - See you on the other side

I just learned today that a good friend of mine from my post-college years in Denver has passed away. The following is from the Harding University alumni newsletter:
Rhonda Frazier (’94)

A celebration of life will be held at 1 p.m. Saturday, Dec. 16, at the Church of Christ in Prineville, Oregon, for Rhonda Gaylene Frazier of Madras, formerly of Lane County, who died Dec. 8 of Alzheimer's disease. She was 34.

She was born Aug. 23, 1972, in Heidelberg, Germany, to Doug and Dawn Kimball Frazier. She worked in fashion merchandising. She was a member of Chi Omega Phi and the women’s track team at Harding University.

Survivors include her parents; a son, Clay; a grandmother, Azalea Kimball Hatfield; a sister, Janelle Strong of Prineville; and a brother, Justin of Eugene.

Arrangements by Autumn Funerals in Redmond. Remembrances to Clay Frazier Trust Fund at Mid-Oregon Credit Union.

Rhonda and I were buddies when I graduated from Harding in '96 and moved back to Denver. She had also made the move to Denver after graduating from Harding, and she lived in the same apartment complex as Mark Story, another good friend of mine. I was probably at either her place or Mark's at least once a week.

I lost touch with Rhonda after I moved back to Searcy and she moved to Oregon. One Christmas she sent me a Christmas card with a photo of her and her son, but I haven’t talked to her in years. She was a Christian and a very thoughtful friend, and I pray that the Lord takes her home and provides comfort to her family.

Here’s a photo of Rhonda and I playing in a mud volleyball tournament in 1996.



This photo was taken at a Christmas party in 1996. Rhonda is on the far left.

Monday, December 11, 2006

Link rot in CACM

I was really surprised this afternoon to see an article in the Communications of the ACM that cited a cached URL from the MSN search engine in place of a missing web page:
The link to the U.S. Secret Service “Operation 4-1-9” report at www.secretservice.gov/alert419.htm appeared to be broken when this column was written, but cached copies remain available (for example, cc.msnscache.com/cache.aspx?q=3910458378891〈=en-US).
Communications of the ACM, Volume 49, Number 12 (2006), Page 18.
The editors of CACM may not be aware of this, but search engines do not keep cached copies of pages long. In fact, they will often purge their caches of any web page that returns a 404 when crawling. (You can read my paper on an experiment which illustrates this.) Citing a cached page from a search engine should never be done in academic writing. Instead of citing just one broken URL, CACM has now cited two.

If you are interested in learning more about link rot and how to combat it, check out this Wikipedia article that I contribute to.

Speaking of link rot, Baden Hughes of the University of Melbourne has recently published a study entitled Link? Rot. URI Citation Durability in 10 Years of AusWeb Proceedings. (Not sure why he used a question mark in his title.) He used many of the methodologies that I used when examining link rot in D-Lib Magazine last year. Turns out AusWeb URLs have a much lower half-life (6 years) than D-Lib article URLs (10 years). This is probably because authors of D-Lib articles are more aware of link rot than authors in other professions.

I’m curious if any other on-line magazine or journal can beat D-Lib’s 10 year half-life. I have a suspicion JMIR articles could since many of them use WebCite.

Friday, December 08, 2006

Agassi and Blake in Norfolk

Last night I finally got to see Andre Agassi in person at Anthem LIVE! at the Constant Center. Anthem LIVE! was hosted by James Blake to raise money for cancer research (James’ father died of cancer several years ago). Becky surprised me with a ticket to the event a few weeks ago. She couldn’t go with me though since this weekend she was in her friend Amy’s wedding in Memphis.


This is only the second time I’ve seen a live professional tennis match- the first was years ago in Denver when I saw an exhibition with Jimmy Connors and David Wheaton. And then there was the time I ran into Steffi Graf (now Andre’s wife) years ago in an elevator in San Antonio (nice legs ), but she wasn't exactly swining a racquet. So I was really excited about the event.

Anthem LIVE! opened with Boyd Tinsley (of the Dave Matthews Band) and the Blake brothers playing doubles against the Bryan brothers. Bob and Mike Bryan are currently the number 1 ranked doubles team in the world. After 3 games, James and his brother Thomas took on the Bryan brothers in a lively 8 game set with the Bryan brothers ending up on top.

After the doubles match, I ran down to the entrance way where Agassi was going to enter and took some snaps with my phone (Becky took our digital camera with her to Memphis). The crowd went nuts when Agassi came out of the tunnel. I was about 10 feet away despite the usher who kept yelling at me to return to my seat. Common- I’m not going to stab the guy. smile

Andre and James played an entertaining two set match, ending in a win for James after a third set tie-breaker. There were some great shots by both players, and the crowd loved it. Most people stayed to the very end (around 10:30 pm).

An interesting fact I learned last night: Agassi and Graf are the only two tennis players to have won every Grand Slam tournament and a gold metal in tennis. Can you image the genes their kids must have?!

I may not have gotten to see Agassi play in the US Open or Wimbledon, but this was pretty awesome. Now if I can just get Becky to let me name our son Agassi...

Saturday, December 02, 2006

Forum posts lost from Beryl Project

This week I received an email from Paul Dorman who was wanting to use Warrick to recover http://forums.beryl-project.org. Beryl is a combined window manager and compositing manager that runs on top of Xgl or AIGLX. According to Paul, the site was lost when a hard drive crashed and no backups were available. There is some discussion of recovering the website here.

Apparently one of the forum members named TreviƱo used Warrick to recover quite a few of the pages and has them hosted here. One of the forum members praised Google for "their excellent off site backup that is watchin us all." Kinda gives you a warm cozy feeling.

Friday, December 01, 2006

Hasta la Vista

Microsoft finally released Windows Vista yesterday. Robert Vamosi has written about the Five reasons to love (and hate) Windows Vista, and Joel Spolsky has written an interesting bit about the complexity of the Vista's off button.

I’m looking forward to giving Vista a try, but not anytime soon… my current laptop runs like a dog with XP, much less with Vista. Better to purchase a new computer in 2007 with the OS already installed.