Wednesday, May 20, 2009
Java Sitemap Parser
The Java Sitemap Parser was the final project for my Search Engine Development class. I talked about the project a few weeks ago and how prevalent Sitemaps are becoming. Originally we wanted to add Sitemap support to Nutch, but developing just the parser proved to be quite a task. By releasing it as an independent project, I'm hoping Nutch, Heritrix, and other open-source crawlers will integrate it into their systems.
Wednesday, April 22, 2009
Nutch, Sitemaps, and Google's findings
The reason I mention our Sitemap project is that WWW 2009 is meeting in Madrid this week, and a paper entitled Sitemaps: Above and Beyond the Crawl of Duty is being presented today by Uri Schonfeld (UCLA) and Narayanan Shivakumar (Google). This is the first paper to report on widespread usage of Sitemaps in the Web using Google's crawling history.
Schonfeld & Shivakumar report that Sitemaps were used by approximately 35 million websites in late 2008, exposing several billion URLs. 58% of the URLs included last modification dates, 7% included change frequency, and 61% a priority. About 76.8% of Sitemaps used XML formatting, and only 3.4% used plain text. Interestingly, 17.5% of Sitemaps are formatted incorrectly.
The figure below represents how many URLs Google discovered via Sitemaps (red) vs. regular crawling (green) for cnn.com. Notice that on any given day, more URLs could normally be discovered via Sitemaps.

Another interesting figure (below) shows when a URL was discovered via Sitemaps vs. regular web crawling for cnn.com. In most cases URLs were discovered at the same rate, but there are a number of them (dots below the line) that were discovered via Sitemaps much earlier than web crawling.

CNN's website is not typical. Schonfeld & Shivakumar report that in a dataset of 5 billion+ URLs, 78% were discovered via Sitemaps first compared to 22% via web crawling.
The paper also describes an algorithm that can be used by search engines to prioritize URLs discovered via web crawling and Sitemaps as well. I've covered the high-lights, but I recommend you read the paper if you're interested in some of the finer details.
Monday, February 16, 2009
Specifying canonical URLs
Essentially, the new attribute value will allow a webmaster to tell a web crawler to ignore a page if it is accessible from another URL. So if a I have a single page that is accessible at URLs A, B, and C, I can tell the web crawler that URLs B and C are pointing to the same content as A by placing the following code in the head element of the page:
<link rel="canonical" href="http://foo.com/A" />
When the web crawler grabs the pages using URLs B or C, it will find the given canonical URL A in the header and therefore ignore the contents of the pages since they duplicate page A.
Of course the entire mechanism requires a willing and competent webmaster to implement it. Webmasters who are very concerned about SEO are likely to use it since it will help bolster the PageRank of certain pages. But the rest of us who aren't concerned about our rankings can safely ignore this new functionality.
See also rel="nofollow".
Monday, July 21, 2008
Webcrawlers wanted
I'm not a huge fan of advertisements, but this ad on Facebook, sponsored by AddGooro, is just too cool. If any of my search engine students are reading this... they're hiring.
Friday, July 04, 2008
Fav5
- The Deep Web is getting a little shallower: Google has recently improved their ability to crawl Flash content. It used to be that a website with a Flash interface was practically invisible to search engine crawlers, but now Google can find the links to other pages within the Flash program and even indexes the program's textual content.
- If you were late signing up for a free Yahoo email address, you might have ended up with frank1837abc@yahoo.com since all the good names were already taken. But now Yahoo has opened up two new domains under "ymail" and "rocketmail". Although Yahoo is still the email market leader with 266 million worldwide users, they want to stay ahead of Microsoft who is a close second with 264 million.
- Are Google, Yahoo, and Microsoft censoring themselves more than they must in China? According to a report by Citizen Lab, some sites are blocked by some search engines and others let them through. In a test to see how many questionable sites were blocked, Google had censored 15.2% of the sites tested, Microsoft censored 15.7%, Yahoo 20.8%, and Baidu (the most popular Chinese search engine) 26.4%.
- A Harvard marketing professor takes on Chris Anderson's "Long Tail" theory by analyzing real data about online video rentals and song purchases. Her conclusion is that maybe we aren't as individualistic in our taste as Anderson suggested.
- In the spirit of July 4, read about how LANL scientists are making fireworks a little "greener".
Friday, April 11, 2008
Fav5
- Google recently announced their newest attempt at crawling the Deep Web. When their crawler sees a form on a "high-quality site", it enters words and selects various checkboxes and radio buttons. The form is then submitted, and the resulting page is crawled. Very cool. Google also announced the Google App Engine. It's a service that allows you to run your web applications on Google's infrastructure so it can grow to accommodate a large amount of traffic. It's a free service until it exceeds a set disk space and bandwidth quota. Unfortunately, there's a limit of 10,000 developers, and I wasn't quick enough to sign up, so I'll have to wait until they increase the limit.
- Joel De Young, one of our Harding CS graduates, was recently interviewed about his company's newest game: Penny Arcade Adventures: On the Rain-slick Precipice of Darkness. Sounds like Joel's HotheadGames is making some buzz.
- A study by Univ of Minnesota researchers shows there are two indicators that determine how good of an answer you will receive when posting to question-and-answer sites like Yahoo! Answers: how much you pay, and how many responses you receive.
- Jansen, Booth, and Spink have developed a system which can automatically classify a search query as navigational, transactional, and informational. They used their classifier on a large dataset of search engine queries and were accurate 74% of the time. Based on their results, approximately 80% of all search engine queries are informational, 10% transactional, and 10% informational.
- Congratulations to the Harding Programming Team on their first place finish at CCSC-MS. We showed them Razorbacks how to program...

Friday, March 21, 2008
Fav5
- Technology Review has just released its top 10 emerging technologies of 2008, those technologies that are most likely to make a huge impact on the way we live. I personally find offline Web applications to be most interesting. This is an area Adobe and Google are vigorously pursuing.
- Joan Smith (my former office-mate in graduate school) and Michael Nelson (my former Ph.D. adviser and pioneer of digital preservation
) just published an article in D-Lib Magazine showing a year's worth of crawling activity on some specially designed websites. Make sure you check out the animated images in Tables 3 and 6- they show the crawling behavior of an aggressive Google and timid Yahoo and MSN. - Ethics 101: You're a popular search engine that allows users to see emergent news from a number of sources. But you're housed in a country where the authorities often seek to squelch information that doesn't hold to the party line. If you show news stories that make the authorities look bad, your search engine will be shut down or blocked. What do you do?
- Joel Spolsky writes an excellent article about IE 8 and the many difficulties of creating software that supports "standards". This is required reading for all you software engineers and CS students. (The article is a little long, but it's worth it.)
- While we're on the topic of web browsers, almost a year ago Steve Jobs announced that the next version of Safari (Apple's web browser) would run on Windows. Now Apple is using iTunes to push it out to the Microsoft world. My guess is that many unsuspecting users are going to think they need Safari to run iTunes since it appears in the same dialog box they always see when needing to install an iTunes update. Many users are going to downloaded it and puzzle over the new icon on their desktop. This should be interesting...
Thursday, February 28, 2008
Fav5
Abilene Christian University, one of our sister schools, has decided to give all incoming freshmen next year an iPod. Is this a gimmick or the future of secondary education? Let's hope the move helps boost their enrollment and shrink their $3 million dollar budget deficit.- Google Sites has just launched as part of Google Apps. Essentially it allows you to create a web page or website via wiki tools that anyone in your group or organization can modify. I haven't tried it yet, but it looks promising.
- Matt Cutts always has some interesting items to read: Subscribed links, how search engines should deal with NOINDEX, and Blogger Play.
- Good news for those of us that work in IT: salaries are expected to grow in 2008. Will more money start getting freshmen to major in CS?
- Java or .NET? According to the Tiobe Programming Community Index, Java is still ranked number 1. But in a recent survey by Info-Tech of 1,900 companies, 12% of enterprises focus exclusively on .NET , and 49% center primarily on .NET. This compares to 3% and 20% of companies that are focused exclusively or primarily on Java.
Saturday, December 08, 2007
Fav5
- Good news for the CS majors: According to Bureau of Labor Statistics, computing-based jobs are expected to grow 24% over the next decade, the highest rate of growth of any "professional and related occupations."
- Maybe the grass isn't greener on Mac's side of the fence.
- Google has just released an impressive REST-based chart API. It allows 50K queries per day and allows you to create pie graphs, line graphs, scatterplots, bar charts, and Venn diagrams.
Just for fun-- my Family League fantasy football record (8-5) as of today: - The Heritrix 2.0 web crawler is here. I'll be using this crawler (most likely) in my search engine class next semester.
- The planet's latest enemies: divorced observers of Hanukkah. Wow... no one is safe blame these days.
Monday, October 08, 2007
Crawling behavior of Google, MSN, and Yahoo
The photo below is a visualization of Yahoo's crawling behavior (Yahoo was the most active of the three crawlers). You can see an animation of the tree growing each day here.

We performed a similar experiment at ODU back in 2005 except we removed web pages every day from our collections to see how long they would stay cached by the search engines. And Joan, a fellow Ph.D. student at ODU who worked on the previous experiment, has been performing another experiment over the last several months with a very deep and wide collection of pages. (You can find links to the pages at the bottom-right corner of this page labeled Joan1-4).
What I found really interesting about this experiment was the methodology used to created the web pages. They altered the pages by appending the crawl log of the pages and allowing anyone (including spam bots) to make comments that appeared on the pages. This kept the page contents unique and changing to entice more crawls.
I really hope the authors of the study will submit their findings for publication in a peer-reviewed conference or journal.
Friday, October 06, 2006
Become.com's web crawler
Today a member of the Heritrix list serve pointed everyone to an article on Sun’s website that discusses Become.com’s web crawler. The article dates back to August of 2005, so it’s a little dated. I couldn’t find any updated information on the crawler, but apparently it is proprietary, and the source code will likely never see the light of day.Become.com actually developed 2 crawlers in 2004- one written entirely in Java and the other mostly Java with some C++. The article states that the crawlers "may be the most sophisticated, massively scaled Java technology application in existence."
The article doesn’t mention anything about Heritrix, a crawler which is also completely written in Java. Although Heritrix doesn’t currently have a distributed architecture, it could still be deployed in such an environment. It would be really interesting to see the two crawlers compete at the National Java Crawling Championships.
Wednesday, August 09, 2006
Crawling the Web is very, very hard…
- The website’s root URL not redirect the crawler to a URL that is on a different host. If it does, the new URL should replace the old website URL.
- The website’s root URL should not appear to be a splash page for a website on a different host or indicate that the website has moved to a different host. If it does, the new URL should replace the old website URL.
- The website should not have all of its contents blocked by robots.txt. If some directories are blocked, that’s ok.
- The website’s root URL should not have a noindex/nofollow meta tag which would prevent a crawler from grabbing anything else on the website.
- The website should not have any more than 10K resources.
The restrictions seem very straightforward, but in practice they are very time consuming to enforce. Requirement 2 requires me to manually visit the site. Did I mention not all the sites are in English? That makes it even more difficult. Requirement 3 means I have to manually examine the robots.txt. Req. 4 requires manually examination of the root page, and req. 5 means I have to waste several days crawling a site before I know if it is too large or not.
I guess I could build a tool for req 3 and 4, but I’m not in the mood.
Anyway, I ended up making about 50 replacements (at least) and starting my crawls over again. Now I finally have 300 websites that meet my requirements.In the past I’ve used Wget to do my crawling, but I’ve decided to use Heritrix since it has quite a few useful features missing from Wget. But Heritrix isn’t perfect. I made a suggestion that Heritrix show the number of URLs from each host remaining in the frontier when examining a crawl report:
http://sourceforge.net/tracker/index.php?func=detail&amp;aid=1533116&group_id=73833&atid=539102
The other difficulty with Heritrix is in extracting what you have crawled. I will need to write a script that will build an index into the ARC files so I can quickly extract data for a website. Since all the crawl data is merged into a series of ARC files, it is really difficult to throw away crawl data for a website you aren’t interested in. I could write a script to do it, but at this point it’s not worth my time.
Friday, June 23, 2006
Heritrix - An archival quality crawler
- limiting the size of each file downloaded
- allowing a crawl to be paused and the frontier to be examined and modified
- following links in CSS and Flash
- crawling multiple sites at the same time without invoking multiple instances of the crawler
- storing crawls in an Arc file
The interface is not exactly intuitive, and a near complete reading of the entire manual is required to put together a decent crawl. Of course if you want to use sophisticated open-source software, you usually have to put in some significant effort to get it to work right. Thankfully, several of the developers (Michael Stack, Igor Ranitovic, and Gordon Mohr) have been very helpful in answering some of my newbie questions on the Heritrix list serve.
In learning about Heritrix, I’ve put together a page on Wikipedia. Hopefully the entry will drum up more general interest in Heritrix as well. I was really surprised no one had created the page before.
Thursday, April 13, 2006
URL Canonicalization
http://www.Harding.edu/USER/dsteil/www/abc/../index.htm
could be normalized to produce the canonical URL:
http://www.harding.edu/user/dsteil/www/
Search engines typically use different URL canonicalization policies which makes it difficult for Warrick to tell if URL x from MSN is the same as URL y from Google. I’ve noted some peculiarities in my blog here, here and here. Matt Cutts at Google also discussed some of their canonicalization policies back in Jan 2006.
I have not found much work in the literature about URL canonicalization/normalization. RFC 3986 has some standard normalization procedures that should be done. Pant et al. (2004) has a section about it in their chapter Crawling the Web from the book Web Dynamics. The first paper I’ve seen that deals with the issue head-on is by Sang Ho Lee et al. (2005) "On URL normalization".
I also checked Wikipedia and didn’t find anything about URL canonicalization. I decided to create a page about it and added a reference to it from the web crawler page. That was the first page I ever created on Wikipedia. Proverbs 25:2 – “It is the glory of God to conceal a thing; but the glory of kings is to search out a matter.” I’m no king, but I think God actually delights in our effort to learn about the great world He has created, and I appreciate Wikipedia providing a unique resource for us to consolidate and share our learning.
Friday, February 10, 2006
Some thoughts on robots.txt
There are numerous limitations to robots.txt that have been noted (see Martijn Koster’s article). A few things that are lacking: ability to specify how frequently server requests should be made, the ideal times to make automated requests, permissions to visit vs. index vs. cache (make available in search engine caches).
According to Matt Cutts, Google supports the “Allow:” directive and wildcards (*) which are not part of the standard. The Google Sitemap team even developed a tool that can be used to ensure compliance with their robots.txt non-standard standard. Matt went on to comment that Google does not support a time delay between requests because some webmasters use values that would only allow Google to crawl 15-20 URLs in a day. Yahoo and MSN support this feature using a “Crawl-delay: XX” directive.
Well, I'm out of thoughts. :) Stay tuned…
Friday, January 27, 2006
Yahoo Reports URLs with No Slash
1) http://privacy.getnetwise.org/sharing/tools/ns6
2) http://privacy.getnetwise.org/browsing/tools/profiling
among others. Are ns6 and profiling directories or dynamic pages? You can’t tell by just looking at the URLs… Yahoo strips off the slash (`/`) from the end of URLs that are directories. The only way to tell is to actually visit the URL. URL 1 will return a 301 code (moved permanently) along with the correct URL:
http://privacy.getnetwise.org/sharing/tools/ns6/
URL 2 will respond with a 200 code because it is a dynamic page. This is no big deal for the user looking for search results, but it is a big deal for an application like Warrick which needs to know if a URL is pointing to a directory or not without actually visiting the URL.
I’ve contacted Yahoo about the "problem" but did not receive a response:
http://finance.groups.yahoo.com/group/yws-search-web/message/309
Google and MSN don’t have this problem.
Tuesday, January 10, 2006
Case Insensitive Crawling
http://foo.org/bar.html
http://foo.org/BAR.html
?
A very smart crawler would recognize that these URLs point to the same resource by comparing the ETags or even better, by examining the “Server:” HTTP header in the response. If the “Server:” contains
Server: Microsoft-IIS/5.1
or
Server: Apache/2.0.45 (Win32)
then it would know that the URLs refer to the same resource since Windows uses a case insensitive filesystem. If *nix is the web server’s OS then it can be safely assumed that the URLs are case sensitive. Other operating systems like Darwin (the open source UNIX-based foundation of Mac OS X) may use HFS+ which is a case insensitive filesystem. A danger of using such a filesystem with Apache’s directory protection is discussed here. Mac OS X could also use UFS which is case sensitive, and therefore a crawler cannot make a general rule for this OS to ignore case sensitive URLs.
Finding URLs in Google, MSN, and Yahoo from case insensitive websites is problematic. Consider the following URL:
http://www.harding.edu/user/fmccown/www/comp250/syllabus250.html
Searching for this URL in Google (using the “info:” parameter) will return back a “not found” page. But if the URL
http://www.harding.edu/USER/fmccown/WWW/comp250/syllabus250.html
is searched for, it will be found. Yahoo performs similarly. It will find the all-lowercase version of the URL but not the mixed-case version.
MSN takes the most flexible approach. All “url:” queries are performed in a case insensitive manner. They appear to take this strategy regardless of the website’s OS since a search for
url:http://www.cs.odu.edu/~FMCCOwn/
and
url:http://www.cs.odu.edu/~fmccown/
are both found by MSN even though the first URL is not valid (Unix filesystem). The disadvantage of this approach is what happens when bar.html and BAR.html are 2 different files on the web server. Would MSN only index one of the files?
The Internet Archive, like Google and Yahoo, is pinicky about case. The following URL is found:
http://web.archive.org/web/*/www.harding.edu/USER/fmccown/WWW/comp250/syllabus250.html
but this is not:
http://web.archive.org/web/*/www.harding.edu/user/fmccown/www/comp250/syllabus250.html
Update 3/30/06:
If you found this information interesting, you might want to check out my paper Evaluation of Crawling Policies for a Web-Repository Crawler which discusses these issues.