Showing posts with label web crawler. Show all posts
Showing posts with label web crawler. Show all posts

Wednesday, May 20, 2009

Java Sitemap Parser

I've just released the Java Sitemap Parser on SourceForge.net. The software is capable of reading Sitemaps in XML, Atom, RSS, and text format. As far as I can tell, this is the first open source Sitemap-parsing software available on the Web.

The Java Sitemap Parser was the final project for my Search Engine Development class. I talked about the project a few weeks ago and how prevalent Sitemaps are becoming. Originally we wanted to add Sitemap support to Nutch, but developing just the parser proved to be quite a task. By releasing it as an independent project, I'm hoping Nutch, Heritrix, and other open-source crawlers will integrate it into their systems.

Wednesday, April 22, 2009

Nutch, Sitemaps, and Google's findings

My search engine class is winding down, but our final project is to implement a Sitemap Protocol parser for Nutch, a popular open-source search engine. I mentioned a while back that Nutch is not for wimps... my students would certainly vouch for the huge learning curve to making code modifications. I've even had to scale back how much work my students do because of the complexity of changes required. I'm going to do the difficult part of integrating their code with the innards of Nutch sometime in the next few weeks.

The reason I mention our Sitemap project is that WWW 2009 is meeting in Madrid this week, and a paper entitled Sitemaps: Above and Beyond the Crawl of Duty is being presented today by Uri Schonfeld (UCLA) and Narayanan Shivakumar (Google). This is the first paper to report on widespread usage of Sitemaps in the Web using Google's crawling history.

Schonfeld & Shivakumar report that Sitemaps were used by approximately 35 million websites in late 2008, exposing several billion URLs. 58% of the URLs included last modification dates, 7% included change frequency, and 61% a priority. About 76.8% of Sitemaps used XML formatting, and only 3.4% used plain text. Interestingly, 17.5% of Sitemaps are formatted incorrectly.

The figure below represents how many URLs Google discovered via Sitemaps (red) vs. regular crawling (green) for cnn.com. Notice that on any given day, more URLs could normally be discovered via Sitemaps.



Another interesting figure (below) shows when a URL was discovered via Sitemaps vs. regular web crawling for cnn.com. In most cases URLs were discovered at the same rate, but there are a number of them (dots below the line) that were discovered via Sitemaps much earlier than web crawling.


CNN's website is not typical. Schonfeld & Shivakumar report that in a dataset of 5 billion+ URLs, 78% were discovered via Sitemaps first compared to 22% via web crawling.

The paper also describes an algorithm that can be used by search engines to prioritize URLs discovered via web crawling and Sitemaps as well. I've covered the high-lights, but I recommend you read the paper if you're interested in some of the finer details.

Monday, February 16, 2009

Specifying canonical URLs

Last week the big three search engines (Google, Yahoo, and Live Search) announced their support for a new HTML attribute value which will help prevent search engines from indexing duplicate content. Search engines naturally want to avoid crawling and indexing duplicate content because it lessens the quality of search result pages. Google's Webmaster Central Blog has a good write-up about the new rel="canonical" attribute value.

Essentially, the new attribute value will allow a webmaster to tell a web crawler to ignore a page if it is accessible from another URL. So if a I have a single page that is accessible at URLs A, B, and C, I can tell the web crawler that URLs B and C are pointing to the same content as A by placing the following code in the head element of the page:

<link rel="canonical" href="http://foo.com/A" />

When the web crawler grabs the pages using URLs B or C, it will find the given canonical URL A in the header and therefore ignore the contents of the pages since they duplicate page A.

Of course the entire mechanism requires a willing and competent webmaster to implement it. Webmasters who are very concerned about SEO are likely to use it since it will help bolster the PageRank of certain pages. But the rest of us who aren't concerned about our rankings can safely ignore this new functionality.

See also rel="nofollow".

Monday, July 21, 2008

Webcrawlers wanted

I'm not a huge fan of advertisements, but this ad on Facebook, sponsored by AddGooro, is just too cool. If any of my search engine students are reading this... they're hiring.

Friday, July 04, 2008

Fav5

Happy Independence Day! My pick of the week's top 5 items of interest:

  1. The Deep Web is getting a little shallower: Google has recently improved their ability to crawl Flash content. It used to be that a website with a Flash interface was practically invisible to search engine crawlers, but now Google can find the links to other pages within the Flash program and even indexes the program's textual content.

  2. If you were late signing up for a free Yahoo email address, you might have ended up with frank1837abc@yahoo.com since all the good names were already taken. But now Yahoo has opened up two new domains under "ymail" and "rocketmail". Although Yahoo is still the email market leader with 266 million worldwide users, they want to stay ahead of Microsoft who is a close second with 264 million.

  3. Are Google, Yahoo, and Microsoft censoring themselves more than they must in China? According to a report by Citizen Lab, some sites are blocked by some search engines and others let them through. In a test to see how many questionable sites were blocked, Google had censored 15.2% of the sites tested, Microsoft censored 15.7%, Yahoo 20.8%, and Baidu (the most popular Chinese search engine) 26.4%.

  4. A Harvard marketing professor takes on Chris Anderson's "Long Tail" theory by analyzing real data about online video rentals and song purchases. Her conclusion is that maybe we aren't as individualistic in our taste as Anderson suggested.

  5. In the spirit of July 4, read about how LANL scientists are making fireworks a little "greener".

Friday, April 11, 2008

Fav5

Today I'm stuck at home trying to recover from the latest disease Ethan has brought home. Here's my 5 items of interest for the week:
  1. Google recently announced their newest attempt at crawling the Deep Web. When their crawler sees a form on a "high-quality site", it enters words and selects various checkboxes and radio buttons. The form is then submitted, and the resulting page is crawled. Very cool. Google also announced the Google App Engine. It's a service that allows you to run your web applications on Google's infrastructure so it can grow to accommodate a large amount of traffic. It's a free service until it exceeds a set disk space and bandwidth quota. Unfortunately, there's a limit of 10,000 developers, and I wasn't quick enough to sign up, so I'll have to wait until they increase the limit.

  2. Joel De Young, one of our Harding CS graduates, was recently interviewed about his company's newest game: Penny Arcade Adventures: On the Rain-slick Precipice of Darkness. Sounds like Joel's HotheadGames is making some buzz.

  3. A study by Univ of Minnesota researchers shows there are two indicators that determine how good of an answer you will receive when posting to question-and-answer sites like Yahoo! Answers: how much you pay, and how many responses you receive.

  4. Jansen, Booth, and Spink have developed a system which can automatically classify a search query as navigational, transactional, and informational. They used their classifier on a large dataset of search engine queries and were accurate 74% of the time. Based on their results, approximately 80% of all search engine queries are informational, 10% transactional, and 10% informational.

  5. Congratulations to the Harding Programming Team on their first place finish at CCSC-MS. We showed them Razorbacks how to program... wink

Friday, March 21, 2008

Fav5

It's Spring Sing weekend at Harding, and the sun is shining. Here's my pick of the week's top 5 items of interest:
  1. Technology Review has just released its top 10 emerging technologies of 2008, those technologies that are most likely to make a huge impact on the way we live. I personally find offline Web applications to be most interesting. This is an area Adobe and Google are vigorously pursuing.

  2. Joan Smith (my former office-mate in graduate school) and Michael Nelson (my former Ph.D. adviser and pioneer of digital preservation wink) just published an article in D-Lib Magazine showing a year's worth of crawling activity on some specially designed websites. Make sure you check out the animated images in Tables 3 and 6- they show the crawling behavior of an aggressive Google and timid Yahoo and MSN.

  3. Ethics 101: You're a popular search engine that allows users to see emergent news from a number of sources. But you're housed in a country where the authorities often seek to squelch information that doesn't hold to the party line. If you show news stories that make the authorities look bad, your search engine will be shut down or blocked. What do you do?

  4. Joel Spolsky writes an excellent article about IE 8 and the many difficulties of creating software that supports "standards". This is required reading for all you software engineers and CS students. (The article is a little long, but it's worth it.)

  5. While we're on the topic of web browsers, almost a year ago Steve Jobs announced that the next version of Safari (Apple's web browser) would run on Windows. Now Apple is using iTunes to push it out to the Microsoft world. My guess is that many unsuspecting users are going to think they need Safari to run iTunes since it appears in the same dialog box they always see when needing to install an iTunes update. Many users are going to downloaded it and puzzle over the new icon on their desktop. This should be interesting...

Thursday, February 28, 2008

Fav5

This one is coming a little early, because I'm heading out to California in the morning. I'm attending a program committee meeting at Stanford University over the weekend. This is my first trip out there, so I'm pretty excited. I'll be touring the Googleplex and the Internet Archive on Monday.
  1. Abilene Christian University, one of our sister schools, has decided to give all incoming freshmen next year an iPod. Is this a gimmick or the future of secondary education? Let's hope the move helps boost their enrollment and shrink their $3 million dollar budget deficit.

  2. Google Sites has just launched as part of Google Apps. Essentially it allows you to create a web page or website via wiki tools that anyone in your group or organization can modify. I haven't tried it yet, but it looks promising.

  3. Matt Cutts always has some interesting items to read: Subscribed links, how search engines should deal with NOINDEX, and Blogger Play.

  4. Good news for those of us that work in IT: salaries are expected to grow in 2008. Will more money start getting freshmen to major in CS?

  5. Java or .NET? According to the Tiobe Programming Community Index, Java is still ranked number 1. But in a recent survey by Info-Tech of 1,900 companies, 12% of enterprises focus exclusively on .NET , and 49% center primarily on .NET. This compares to 3% and 20% of companies that are focused exclusively or primarily on Java.

Saturday, December 08, 2007

Fav5

Only one week of finals before the semester is out! My pick of the week's top 5 items of interest:
  1. Good news for the CS majors: According to Bureau of Labor Statistics, computing-based jobs are expected to grow 24% over the next decade, the highest rate of growth of any "professional and related occupations."

  2. Maybe the grass isn't greener on Mac's side of the fence.

  3. Google has just released an impressive REST-based chart API. It allows 50K queries per day and allows you to create pie graphs, line graphs, scatterplots, bar charts, and Venn diagrams.

    Just for fun-- my Family League fantasy football record (8-5) as of today:

    Fantasy football standings

  4. The Heritrix 2.0 web crawler is here. I'll be using this crawler (most likely) in my search engine class next semester.

  5. The planet's latest enemies: divorced observers of Hanukkah. Wow... no one is safe blame these days.

Monday, October 08, 2007

Crawling behavior of Google, MSN, and Yahoo

Thanks, Marko, for sending me a link to this study analyzing the crawling behavior of the big three search engines. The experiment involved setting up a synthetic collection of web pages arranged as a binary search tree. They monitored the crawl log and performed search engine queries for a period of one year (2005-4-13 to 2006-4-13).

The photo below is a visualization of Yahoo's crawling behavior (Yahoo was the most active of the three crawlers). You can see an animation of the tree growing each day here.


We performed a similar experiment at ODU back in 2005 except we removed web pages every day from our collections to see how long they would stay cached by the search engines. And Joan, a fellow Ph.D. student at ODU who worked on the previous experiment, has been performing another experiment over the last several months with a very deep and wide collection of pages. (You can find links to the pages at the bottom-right corner of this page labeled Joan1-4).

What I found really interesting about this experiment was the methodology used to created the web pages. They altered the pages by appending the crawl log of the pages and allowing anyone (including spam bots) to make comments that appeared on the pages. This kept the page contents unique and changing to entice more crawls.

I really hope the authors of the study will submit their findings for publication in a peer-reviewed conference or journal.

Friday, October 06, 2006

Become.com's web crawler

Today a member of the Heritrix list serve pointed everyone to an article on Sun’s website that discusses Become.com’s web crawler. The article dates back to August of 2005, so it’s a little dated. I couldn’t find any updated information on the crawler, but apparently it is proprietary, and the source code will likely never see the light of day.

Become.com actually developed 2 crawlers in 2004- one written entirely in Java and the other mostly Java with some C++. The article states that the crawlers "may be the most sophisticated, massively scaled Java technology application in existence."

The article doesn’t mention anything about Heritrix, a crawler which is also completely written in Java. Although Heritrix doesn’t currently have a distributed architecture, it could still be deployed in such an environment. It would be really interesting to see the two crawlers compete at the National Java Crawling Championships.

Wednesday, August 09, 2006

Crawling the Web is very, very hard…

I’ve spent the past couple of weeks trying to randomly select 300 websites from a dmoz.org. There were only a couple of restrictions I placed on the selection:
  1. The website’s root URL not redirect the crawler to a URL that is on a different host. If it does, the new URL should replace the old website URL.
  2. The website’s root URL should not appear to be a splash page for a website on a different host or indicate that the website has moved to a different host. If it does, the new URL should replace the old website URL.
  3. The website should not have all of its contents blocked by robots.txt. If some directories are blocked, that’s ok.
  4. The website’s root URL should not have a noindex/nofollow meta tag which would prevent a crawler from grabbing anything else on the website.
  5. The website should not have any more than 10K resources.

The restrictions seem very straightforward, but in practice they are very time consuming to enforce. Requirement 2 requires me to manually visit the site. Did I mention not all the sites are in English? That makes it even more difficult. Requirement 3 means I have to manually examine the robots.txt. Req. 4 requires manually examination of the root page, and req. 5 means I have to waste several days crawling a site before I know if it is too large or not.

I guess I could build a tool for req 3 and 4, but I’m not in the mood.

Anyway, I ended up making about 50 replacements (at least) and starting my crawls over again. Now I finally have 300 websites that meet my requirements.

In the past I’ve used Wget to do my crawling, but I’ve decided to use Heritrix since it has quite a few useful features missing from Wget. But Heritrix isn’t perfect. I made a suggestion that Heritrix show the number of URLs from each host remaining in the frontier when examining a crawl report:

http://sourceforge.net/tracker/index.php?func=detail&amp;amp;aid=1533116&group_id=73833&atid=539102

Right now it is very difficult to tell if a host has been completely crawled or not. I would love to work on this myself, but I just don’t have the time right now. Maybe I'll get a student to work on this next time I'm teaching. ;)

The other difficulty with Heritrix is in extracting what you have crawled. I will need to write a script that will build an index into the ARC files so I can quickly extract data for a website. Since all the crawl data is merged into a series of ARC files, it is really difficult to throw away crawl data for a website you aren’t interested in. I could write a script to do it, but at this point it’s not worth my time.

Anyway, crawling is a very error-prone and difficult process, but I can’t wait to teach about it when I return to Harding! (I'm really getting the itch to get back in the classroom.)

Friday, June 23, 2006

Heritrix - An archival quality crawler

This week I’ve been experimenting with Heritrix, the Internet Archive’s web crawler. It has some functionality that Wget doesn’t provide including:
  • limiting the size of each file downloaded
  • allowing a crawl to be paused and the frontier to be examined and modified
  • following links in CSS and Flash
  • crawling multiple sites at the same time without invoking multiple instances of the crawler
  • storing crawls in an Arc file
Since Heritrix was built with Java and was pre-configured to run on a Linux system, I didn’t have to expend much effort to get it to run on Solaris. I untarred the distribution file, set a couple of environment variables, started the web server interface, and boom it was working.

The interface is not exactly intuitive, and a near complete reading of the entire manual is required to put together a decent crawl. Of course if you want to use sophisticated open-source software, you usually have to put in some significant effort to get it to work right. Thankfully, several of the developers (Michael Stack, Igor Ranitovic, and Gordon Mohr) have been very helpful in answering some of my newbie questions on the Heritrix list serve.

In learning about Heritrix, I’ve put together a page on Wikipedia. Hopefully the entry will drum up more general interest in Heritrix as well. I was really surprised no one had created the page before.

Thursday, April 13, 2006

URL Canonicalization

The term URL canonicalization (also frequently called URL normalization) refers to the process that is performed on a URL to make it easier to tell if two syntactically different URLs are the same. For example, the URL

http://www.Harding.edu/USER/dsteil/www/abc/../index.htm

could be normalized to produce the canonical URL:

http://www.harding.edu/user/dsteil/www/

Search engines typically use different URL canonicalization policies which makes it difficult for Warrick to tell if URL x from MSN is the same as URL y from Google. I’ve noted some peculiarities in my blog here, here and here. Matt Cutts at Google also discussed some of their canonicalization policies back in Jan 2006.

I have not found much work in the literature about URL canonicalization/normalization. RFC 3986 has some standard normalization procedures that should be done. Pant et al. (2004) has a section about it in their chapter Crawling the Web from the book Web Dynamics. The first paper I’ve seen that deals with the issue head-on is by Sang Ho Lee et al. (2005) "On URL normalization".

I also checked Wikipedia and didn’t find anything about URL canonicalization. I decided to create a page about it and added a reference to it from the web crawler page. That was the first page I ever created on Wikipedia. Proverbs 25:2 – “It is the glory of God to conceal a thing; but the glory of kings is to search out a matter.” I’m no king, but I think God actually delights in our effort to learn about the great world He has created, and I appreciate Wikipedia providing a unique resource for us to consolidate and share our learning.

Friday, February 10, 2006

Some thoughts on robots.txt

The Robots Exclusion Protocol has been around since June 1994, but there is no official standards body or RFC for the protocol. That leaves others free to tinker with it and add their own bells and whistles.

There are numerous limitations to robots.txt that have been noted (see Martijn Koster’s article). A few things that are lacking: ability to specify how frequently server requests should be made, the ideal times to make automated requests, permissions to visit vs. index vs. cache (make available in search engine caches).

According to Matt Cutts, Google supports the “Allow:” directive and wildcards (*) which are not part of the standard. The Google Sitemap team even developed a tool that can be used to ensure compliance with their robots.txt non-standard standard. Matt went on to comment that Google does not support a time delay between requests because some webmasters use values that would only allow Google to crawl 15-20 URLs in a day. Yahoo and MSN support this feature using a “Crawl-delay: XX” directive.

Well, I'm out of thoughts. :) Stay tuned…

Friday, January 27, 2006

Yahoo Reports URLs with No Slash

Yahoo does not properly report URLs that end in a directory with a slash at the end. For example, the query for "site:privacy.getnetwise.org" will yield the following URLs:

1) http://privacy.getnetwise.org/sharing/tools/ns6
2) http://privacy.getnetwise.org/browsing/tools/profiling

among others. Are ns6 and profiling directories or dynamic pages? You can’t tell by just looking at the URLs… Yahoo strips off the slash (`/`) from the end of URLs that are directories. The only way to tell is to actually visit the URL. URL 1 will return a 301 code (moved permanently) along with the correct URL:

http://privacy.getnetwise.org/sharing/tools/ns6/

URL 2 will respond with a 200 code because it is a dynamic page. This is no big deal for the user looking for search results, but it is a big deal for an application like Warrick which needs to know if a URL is pointing to a directory or not without actually visiting the URL.

I’ve contacted Yahoo about the "problem" but did not receive a response:
http://finance.groups.yahoo.com/group/yws-search-web/message/309

Google and MSN don’t have this problem.

Tuesday, January 10, 2006

Case Insensitive Crawling

What should a Web crawler do when it is crawling a website that is housed on a Windows web server and it comes across the following URLs:

http://foo.org/bar.html
http://foo.org/BAR.html

?

A very smart crawler would recognize that these URLs point to the same resource by comparing the ETags or even better, by examining the “Server:” HTTP header in the response. If the “Server:” contains

Server: Microsoft-IIS/5.1

or

Server: Apache/2.0.45 (Win32)

then it would know that the URLs refer to the same resource since Windows uses a case insensitive filesystem. If *nix is the web server’s OS then it can be safely assumed that the URLs are case sensitive. Other operating systems like Darwin (the open source UNIX-based foundation of Mac OS X) may use HFS+ which is a case insensitive filesystem. A danger of using such a filesystem with Apache’s directory protection is discussed here. Mac OS X could also use UFS which is case sensitive, and therefore a crawler cannot make a general rule for this OS to ignore case sensitive URLs.

Finding URLs in Google, MSN, and Yahoo from case insensitive websites is problematic. Consider the following URL:

http://www.harding.edu/user/fmccown/www/comp250/syllabus250.html

Searching for this URL in Google (using the “info:” parameter) will return back a “not found” page. But if the URL

http://www.harding.edu/USER/fmccown/WWW/comp250/syllabus250.html

is searched for, it will be found. Yahoo performs similarly. It will find the all-lowercase version of the URL but not the mixed-case version.

MSN takes the most flexible approach. All “url:” queries are performed in a case insensitive manner. They appear to take this strategy regardless of the website’s OS since a search for

url:http://www.cs.odu.edu/~FMCCOwn/

and

url:http://www.cs.odu.edu/~fmccown/

are both found by MSN even though the first URL is not valid (Unix filesystem). The disadvantage of this approach is what happens when bar.html and BAR.html are 2 different files on the web server. Would MSN only index one of the files?

The Internet Archive, like Google and Yahoo, is pinicky about case. The following URL is found:

http://web.archive.org/web/*/www.harding.edu/USER/fmccown/WWW/comp250/syllabus250.html


but this is not:

http://web.archive.org/web/*/www.harding.edu/user/fmccown/www/comp250/syllabus250.html


Update 3/30/06:

If you found this information interesting, you might want to check out my paper Evaluation of Crawling Policies for a Web-Repository Crawler which discusses these issues.