- The Web 2.0 Summit 2008 concludes today. Winners: web apps that save the world. Losers: Yahoo's Jerry Wang.
- Firefox for your cell phone?
- Have you heard of the Year 2038 Problem?
- Flash or Java for game development?
- Whether you're happy about Obama being elected or not, we've just experienced something monumental in American history. And it is refreshing to know that race is significantly less of an issue today.
Someone with a lot of time on his hands also compared the web presence of McCain and Obama in Google, Facebook, Twitter, etc.; Obama won hands-down.
Showing posts with label yahoo. Show all posts
Showing posts with label yahoo. Show all posts
Friday, November 07, 2008
Fav5
My pick of the week's top 5 items of interest:
Labels:
fav5,
games,
web browsers,
yahoo
Saturday, July 19, 2008
Fav5
My pick of the week's top 5 items of interest:
- With the growing popularity of the iPhone, website designers are paying much more attention to how their sites look on devices with very small real estate.
- Will Microsoft bite at $33 a share for Yahoo?
- We're getting just a little closer to realize quantum computing.
- This will be of interest to many of my students: How much do game programmers make?
- Last fall, Harding Univ made the move to Google Mail, freeing us from the burden of maintaining our own email system. Unfortunately, the IT guys left all of our old emails on the old system and didn't forward them on to our new GMail accounts. I just received an email saying in the next few weeks the IT guys will purge all of our emails from the old system. Anyone want to bet some very important email is about to be erased forever? (I did move my email over, but it took a few hours... I would hate to have lost emails from when Becky and I were dating.)
Friday, July 11, 2008
Fav5
My pick of the weeks' top 5 items of interest:
- Good news: the number of people working in IT jobs in the US has hit a record 4 million, and the unemployment rate is a meager 2.3%. Advice to incoming freshmen: Give computer science a try.
- Chill. It's just a phone.
- Neil McAllister asks an intriguing question: "Is the Web still the Web?"
- Digg has rolled out a new recommendation system based on the wisdom of crowds.
- Some cool new stuff: Protocol Buffers from Google and the BOSS API from Yahoo (see my post from yesterday).
Friday, July 04, 2008
Fav5
Happy Independence Day! My pick of the week's top 5 items of interest:
- The Deep Web is getting a little shallower: Google has recently improved their ability to crawl Flash content. It used to be that a website with a Flash interface was practically invisible to search engine crawlers, but now Google can find the links to other pages within the Flash program and even indexes the program's textual content.
- If you were late signing up for a free Yahoo email address, you might have ended up with frank1837abc@yahoo.com since all the good names were already taken. But now Yahoo has opened up two new domains under "ymail" and "rocketmail". Although Yahoo is still the email market leader with 266 million worldwide users, they want to stay ahead of Microsoft who is a close second with 264 million.
- Are Google, Yahoo, and Microsoft censoring themselves more than they must in China? According to a report by Citizen Lab, some sites are blocked by some search engines and others let them through. In a test to see how many questionable sites were blocked, Google had censored 15.2% of the sites tested, Microsoft censored 15.7%, Yahoo 20.8%, and Baidu (the most popular Chinese search engine) 26.4%.
- A Harvard marketing professor takes on Chris Anderson's "Long Tail" theory by analyzing real data about online video rentals and song purchases. Her conclusion is that maybe we aren't as individualistic in our taste as Anderson suggested.
- In the spirit of July 4, read about how LANL scientists are making fireworks a little "greener".
Labels:
google,
los alamos,
microsoft,
web crawler,
yahoo
Saturday, June 07, 2008
Fav5
My pick of the week's top five items of interest:
- How valuable is your old high school or college yearbook? At Purdue University, this year's yearbook will be their last. In fact there are only 80 US colleges that still produce yearbooks, down from 100 last year. Interest is declining with part of the blame on social networking sites like Facebook. What today's students don't realize is that 20 years from now, you may not have access to the memories you have now... free services like Facebook have no obligation to retain them indefinitely for you.
- Want to save some electricity and don't mind black? Try out Blackle. Blackle displays search results straight from Google, but because they are displaying search results on a primarily black screen, they are using considerably less power than Google which uses white. A similar search engine is called Blaxel, but I don't think the two are affiliated.

Update 6-9-08:
My Finnish amigo has pointed out the error in my post. Apparently only CRTs save energy by displaying black, but the newer LCD monitors which most everyone is using today actually take more energy to display black. So Blackle and Blaxel are actually wasting more energy! Thanks for the tip, Timo. - The Digital Lives research project is conducting a quick survey of how individuals store personal computer files, find them in the future, and archive them. If you have 10 minutes and want to contribute to digital preservation research (and possibly win £200 in British Library gift vouchers), please take the survey.
- This is kinda cool: Yahoo has opened up its search results page to developers using a new platform called SearchMonkey. They've also developed a listing of numerous SearchMonkey plug-ins in their Yahoo Search Gallery. Do you want to see details of a movie when searching Yahoo? Download the IMDB presentation enhancement.
- I'm not totally sure what to make of this: a new search engine called RushmoreDrive tailors search results for the black community only. Apparently Google is too white; African Americans want different search results than European Americans, Asian Americans, etc. While I agree that web search that takes into account the user's profile (e.g., interests, age, gender, location, etc.) are likely to produce better search results, creating a search engine that tailors only to one racial group smacks of racism. While I'm sure this isn't RuchmoreDrive's intention, wouldn't we all agree that a search engine called Whitey.com that was built for the white community only was racist?
Labels:
digital preservation,
fav5,
search engines,
web search,
yahoo
Thursday, May 15, 2008
Baseball and search engine queries
This is a really interesting visualization from Spatial Variation in Search Engine Queries, a WWW08 paper by Backstrom & Kleinberg (Cornell) and Kumar & Novak (Yahoo!).

What you are seeing are the "spheres of influence" for various Major League baseball teams. To create this graph, the authors mined the search engine logs from Yahoo! and plotted where the most popular baseball team queries originated from (based on IP address).
The result is a geographical break-down of each team's fan base. The authors point out that the boundaries even follow state boundaries closely: "For instance, in Michigan, across the lake from Chicago but far from Detroit, it is the Tigers, not the Cubs who have the largest following."
It's also interesting to note there are some places like Arkansas where there is no single winner. That's what happens when you live in a state with zero professional teams. Oh, how I miss my childhood years living in Denver...

What you are seeing are the "spheres of influence" for various Major League baseball teams. To create this graph, the authors mined the search engine logs from Yahoo! and plotted where the most popular baseball team queries originated from (based on IP address).
The result is a geographical break-down of each team's fan base. The authors point out that the boundaries even follow state boundaries closely: "For instance, in Michigan, across the lake from Chicago but far from Detroit, it is the Tigers, not the Cubs who have the largest following."
It's also interesting to note there are some places like Arkansas where there is no single winner. That's what happens when you live in a state with zero professional teams. Oh, how I miss my childhood years living in Denver...
Labels:
baseball,
visualization,
web search,
yahoo
Saturday, June 30, 2007
Save your Yahoo! Photos
I just received this email yesterday from Yahoo! Photos, telling me that my photos were going to disappear unless I took action. I couldn't help but wonder what if I my spam filter had eaten the message, or what if I wasn't using this email account anymore? Considering that just 57% of individuals actually backup their personal data, how many people do you think are depending on Yahoo! to preserve their photos indefinitely? You can image the horrible feeling of logging in and realizing all those photos of your newborn child have ended up in the big bit bucket in the sky. Sigh...
Dear Yahoo! Photos user,
...
We will officially close Yahoo! Photos on Thursday, September 20, 2007, at 9 p.m. PDT. Until then, we are offering you the opportunity to move to another photo sharing service (Flickr, KODAK Gallery, Shutterfly, Snapfish, or Photobucket), download your original-resolution photos back to your computer, or buy an archive CD from our featured partner (for users of the New Yahoo! Photos only). All you need to do is tell us what to do with your photos before we close, after which any photos remaining on Yahoo! Photos will be deleted and no longer accessible.
...
Please give us your decision by Thursday, September 20, 2007, at 9 p.m. PDT. After that time, any photos remaining in Yahoo! Photos will be deleted. Click here to make your decision, or review a list of our frequently asked questions.
Friday, May 04, 2007
First Fav5
I'm starting a new weekly installment on my blog entitled "Fav5" where I list 5 items of interest for the week. I'm not putting any restrictions on the list, so anything may appear here. You may notice this week's list is a little heavy on the search engine side.- I recently found out that Ask, MSN, and Yahoo are now supporting Google's Sitemap Protocol. The announcement was made back in Nov. 2006 although Ask apparently joined a little later, and MSN still hasn't implemented it yet. The "official" website for sitemaps is now http://www.sitemaps.org/.
Just last month an autodiscovery method was announced which makes it easy to notify all the search engines of your sitemap by naming it in your robots.txt file. Alternatively you can use the search engine's interface or ping a search engine via an HTTP request to let them know the sitemap file location. - If you didn't already know, Google posts their Tech Talks on Google Video. There are some top-notch presentations out there if you have an hour to burn per talk.
The most recent talk is The Internet of Things: What is a Spime and why is it useful?
by Science Fiction writer and futurist Bruce Sterling. Early in the talk, Bruce rightly lauds Vannevar Bush's prediction of a memex device as one of the most "brilliant acts of technological forecasting, ever." - Google announced their new Web History feature a few weeks ago:
Today, we're pleased to announce the launch of Web History, a new feature for Google Account users that makes it easy to view and search across the pages you've visited. If you remember seeing something online, you'll be able to find it faster and from any computer with Web History. Web History lets you look back in time, revisit the sites you've browsed, and search over the full text of pages you've seen. It's your slice of the web, at your fingertips.
Matt Cutts shows some use cases for Web History and also ties it in to Bush's memex device. - Joel points out a very interesting usability problem posed by the elevators installed in the new World Trade Center highrise.
- And on a personal note, I'm going to be presenting a poster Search Engines and their Public Interfaces: Which APIs are the Most Synchronized? at the World Wide Web conference in Banff, Alberta next week. You can see a PDF of my poster here. Please come by my poster and say hello if you're at the event.
I just found out Johan (ex-ODU professor) is going to be at WWW too. It's too bad his poster is kinda lame.
Thursday, August 03, 2006
Yahoo transforming FRAME tags
The past several months I’ve been ramping-up for a huge experiment where I’ll be reconstructing several hundred websites. I’ve been learning to use Heritrix and process ARC files, and I’ve been periodically tweaking Warrick. Today I found out that Yahoo has changed the way it caches HTML pages that contain frames.
For example, the page at http://www.harding.edu/comp/ contains the following HTML:
<FRAMESET COLS="195,*" FRAMEBORDER=no FRAMESPACING=0>
<FRAME SRC=menu.html NAME="MENU" MARGINWIDTH=0 MARGINHEIGHT=0>
<FRAME SRC=welcome.html NAME="MAIN">
</FRAMESET>
In Yahoo’s cached page for this URL, the FRAME tags are converted to the following (I’ve added some white space for readability):
<frameset rows="200,*"><frame scrolling="no" noresize="" frameborder="0" marginwidth="0" marginheight="0" src="http://216.109.125.130/search/cache?.intl=us&u=www.harding.edu%2fcomp%2f&
w=%22harding+.edu%22&d=XS7fRmP9NNYx&origargs=p%3durl%253Ahttp%253A%252F%252F
www.harding.edu%252Fcomp%252F%26toggle%3d1%26ei%3dUTF-8%26_intl%3dus&frameid=-1">
<FRAMESET COLS="195,*" FRAMEBORDER=no FRAMESPACING=0>
<frame security="restricted" MARGINHEIGHT="0" MARGINWIDTH="0" NAME="MENU" SRC="http://216.109.125.130/search/cache?.intl=us&u=www.harding.edu%2fcomp%2f&
w=%22harding+.edu%22&d=XS7fRmP9NNYx&origargs=p%3durl%253Ahttp%253A%252F%252F
www.harding.edu%252Fcomp%252F%26toggle%3d1%26ei%3dUTF-8%26_intl%3dus&frameid=1" >
<frame security="restricted" NAME="MAIN" SRC="http://216.109.125.130/search/cache?.intl=us&u=www.harding.edu%2fcomp%2f&
w=%22harding+.edu%22&d=XS7fRmP9NNYx&origargs=p%3durl%253Ahttp%253A%252F%252F
www.harding.edu%252Fcomp%252F%26toggle%3d1%26ei%3dUTF-8%26_intl%3dus&frameid=2" >
</frameset></FRAMESET>
Yahoo is placing their own FRAMESET tags around mine and loading the two column frames with pages directly from their cache. Notice the use of security="restricted" within the FRAME tag which tells the browser to place security constraints on the frame sources; this disables any JavaScript in my pages.
While this conversion of FRAME tags makes the page easier to view from their cache, it completely destroys the original HTML. There’s no way I can even parse through the arguments to tell what URL used to be in the SRC attribute. ARG! Now I’m going to have to add a rule to Warrick that tells it to ignore Yahoo cached pages that contain FRAME tags. Google and MSN have yet to implement this “trick”, and hopefully they never do.
For example, the page at http://www.harding.edu/comp/ contains the following HTML:
<FRAMESET COLS="195,*" FRAMEBORDER=no FRAMESPACING=0>
<FRAME SRC=menu.html NAME="MENU" MARGINWIDTH=0 MARGINHEIGHT=0>
<FRAME SRC=welcome.html NAME="MAIN">
</FRAMESET>
In Yahoo’s cached page for this URL, the FRAME tags are converted to the following (I’ve added some white space for readability):
<frameset rows="200,*"><frame scrolling="no" noresize="" frameborder="0" marginwidth="0" marginheight="0" src="http://216.109.125.130/search/cache?.intl=us&u=www.harding.edu%2fcomp%2f&
w=%22harding+.edu%22&d=XS7fRmP9NNYx&origargs=p%3durl%253Ahttp%253A%252F%252F
www.harding.edu%252Fcomp%252F%26toggle%3d1%26ei%3dUTF-8%26_intl%3dus&frameid=-1">
<FRAMESET COLS="195,*" FRAMEBORDER=no FRAMESPACING=0>
<frame security="restricted" MARGINHEIGHT="0" MARGINWIDTH="0" NAME="MENU" SRC="http://216.109.125.130/search/cache?.intl=us&u=www.harding.edu%2fcomp%2f&
w=%22harding+.edu%22&d=XS7fRmP9NNYx&origargs=p%3durl%253Ahttp%253A%252F%252F
www.harding.edu%252Fcomp%252F%26toggle%3d1%26ei%3dUTF-8%26_intl%3dus&frameid=1" >
<frame security="restricted" NAME="MAIN" SRC="http://216.109.125.130/search/cache?.intl=us&u=www.harding.edu%2fcomp%2f&
w=%22harding+.edu%22&d=XS7fRmP9NNYx&origargs=p%3durl%253Ahttp%253A%252F%252F
www.harding.edu%252Fcomp%252F%26toggle%3d1%26ei%3dUTF-8%26_intl%3dus&frameid=2" >
</frameset></FRAMESET>
Yahoo is placing their own FRAMESET tags around mine and loading the two column frames with pages directly from their cache. Notice the use of security="restricted" within the FRAME tag which tells the browser to place security constraints on the frame sources; this disables any JavaScript in my pages.
While this conversion of FRAME tags makes the page easier to view from their cache, it completely destroys the original HTML. There’s no way I can even parse through the arguments to tell what URL used to be in the SRC attribute. ARG! Now I’m going to have to add a rule to Warrick that tells it to ignore Yahoo cached pages that contain FRAME tags. Google and MSN have yet to implement this “trick”, and hopefully they never do.
Monday, June 05, 2006
Getting external backlinks from Google, Yahoo, and MSN
It’s often useful to know how many external backlinks are pointing to a particular URL. This metric can be used to partly determine a page’s popularity on the Web. A good tool for automating this process is the Backlink Analyzer which uses the Google, MSN, and Yahoo APIs using the "link:" command. The software allows a user to specify sites to ignore in the backlink counts, a useful function since the link: command returns backlinks from external and internal links for all three search engines.
I don’t know for sure if Backlink Analyzer is doing this or not, but for Yahoo and MSN, it is possible to perform a single command to give only external backlinks using a combination of link: and -site: parameters. For example, the following query will show all the pages pointing to my Warrick page for Yahoo and MSN:
link:http://www.cs.odu.edu/~fmccown/research/lazy/warrick.html -site:www.cs.odu.edu
Google will not handle the –site: parameter successfully in this query although it does handle it in other types of queries.
I don’t know for sure if Backlink Analyzer is doing this or not, but for Yahoo and MSN, it is possible to perform a single command to give only external backlinks using a combination of link: and -site: parameters. For example, the following query will show all the pages pointing to my Warrick page for Yahoo and MSN:
link:http://www.cs.odu.edu/~fmccown/research/lazy/warrick.html -site:www.cs.odu.edu
Google will not handle the –site: parameter successfully in this query although it does handle it in other types of queries.
Tuesday, May 09, 2006
Server encoding caching experiment
To determine if my server-side component encodings could be inserted into indexable/cacheable HTML files, I ran a little experiment. I created 3 HTML files that contained encoded chunks in HTML comments at the base of each file:
html_encoded1.html - 2 KB
html_encoded2.html - 45 KB
html_encoded3.html - 99 KB
If you view the source of the pages, you’ll see something like this at the end:
Google – cached all three
MSN – cached 1 and 2
Yahoo – indexed 2 only (not available in their cache)
Ask – nada
To see if Google can handle any more, I have created 4 new files of 150, 200, 250, and 300 KB. Looks like 99 KB is too large for MSN. Yahoo’s cache is really inconsistent- maybe 2 is in there, maybe it’s not. Why didn’t they grab 1?
I’ll check back in a couple of weeks and see if anything else has been cached.
Update: 6/20/06
Google and MSN have cached all files that range up to 300 KB. Yahoo has only indexed the first 3 (none are cached), and Ask has nothing.
Now I'm going to create a 400 KB, 500 KB, and 1 MB file and see what happens.
Update: 2/21/07
The cache limits for the search engines appear to be the following: Google - 977 KB, Yahoo - 214 KB, and MSN - 1 MB. I still cannot tell for sure what Ask's limit is, but I ran an experiment where I found 984 KB cached for a document that was 1.6 MB. Google's limit has been confirmed by others.
html_encoded1.html - 2 KB
html_encoded2.html - 45 KB
html_encoded3.html - 99 KB
If you view the source of the pages, you’ll see something like this at the end:
<!-- BEGIN_FILERECOVERYI placed these files in my public_html folder on April 19, and linked to them from my index.html page. Today I checked Google, MSN, Yahoo, and Ask to see if any of them were cached. Here’s the results:
chunks = 4
filename = xor.o
recover = 2
orig_size = 1105
block_size = 554
block_num = 3
fY/xaGQn0V5MOOpLnM1WIsIUMirrVBQ2XNhidvc5yjL9tEyKTmNjNPjcrJzcPWvs INxxHl1Gt5lKQAYoNi1DXOhFI5ExBm15Nxx1T/hFCwVvsyaHsQQdd3lcqWJl+WTw BTlkiI8yWcPPoy38dqgTVnc4aSNd+0YQWW0bDl67/6XTnych3rSXn5YEYhVMU2eS LCR/0N4pAhKgeMb7SXtdJNQ6WykqDXYJAjtTOIrT2CLaPNRdKbU/ydsvUSDenSt+
Etc…
END_FILERECOVERY -->
Google – cached all three
MSN – cached 1 and 2
Yahoo – indexed 2 only (not available in their cache)
Ask – nada
To see if Google can handle any more, I have created 4 new files of 150, 200, 250, and 300 KB. Looks like 99 KB is too large for MSN. Yahoo’s cache is really inconsistent- maybe 2 is in there, maybe it’s not. Why didn’t they grab 1?
I’ll check back in a couple of weeks and see if anything else has been cached.
Update: 6/20/06
Google and MSN have cached all files that range up to 300 KB. Yahoo has only indexed the first 3 (none are cached), and Ask has nothing.
Now I'm going to create a 400 KB, 500 KB, and 1 MB file and see what happens.
Update: 2/21/07
The cache limits for the search engines appear to be the following: Google - 977 KB, Yahoo - 214 KB, and MSN - 1 MB. I still cannot tell for sure what Ask's limit is, but I ran an experiment where I found 984 KB cached for a document that was 1.6 MB. Google's limit has been confirmed by others.
Yahoo Site Explorer
I just discovered Yahoo’s Site Explorer which was apparently released in September 2005. The tool allows you to see which pages of a site are currently indexed by Yahoo and the inlinks to a particular page. For example, I can see that Yahoo currently has around 1400 URLs indexed from my ODU website, and there are 19 inlinks pointing to the Warrick page. There is an API for accessing the service so page scraping is unnecessary. Now if only we can get Google to provide a similar service!
Tuesday, February 21, 2006
Ghostsites
Steve Baldwin has a really nice blog called Ghostsites which is dedicated to really old web pages and sites. A few entries that caught my eye:
- Is This the World’s Oldest Active Web Page?
- This Site Proudly Optimized for Netscape 2.0!
- The Museum of E-Failure
Labels:
google,
internet archive,
warrick,
yahoo
Friday, January 27, 2006
Yahoo Reports URLs with No Slash
Yahoo does not properly report URLs that end in a directory with a slash at the end. For example, the query for "site:privacy.getnetwise.org" will yield the following URLs:
1) http://privacy.getnetwise.org/sharing/tools/ns6
2) http://privacy.getnetwise.org/browsing/tools/profiling
among others. Are ns6 and profiling directories or dynamic pages? You can’t tell by just looking at the URLs… Yahoo strips off the slash (`/`) from the end of URLs that are directories. The only way to tell is to actually visit the URL. URL 1 will return a 301 code (moved permanently) along with the correct URL:
http://privacy.getnetwise.org/sharing/tools/ns6/
URL 2 will respond with a 200 code because it is a dynamic page. This is no big deal for the user looking for search results, but it is a big deal for an application like Warrick which needs to know if a URL is pointing to a directory or not without actually visiting the URL.
I’ve contacted Yahoo about the "problem" but did not receive a response:
http://finance.groups.yahoo.com/group/yws-search-web/message/309
Google and MSN don’t have this problem.
1) http://privacy.getnetwise.org/sharing/tools/ns6
2) http://privacy.getnetwise.org/browsing/tools/profiling
among others. Are ns6 and profiling directories or dynamic pages? You can’t tell by just looking at the URLs… Yahoo strips off the slash (`/`) from the end of URLs that are directories. The only way to tell is to actually visit the URL. URL 1 will return a 301 code (moved permanently) along with the correct URL:
http://privacy.getnetwise.org/sharing/tools/ns6/
URL 2 will respond with a 200 code because it is a dynamic page. This is no big deal for the user looking for search results, but it is a big deal for an application like Warrick which needs to know if a URL is pointing to a directory or not without actually visiting the URL.
I’ve contacted Yahoo about the "problem" but did not receive a response:
http://finance.groups.yahoo.com/group/yws-search-web/message/309
Google and MSN don’t have this problem.
Tuesday, January 10, 2006
Case Insensitive Crawling
What should a Web crawler do when it is crawling a website that is housed on a Windows web server and it comes across the following URLs:
http://foo.org/bar.html
http://foo.org/BAR.html
?
A very smart crawler would recognize that these URLs point to the same resource by comparing the ETags or even better, by examining the “Server:” HTTP header in the response. If the “Server:” contains
Server: Microsoft-IIS/5.1
or
Server: Apache/2.0.45 (Win32)
then it would know that the URLs refer to the same resource since Windows uses a case insensitive filesystem. If *nix is the web server’s OS then it can be safely assumed that the URLs are case sensitive. Other operating systems like Darwin (the open source UNIX-based foundation of Mac OS X) may use HFS+ which is a case insensitive filesystem. A danger of using such a filesystem with Apache’s directory protection is discussed here. Mac OS X could also use UFS which is case sensitive, and therefore a crawler cannot make a general rule for this OS to ignore case sensitive URLs.
Finding URLs in Google, MSN, and Yahoo from case insensitive websites is problematic. Consider the following URL:
http://www.harding.edu/user/fmccown/www/comp250/syllabus250.html
Searching for this URL in Google (using the “info:” parameter) will return back a “not found” page. But if the URL
http://www.harding.edu/USER/fmccown/WWW/comp250/syllabus250.html
is searched for, it will be found. Yahoo performs similarly. It will find the all-lowercase version of the URL but not the mixed-case version.
MSN takes the most flexible approach. All “url:” queries are performed in a case insensitive manner. They appear to take this strategy regardless of the website’s OS since a search for
url:http://www.cs.odu.edu/~FMCCOwn/
and
url:http://www.cs.odu.edu/~fmccown/
are both found by MSN even though the first URL is not valid (Unix filesystem). The disadvantage of this approach is what happens when bar.html and BAR.html are 2 different files on the web server. Would MSN only index one of the files?
The Internet Archive, like Google and Yahoo, is pinicky about case. The following URL is found:
http://web.archive.org/web/*/www.harding.edu/USER/fmccown/WWW/comp250/syllabus250.html
but this is not:
http://web.archive.org/web/*/www.harding.edu/user/fmccown/www/comp250/syllabus250.html
Update 3/30/06:
If you found this information interesting, you might want to check out my paper Evaluation of Crawling Policies for a Web-Repository Crawler which discusses these issues.
http://foo.org/bar.html
http://foo.org/BAR.html
?
A very smart crawler would recognize that these URLs point to the same resource by comparing the ETags or even better, by examining the “Server:” HTTP header in the response. If the “Server:” contains
Server: Microsoft-IIS/5.1
or
Server: Apache/2.0.45 (Win32)
then it would know that the URLs refer to the same resource since Windows uses a case insensitive filesystem. If *nix is the web server’s OS then it can be safely assumed that the URLs are case sensitive. Other operating systems like Darwin (the open source UNIX-based foundation of Mac OS X) may use HFS+ which is a case insensitive filesystem. A danger of using such a filesystem with Apache’s directory protection is discussed here. Mac OS X could also use UFS which is case sensitive, and therefore a crawler cannot make a general rule for this OS to ignore case sensitive URLs.
Finding URLs in Google, MSN, and Yahoo from case insensitive websites is problematic. Consider the following URL:
http://www.harding.edu/user/fmccown/www/comp250/syllabus250.html
Searching for this URL in Google (using the “info:” parameter) will return back a “not found” page. But if the URL
http://www.harding.edu/USER/fmccown/WWW/comp250/syllabus250.html
is searched for, it will be found. Yahoo performs similarly. It will find the all-lowercase version of the URL but not the mixed-case version.
MSN takes the most flexible approach. All “url:” queries are performed in a case insensitive manner. They appear to take this strategy regardless of the website’s OS since a search for
url:http://www.cs.odu.edu/~FMCCOwn/
and
url:http://www.cs.odu.edu/~fmccown/
are both found by MSN even though the first URL is not valid (Unix filesystem). The disadvantage of this approach is what happens when bar.html and BAR.html are 2 different files on the web server. Would MSN only index one of the files?
The Internet Archive, like Google and Yahoo, is pinicky about case. The following URL is found:
http://web.archive.org/web/*/www.harding.edu/USER/fmccown/WWW/comp250/syllabus250.html
but this is not:
http://web.archive.org/web/*/www.harding.edu/user/fmccown/www/comp250/syllabus250.html
Update 3/30/06:
If you found this information interesting, you might want to check out my paper Evaluation of Crawling Policies for a Web-Repository Crawler which discusses these issues.
Subscribe to:
Posts (Atom)