Showing posts with label web search. Show all posts
Showing posts with label web search. Show all posts

Friday, December 12, 2008

Fav5

My pick of the week's top 5 items of interest:
  1. Google has released its list of popular searches in their 2008 Google Zeitgeist. The most popular "what is..." query? What is love. Baby, don't hurt me.

  2. According to J.L. Needham, Google's manager of public-sector content partnerships, approximately 1,000 federal government websites are inaccessible to search engines -- that is, they lie in the Deep Web.

  3. How well are we preserving our video game heritage? The paper 'Grand Theft Archive': A quantitative Analysis of the State of Computer Game Preservation attempts to answer this question. Their findings aren't very positive.

  4. Although this is three years old, I just came across it this week. Charles Petzold asks, Does Visual Studio Rot the Mind? Petzold's Windows programming books have been a staple of my GUI class since 1997.

  5. Some Harding students have developed this very creative and entertaining Christmas card. Enjoy.

Saturday, November 22, 2008

Google unveils SearchWiki

On Thursday, Google unveiled their newest attempt to deliver the most relevant search results. SearchWiki allows a logged-in user to move results up and down, delete results they don't like, and even leave comments attached to results.

The screenshot below shows a search for "google searchwiki". I've highlighted the Promote, Remove, and Comment icons that appear next to each result. Note that the news result (ranked first) doesn't have any icons- Google always wants that result to be at the top.



I decided to leave a comment on the YouTube video, and as I entered the comment, Google warned me it would be visible to everyone. I also chose to promote the result. Now I can see my comment as well as 16 people who clicked to promote the result and 3 who wanted it removed:



This ability to give personal input to search results is only available to logged-in users. If you are not logged in (you don't see a "Sign out" link in the upper-right corner of the results page), you won't be able to vote.

It's not clear how Google will use personal votes to promote or degrade results, but you have to believe it will be a factor when ranking results in the future, a Human PageRank if you will. It's certainly something other search engines have been thinking about.

Monday, November 17, 2008

Search Engine Development offered this spring

I'll be offering a special class entitled Search Engine Development (COMP 475) this spring semester. This is the second time I've taught this class, but it will likely not be offered again for another couple of years. The pre-requisites are COMP 245 (Data Structures) and 250 (Internet Development I).

We'll be using Java to build a web crawler, indexer, and rank results. At the end of the semester, we'll have a fully-functional web search engine.

Here are some of the topics we'll be covering:
  • Web characterization
  • History of web search
  • Information retrieval (IR)
  • Web crawling
  • Deep web
  • Content indexing
  • Query processing
  • Search results ranking (e.g., PageRank and HITS)
  • Search engine optimization (SEO)
  • Adversarial IR
  • Personalization of search results

Below are some slides I shared last Fri morning in our Computing Seminar on adversarial information retrieval on the Web, one of the topics we'll cover in class.

Friday, August 29, 2008

Fav5

We just completed our first week of 16 for the fall semester. Only 15 more to go! ;-) And with that, my pick of the week's top 5 items of interest:
  1. Microsoft has recently released Photosynth to the public. This technology was developed jointly by Microsoft Live Labs and the University of Washington a few years ago. It takes photos from various view points and synthesizes them into a 3D object when can be rotated. Warning: leaving your browser pointing to a Photosynth page will use up 50% of your processor (at least it did for me), slowing everything down.



  2. The Internet Archive, Library of Congress, and a few other partners are going to archive many of the government's websites at the end of the Bush administration. In the past this has been done by the National Archives and Records Administration (NARA), but they decided it wasn't worth the effort this time around. I'm glad someone else thinks it is.

  3. After less than one year, Yahoo Mash is no more. Wonder how many people are losing a year's worth of social interaction? Someday the same thing is going to happen to MySpace or Facebook, and there are going to be revolts in the streets.

  4. Some interesting research done by facesaerch reveals that individuals usually google Windows and Linux on weekdays and Apple on weekends. Does this mean we "work and suffer with MS and Linux" during the week and "relax with Apple on the weekends"? Or are people just more interested in vegetation on weekends? ;-)

  5. Internet Explorer 8 Beta is available. Like many others, I've made the switch to Firefox, mainly because I liked the add-on features. Apple's sly Safari install didn't convince me to switch. But I'm very tempted to give IE 8 a try. Has anyone tried it out yet?

Saturday, August 02, 2008

Fav5

My pick of the week's top 5 items of interest:
  1. In The 'Anti-Java' Professor and the Jobless Programmers, CS professor Robert Dewar complains of the Java-centric mentality that many American CS graduates have today. Here are two questions Dewar might ask a recent graduate in an interview in order to separate the wheat from the chaff:
    1) You begin to suspect that a problem you are having is due to the compiler generating incorrect code. How would you track this down? How would you prepare a bug report for the compiler vendor? How would you work around the problem?

    2) You begin to suspect that a problem you are having is due to a hardware problem, where the processor is not conforming to its specification. How would you track this down? How would you prepare a bug report for the chip manufacturer, and how would you work around the problem?

  2. In Google Still Not Indexing Hidden Web URLs, Kat Hagedorn and Joshua Santelli follow-up on a paper I published two years ago. Apparently Google is still not doing a good job indexing the OAI-PMH corpus; only 44% of the URLs tested were indexed by Google.

  3. Carnegie Mellon has just introduced a new masters degree: Master of
    Tangible Interaction Design
    . The one year degree combines computer science and architecture.

  4. The Google Blog has a series of posts discussing how Google ranks their results. The discussion is understandable for non-techies and delves into the psychology of web search.

  5. Just for fun: Super Mario Bros. in 20 lines of JavaScript.

Friday, July 25, 2008

Fav5

My pick of this week's top 5 items of interest:
  1. It's official: Google's Knol, a Wikipedia-like source of user-contributed information, is now available to the public. Lots of people are talking about it. I found the interface a little lacking... there's a search box and a list of some randomly selected Knols, but I can't figure out how to browse by subject (of course it's difficult to do this in Wikipedia as well). Also I can't tell how many knols there are yet; a search for "a" shows about 80 knols, and most appear to be health-related. My guess is the tech guys may stick with Wikipedia.

  2. At this week's SIGIR'08 conference, Microsoft Research presented a paper called BrowseRank: Letting Web Users Vote for Page Importance. The paper introduces a new relevance ranking algorithm called BrowseRank. Instead of relying on the web graph to assign web page importance as PageRank does, BrowseRank assigns importance based on users' browsing behavior. A good summary of the paper is at CNET News.

  3. Facebook will soon be using Microsoft's web search technology to give search results and sponsored ads. Currently Google is powering MySpace.

  4. According to a new report from the antivirus company Sophos, they have detected over 16,000 malicious web pages each day in the first half of 2008, most using SQL-injection techniques. Blogspot.com hosts the largest number of malicious web pages, mostly because of how easy it is to setup a blog with this service and to inject

  5. You'd better hurry: The domain ☼.com is still available! Oh, and Google has discovered at least 1 trillion pages on the Web.

Friday, June 13, 2008

Fav5

My pick of the week's top 5 items of interest:
  1. According to a RAND study, the US is still tops in science and technology. According to the report:
    The United States accounts for 40 percent of the total world's spending on scientific research and development, employs 70 percent of the world's Nobel Prize winners and is home to three-quarters of the world's top 40 universities.
    How long will we remain #1? Not very if we continue to make it difficult for foreigners to study here. Americans are just not majoring in technical fields like computer science like they used to.

  2. Someone has actually made a play about the publicly released AOL search queries from 2006. The play focuses on User 927, one of the "anonymous" users whose queries ranged from innocuous to downright deviant. hmmm.... I think I'll skip this one.

  3. Just this week, an Illinois public official dropped his requests to force MySpace to unveil the creators of several "defamatory" profiles that spoofed his identity. Stinks having web-savy enemies.

  4. Matt Cutts confirmed that the file extension of your web pages is very important to Google. The Googlebot will not crawl pages with extensions for binary content like .exe, .dll, .tar, etc. (at least not yet). Matt gives an interesting example: a URL ending with "/web2.0" will be rejected by Googlebot, but "/web2.0/" will be accepted.

  5. Just for fun: try the Bird Flocking Behavior Simulator. The Java applet models the flocking behavior of birds (blue and green flocks). You can place obstacles in their path (left click), give them food (right click), and add predators (red birds which eat the blue and green birds). I had a little fun trapping several of them inside a barrier. (Yes, I do have better things to do with my time. wink)


Saturday, June 07, 2008

Fav5

My pick of the week's top five items of interest:
  1. How valuable is your old high school or college yearbook? At Purdue University, this year's yearbook will be their last. In fact there are only 80 US colleges that still produce yearbooks, down from 100 last year. Interest is declining with part of the blame on social networking sites like Facebook. What today's students don't realize is that 20 years from now, you may not have access to the memories you have now... free services like Facebook have no obligation to retain them indefinitely for you.

  2. Want to save some electricity and don't mind black? Try out Blackle. Blackle displays search results straight from Google, but because they are displaying search results on a primarily black screen, they are using considerably less power than Google which uses white. A similar search engine is called Blaxel, but I don't think the two are affiliated.



    Update 6-9-08:

    My Finnish amigo has pointed out the error in my post. Apparently only CRTs save energy by displaying black, but the newer LCD monitors which most everyone is using today actually take more energy to display black. So Blackle and Blaxel are actually wasting more energy! Thanks for the tip, Timo.

  3. The Digital Lives research project is conducting a quick survey of how individuals store personal computer files, find them in the future, and archive them. If you have 10 minutes and want to contribute to digital preservation research (and possibly win £200 in British Library gift vouchers), please take the survey.

  4. This is kinda cool: Yahoo has opened up its search results page to developers using a new platform called SearchMonkey. They've also developed a listing of numerous SearchMonkey plug-ins in their Yahoo Search Gallery. Do you want to see details of a movie when searching Yahoo? Download the IMDB presentation enhancement.

  5. I'm not totally sure what to make of this: a new search engine called RushmoreDrive tailors search results for the black community only. Apparently Google is too white; African Americans want different search results than European Americans, Asian Americans, etc. While I agree that web search that takes into account the user's profile (e.g., interests, age, gender, location, etc.) are likely to produce better search results, creating a search engine that tailors only to one racial group smacks of racism. While I'm sure this isn't RuchmoreDrive's intention, wouldn't we all agree that a search engine called Whitey.com that was built for the white community only was racist?

Thursday, May 15, 2008

Baseball and search engine queries

This is a really interesting visualization from Spatial Variation in Search Engine Queries, a WWW08 paper by Backstrom & Kleinberg (Cornell) and Kumar & Novak (Yahoo!).


What you are seeing are the "spheres of influence" for various Major League baseball teams. To create this graph, the authors mined the search engine logs from Yahoo! and plotted where the most popular baseball team queries originated from (based on IP address).

The result is a geographical break-down of each team's fan base. The authors point out that the boundaries even follow state boundaries closely: "For instance, in Michigan, across the lake from Chicago but far from Detroit, it is the Tigers, not the Cubs who have the largest following."

It's also interesting to note there are some places like Arkansas where there is no single winner. That's what happens when you live in a state with zero professional teams. Oh, how I miss my childhood years living in Denver... wink

Friday, April 25, 2008

Fav5

My 5 items of interest for the week:
  1. This is very cool: Google News now shows quotes taken from news sources. What has John McCain been saying recently?

  2. Udi Manber, VP of search quality at Google, answers 20 questions about web search. (I made my search engine class read this.)

  3. Researchers at University of California, Santa Cruz are working on long-term archiving using disks rather than tape. Their system, Pergamum, "is a distributed network of intelligent, disk-based, storage appliances that stores data reliably and energy-efficiently."

  4. Video games are becoming a big business ($9.5 billion last year in the US), and universities are listening, creating degree programs for game developers.

  5. Microsoft worked so hard to get their Office Open XML (OOXML) standardize accepted by ISO. And now the bad news: Office 2007 doesn't produce documents that adheres to their standard. My guess is the next service pack changes the file format so it does conform.

Friday, April 11, 2008

Fav5

Today I'm stuck at home trying to recover from the latest disease Ethan has brought home. Here's my 5 items of interest for the week:
  1. Google recently announced their newest attempt at crawling the Deep Web. When their crawler sees a form on a "high-quality site", it enters words and selects various checkboxes and radio buttons. The form is then submitted, and the resulting page is crawled. Very cool. Google also announced the Google App Engine. It's a service that allows you to run your web applications on Google's infrastructure so it can grow to accommodate a large amount of traffic. It's a free service until it exceeds a set disk space and bandwidth quota. Unfortunately, there's a limit of 10,000 developers, and I wasn't quick enough to sign up, so I'll have to wait until they increase the limit.

  2. Joel De Young, one of our Harding CS graduates, was recently interviewed about his company's newest game: Penny Arcade Adventures: On the Rain-slick Precipice of Darkness. Sounds like Joel's HotheadGames is making some buzz.

  3. A study by Univ of Minnesota researchers shows there are two indicators that determine how good of an answer you will receive when posting to question-and-answer sites like Yahoo! Answers: how much you pay, and how many responses you receive.

  4. Jansen, Booth, and Spink have developed a system which can automatically classify a search query as navigational, transactional, and informational. They used their classifier on a large dataset of search engine queries and were accurate 74% of the time. Based on their results, approximately 80% of all search engine queries are informational, 10% transactional, and 10% informational.

  5. Congratulations to the Harding Programming Team on their first place finish at CCSC-MS. We showed them Razorbacks how to program... wink

Tuesday, March 11, 2008

Search engine text snippet

Recently my search engine class was discussing where search engines like Google get the text snippet they display on the search results pages (SERPs). For example, a search for frank mccown using Live.com shows the following SERP:


The text snippet circled in red might be from the full text of the web page, from a description from ODP (see this example), or, as in this case, it might come from the meta description tag like the one I have in my web page:
<meta name="description" content="Frank McCown, Assistant Professor of Computer Science">
Just for fun, I changed this tag a few weeks ago. Although Live has re-crawled my page since then, they have not re-indexed the meta tag. But Google (and Yahoo) have:


This of course is a very misleading description of my web page, but Google doesn't appear to care. A Google search for this description will not make my page appear in the SERP, so the misleading description is only hurting me (and those trying to find me).

I'm a little surprised Google doesn't at least search for the terms from the description in the accompanying web page just to make sure the description is not misleading. I suppose it's better to err on the side of giving your web page author the benefit of the doubt than to deal with all the complaints of those whose tags might be ignored.

Matt Cutts from Google discusses text snippets on YouTube.

Friday, February 01, 2008

Fav5

It's been a few weeks since I posted a Fav5, but it's back. My pick of the week's top 5 items of interest:
  1. Google has just released the Social Graph API. The API will potentially make it much easier to port social graphs from one social network to another.

  2. Here's a quote I like to see:
    "Technology workers remain among the highest paid employees, especially those with management experience and hard-to-find skills."
    - Dice Holdings CEO Scot Melland
    Now if we can just narrow that gender gap a little in the IT field.

  3. Microsoft wants to buy Yahoo for $31 a share. I don't think this is going to pass muster, but if such a deal did take place, Google would seriously need to watch out.

  4. In an interview with VentureBeat, Marissa Mayer (Google’s leading VP in search) discussed the possibility of using social networks for searching the Web.

  5. A few weeks ago I was talking about Googlebombing in my search engine class. Philipp Lenssen writes that although Google has supposedly figured out a way to diffuse the bomb, a search for dangerous cult shows the Scientology website first. I don't think this is technically a googlebomb since a large number of people didn't jointly conspire to create such a link; I think intent is an important criteria in creating a googlebomb.