Wednesday, April 29, 2009

Upload an image in PHP

I created this function for my Internet Development students which saves a single uploaded image to disk. Example:
// Assuming the web server has write permissions to /mydir
SaveUploadedImage("/mydir/myimage.png");

The function can easily be modified to handle multiple filenames (change the parameter to accept an array of filenames and modify the final foreach block). Note that this is modified code from the webdeveloper.com forum. If you want to know more about uploading files in PHP, check out the PHP - File Upload tutorial.


// Return empty string if uploaded image is successfully saved as
// $image_filename or an error message. $image_filename should be
// saved in a directory that the web server can write to.
function SaveUploadedImage($image_filename)
{
// This function is greatly modified code from
// http://www.webdeveloper.com/forum/showthread.php?t=101466


// Possible PHP upload errors
$errors = array(1 => 'php.ini max file size exceeded',
2 => 'html form max file size exceeded',
3 => 'file upload was only partial',
4 => 'no file was attached');

// Store nonempty files in the active_keys array
$active_keys = array();
foreach ($_FILES as $key => $file)
{
if (!empty($file['name']))
$active_keys[] = $key;
}

// Check at least one file was uploaded
if (count($active_keys) == 0)
return 'No files were uploaded';

// Check for standard uploading errors
foreach ($active_keys as $key)
{
if ($_FILES[$key]['error'] > 0)
return $_FILES[$key]['tmp_name'] . ': ' . $errors[$_FILES[$key]['error']];
}

// See if the file we are working on really was an HTTP upload
foreach ($active_keys as $key)
{
if (!is_uploaded_file($_FILES[$key]['tmp_name']))
return $_FILES[$key]['tmp_name'] . ' not an HTTP upload';
}

// Make sure the image uploaded appears to be an actual image
foreach ($active_keys as $key)
{
if (!getimagesize($_FILES[$key]['tmp_name']))
return $_FILES[$key]['tmp_name'].' is not an image';
}


// Save every uploaded file to the same filename (normally we'd want to
// save each file with its own unique name, but we are assuming there
// is only one file).
foreach ($active_keys as $key)
{
if (!move_uploaded_file($_FILES[$key]['tmp_name'], $image_filename))
return 'receiving directory (' . $image_filename . ') has insufficient permission';
}

// If you got this far, everything has worked and the file has been successfully saved.

return '';
}

Wednesday, April 22, 2009

Nutch, Sitemaps, and Google's findings

My search engine class is winding down, but our final project is to implement a Sitemap Protocol parser for Nutch, a popular open-source search engine. I mentioned a while back that Nutch is not for wimps... my students would certainly vouch for the huge learning curve to making code modifications. I've even had to scale back how much work my students do because of the complexity of changes required. I'm going to do the difficult part of integrating their code with the innards of Nutch sometime in the next few weeks.

The reason I mention our Sitemap project is that WWW 2009 is meeting in Madrid this week, and a paper entitled Sitemaps: Above and Beyond the Crawl of Duty is being presented today by Uri Schonfeld (UCLA) and Narayanan Shivakumar (Google). This is the first paper to report on widespread usage of Sitemaps in the Web using Google's crawling history.

Schonfeld & Shivakumar report that Sitemaps were used by approximately 35 million websites in late 2008, exposing several billion URLs. 58% of the URLs included last modification dates, 7% included change frequency, and 61% a priority. About 76.8% of Sitemaps used XML formatting, and only 3.4% used plain text. Interestingly, 17.5% of Sitemaps are formatted incorrectly.

The figure below represents how many URLs Google discovered via Sitemaps (red) vs. regular crawling (green) for cnn.com. Notice that on any given day, more URLs could normally be discovered via Sitemaps.



Another interesting figure (below) shows when a URL was discovered via Sitemaps vs. regular web crawling for cnn.com. In most cases URLs were discovered at the same rate, but there are a number of them (dots below the line) that were discovered via Sitemaps much earlier than web crawling.


CNN's website is not typical. Schonfeld & Shivakumar report that in a dataset of 5 billion+ URLs, 78% were discovered via Sitemaps first compared to 22% via web crawling.

The paper also describes an algorithm that can be used by search engines to prioritize URLs discovered via web crawling and Sitemaps as well. I've covered the high-lights, but I recommend you read the paper if you're interested in some of the finer details.

Friday, April 17, 2009

Looks can be deceiving

It's been busy around here... Spring Sing, Easter, Tax Day, etc.

This morning Steve Baber presented a devotional thought at our computing seminar that I thought I'd share with you all. He talked about how easily our eyes can be deceived. Are you seeing a man on the left or the word Liar?

This is especially true when it comes to how we perceive others. How often do you catch yourself judging someone by their looks, their clothes, the house they are living in and the car they are driving?

James warns against this practice in James 2:1-4:
My brothers, as believers in our glorious Lord Jesus Christ, don't show favoritism. Suppose a man comes into your meeting wearing a gold ring and fine clothes, and a poor man in shabby clothes also comes in. If you show special attention to the man wearing fine clothes and say, "Here's a good seat for you," but say to the poor man, "You stand there" or "Sit on the floor by my feet," have you not discriminated among yourselves and become judges with evil thoughts?
Here are three contestants from Britain's Got Talent that feature some contestants whose appearance is misleading: Susan Boyle, Paul Potts, and Andrew Johnston.

Friday, April 03, 2009

Day 2 at DigCCurr 2009

This was a full day of presentations. One of my favorite panels was on personal digital archiving with Jeremy John, Cathy Marshall, David Pearson, and Andreas Rauber. My presentation seemed to go well... the room was packed with people sitting on the floor. Overall I was very pleased with the conference and met a good number of interesting people.

After the conference ended, I took advantage of the 70 degree weather and took a walk around the UNC campus. I then headed up to Franklin St. where a mass of well-dressed college students were gathering. (Franklin St. is where all the cool places to hang out are located. It's also the place where students jump over bonfires after big UNC wins.) The 2 mile walk back to the hotel was fantastic... the homes on Franklin St. are some of the most beautiful and unique homes I've seen.

Now I'm sitting in my hotel room (11 pm) missing my family while a number of college students stand outside my window talking as if no one else but them were at the hotel. It'll be a lot worse tomorrow night if UNC wins, but I'll be back in Searcy by then.

Update:

Looks like UNC is going to the championship, and the partying on Franklin Street continues.

Thursday, April 02, 2009

I'm at DigCCurr 2009

I flew into Raleigh/Durham late last night, and today I am attending the DigCCurr 2009 conference (Digital Curation Practice, Promise and Prospects) in Chapel Hill, NC. Tomorrow I'll be presenting a paper based on my summer work at LANL: Everyone is a Curator: Human-Assisted Preservation for ORE Aggregations. This was work I did with Herbert Van de Sompel and Michael Nelson (my former adviser at ODU).

There are 270 people registered for DigCCurr, but I only know a handful of them. So it was good to see Michael Nelson this morning getting coffee in the lobby... I had no idea he'd be here. Of course now I have to spruce up my presentation and remove my disparaging remarks about Herbert and Michael. wink

Tuesday, March 31, 2009

Happy birthday, Ethan!

My son is 2 today! He's becoming a big boy, and I'm very proud of him.

Monday, March 30, 2009

Enrollment in computer science is finally increasing

Good news: A survey by the Computing Research Association shows that enrollment in computer science courses in the 2007-2008 academic year was up 6.2%, the first increase since the dot-com bust six years ago. The number of new undergraduates majoring in CS is up 9.5%.

Bad news: Women still only receive 11.8% of CS degrees.

So why are enrollments increasing? Fear of the bad economy? The coolness of the iPhone?

We haven't yet seen an increase here at Harding, but I'm betting we will soon.

Wednesday, March 25, 2009

The Web in a box

The Internet Archive and Sun Microsystems have just announced the launching of a new data center that stores IA's entire web archive and serves the Wayback Machine. According to Brewster Kahle:
This 3 Petabyte (3 million gigabyte) datacenter will handle the 500 requests per second as it takes over the full Wayback load.

Tuesday, March 17, 2009

Workshop on Innovation in Digital Preservation (InDP 2009)

In conjunction with JCDL 2009, Andreas Rauber and I will be hosting the first workshop on Innovation in Digital Preservation (InDP 2009). We are soliciting full and short research papers as well as position papers. Read more about the workshop below and on the workshop website.

Digital Preservation (DP) research is often driven by traditional needs and approaches to solve the challenges arising. This is partially due to the rather traditional settings in which the challenges of obsolescence of digital objects have first been identified and dealt with, as well as partially due to the high levels of quality and auditability that these mostly very professional settings require.

But increasingly we are facing non-traditional DP challenges, ranging from non-traditional data collection (such as the Web, especially Web 2.0) to non-traditional institutions and actors, such as SMEs or private/home users. Additionally, non-traditional approaches to maintain digital objects, such as retargetable binary code or self-aware objects are gaining momentum.

This full-day workshop aims to provide a forum where researchers can share and discuss the latest innovations in DP by non-traditional methods. Topics include but are not limited to:

  • Personal archiving and personal information management
  • Archiving Web 1.0, 2.0, and Deep Web
  • Innovative approaches to preservation actions
  • Self-aware objects
  • Archiving solutions for small institutions
  • Binary retargetable code
  • Disaster recovery
  • Theoretical models of information preservation
  • Value of information and forgetting
  • Preserving electronic art

InDP 2009 will be held in conjunction with the ACM/IEEE-CS Joint Conference on Digital Libraries (JCDL) in Austin, Texas on June 19, 2009.

Saturday, March 14, 2009

Map of Science

This past summer I worked with the Digital Library Research and Prototyping Team at LANL. One of the projects they were working on, a "map of science", was just featured on the Wired Science Blog. The graph below is based on a massive collection of scholarly usage data (click on the image to get a more detailed look at it). Basically, the graph shows how users accessing scholarly work in one field (e.g., Cognitive Science) may also access work in another related field (e.g., Sports Medicine).



You can find more information about this work in:

Clickstream Data Yields High-Resolution Maps of Science by Johan Bollen, Herbert Van de Sompel, Aric Hagberg, Luis Bettencourt, Ryan Chute, Marko A. Rodriguez, Lyudmila Balakireva. Public Library of Science ONE, March 11, 2009.

Wednesday, March 11, 2009

Harding and the Economy

An article just published in the Christian Chronicle reports how the economy is affecting Harding University and some of our other sister institutions. Cascade College is closing their doors at the end of the spring semester (see the screen shot from their website below). Pepperdine University is eliminating 50 full- and part-time staff members as well as men's track and the women's swimming and diving program.


Things at Harding aren't quite as dim. The endowment is down, but no staff or faculty jobs are being cut. What the CC didn't note is that next year's enrollment numbers look really good. There is a budget freeze, and it's been rumored we will not receive any pay raises next year, but there's nothing to loose sleep over.

In general, post-secondary education is usually a winner in tough economic times. Individuals out of work will re-tool to make themselves more competitive. Some states are even putting more money into higher education, realizing the positive, long-term economic impact it can have.

At $423 a credit hour, Harding is not cheap, but it is less expensive than many private universities and many state universities. That's going to help us weather the storm.

Everyone is going to need to tighten their belts a little, but Lord willing, Harding is going to emerge from this economic downturn intact. I pray the same is true for our sister institutions.

Sunday, March 08, 2009

Back from SIGCSE 2009

Scott and I returned to Searcy yesterday after attending the morning paper presentations. We both agreed that this was one of the better academic conferences we had attended.

As I mentioned before, collaborative learning was a huge theme at SIGCSE. Some of the big curriculum pushes included game programming, robotics, and parallel programming. A number of the presentations stressed how introducing games, graphics, and robotics in CS1 would probably help in retention and expanding interest to CS minorities (mostly women). However, many of the presentations also failed to show conclusive evidence that this was true. In fact one of the presenters admitted that he spent so much time discussing peripheral concepts in regards to game programming that there was no time left to teach some of the core concepts.

Ideally, I think it would be very worthwhile for our department to offer intro to programming courses that use either robotics, graphics, or games in their approach. Then incoming freshmen could pick the course that most interested them. Certainly it would be a good recruitment tool. The only problem is we don't have enough majors or teachers to offer so many courses. I may at least try to mix in more of these attention getters in my CS1 course next fall.

Something else I heard repeatedly was how we should be using Python in CS1. I can certainly see some benefits of doing so, but there's a number of benefits to teaching C++ first. One presentation showed that using Python in CS1 was no better than using C++ at preparing students for CS2. Until we see some solid research showing that Python is in fact better at increasing retention, I think we should stay where we are.

Next year's SIGCSE is in Milwaukee where the average high is 34 F this time of the year. Brr.

Friday, March 06, 2009

I'm at SIGCSE 2009

My spring break started a little early this year. Scott Ragsdale and I drove down to Chattanooga, TN, on Wed for SIGCSE 2009. This is the first time I’ve attended the conference, and so far I’ve been very impressed, both with the conference and with Chattanooga.

SIGCSE brings together computer science educators from around the globe to share and discuss the latest in computing education. There are around 1200 participants this year and numerous talks and workshops to choose from.

On Wed night Scott and I attended a workshop entitled Engaging Student Learning Through Virtual World Programming. It was mainly about introducing the world of Second Life. We created avatars and then learned how to navigate the virtual world, create objects, and use the Linden Language scripting language. I wasn’t very impressed with Second Life... it felt like a very dysfunctional, overly sexualized place that I didn’t want to be in for very long (although flying is kinda fun). It’s hard for me to imagine my students liking it much either.

Today’s favorite buzz word was “collaborative learning”. Most presenters felt obliged to use it at least twice in their talk. Despite the overuse, I was quite convinced that students do learn more effectively when they are teaching each other. I’m also convinced that I need to make some changes to my intro to programming classes that makes better use of this fact.

At one of the sessions, I learned how I will not be able to teach iPhone development to my GUI students next fall. I was hoping to teach Objective C and iPhone programming in the final five weeks of the course, but the learning curve is just too steep to teach effectively in a 5 week period, especially when compared with Windows Mobile programming.

I’m too exhausted to list everything I saw today, but it was very worthwhile. And tonight’s reception at the Tennessee Aquarium was fantastic.

(This entry was written Thurs night.)

Saturday, February 28, 2009

My first knol

As I mentioned yesterday, I just wrote my first knol entitled Introduction to Web Search Engines. This article was originally meant for my Internet Development class. I wanted to have them understand some of the technical issues of how search engines work because it affects how a website should be developed to make it Google-friendly. At the same time, I didn't want the article to be overly technical... it needed to convey just enough technical information so my students would get the big picture. Whether I hit that sweet-spot or not is debatable.

I couldn't find any similar articles on the Web (this is close), so I thought it would be useful to put one out there, and I've been eager to try out Google's new knol service. Despite my many frustrations yesterday, it wasn't too bad. I didn't find myself needing to manipulate the raw HTML too much, and the ability to add and manage references was very intuitive.

Yes, the amazing graphics are my own. I know I got skillz. wink

Friday, February 27, 2009

Why I (sometimes) hate cloud computing

Cloud computing offers a lot of positive benefits, namely access to data from anywhere.

But today I hate it.

I realize "hate" is a strong word, and I rarely break it out, but today I must.

I decided to write my first Knol yesterday. Google's Knol system uses an online editor and saves your data in their systems (i.e., in the cloud). After working on my article for some time, the system started having problems saving, but a warning message at the top of the screen warned that Knol would be down for about an hour. So after waiting the hour, I returned to the system and was able to make a lot of progress.

Or so I thought.

This morning when I returned to my article, only the first two paragraphs remained. I accessed the revision system to see if there were previous versions that contained my complete article, but every revision was the same: just two paragraphs. This really jolted me because I had repeatedly saved every so often just so something like this wouldn't happen. I never received a single error message after pressing Save.

Thankfully I had been smart and saved a copy of my article in Microsoft Word... my previous experience with cloud computing told me such a move would be smart. So I didn't loose my cool too much... I just copy and pasted my stuff back into Knol and then worked on some formatting issues.

About 10 minutes after I started editing, I got this error message warning me that my session had expired:


The message suggested I refresh the page. Knowing what I know, this is usually not a good way to repair an "expired" session. But there was nothing else I could do. Sure enough, after refreshing the page, my content was all gone. Back to two paragraphs.

Now I'm hot.

After taking a little walk to calm down, I returned to my office and decided to persist. I think I'm almost done with my Knol now, but I'm still feeling raw. If this is what cloud computing is going to be like, I'd rather stay on the ground. I've never had Word lose my document because my session expired.

Now you may be thinking this is an isolated issue with Knol, but I have had similar losses with Blogger (losing entire blog posts when their system was temporarily inaccessible) and Google Calendar (losing a number of appointments I had typed in but were apparently never saved). Maybe this is more a Google problem than a cloud problem, but if Google can't seem to get it right, who will?

Friday, February 20, 2009

Happenings at Harding

There's a lot going on around here at Harding, so a quick post to bring you up to date:
  1. This weekend in the Benson Auditorium, Hal Runkel will be presenting ScreamFree Parenting. Read more about it here.


  2. Jimmy Allen has retired from teaching. If you are a Harding alum, there's a good chance you took his class on Romans. But Dr. Allen is still around... I played basketball with him just a few weeks ago.


  3. Construction of a new Pizza Hut started a few weeks ago, about 50 yards from the Beebe-Capps entrance into campus. Why Harding was unable to purchase the land, I don't know. It's gonna look tacky, but at least it isn't a used car lot.


  4. If you haven't been spammed by the Alumni Office recently, here's your chance to win an all-expense paid Homecoming weekend at Harding. All you have to do is give the Alumni Office the email address of 5 "missing" alumni. The campaign is called Six Degrees of Harding University.


  5. A group of Harding students joined with others in helping storm victims in north-west Arkansas.


  6. The Harding programming team headed by David Farrow smashed the other business programming teams this past weekend in the Axiom programming contest held in Conway, AR. David was actually the lone programmer since his two teammates were just buddies who were there for moral support. David confirmed my suspicion that a CS-trained programmer is 10 times more effective than a business major that knows how to program. wink


  7. Ethan saw his first Harding basketball game last night. He cheered on the Lady Bisons as they won their 6th consecutive win.

Monday, February 16, 2009

How *not* to implement online security

I have an online account with a bank which shall remain nameless. Let's just call them Amgirl Direct. They use a really "sophisticated" security system which they have apparently leased from a third party named Information Technology, Inc.

Here's how I login to my account:
  1. First I must enter my 9 digit number account which I have not been able to memorize because I use it once a month. So I have to search for it in my email each time.

  2. I then am told I need to answer a security question because the bank doesn't recognize my IP address. (Of course it doesn't... my home computer is assigned a new one periodically by my ISP.) The security question is always the same:

    What is your high school mascot?


    I went to two high schools, and I have no idea which mascot I entered originally. But it doesn't matter... if I type in the mascot of either high school, the answer is always wrong.

  3. After I answer the first security question wrong twice, I'm finally asked for my mother's middle name. Thankfully it recognizes my answer to this question.

  4. Next I'm asked to enter my password. But supposedly my password has something to do with an "authentication image" which is always a white vase. I have no idea why. There's no link to an explanation. It's always the same image, and I have only one password, so I'm left wondering what-in-the-world this vase has to do with anything.

    (Note Information Technology, Inc.'s proud declaration of ownership for their system.)

  5. After entering my password, I'm finally logged in (usually). But be careful! If you ever click the back or forward browser button at any time, you are presented with this most unhelpful error message:

    Error

    A Security Error Has Occurred. Your Online Session Has Expired.
    Possible Reasons Include Double Clicking A Link Or Pressing The Browser's Back Forward Or Refresh Buttons.
    Return To The Login Page To Continue Your Session.


    They "expire" my session for using navigation buttons that most users are accustomed to using. And there is no link to a login page... you just have to re-type Amgirl Direct's original URL and proceed through the steps above once again.

I keep asking myself, is using an online bank with this lousy of a system really worth the 2.25% APY?

Specifying canonical URLs

Last week the big three search engines (Google, Yahoo, and Live Search) announced their support for a new HTML attribute value which will help prevent search engines from indexing duplicate content. Search engines naturally want to avoid crawling and indexing duplicate content because it lessens the quality of search result pages. Google's Webmaster Central Blog has a good write-up about the new rel="canonical" attribute value.

Essentially, the new attribute value will allow a webmaster to tell a web crawler to ignore a page if it is accessible from another URL. So if a I have a single page that is accessible at URLs A, B, and C, I can tell the web crawler that URLs B and C are pointing to the same content as A by placing the following code in the head element of the page:

<link rel="canonical" href="http://foo.com/A" />

When the web crawler grabs the pages using URLs B or C, it will find the given canonical URL A in the header and therefore ignore the contents of the pages since they duplicate page A.

Of course the entire mechanism requires a willing and competent webmaster to implement it. Webmasters who are very concerned about SEO are likely to use it since it will help bolster the PageRank of certain pages. But the rest of us who aren't concerned about our rankings can safely ignore this new functionality.

See also rel="nofollow".

Friday, February 13, 2009

Feel the Nutch burn...

This week I spent all 3 hours of class time showing my Search Engine students how to install, run, and modify Nutch, an open-source search engine written in Java. Since Nutch is new to me as well, I spent several hours last week trying to get familiar enough to walk my students through the time-consuming, error-prone, and laborious process of getting Nutch to run on Windows and in Eclipse.

I have labeled my newbie experience the "Nutch burn." And boy does it.

I followed a couple of tutorials that were pretty helpful, but I ran into several problems that required me to scour the Web looking for solutions. After much trial and effort, I was able to overcome and make some modifications to Nutch in Eclipse. I also got my modifications to run from the command line.

The barrier to entry is so high and the learning curve so steep that it makes me wonder... there's got to be a better way. The goal is for my class to make a major contribution to Nutch. Maybe our contribution could be to make the initial install/edit process just a little easier.

Tuesday, February 10, 2009

Ben Stein speaking at Harding University tonight

Ben Stein will be presenting his thoughts on the economy, etc. tonight at Harding University (7:30pm in the Benson Auditorium). Stein is well-known as an author, entertainer, and humorist. He is especially well-known for the hilarious "Bueller...? Bueller...?" scene in Ferris Bueller's Day Off and more recently for the controversial movie Expelled.

In an ironic twist, Stein was recently uninvited as the commencement speaker at the University of Vermont (technically he uninvited himself). Apparently strong opposition arose from some in the UVA academic community because of Stein's stance in Expelled. The theme of Expelled is that the academic community will shut you out for offering an opinion that differs from the status quo.

Disclaimer: I have not seen Expelled and have no opinion for or against the movie.

Update:

After seeing the talk, here are my impressions: Smart & witty. Loves Sonic. Extremely conservative. Fiercely loyal to Nixon. Not scientifically inclined.