Meteorologists debate whether recent Chicago snowstorm was 3rd or 4th largest on record

Headlines after the recent Chicago blizzard suggested that the storm had the third largest amount of snow in Chicago history. But when this was later changed to the 4th largest storm, an argument erupted among meteorologists about what exactly counted as part of this particular storm:

After a brief drop to No. 4, the Blizzard of 2011 has now been put back in its rightful spot as the No. 3 worst blizzard in Chicago history.

Earlier in the day, the National Weather Service downgraded the Ground Hog Day Blizzard to 20 inches, taking away .2 inches of snow they say fell hours before the actual blizzard hit. At the same time, they decided that the 1979 storm lasted three days, not the two generally cited. That upped the storm’s total to 20.3 from the 18.8 inches generally credited to the storm…

But during a teleconference with meteorologists from Chicago area media outlets, there was such outcry over the weather service’s decision to lower the total snowfall from this year’s blizzard that the decision was reversed.

“You really are getting into hazardous territory,” WGN meteorologist Tom Skilling warned National Weather Service officials during the teleconference. “To downgrade this storm in any way shape or form is highly subjective. You guys are the arbiters of this, but I don’t agree with it.”…

Allsopp emphasized that these storm totals are more for the public’s benefit than for the record books. The official snow records are listed by calendar days.

Even the weather, data we might consider “hard data,” is open to different interpretations. It is interesting that the final decision went the way of the local forecasters. While Skilling is right to suggest that the decision to downgrade the storm was subjective, wasn’t ranking the storm 3rd also subjective?

Perhaps the key is the final statement in the article: this is for the public, not the record books. In the long run, does it make Chicago area residents feel better or more proud to know that the recent storm was the 3rd largest? If we went by the official snowfall by calendar day, this website suggests the record was 18.6 inches on January 2, 1999.

Trying to count the people on the streets in Cairo

This is a problem that occasionally pops up in American marches or rallies: how exactly should one estimate the number of people in the crowd? This has actually been quite controversial at points as certain organizers of rallies have produced larger figures than official government or media estimates. And with the ongoing protests taking place in Cairo, the same question has arisen: just how many Egyptians have taken to the streets in Cairo? There is a more scientific process to this beyond a journalist simply making a guess:

To fact-check varying claims of Cairo crowd sizes, Clark McPhail, a sociologist at the University of Illinois and a veteran crowd counter, started by figuring out the area of Tahrir Square. McPhail used Google Earth’s satellite imagery, taken before the protest, and came up with a maximum area of 380,000 square feet that could hold protesters. He used a technique of area and density pioneered in the 1960s by Herbert A. Jacobs, a former newspaper reporter who later in his career lectured at the University of California, Berkeley, as chronicled in a Time Magazine article noting that “If the crowd is largely coeducational, he adds, it is conceivable that people might press closer together just for the fun of it.”

Such calculations of capacity say more about the size of potential gathering places than they do about the intensity of the political movements giving rise to the rallies. A government that wants to limit reported crowd sizes could cut off access to its cities’ biggest open areas.

From what I have read in the past on this topic, this is the common approach: calculate how much space is available to protesters or marchers, calculate how much space an individual needs, and then look at photos to see how much of that total space is used. The estimates can then vary quite a bit depending on how much space it is estimated each person wants or needs. These days, the quest to count is aided by better photographs and satellite images:

That is because to ensure an accurate count, some computerized systems require multiple cameras, to get high-resolution images of many parts of the crowd, in case density varies. “I don’t know of real technological solutions for this problem,” said Nuno Vasconcelos, associate professor of electrical and computer engineering at the University of California, San Diego. “You will have to go with the ‘photograph and ruler’ gurus right now. Interestingly, this stuff seems to be mostly of interest to journalists. The funding agencies for example, don’t seem to think that this problem is very important. For example, our project is more or less on stand-by right now, for lack of funding.”

Without any such camera setup, many have turned to some of the companies that collect terrestrial images using satellites, but these companies have collected images mostly before and after the peak of protests this week. “GeoEye and its regional affiliate e-GEOS tasked its GeoEye-1 satellite on Jan. 29, 2011 to collect half-meter resolution imagery showing central Cairo, Egypt,” GeoEye’s senior vice president of marketing, Tony Frazier, said in a written statement. “We provided the imagery to several customers, including Google Earth. GeoEye normally relies on our partners to provide their expert analysis of our imagery, such as counting the number of people in these protests.” This image was taken before the big midweek protests. DigitalGlobe, another satellite-imagery company, also didn’t capture images of the protests, according to a spokeswoman, but did take images later in the week.

Because these images are difficult to come by in Egypt, it is then difficult to make an estimate. As the article notes, this is why you will get vague estimates for crowd sizes in news stories like “thousands” or “tens of thousands.”

Since this is a problem that does come up now and then, can’t someone put together a better method for making crowd estimates? If certain kinds of images could be obtained, it seems like an algorithm could be developed that would scan the image and somehow differentiate between people.

One chart that situates the 2012 Republican presidential contenders

One of the key purposes of a chart or graph is to distill a lot of complicated information into a simple graphic so readers can quickly draw conclusions. In the midst of a crowded field of people who may (or may not) be vying to be the Republican candidate for president in 2012, one chart attempts to do just that.

This chart has two axes: moderate to conservative and insider to outsider. While these may be fuzzy concepts, creator Nate Silver suggests these axes give us some important information:

With that said, it is exceptionally important to consider how the candidates are positioned relative to one another. Too often, I see analyses of candidates that operate through what I’d call a checkbox paradigm, tallying up individual candidates’ strengths and weaknesses but not thinking deeply about how they will compete with one another for votes.

Silver then goes on to explain two other pieces of information for each candidate that is part of the circle used to place each candidate on the graph: the color indicates the region and the size of the circle represents their relative stock on Intrade.

Based on this chart, it looks like we have a diagonal running from top left to bottom right, from moderate insider (Mitt Romney) to conservative outsider (Sarah Palin) with Tim Pawlenty and Mike Huckabee trying to straddle the middle. We will have to see how this plays out.

But as a statistics professor who is always on the lookout for cool ways of presenting information, this is an interesting graphic.

Exactly how many American homes are vacant?

Two bloggers have a disagreement about how many vacant homes there are in the United States. Check out the debate and the comments below.

The moral of the story: one still needs to interpret statistics and what exactly they are measuring. The different between 11% and 2% is quite a lot: the first figure suggests 1 out of 10 housing units are vacant while the second figure suggests it is 1 out of 50. If you look at Table 1 of this Census Bureau release regarding housing figures from Quarter 4, it looks like the vacancy rate is 2.7%. But there may be confusion based on Table 3 which suggests the vacancy for all housing units is roughly 11% for year-round units. And later in the release, page 11 of the document, gives the formula for the vacancy calculation and an explanation: “The homeowner vacancy rate is the proportion of the homeowner inventory that is vacant for sale.”

There are some other figures of note in this document. Table 4 shows that the homeownership rate is at 66.5%, down from a peak of 69.2% in the fourth quarter of 2004. (It is interesting to note that this rate peaked a couple of years before the housing market is popularly thought to have gone downhill. What happened between Q4 2004 and the start of the larger economic crisis? Table 7 has homeownership rates by race: the white rate has dropped 1.1% since 1Q 2007 while Blacks and Latinos have seen bigger drops (3.2% and 3.3%).

An example of statistics in action: measuring faculty performance by the grades students receive in subsequent courses

Assessment, whether it is for student or faculty outcomes,  is a great area in which to find examples of statistics. This example comes from a discussion of assessing faculty by looking at how students do in subsequent courses:

[A]lmost no colleges systematically analyze students’ performance across course sequences.

That may be a lost opportunity. If colleges looked carefully at students’ performance in (for example) Calculus II courses, some scholars say, they could harvest vital information about the Calculus I sections where the students were originally trained. Which Calculus I instructors are strongest? Which kinds of homework and classroom design are most effective? Are some professors inflating grades?

Analyzing subsequent-course preparedness “is going to give you a much, much more-reliable signal of quality than traditional course-evaluation forms,” says Bruce A. Weinberg, an associate professor of economics at Ohio State University who recently scrutinized more than 14,000 students’ performance across course sequences in his department.

Other scholars, however, contend that it is not so easy to play this game. In practice, they say, course-sequence data are almost impossible to analyze. Dozens of confounding variables can cloud the picture. If the best-prepared students in a Spanish II course come from the Spanish I section that met at 8 a.m., is that because that section had the best instructor, or is it because the kind of student who is willing to wake up at dawn is also the kind of student who is likely to be academically strong?

It sounds like the relevant grade data for this sort of analysis would not be difficult. The hard part is making sure the analysis includes all of the potentially relevant factors, “confounding variables,” that could influence student performance.

One way to limit these issues is to limit student choice regarding sections and instructors. Interesting, this article cites studies done at the Air Force Academy, where students don’t have many options in the Calculus I-II sequence. In summary, this setting means “the Air Force Academy [is] a beautifully sterile environment for studying course sequences.”

Some interesting findings both from the Air Force Academy and Duke: students who were in introductory/earlier classes that they considered more difficult or stringent did better in subsequent courses.

Finding the right model to predict crime in Santa Cruz

Science fiction stories are usually the setting when people talk about predicting crimes. But it appears that the police department in Santa Cruz is working with an academic in order to forecast where crimes will take place:

Santa Cruz police could be the first department in Northern California that will deploy officers based on forecasting.

Santa Clara University assistant math professor Dr. George Mohler said the same algorithms used to predict aftershocks from earthquakes work to predict crime.”We started with theories from sociological and criminological fields of research that says offenders are more likely to return to a place where they’ve been successful in the past,” Mohler said.

To test his theory, Mohler plugged in several years worth of old burglary data from Los Angeles. When a burglary is reported, Mohler’s model tells police where and when a so-called “after crime” is likely to occur.

The Santa Cruz Police Department has turned over 10 years of crime data to Mohler to run in the model.

I wonder if we will be able to read about the outcome of this trial, regardless of whether the outcome is good or bad. If the outcome is bad, perhaps the police department or the academic would not want to publicize the results.

On one hand, this simply seems to be a problem of getting enough data to make accurate enough predictions. On the other hand, there will always be some error in the predictions. For example, how could a model predict something like what happened in Arizona this past weekend? Of course, one could include some random noise into the model – but these random guesses could easily be wrong.

And knowing the location of where crime would happen doesn’t necessarily mean that the crime could be prevented.

US land use statistics from the 2011 Statistical Abstract of the United States

I have always enjoyed reading or looking through almanacs or statistical abstracts: there is so much interesting information from crop production to sports results to country profiles and more. Piquing my interest, the New York Times has a small sampling of statistics from 2011 Statistical Abstract of the United States.

One reported statistic struck me: “The proportion of developed land reached a record high: 5.6 percent of all land in the continental U.S.” At first glance, I am not surprised: a number of the car trips I make to visit family in different locations includes a number of hours of driving past open fields and forests. Even with all the talk we hear of sprawl, there still appears to be plenty of land that could be developed.

But the Statistical Abstract allows us to dig deeper: how exactly is American land used? According to 2003 figures (#363, Excel table), 71.1% of American land is rural with 19% total and 20.9% total being devoted to crops and “rangeland,” respectively. While developed land may have reached a record high (5.6%), Federal land is almost four times larger (20.7%).

Another factor here would have to be how much of the total land could actually be developed. How much of that rural land is inaccessible or would require a large amount of work and money to improve?

So whenever there is a discussion of developable land and sprawl, it seems like it would be useful to keep these statistics in mind. How much non-developed land do we want to have as a country and should it be spread throughout the country? How much open land is needed around cities or in metropolitan regions? And what should this open land be: forest preserve, state park, national park, open fields, farmland, or something else?

A humorous yet relevant comment from the scientific past: “Oh, well, nobody is perfect.”

When I look at sociology journal articles from the past, a few things strike me: the lack of high-powered statistics and a simplicity in explanation and research design. In the current world of publishing demands and the push for always high-quality, ground-breaking work, these earlier articles look like they were from a more innocent era.

I was reminded of this by a recent Wired post. In this case, a geology journal had published an article in the early 1960s and another scholar had responded in print to this article by pointing out a mistake on the part of the original authors. This is not uncommon. What does look particularly uncommon is the response by the original authors: “Oh, well, nobody is perfect.”

In a perfect world, isn’t this how science is supposed to work: just admit your mistakes, don’t repeat them, and move on? But I can’t imagine that many current scholars could give such a reply, perhaps in fear that their career or reputation would be in jeopardy. And in the world of scientific journals, is this sort of back and forth (with candidness) even possible much of the time?

I also infer a sense of humility on the part of the original authors. Instead of going on for pages about how their mistake was defensible or trying to pass the blame, a quick one-liner admits the mistake, diffuses the situation, and everyone can move on.

Trying to understand China’s economy with a lack of statistics

Megan McArdle writes about the issue of a lack of comprehensive data to understand what is happening with China’s economy:

But central planners badly need good, comprehensive data.  Once you limit the autonomy of local nodes to make decisions, you need some sort of massive data set to overcome information loss as decisions move up the hierarchy.

Libertarians often use this to argue against any sort of central planning, but that’s not the point of this post.  All modern economies engage in some level of planning, whether it is monetary policy or infrastructure construction.  It was in response to the problems of managing production during World War I that economists first conspired to create US economic statistics.

The Chinese government is extremely enthusiastic about managing their economy, and they put a lot of thought into it.  But the lack of good statistics on economic performance makes an already near-impossible challenge even more daunting.

It is remarkable to recognize how much data there is out there these days in the United States. And even with all that data, it is often not always clear what should be done – government officials, investors, journalists, and citizens need to know how to interpret the data and figure out how to respond.

What would it take to get comprehensive data in China?

h/t Instapundit

The statistical calculations used for counting votes

Some might be surprised to hear that “Counting lots of ballots [in elections] with absolute precision is impossible.” Wired takes a brief look at how the vote totals are calculated:

Most laws leave the determination of the recount threshold to the discretion of registrars. But not California—at least not since earlier this year, when the state assembly passed a bill piloting a new method to make sure the vote isn’t rocking a little too hard. The formula comes from UC Berkeley statistician Philip Stark; he uses the error rate from audited precincts to calculate a key statistical number called the P-value. Election auditors already calculate the number of errors in any given precinct; the P-value helps them determine whether that error rate means the results are wrong. A low P-value means everything is copacetic: The purported winner is probably the one who indeed got the most votes. If you get a high value? Maybe hold off on those balloon drops.

A p-value is a key measure in most statistical analysis – it provides a measure of how much error is in the data and whether the obtained results are just by chance or whether we can be fairly sure (95% or more) the statistical estimation represents the whole population.

So what is the acceptable p-value for elections in California?

I would be curious to know whether people might seize upon this information for two reasons: (1) it shows the political system is not exact and therefore, possibly corrupt and (2) they distrust statistics altogether.