Comparing stories and statistics

A mathematician thinks about the differences between stories and statistics and the people who prefer one side over another:

Despite the naturalness of these notions, however, there is a tension between stories and statistics, and one under-appreciated contrast between them is simply the mindset with which we approach them. In listening to stories we tend to suspend disbelief in order to be entertained, whereas in evaluating statistics we generally have an opposite inclination to suspend belief in order not to be beguiled. A drily named distinction from formal statistics is relevant: we’re said to commit a Type I error when we observe something that is not really there and a Type II error when we fail to observe something that is there. There is no way to always avoid both types, and we have different error thresholds in different endeavors, but the type of error people feel more comfortable may be telling. It gives some indication of their intellectual personality type, on which side of the two cultures (or maybe two coutures) divide they’re most comfortable. I’ll close with perhaps the most fundamental tension between stories and statistics. The focus of stories is on individual people rather than averages, on motives rather than movements, on point of view rather than the view from nowhere, context rather than raw data. Moreover, stories are open-ended and metaphorical rather than determinate and literal…

I’ll close with perhaps the most fundamental tension between stories and statistics. The focus of stories is on individual people rather than averages, on motives rather than movements, on point of view rather than the view from nowhere, context rather than raw data. Moreover, stories are open-ended and metaphorical rather than determinate and literal.

This is a good discussion and one that I think about often while teaching statistics or research methods. Stories are often easy for students to grab unto, particularly if told from an interesting point of view. In the end, these stories (particularly the “classics”) have the ability to illuminate the human condition or interesting concerns but don’t have the same ability to offer more concrete overviews of the typical or common experience. Statistics do offer a different lens for viewing the world, one where individual experiences are muted in favor of data about larger groups. Both can miss important features of the reality around us but offer different angles for tackling similar concerns.

Both have their place and I would suggest both are necessary.

The presence of error in statistics as illustrated by basketball predictions

TrueHoop has an interesting paragraph from this afternoon illustrating how there is always error in even complicated statistical models:

A Laker fan wrings his hands over the fact that advanced stats prefer the Heat and LeBron James to the Lakers and Kobe Bryant. It’s pitched as an intuition vs. machine debate, but I don’t see the stats movement that way at all. Instead, I think everyone agrees the only contest that matters takes place in June. In the meantime, the question is, in clumsily predicting what will happen then (and stats or no, all such predictions are clumsy) do you want to use all of the best available information, or not? That’s the debate about stats in the NBA, if there still is one.

By suggesting that predictions are clumsy, Abbott is highlighting an important fact about statistics and statistical analysis: there is always some room for error. Even with the best statistical models, there is always a chance that a different outcome could result. There are anomalies that pop up, such as a player who has an unexpected breakout year or a young star who suffers an unfortunate injury early in the season. Or perhaps an issue like “chemistry,” something that I imagine is difficult to model, plays a role. The better the model, meaning the better the input data and the better the statistical techniques, the more accurate the predictions.

But in the short term, there are plenty of analysts (and fans) who want some way to think about the outcome of the 2010-2011 NBA season. Some predictions are simply made on intuition and basketball knowledge. Other predictions are made based on some statistical model. But all of these predictions will serve as talking points during the NBA season to help provide some overarching framework to understand the game by game results. Ultimately, as Gregg Easterbrook has pointed out in his TMQ column during the NFL off-season, many of the predictions are wrong – though the makers of the predictions are not often punished for poor results.

Comparing greatness of players past and present an enjoyable part of sports fandom

As the NBA season approaches, discussion this week has centered on the relative status of several players: Kobe Bryant, Kevin Durant, LeBron James, and Michael Jordan. While the first three players in this list were involved in a question about who is the best current player and potential MVP, Jordan also has been inserted in the discussion due to his starring role in NBA2K11 and comments he made about the number of points he could score if he played today when more fouls are called.

Several quick thoughts come to mind:

1. The new era of statistics in sports offers more opportunities to make comparisons of players across different eras, particularly if you can control for certain features of the game at each time period (like the average pace in basketball).

2. I wonder how much current players think about issues like these. Fans seems to like these discussions. It allows the average guy sitting on the couch to say, “my guy, whoever that may be, can match up or beat your guy.”

3. Jordan, like some other old players, still likes to be part of these discussions.

4. All of these discussions are magnified by the non-stop media attention for sports these days. I can hear it on local sports talk radio which all sound like the CNN of the radio airwaves; stories are repeated all day long with slightly different interpretations.

Many Americans not optimistic that their childrens’ lives will be better

A recent Bloomberg poll asked Americans whether they felt the future would be better for their children. The results:

What optimism there is about the immediate future doesn’t carry over to the longer term. Pluralities of those polled say they’re not hopeful they will have enough money in retirement and expect they will have to keep working to make up the difference. More than 50 percent aren’t confident or are just somewhat confident their children will have better lives than they have.

This belief has been an important part of the American Dream for decades. American parents seem willing to sacrifice much for their children to help insure this. Americans are usually quite optimistic about the future and tend to believe American ingenuity and progress will lead the way.

A question I would like to ask on these surveys: would it be okay for your children to have the same quality of life as you have experienced? If not, why not?

(A note about the reporting: many numerical statistics from the survey are thrown out. However, there is little context. The author tries to throw in some commentary about how these statistics link up with what is going on in the country or with a few quotes but this doesn’t add much. Additionally, let’s break down the numbers a bit more: do they differ by gender, race, region, political party, etc.?)

A word cloud as an accurate information graphic

There are many ways to visually present data or statistics. One issue can arise when parts of graphs or images are not displayed in the correct proportions. Does using a word cloud fall into these difficulties?

Gallup has put together a word cloud of American’s perceptions about the federal government. Some phrases, such as “too big” and “corrupt” are much bigger. Some words are on their sides such as “good” and “terrible.”

Overall, I would say the word cloud is probably not the best choice in this situation. It is hard to judge the most popular responses and the relative proportions of each response. While one can quickly pick up that the majority of responses were negative, it is not a very precise graphic.

Reviewing “American Grace”: it is readable!

The book American Grace: How Religion Divides and Unites Us was released this past week. In addition to being co-authored by Robert Putnam (author of well-known Bowling Alone), the study has been hailed by several sources as a (and perhaps the) comprehensive look at religion in American society.

But a feature of a positive review written by a historian in the San Francisco Chronicle struck me as intriguing:

Among the great virtues of this volume is its combination of two features that are all too rarely found in close proximity. One is a commitment to the most rigorous standards of contemporary social science, bolstered by statistical sophistication. Do you like multiple regression analysis? You’ll find lots of it here. The other feature is a commitment to get their message across to educated readers who are put off by the excessive jargon and abstraction of most sociological studies. Only such a combination could make a 673-page tome worth the attention “American Grace” deserves.

Reading between the lines, here is what is being said: sociologists are not often able to combine statistical evidence (regression analysis of survey results is the gold standard for studies like this that claim to be comprehensive looks at American society) and winsome writing. Essentially, the book is “readable.”

A few thoughts come to mind:

1. What exactly about it makes it “readable” or “understandable”?

2. When reading a book using regression analysis, how much should the “typical educated reader” know about this kind of analysis? This might say more about general statistical knowledge, even among the educated, than it does about the book.

3. This is a valid concern for a book that hopes to be read by many people – writers should always consider their audience. However, it still strikes me as a lower-level priority: isn’t the argument of the book much more important than how it was written? The style of writing can detract from the argument but what we should grapple with are Putnam and Campbell’s conclusions.

Making the case for reputational rankings

A statistician argues that the National Research Council’s study of doctoral programs released earlier this week should have included reputational rankings:

Mr. Stigler says that it was a mistake for the NRC to so thoroughly abandon the reputational measures it used in its previous doctoral studies, in 1982 and 1995. Reputational surveys are widely criticized, he says, but they do provide a check on certain kinds of qualitative measures. When the new NRC counts faculty publication rates, it does not offer any information about whether scholars in the field believe those publications are any good. (That’s especially true in humanities fields, where the NRC report does not include citation counts.)

“Everybody involved in this was trying hard, and with good intentions and high integrity,” Mr. Stigler says. “But once they decided to rule out reputation, they cut off what I consider to be the most useful measure from all past surveys.”

In an e-mail message to The Chronicle this week, Mr. Ostriker declined to reply to Mr. Stigler’s specific statistical criticisms. But he pointed out that the National Academies explicitly instructed his committee not to use reputational measures.

I was curious about this when I looked at the list of sociology doctoral programs. Perhaps several of the schools that were lower than I expected, such as the University of California – Berkeley, were lower because of this.

Stigler defends reputational measures but I’ve seen others argue that they prohibit “true” rankings within fields because certain schools retain a reputation even without the necessary output (research, good grad students, etc.). This particular discussion is part of a larger one where it will need to be decided whether reputational rankings should be used or not.

“Most romantic ‘L’ station” analysis with faulty generalization

Among other features, Craigslist offers a “missed connections” page where people can try to identify and track down people they ran into in public life. Based on this data, Craigslist recently released a list of the “most romantic” spots in Chicago’s CTA system:

Turns out it’s Belmont. The stop on the CTA’s Red, Brown and Purple lines won the title of Chicago’s most romantic ‘L’ station in a report the Web site released Thursday. The crown for most romantic train line went to the Red Line.

The site did a four-week study this summer of more than 250 missed connections postings (read, “potential hookups”) in Chicago and ranked stations based on a scale called the Train Romance Index Score Total — or TRIST for short.

The TRIST is calculated by dividing the number of missed connections that mention a CTA station or line by the number of riders a year who use that station or line. Then, that number is multiplied by 10 to get a whole number and rounded to two decimal places.

The romantic train line is defined as having the best odds of a passenger spotting another rider across a crowded train or platform and then posting a missed connections listing to get in touch.

The data is limited (a pretty small sample) but the generalization is the biggest issue: does this really reveal what is the most romantic spot on the CTA rail system? It is probably much more indicative of who uses Craigslist (young North-siders?).

My guess is that this was simply meant to be fun and promote Craigslist. But sometimes statistics and arguments like this take on a life of their own…

Sorting the good from the bad statistics about Evangelicals

Sociologist Bradley Wright talks with Christianity Today about his latest book: Christians Are Hate-Filled Hypocrites…and Other Lies You’ve Been Told: A Sociologist Shatters Myths From the Secular and Christian Media. Here is CT’s quick summary of the argument:

Young people are not abandoning church. Evangelical beliefs and practices get stronger with more education. Prayer, Bible reading, and evangelism are up. Perceptions about evangelicals have improved dramatically. The data are clear on these matters, says University of Connecticut sociologist Bradley Wright, but evangelicals still want to believe the worst statistics about themselves.

One question to then ask is why Evangelicals buy into these negative statistics. The subculture argument, when applied to evangelicals, might suggest that these numbers help keep people fired up by reminding them that the group could lose its distinctiveness if drastic action is not taken.

Wright suggests his goal is to encourage Evangelicals:

This is not a call for complacency but for encouragement. Why not say, “We’re reading our Scriptures more than most other religious traditions; let’s do even better”? Instead, what we hear is, “Christianity’s going to fail. You’re all a bunch of failures. But if you buy my book, listen to my sermon, or go to my conference, I’ll solve everything.” These fear messages demoralize people, hinder the message of the church, and hide real problems.

I would like to see exactly what statistics he looks at and debunks. Wright is not the first to suggest Evangelicals have some issues with statistics.

The online “mega-reviewers”

One of the innovations of online stores is the ability for users to rate what they like and then for other users to base decisions or comment on those previous ratings. A site like Amazon.com is amazing in this regard; within a few minutes, a reader can get a much better idea about a product.

But statistics from Netflix, another site that allows user reviews, indicate that many users don’t rate anything while there is a small percentage of people who might be called “mega-reviewers”:

About a tenth of one percent (0.07%) of Netflix users — more than 10,000 people —  have rated more than 20,000 items. And a full one percent, or nearly 150,000 Netflixers, have rated more than 5,000 movies. By contrast, only 60 percent of Netflix users rate any movies at all, and the typical person only gives out 200 starred grades.

This rating pattern might fit a Poisson or a negative binomial regression where many people rate none or very few movies while there is a smaller percentage who rate a lot. (A useful statistic in helping to figure out the shape of the curve: while there is 40% that doesn’t rate anything, of the 60 percent who rate any movies at all, what is that median?)

The Atlantic talks to two of mega-reviewers who seem to motivated by seeing what the system would recommend to them after having all of their input. Interestingly, they suggest Netflix still recommends movies to them that they don’t like after watching them.