On Facebook, it’s not 6 degrees of separation but rather 4.74 degrees

One effect of globalization is that people are more aware of world events and are better connected to others. A new study using Facebook data suggests the average user is separated from any other user in the world by just 4.74 degrees:

On Facebook, however, the average user is only 4.74 degrees away from any other Facebooker…

That conclusion comes from a non-peer-reviewed study of 721 million active Facebook users, released by Facebook in collaboration with the Università degli Studi di Milano, the blog post says…

The Palo Alto, California, company says 99.6% of all Facebook users studied were separated by five degrees or less from any other Facebook user; 92% were separated by only four degrees…

“The average distance in 2008 was 5.28 hops, while now it is 4.74,” Facebook says.

While this is indeed an interesting finding (particularly since it is related to Stanley Milgram’s six degree studies decades ago), there are bigger questions at stake here. With people 4.74 connections away, how exactly does this impact a user’s life or positively influence their life? We know that information and culture passes through networks but how exactly does this work on Facebook? Can the life of a user in Siberia really affect the life of a college student in Arizona?

One issue here is that Facebook itself currently allow for limiting connectability between users. Sources like The Facebook Effect suggest that Mark Zuckerberg would really like a more open network where people could see each other’s information and actually interact with others beyond the “friends” structure. However, it doesn’t appear that most users would want this at this time – most Facebook friends are people users are already know and there are concerns about privacy. How does the company move people into accepting a more open network so that users can openly take advantage of those chains 4.74 people long?

Also, who tend to be the people in the networks that help connect people the most? College students? People who live in larger metropolitan areas? People with the most friends? People with the most diversity in their own friend lists?

What is “The Big Data Boom”on the Internet good for?

The Internet is a giant source of ready-to-use data:

Today businesses can measure their activities and customer relationships with unprecedented precision. As a result, they are awash with data. This is particularly evident in the digital economy, where clickstream data give precisely targeted and real-time insights into consumer behavior…

Much of this information is generated for free, by computers, and sits unused, at least initially. A few years after installing a large enterprise resource planning system, it is common for companies to purchase a “business intelligence” module to try to make use of the flood of data that they now have on their operations. As Ron Kohavi at Microsoft memorably put it, objective, fine-grained data are replacing HiPPOs (Highest Paid Person’s Opinions) as the basis for decision-making at more and more companies.

The wealth of data also makes it easy to run experiments:

Consider two “born-digital” companies, Amazon and Google. A central part of Amazon’s research strategy is a program of “A-B” experiments where it develops two versions of its website and offers them to matched samples of customers. Using this method, Amazon might test a new recommendation engine for books, a new service feature, a different check-out process, or simply a different layout or design. Amazon sometimes gets sufficient data within just a few hours to see a statistically significant difference…

According to Google economist Hal Varian, his company is running on the order of 100-200 experiments on any given day, as they test new products and services, new algorithms and alternative designs. An iterative review process aggregates findings and frequently leads to further rounds of more targeted experimentation.

This sounds like a social scientist’s dream – if we could get our hands on the data.

My big question about all of this data is this: what should be done with it? This article, and others I’ve seen, have said that it will transform business. If this is just a way for businesses to become more knowledgeable, more efficient, and ultimately, more profitable, is this enough? Occasionally, we hear of things like discovering and/or tracking epidemics by looking at search queries or tools like the “mechanical turk” to crowdsource small but needed work. On the whole, does the data from the Internet advance human flourishing or concentrate some benefits in the hands of a few or even hinder flourishing? Does this data give us insights into health and medicine, international relations, and social interactions or does it primarily give entrepreneurs and established companies the chance to make more money? Are these questions that anyone really asks or cares about?

ASA pushing for better sociology Wikipedia entries

This news came out earlier this week in the American Sociological Association’s Footnotes: the ASA is hoping sociologists and sociology students will help improve Wikipedia pages pertaining to sociology.

In an essay on the association’s online newsletter (scheduled to be included in the next edition of its print newsletter), Wright this week announced the Sociology in Wikipedia Initiative: a formal call to sociologists to help improve and expand Wikipedia entries that might benefit from their expertise and consider assigning their students to do the same.

“Wikipedia has become an important global public good,” Wright writes in the essay. “Since it is a reference source for sociologically relevant ideas and knowledge that is widely used by both the general public and students, it is important that the quality of sociology entries be as high as possible. This will only happen if sociologists themselves contribute to this public good.”

Not only might Wikipedia benefit from contributions by students steeped in academic research methods, but the exercise might help students learn how to read the crowd-sourced encyclopedia in the proper context, said Wright.

“What better way to get students to understand that it’s actual people like them who have written this stuff, than for them to write this stuff?” he said.

Is this “public sociology” at work? I don’t mind this call as it would help ensure that Wikipedia has accurate and in-depth sociology information rather than just a bare bones outline. Actually, I’ve thought the sociology Wikipedia entries weren’t that bad already, particularly compared to other disciplines. For example, the statistics pages on Wikipedia are technically correct but it is very difficult for a layperson to understand what is going on.

But how many sociology faculty will spend much time with this since there aren’t many professional incentives? Even publishing in online journals as opposed to more traditional print journals is not well-regarded so what’s the point of helping improve Wikipedia entries? This may seem like a move toward embracing technology and toward a younger generation of sociologists but the discipline has a long way to go.

At least a few leaders of major academic groups are admitting that they use Wikipedia as a source. Not too long, admitting this would not have been good for one’s status. How far away are we from Wikipedia being an acceptable source?

One of the new research frontiers: studying dating online

There are now a number of academics studying online dating sites as they allow insights into relationship formation that are difficult to observe elsewhere in large numbers:

Like contemporary Margaret Meads, these scholars have gathered data from dating sites like Match.com, OkCupid and Yahoo! Personals to study attraction, trust, deception — even the role of race and politics in prospective romance…

“There is relatively little data on dating, and most of what was out there in the literature about mate selection and relationship formation is based on U.S. Census data,” said Gerald A. Mendelsohn, a professor in the psychology department at the University of California, Berkeley…

Andrew T. Fiore, a data scientist at Facebook and a former visiting assistant professor at Michigan State University, said that unlike laboratory studies, “online dating provides an ecologically valid or true-to-life context for examining the risks, uncertainties and rewards of initiating real relationships with real people at an unprecedented scale.”…

Of the romantic partnerships formed in the United States between 2007 and 2009, 21 percent of heterosexual couples and 61 percent of same-sex couples met online, according to a study by Michael J. Rosenfeld, an associate professor of sociology at Stanford. (Scholars said that most studies using online dating data are about heterosexuals, because they make up more of the population.)

The rest of the article has some research findings about appearance, race, and political ideology derived from studies of online dating site members.

Researchers will go wherever the research subjects are so if the people are expanding their dating pools online, that is where the research to go. It would be interesting to hear if any of these researchers have received pushback from people within their own fields who scoff at online dating sites or ask them to demonstrate the worthiness of studying online behavior.

 

The NFL says the “All-22” camera angle is proprietary information

The NFL is a TV ratings powerhouse and makes billions each year on selling television rights. However, fans don’t see the same action that the league and teams watch because the league claims its “All-22” view is proprietary information:

If you ask the league to see the footage that was taken from on high to show the entire field and what all 22 players did on every play, the response will be emphatic. “NO ONE gets that,” NFL spokesman Brian McCarthy wrote in an email. This footage, added fellow league spokesman Greg Aiello, “is regarded at this point as proprietary NFL coaching information.”

For decades, NFL TV broadcasts have relied most heavily on one view: the shot from a sideline camera that follows the progress of the ball. Anyone who wants to analyze the game, however, prefers to see the pulled-back camera angle known as the “All 22.”

While this shot makes the players look like stick figures, it allows students of the game to see things that are invisible to TV watchers: like what routes the receivers ran, how the defense aligned itself and who made blocks past the line of scrimmage.

By distributing this footage only to NFL teams, and rationing it out carefully to its TV partners and on its web site, the NFL has created a paradox. The most-watched sport in the U.S. is also arguably the least understood. “I don’t think you can get a full understanding without watching the entirety of the game,” says former head coach Bill Parcells. The zoomed-in footage on TV broadcasts, he says, only shows a “fragment” of what happens on the field.

Why does the NFL do this? Here are a few plausible scenarios:

1. It can do it so it will. The NFL won’t be bullied into doing something it doesn’t want to do. As long as the money keeps pouring in for TV rights, there is little pressure the public can put on the league for this footage.

1a. If enough fans and commentators picked up on this, could they force the NFL’s hand? It seems unlikely.

2. The NFL makes billions on TV rights and perhaps wants to package this video in a certain way. A later part of the story suggests the NFL has quietly floated the idea of selling access to this footage.

3. The league is worried about legitimate football competitors. There are not currently any viable threats but this could pop up again.

4. The league thinks this is the core data of the NFL, what actually happens on all plays, and will go to great lengths to protect its “intellectual property.” I find this a little hard to believe: aren’t there plenty of people who could understand and scheme what happens on a football field even if the primary camera angle doesn’t show it? Are teams really that worried about what the public might see or that other teams are missing things in the video?

Measuring how much the Internet is worth: $8 trillion?

A recent report by McKinsey puts the value of the Internet at $8 trillion. Here are a few other fun facts:

There is a lot of Internet to measure, with two billion global consumers and $8 trillion in total revenue. So McKinsey’s report limited its scope to the online economy in the G-8 countries plus five more: Brazil, China, India, South Korea and Brazil. It defined Internet activities as private consumption (electronic equipment, e-commerce, broadband subscriptions, mobile Internet, and hardware and software consumption); private investment (from the telecommunications industry and the maintenance of extranet, intranet, and Web sites); public expenditure (spending and buying by government in software hardware and services); and trade (which accounts for exports of Internet equipment plus business-to-business services with overseas companies)…

As an industry, the Internet contributes more to the typical developed economy than mining, utilities, agriculture, or education. In Sweden, fully one-third of economic growth in the five years leading up to the recession came from Internet activities. For the entire G-8, the average was 21 percent. In an analysis of France since the mid-1990s, McKinsey found that the Internet created more than twice the number of jobs it destroyed.

Much of the Internet’s contribution to our lives is nearly impossible to measure. For example, I use email. How much is that worth to me? I can’t even begin to say. I read hundreds of news sources a day. What is that worth to me, or to the news organizations? Pricing this kind of thing is exhausting to think about. But since analyzing what the rest of us find “exhausting to think about” is McKinsey’s job, their researchers looked at the “consumer surplus” of the Internet, concluding that the total annual benefit to the United States comes out to $64 billion…

The United States is the world leader in the online industry, grabbing 30 percent of global Internet revenues. But the UK is the world leader in online retail. The British spent $2,535 on e-stuff in 2009, more than twice the average of the world’s largest countries and still 1.4 times the amount of the typical U.S. shopper. Sweden leads the world in Internet’s contribution to GDP. Fully 6.3 of the country’s economy is online — twice Germany, France or India. In Russia, the Internet contributes not even one percent of GDP.

Some interesting stuff here:

1. I appreciate the emphasis on the difficulty of measuring this topic. In addition to simply thinking about the economic benefits, we could spend a lot of time discussing how it has altered social interaction, private practices, and democracy. I wonder what the margin of error is on the estimates.

2. There is some indication of the splits between the Internet haves and have-nots. If the Internet is so valuable, should this be a leading component of aid to poorer countries? It does require a decent investment in infrastructure but it would allow people to easily connect to first-world countries and industries. For example, what is the impact of the less than $100 laptop that was touted for years?

3. With all of this money (and value floating around), it is a reminder why so many states want to get their hands on sales tax revenues from Internet sales. Do European countries like Britain have a similar system? I have bought a few things from Amazon.co.uk in the past and I don’t recall the experience being much different.

4. I would be interested to know the future prospects for the Internet’s growth: how quickly will it grow? How much will it expand? Is most of the growth within developed countries or in opening or expanding newer markets (China and India plus others)?

Viewing the insides of stores on Google Maps

Adding to its Street View capabilities, Google also will allow browsers to see inside some retail establishments that allowed Google to photograph their interiors:

A test program launched in April of last year was bearing fruit in a growing array of panoramic images taken inside businesses that volunteered to be part of the project.

“We’ve been seeing renewed interest in the past few days because, as promised, we’re getting more imagery online,” Google spokeswoman Deanna Yick told AFP on Monday…

Small businesses in Japan, Australia, New Zealand, and the United States have been able to invite Street View photographers into their shops or eateries to capture images then served up with Google online maps.

“With this immersive imagery, potential customers can easily imagine themselves at the business and decide if they want to visit in person,” Google Maps product manager Gadi Royz said in a blog post early this year.

My big question: will this actually bring more customers inside the shops? I’m skeptical: how many times would someone be wondering about whether they should visit a store, look up the interior image on Street View, and then make a positive decision. What if the image is actually a negative thing, perhaps due to the lighting (I wonder if they adjusted for this), outdated decor, or, for lack of a better term, a lack of “coolness”?

We could also ask whether Google’s efforts in these areas actually encourage in-person community. If given more information in general through search engines, images, and reviews (with Google recently buying Zagat), will people be more likely to venture out of their homes or away from their internet-enabled devices? Will they become overwhelmed with the choices (like Barry Schwartz argues in The Paradox of Choice) and be less likely to choose any?

In the end, Google must think that providing these interior images are going to help them make money.

A new way to do the college search process: one comprehensive website to match students to colleges

The policy director of an education think tank writes in Washington Monthly, itself a purveyor of college rankings, that the future of college admissions will come in the form of a single, comprehensive website that will match prospective students and colleges:

This is the future of college admissions. The market for matching colleges and students is about to undergo a wholesale transformation to electronic form. When the time comes for Jameel to apply to colleges, ConnectEDU will take all of the information it has gathered and use sophisticated algorithms to find the best colleges likely to accept him—to find a match for Jameel in the same way that Amazon uses millions of sales records to advise customers about what books they might like to buy and Match.com helps the lovelorn find a compatible date. At the same time, on the other side of the looking glass, college admissions officers will be peering into ConnectEDU’s trove of data to search for the right mix of students.

This won’t just help the brightest, most driven kids. Bad matching is a problem throughout higher education, from top to bottom. Among all students who enroll in college, most will either transfer or drop out. For African American students and those whose parents never went to college, the transfer/dropout rate is closer to two-thirds. Most students don’t live in the resource-rich, intensely college-focused environment that upper-middle-class students take for granted. So they often default to whatever college is cheapest and closest to home. Tools like ConnectEDU will give them a way to find something better.

We can think of getting into college like this: students need to be slotted into the appropriate school. At this point, students can do certain things to improve their fit and colleges use certain information (though it often comes in a form of a narrative about students that admissions officers construct – I highly recommend Creating a Class). Our current system is highly dependent on students doing the initial legwork in searching out colleges that might fit them but as this article suggests, there are a number of students, particularly poorer students, who don’t do well in this system.

If this website idea catches on, wouldn’t it create more competition within the college market for students? If so, would middle- and upper-class students start complaining?

Also, while the article suggests a website like this is the answer to helping kids who can’t currently play the college game, doesn’t it rest on the idea that (1) people have equal access to this website and (2) that users have the ability or “cultural capital” to sort through the information the website presents? Neither of these might necessarily be true.

h/t Instapundit

Looking for a new area of study? Try Twitterology

If it is in the New York Times, Twitterology must be a viable area of academic study:

Twitter is many things to many people, but lately it has been a gold mine for scholars in fields like linguistics, sociology and psychology who are looking for real-time language data to analyze.

Twitter’s appeal to researchers is its immediacy — and its immensity. Instead of relying on questionnaires and other laborious and time-consuming methods of data collection, social scientists can simply take advantage of Twitter’s stream to eavesdrop on a virtually limitless array of language in action…

One criticism of “sentiment analysis,” as such research is known, is that it takes a naïve view of emotional states, assuming that personal moods can simply be divined from word selection. This might seem particularly perilous on a medium like Twitter, where sarcasm and other playful uses of language often subvert the surface meaning…

Still, the Twitterologists will continue to have a tough row to hoe in justifying their research to those who think that Twitter is a trivial form of communication. No less a figure than Noam Chomsky has taken Twitter to task recently for its “superficiality.”

For more sociological thoughts about Chomsky’s comments, see this post from a few days ago.

Here is my quick take on Twitterology: it has some potential for gathering quick, on-the-ground information. But there are two big issues that this article doesn’t address:

1. Are Twitter users representative of the whole population? Probably not. Twitter feeds might be good for studying very specific groups and movements.

2. How can one make causal arguments with Twitter data? If we had more information about Twitter users from profiles, this might be doable but Twitter is less about Facebook-style profiles. We then need studies that collect the information about Twitter users as well as their Twitter activity. If we want to ask questions like whether Twitter was instrumental or even helped cause the Arab Spring movements, we need more data.

Twitterology may be trendy at the moment but I think it has a ways to go before we can use it to tackle typical questions that sociologists ask.

Argument: Chomsky wrong to suggest Twitter is “superficial, shallow, evanescent”

Nathan Jurgenson argues that Noam Chomsky’s thoughts about Twitter are misguided:

Noam Chomsky has been one of the most important critics of the way big media crowd out “everyday” voices in order to control knowledge and “manufacture consent.” So it is surprising that the MIT linguist dismisses much of our new digital communications produced from the bottom-up as “superficial, shallow, evanescent.” We have heard this critique of texting and tweeting from many others, such as Andrew Keen and Nicholas Carr. And these claims are important because they put Twitter and texting in a hierarchy of thought. Among other things, Chomsky and Co. are making assertions that one way of communicating, thinking and knowing is better than another…

Claiming that certain styles of communicating and knowing are not serious and not worthy of extended attention is nothing new. It’s akin to those claims that graffiti isn’t art and rap isn’t music. The study of knowledge (aka epistemology) is filled with revealing works by people like Michel Foucault, Jean-François Lyotard or Patricia Hill Collins who show how ways of knowing get disqualified or subjugated as less true, deep or important…

In fact, in the debate about whether rapid and social media really are inherently less deep than other media, there are compelling arguments for and against. Yes, any individual tweet might be superficial, but a stream of tweets from a political confrontation like Tahrir Square, a war zone like Gaza or a list of carefully-selected thinkers makes for a collection of expression that is anything but shallow. Social media is like radio: It all depends on how you tune it…

Chomsky, a politically progressive linguist, should know better than to dismiss new forms of language-production that he does not understand as “shallow.” This argument, whether voiced by him or others, risks reducing those who primarily communicate in this way as an “other,” one who is less fully human and capable. This was Foucault’s point: Any claim to knowledge is always a claim to power. We might ask Chomsky today, when digital communications are disqualified as less deep, who benefits?

Back to a classic question: is it the medium or the message? Is there something inherent about 140 character statements and how they must be put together that makes them different than other forms of human communication? I like that Jurgenson notes historical precedent: these arguments have also accompanied the introduction of radio, television, and the Internet.

But could we tweak Chomsky’s thoughts to make them more palatable? What if Chomsky had said that the average Twitter experience was superficial, would he be incorrect? Perhaps the right comparison is necessary – Twitter is more superficial compared to face-to-face contact? But is it more superficial than no contact since face-to-face time is limited? Jurgenson emphasizes the big picture of Twitter, its ability to bring people together and give people the opportunity to follow others and “tune in.” In particular, Twitter and other social media forms allow the average person in the world to potentially have a voice in a way that was never possible before. But for the average user, how much are they benefiting – are they tuned in to major social movements or celebrity feeds? What their friends are saying right now or progress updates from non-profit organizations? Is this a beneficial public space for the average user?

Additionally, does it matter here if Twitter had advertisements and made a big push to make money off of this versus providing a more democratic space? Is Twitter more democratic and deep than Facebook? How would one decide?

In the end, is this simply a generational split?

(See earlier posts on a similar topic: Malcolm Gladwell on the power of Twitter, how Twitter contributed (or didn’t) to movements in the Middle East, and whether using Twitter in the classroom improves student learning outcomes.)