Potential expansion for domain suffixes

Even amidst discussions that the Internet has run out of addresses, there is talk about expanding the list of available domain suffixes beyond the current 21 options. It sounds like these proposals would allow for all sorts of suffixes and this, inevitably, leads to questions about who would get to control certain domains:

This massive expansion to the Internet’s domain name system will either make the Web more intuitive or create more cluttered, maddening experiences. No one knows yet. But with an infinite number of naming possibilities, an industry of Web wildcatters is racing to grab these potentially lucrative territories with addresses that are bound to provoke.

Who gets to run .abortion Web sites – people who support abortion rights or those who don’t? Which individual or mosque can run the .islam or .muhammad sites? Can the Ku Klux Klan own .nazi on free speech grounds, or will a Jewish organization run the domain and permit only educational Web sites – say, remember.nazi or antidefamation.nazi? And who’s going to get .amazon – the Internet retailer or Brazil?

The decisions will come down to a little-known nonprofit based in Marina del Rey, Calif., whose international board of directors approved the expansion in 2008 but has been stuck debating how best to run the program before launching it. Now, the Internet Corporation for Assigned Names and Numbers, or ICANN, is on the cusp of completing those talks in March or April and will soon solicit applications from companies and governments that want to propose and operate the new addresses.

Sounds like we could have some battles on our hands for particular suffixes. Perhaps the companies or organizations with the most money will win.

But many of the options in this article are set up as “good” options versus “bad” options. If given a choice, how many people would want the .nazi domain to be controlled by the Ku Klux Klan? And some of the other options presented in the story, such as whether someone who wants musicians and agents to be able to get .music addresses while the music industry wants to control this for their larger purposes, are less clear. ICANN, the organization who controls the domains, says they have considered this: “For people who might propose controversial domains – such as .nazi, which ICANN officials have worried about – approval will be based on the applicant’s identity and intentions, and on the grounds of “morality and public order.” How in the world will they be able to do this in a way that is satisfying to multiple parties? Is there a way to decide this before the domains are sold or are we simply in for long rounds of litigation?

Trying to count the people on the streets in Cairo

This is a problem that occasionally pops up in American marches or rallies: how exactly should one estimate the number of people in the crowd? This has actually been quite controversial at points as certain organizers of rallies have produced larger figures than official government or media estimates. And with the ongoing protests taking place in Cairo, the same question has arisen: just how many Egyptians have taken to the streets in Cairo? There is a more scientific process to this beyond a journalist simply making a guess:

To fact-check varying claims of Cairo crowd sizes, Clark McPhail, a sociologist at the University of Illinois and a veteran crowd counter, started by figuring out the area of Tahrir Square. McPhail used Google Earth’s satellite imagery, taken before the protest, and came up with a maximum area of 380,000 square feet that could hold protesters. He used a technique of area and density pioneered in the 1960s by Herbert A. Jacobs, a former newspaper reporter who later in his career lectured at the University of California, Berkeley, as chronicled in a Time Magazine article noting that “If the crowd is largely coeducational, he adds, it is conceivable that people might press closer together just for the fun of it.”

Such calculations of capacity say more about the size of potential gathering places than they do about the intensity of the political movements giving rise to the rallies. A government that wants to limit reported crowd sizes could cut off access to its cities’ biggest open areas.

From what I have read in the past on this topic, this is the common approach: calculate how much space is available to protesters or marchers, calculate how much space an individual needs, and then look at photos to see how much of that total space is used. The estimates can then vary quite a bit depending on how much space it is estimated each person wants or needs. These days, the quest to count is aided by better photographs and satellite images:

That is because to ensure an accurate count, some computerized systems require multiple cameras, to get high-resolution images of many parts of the crowd, in case density varies. “I don’t know of real technological solutions for this problem,” said Nuno Vasconcelos, associate professor of electrical and computer engineering at the University of California, San Diego. “You will have to go with the ‘photograph and ruler’ gurus right now. Interestingly, this stuff seems to be mostly of interest to journalists. The funding agencies for example, don’t seem to think that this problem is very important. For example, our project is more or less on stand-by right now, for lack of funding.”

Without any such camera setup, many have turned to some of the companies that collect terrestrial images using satellites, but these companies have collected images mostly before and after the peak of protests this week. “GeoEye and its regional affiliate e-GEOS tasked its GeoEye-1 satellite on Jan. 29, 2011 to collect half-meter resolution imagery showing central Cairo, Egypt,” GeoEye’s senior vice president of marketing, Tony Frazier, said in a written statement. “We provided the imagery to several customers, including Google Earth. GeoEye normally relies on our partners to provide their expert analysis of our imagery, such as counting the number of people in these protests.” This image was taken before the big midweek protests. DigitalGlobe, another satellite-imagery company, also didn’t capture images of the protests, according to a spokeswoman, but did take images later in the week.

Because these images are difficult to come by in Egypt, it is then difficult to make an estimate. As the article notes, this is why you will get vague estimates for crowd sizes in news stories like “thousands” or “tens of thousands.”

Since this is a problem that does come up now and then, can’t someone put together a better method for making crowd estimates? If certain kinds of images could be obtained, it seems like an algorithm could be developed that would scan the image and somehow differentiate between people.

Untimely end for source of Internet meme about Internet

The NBC employee who released footage of Bryant Gumbel, Katie Couric, and Elizabeth Vargas struggling to define the Internet back in 1994 has been fired. So much for (corporate) information wanting to be free.

But I also don’t quite understand what all the fuss has been about. Sure, their conversation sounds silly to us today. But this was only 16 years ago. If anything, this clip and its popularity demonstrates how quickly the Internet has become an part of everyday life. Back in 1994, the Internet was not used by the common American. My family got AOL in the next year or two, I remember a friend’s family having Prodigy around this time, but most people had no access and realistically, no need for access. Couric, Gumbel, and Vargas were like many Americans: just trying to figure out what this new technology was and how it was used.

More broadly, this released video fits with patterns of more modern people laughing at or commenting on how much better life is now compared to the past. From the vantage point of 2011, we can see the benefits of the Internet and we are bombarded with messages from companies suggesting we need even more of it (in our phones, in our treadmills, etc.). But anytime new technology is introduced, it takes time for the mass public to figure out whether it is a good change or not.

On the hidden or out of the way yet sometimes thriving web forums

This is something I have noticed recently in several sites I visit frequently: there is a little community of consistent posters who have been drawn together and slowly get to know each other. While one of these sites, the Ask Amy column posted on the site of the Chicago Tribune, is not exactly hidden, The Economist discusses some web groups that have formed in really hard to find or unlikely places:

The programming crew had accidentally created a community of the sort that crop up all over the internet. Most online discussions take place in discussion forums designed to allow people to create an identity and interact in threaded, chronological conversations. But the hidden recesses of the web provide enough soil to root entire worlds, too. Wherever one person may post words which more than one other may read and respond to, a world is born.

Read the article to hear how devoted fans of Douglas Adams founded a group in a forum that was an afterthought and how some people unhappy with Sonic Drive-In’s service found each other.

Sounds like a start to a very interesting research study: what exactly motivates people to (1) seek out these spaces and (2) then continue with discussions and getting to know each other. The description of what happens in these settings in out of way parts of the Internet is hilarious:

It’s been thirteen years of hosting an accidental community. It’s somewhat like ignoring the vegetable drawer of your fridge for a year, then opening it to find a bunch of very grateful sentient tomatoes busily working on their third opera.

I would guess that the people who participate in groups like this are a limited number of total web users. I wouldn’t tend to be drawn to such forums: read a comment section of any blog or news story and you would likely find the conversation to be quite tedious or inflammatory. But I can remember the heady early days of AOL when chat rooms were the exciting feature of the Internet (and content took forever to load).

And these groups can be like real-life groups, meaning that they become territorial and protective:

Another surprise is that they will treat growth as a perturbation as well, and they will spontaneously erect barriers to that growth if they feel threatened by it. They will flame and troll and otherwise make it difficult for potential new members to join, and they will invent in-jokes and jargon that makes the conversation unintelligible to outsiders, as a way of raising the bar for membership.

It sounds like there is a starting period when the group might be somewhat fluid as people stumble unto such forums. But once the group coalesces and becomes a collective entity, others are not welcome and sharp boundaries are drawn to limit the influence of outsiders. So if one wants to become part of such a group, does one simply have to be lucky or have good timing?

Another question: what do the users get out of participating in such long digressions?

Found hypocrisy; still searching for clarity

In case you haven’t heard, a few days ago Google started publicly accusing Microsoft’s Bing of stealing its search results.  Juan Carlos Perez over at PCWorld has published an interesting roundup of reactions to Google’s new “strategy” of public accusations:

While the merits of Google’s accusation are up for debate — Microsoft denies the charge — the fact that Google chose to complain in such a loud and agitated manner has become fertile ground for analysis and comment by industry observers.

Opinions range from those who view Google’s actions as hypocritical to others who say the company did the right thing by airing its grievance.

PCWorld’s link to Daniel Eran Dilger reaction over at Roughly Drafted is especially worth checking out.  Personally, I come down on the “Google is being hypocritical” side of things.  It’s hard to have the expansive view of copyright law and fair use that Google embraces for its own activities and then to complain with any legitimacy about Microsoft’s alleged behavior.

Unfortunately, copyright law in general (and fair use in particular) is notoriously unclear, malleable, and subject to judicial whims.  It’s doubtful that Google will actually sue Microsoft over this, so we may never know what the “answer” is.

However, even if a U.S. court upheld Microsoft’s right to copy Google’s search results (assuming that’s what happened here), that would only give us an answer (1) on these specific facts (2) as between parties willing to litigate (and maybe even (3) before that particular judge).  Given the high costs of litigation, most non-Fortune-500 copyright users claiming fair use rights usually find it is in their best interest to settle for a few thousand dollars when saddled with a copyright infringement lawsuit.  Indeed, there are companies based on this very business model that are out there suing people; the number of copyright infringement suits is rising.

This latest spat between Google and Microsoft is, to some extent, a sideshow, but it does highlight some of the problems that uncertainty breeds within copyright law.  I’m not worried about Microsoft’s ability to defend itself:  it’s a multi-billion dollar company with lawyers and PR specialists both in-house and on speed dial.  I am worried about the start ups that are seeking to be the next Google or Microsoft:  they generally can’t afford to get anywhere close to the line because they know that an infringement lawsuit may mean millions in legal fees and damages, so they back off and play it safe.

That’s the real cost of un-clarity in copyright law.

How the John Edwards affair became news

How exactly certain scandals come to light when they do is often an interesting tale. The former editor of the National Enquirer explains how his investigative team put together the story of John Edwards’ affair. The tale involves the use of technology and a profiler who provided insights into how to trap Edwards in his lies:

I knew there was no viable scenario for Edwards to confess to the Enquirer. I faced the bitter realization that another news organization would reap the benefits of our team’s hard work and get the confession, but I also knew that ultimately that confession would validate the Enquirer‘s earlier story as well as the new one.

Behind the scenes we exerted pressure on Edwards, sending word though mutual contacts that we had photographed him throughout the night. We provided a few details about his movements to prove this was no bluff.

For 18 days we played this game, and as the standoff continued the Enquirer published a photograph of Edwards with the baby inside a room at the Beverly Hilton hotel.

Journalists asked if we had a hidden camera in the room. We never said yes or no. (We still haven’t). We sent word to Edwards privately that there were more photos.

He cracked. Not knowing what else the Enquirer possessed and faced with his world crumbling, Edwards, as the profiler predicted, came forward to partially confess. He knew no one could prove paternity so he admitted the affair but denied being the father of Hunter’s baby, once again taking control of the situation.

Perhaps this story isn’t anything unusual – technology makes information gathering a lot easier. Yet it is somewhat shocking to me that plenty of powerful people, like John Edwards or Tiger Woods, think that they can get away with things in the long run. Sure, the National Enquirer had to spend months tracking down this story but in the end, it was doable and effectively changed the public perception of John Edwards forever. Is there something that happens when people are put in powerful positions that changes their perceptions of what they can and can’t get away with?

Is it even possible for the powerful to get away with things like this any more? How many “scandals” are lurking out there somewhere? It is certainly a far cry from the days of the 1950s and before when sportwriters routinely shied away from reporting on what athletes did away from home and political reporters didn’t talk about everything.

The land of 100,000 lawsuits

Some enterprising anonymous researcher has determined that almost 100,000 copyright infringement lawsuits have been filed in the U.S. in the past year:

In the United States the judicial system is currently being overloaded with new cases, but the scope of the issue was never really clear until now. An anonymous TorrentFreak reader has spent months compiling a complete overview of all the mass P2P lawsuits that have been filed in the US since the beginning of 2010, listing all the relevant case documents and people involved in a giant spreadsheet.

The research shows that between 8th January 2010 and 21st January 2011, a total of 99,924 individuals have been sued. The vast majority of the defendants have allegedly used BitTorrent to share copyrighted works but a few hundred ed2k users are also included.

Of the 80 cases that were filed originally, 68 are still active, with 70,914 defendants still in jeopardy.

The raw data is available is spreadsheet form over on Google Docs.

As the disparity between 80 and 70,914 indicates, these types of lawsuits completely overwhelm the courts.  The U.S. justice system is simply not set up to handle this kind of volume, especially for suits as notoriously tricky to argue as copyright infringement.

Find (if ye know how to seek)

It’s a few days old now, but I just ran across a post over on TorrentFreak describing how Google has started removing “torrent”-related results from its auto-complete search results:

Without a public notice Google has compiled a seemingly arbitrary list of keywords for which auto-complete is no longer available. Although the impact of this decision does not currently affect full search results, it does send out a strong signal that Google is willing to censor its services proactively, and to an extent that is far greater than many expected.

Among the list of forbidden keywords are “uTorrent”, a hugely popular piece of entirely legal software and “BitTorrent”, a file transfer protocol and the name of San Fransisco based company BitTorrent Inc. As of today [1/26/2011], these keywords will no longer be suggested by Google when you type in the first letter, nor will they show up in Google Instant.

All combinations of the word “torrent” are also completely banned. This means that “Ubuntu torrent” will not be suggested as a user types in Ubuntu, and the same happens to every other combination ending in the word torrent. This of course includes the titles of popular films and music albums, which is the purpose of Google’s banlist.

This is quite an interesting development.  Personally, I have found Google’s auto-complete functionality very helpful in finding the names of half-remembered items.  It is a disturbing reminder of just how much control Google exerts–not only over what we find, but over what we search for.

Vast worlds of discovery

In case you thought the age of discovery was over, Wired’s Threat Level blog is reporting that a 21-year-old hacker George Hotz who released the PlayStation 3 jailbreak has been ordered to surrender

any and all computer hardware and peripherals containing circumvention devices, technologies, programs, parts thereof, or other unlawful material, including but not limited to code and software, hard disc drives, computer software, inventory of CD-ROMS, computer diskettes, or other material containing circumvention devices, technologies, programs, parts thereof, or other unlawful material.

As Hotz lawyer put it,

The information sought at issue [the jailbreak code] is less than 100 kilobytes of data. Mr. Hotz has terabytes of storage devices….Impounding his computers, it’s like starting a forest fire to cut down a single tree.

Though the court’s order does seem like overkill, it is unfortunately a typically broad discovery request.  Sony may simply be trying to harass Hotz and/or hamper any future work, a theory especially plausible insofar as the court also ordered that Hotz “shall retrieve” the jailbreak he posted.  Given the number of websites that have re-posted Hotz’s original code, this would seem to be impossible.  As Hotz’s lawyer rather cogently quipped, ““Mr. Hotz can’t retrieve the internet.”

Wired has posted the judge’s order here (PDF).