Tuesday, March 06, 2012

The Mystery of Power-Law Distributions

One criticism of sociology, and the macro social sciences more generally (such as political science, anthropology, and economics), is that there are very few "laws" of social reality. There are, however, some sociological regularities that are as yet not fully explained, and which seem bizarre. The most enduring and puzzling of these are power-law distributions (a well-known special case of this is "Zipf's Law"), which is the fact that  "large" instances of things are extremely rare, while "small" occurrences of things are extremely common (where size can refer to frequency in a population, population size, geographic space, and so on). In practice this means that a handful of words are much more frequent than other words (and most words are rarely used), wealth is concentrated in a small number of people (and most people are poor), there are a handful of really popular songs (and a vast number of unpopular tunes), and so on. Even the sizes of sand particles on a beach follow a power-law distribution: how often have you seen a boulder on a beach?

What might explain the ubiquity of power-law distributions? As far as I can tell, nobody is entirely sure, although we have some good guesses. For example, the sociologist Herbert Simon outlined a theory of preferential growth attachment (also known as the "rich get richer" effect), in which songs that are already fairly popular will become more popular, cities that are already large will become even larger, and words already used widely will become even more widely used. Note that this explanation hinges on a positive feedback effect: the probability that any thing gets "larger" is directly proportional to the current "largeness" of the thing; or, to put it another way, large values get amplified rather than cancelled out (as in a normal distribution).

Power-law distributions have important cultural, statistical, and political implications.

Culturally, there are several implications. First, most cultural constructs  are rarely used and only a handful are common among any group of people. To put it another way, the shared part of culture is likely to be relatively small, while the particular part of culture is vast. Second, frequently used cultural constructs are particularly stable over time; that is, 500 years from the word "the" will still be used, while "sesquipedalian" has a more uncertain future. Third, the stability of a cultural system is derived from the more frequently used cultural constructs, while the dyanmism is among the less frequently used constructs. Fourth, initial conditions are extremely important for the frequency and hence durability of cultural constructs: for instance, small, random fluctuations led to the popularity of "the" in the English language. Finally, following from the previous point, the consequences of initial conditions are highly unpredictable; given small initial changes English speakers today might instead be using the word "tha" or "se" instead of "the." 

Statistically, the presence of power-law distributions is a reminder that classical linear regression (based on the normal distribution) is not always the appropriate fit to a scatter plot of two variables, and that summarizing a distribution as a mean or median can be highly misleading.

Politically, power-law distributions have a unique implication for efforts to deal with wealth inequality: one effective way to alter the distribution of wealth is to remove the positive feedback effects from wealth. The desired distribution of wealth would thus be described by a normal rather than power law function. Importantly, removing the positive feedback effects of wealth would not lead to the removal of inequality, but rather a change in the distribution so that the mean, median, and mode are the same. From this perspective, policies should be in place so that (in principle) a person's change in wealth is independent of their current level of wealth. Such policies might include very high taxes on capital gains, restrictions on the influence of wealth in political decision-making, rules specifying equal monetary amounts from promotions for all occupational levels in a firm, and so on.

Monday, March 05, 2012

Visualizing a Correlation Table

Correlation tables are ubiquitous in social science research, but very rarely they are visualized. As I've emphasized in previous posts, I'm a strong advocate for visualizing data and models whenever possible. For example, for my research I graphed correlations using Adrian Mander's plotmatrix command in Stata. Using Mander's package, I could create a graph that clearly shows all the information in a parsimonious way; moreover, unlike a correlation table, correlation patterns are intuitively grasped from the shading of the cells, and there is an implicit emphasis on the correlation size rather than statistical significance.

Sunday, March 04, 2012

Why Models are Not Data

In doing research, sometimes it can be easy to think that the models one is using are in fact the data -- but this is clearly not true. Even the mean of a sample of data is a model of the central tendency of the data, and not the data itself. One clear example of why models are not data is Anscombe's quartet. For example, take the following:

What is remarkable about this quartet is that for all of these scatter plots the mean of x is the same (exactly), the variance of x is the same (exactly), the mean of y is the same (to two decimal places), the variance of y is the same (to three decimal places), the correlation between x and y is the same (to three decimal places), and the linear regression equation is the same (to two or three decimal places). In other words, the models of the data (e.g., mean, variance, correlation, etc.) are the same, but the data are not!

So what's the solution? As I've mentioned in previous posts, graphing the data is crucial, because we're forced to confront the actual data, and not models of the data.

Saturday, March 03, 2012

R versus Stata Redux

I've used both R and Stata for a long time, but these days I use Stata much more frequently than R. While R is useful for some kinds of graphics (especially three-dimensional graphics) and some statistical procedures (for example, finite mixture models), in general I prefer Stata as the go-to statistical program. The reasons are clear: Stata has superior help files for almost all ado files, Stata graphics are excellent (even contour plots are available in Stata), cleaning data is a breeze in Stata but awkward in R, labeling data is much efficient in Stata (in fact, as far as I can tell R does not allow for labeling variable names, while Stata allows for labeling levels of a variable, the variable itself, and the data set), and for many procedures Stata's syntax is much more parsimonious than R's.

Yet, R is worth learning because the 3-D graphics available are often extremely useful for exploring the data, and there will certainly be cases in which R will have statistical procedures that are unavailable or cumbersome in Stata (Bayesian analyses and finite mixture models come to mind, for example).


Friday, March 02, 2012

Culture and Poverty

The New York Times has an article covering the concept of the culture of poverty here. The article is fairly accurate, and does a good job highlighting that the study of culture and poverty had its origins in left-wing Marxists (although I would have mentioned Bowles and Gintis, who emphasized that cultural values and norms of obedience to capitalist ideologies rather than intelligence contribute to the social reproduction of inequality). The author elides the fact that the problem with the concept of the "culture of poverty" is that such a thing does not, and never has, existed: culture is everywhere, not just among the a subset of the economically disadvantaged. The appropriate question, then, is: given that we know that culture is a constituent part of the human experience, how does it matter not just for poverty, but for happiness, well-being, inequality, wealth, and so on?

Thursday, March 01, 2012

Values and Politics

I'm a bit biased, but the front page of the Huffington Post highlighted a fascinating study on education, culture and politics today.

Wednesday, February 29, 2012

Reading the New York Times in Stata

One useful command for taking a break from research is Neal Caren's "nytimes" ado file. This command lists the most recent headlines with brief summaries from the New York Times. Best of all, no subscription is required!

Tuesday, February 28, 2012

Utility Theory as Naive Cultural Theory

Here's a fascinating presentation by the economist Steve Keen on utility theory and neoclassical economics. From the perspective of a cultural sociologist, what is of particular interest is that the utility theory underlying neoclassical economics has the appearance of a naive cultural theory. Specifically, the indifference curves that constitute supply and demand curves in neoclassical analysis are based on strong, disproved assumptions about how people value things in the world: first, completeness (i.e., that the individual knows their evaluative ranking of all combinations of things); second, transitivity (i.e., if thing A is valued to B, and B to C, then A is valued over C); third, non-satiation (i.e., more things are always valued to less); fourth, convexity (i.e., for each thing, additional value falls); fifth, structural independence from culture (i.e., what an individual values is independent of how much income the have); finally, no curse of dimensionality (i.e., information processing abilities are unlimited). No cultural theory  in sociology has even approached the disbelief required for these kinds of assumptions. Fortunately, some sociologists (for example Michael Hechter), have sought to correct this naive cultural theory, and have advocated eloquently and convincingly for a richer understanding of values in economic models of human behavior.

Monday, February 27, 2012

The Phil Gramm Effect

I recently re-read Andrew Abbott's brilliant article on the problems with classical linear regression. One of the most persuasive criticisms is that statistical models are extremely difficult to use for examining small changes with big effects (but big changes with small effects can be modeled). I like to call this the "Phil Gramm Effect" because arguably one of the most important causes of the 2008 financial crisis (an undoubtedly big effect) was Phil Gramm (a small change), since he was the driving force for gutting the Glass-Steagall Act and shifting government regulations in favor of private companies (often called "deregulation," but more accurately termed "re-regulation").

Sunday, February 26, 2012

Big Science in Sociology

The search for the Higgs Boson particle has captivated a wide range of people all over the world, and the construction of the Large Hadron Collider is the reason for this widespread interest. Is such a "big science" approach possible in the social sciences, including sociology? Although the details to me seem obscure, researchers in Europe have developed a proposal for what they call the FuturICT, a "big science" project for the social sciences (ICT stands for "Information and Communication Technology") in the mode of the Manhattan Project, Apollo Project, and Large Hadron Collider. But what is it, exactly, that they are proposing? I get the sense it's a giant computer simulation, but it doesn't seem entirely clear.