Zombies are now a common topic of discussion. In fact, the data we have available from Google Trends (for the phrase "zombie attack") strongly suggest an increasing risk of zombification across the world:
However, academic research on zombies is limited (i.e,. non-existent), mainly because of the lack of high quality data. For those interested in studying zombies, I refer readers to Andrew Gelman's paper (co-written, apparently, by the great zombie film director George Romero) on how to measure zombie outbreaks via indirect survey techniques. You can find his article here. Even if you're not interested in zombies, his paper offers some good ideas on how to sample difficult-to-reach populations more generally.
Showing posts with label Data. Show all posts
Showing posts with label Data. Show all posts
Thursday, May 31, 2012
Tuesday, March 13, 2012
Scatter Plot Matrix in R
Stata has a large number of graphics capabilities (and I highly recommend Stata over other statistical packages for a variety of reasons), but in a few instances R is more useful. In particular, I find R useful for creating beautiful scatter plot matrices and 3-D graphical displays. To my knowledge, currently these kinds of graphics are very difficult (if not impossible) to create in Stata 12. What I like about scatter plot matrices is that can have a high data-to-ink ratio, packing together fitted lines, scattered data, histograms, correlations (proportional to the size of the correlation), and statistical significance "stars" (since reviewers seem to like them). Moreover, I like that all the information effectively puts the "stars" associated with statistical significance in appropriate context: there is an incredible amount of variability in the size of correlations and distribution of data among all the "three-star" correlations, underscoring the limited usefulness of statistical significance as a tool for understanding the social reality given to us by data.
Friday, February 24, 2012
3-D Bar Graph "Masterpiece"
I encountered this post on how to turn a "boring" bar graph into a 3-D "masterpiece." What's striking to me is that most of the people commenting actually want to replicate this graph, even though it violates the basic principles of effective statistical graphics, according to Tufte and others. For example, the 3-D effect distorts the information displayed by the "boring" bar graph, making comparisons difficult, and the visualization effects distract from the underlying data as conveyed by the differing heights of the bars. Here's the "masterpiece" in its full glory:
Saturday, February 11, 2012
Era of Big Data
The New York Times has a great article discussing the era of big data. This might have a Kurzweil-esque ring to it, but due to technology change big data is becoming increasingly available and ready for analysis: in fact, there are more data sets out there than brains to analyze them, especially when one notes the incredible number of combinations of analyses that could be conducted even on a single data set with 100 variables (in what is known as the curse of dimensionality). However, one problem with big data is that, since so much of the data are collected by private entities, much of it may not be available to academics and independent researchers.
Thursday, January 12, 2012
Inequality versus Dispersion
I'm glad to see that Alan Krueger, chairman of the Council of Economic Advisers (a fancy name for a panel of three economists), discussed the problems with inequality in his address today. You can find his remarks and graphs here. I liked his graphs, and he shows convincingly many of the standard findings in sociology and political science on politics and inequality in the United States. However, I found the following comments puzzling:
Although I have done much research in my career on inequality, I used to have an aversion to using the term inequality. The Wall Street Journal ran an article in the mid-1990s that noted that I prefer to use the term “dispersion.” But the rise in income dispersion – along so many dimensions – has gotten to be so high, that I now think that inequality is a more appropriate term.The mixing of the statistical concept of dispersion with the sociological concept of inequality muddles the discussion. It's true that any distribution is often described by some measure of dispersion (e.g., standard deviation) and central tendency (e.g., mean or mode). But inequality encompasses a concept of equity, as well as some concept of disparity (or disparities), neither of which is analogous to the statistical concept of dispersion. Moreover, if we use Krueger's logic it's unclear at what threshold "dispersion" is labeled "inequality"; for instance, his comments imply that Sweden currently has dispersion, while the United States has inequality, although many Swedes would probably disagree.
Tuesday, January 03, 2012
Congratulations to the Digging into Data Recipients
The list of the round two award recipients for the 2011 Digging into Data challenge are listed here.
Tuesday, July 19, 2011
Cultural Contradictions of Pop Economics
Andy Gelman has a fascinating post on the apparent contradictions of pop micro-economists today. I highly recommend reading it, as well as the comments. In essence, Gelman argues that many pop economists take one of two positions, depending on the circumstance: first, people are rational and respond to incentives (and thus behavior that appears irrational is actually rational once you take the perspective of an economist); second, people are irrational and don't respond to incentives (and thus, they need economists, with their open minds, to show them how to be rational). The problem, argues, Gelman is that these positions are entirely contradictory, and that pop-economics plays hopscotch with these viewpoints, switching from one to the other.
Friday, December 11, 2009
Do Social Networks Affect Health?
In a recent series of ground-breaking articles over the past several years, Nicholas Christakis (a sociologist here at Harvard) and James Fowler (a political scientist at UC Davis) have shown that health behaviors seem to flow through social networks. Using new data from the Framingham Heart Study, they've shown apparent contagion effects for obesity, smoking, depression, and even isolation. Recently, however, Ethan Cohen-Cole and Jason M. Fletcher (two economists) have used data from Add Health to show apparent "implausible social network effects" for acne, height, and headaches. The economists also demonstrate that after controlling for "environmental confounders" the effects disappear, suggesting that contagion through networks are really just capturing similar structural conditions (e.g., the fact that I get obese after you are obese is really because we both live in a neighborhood with another fast food restaurant).
Although Cohen-Cole's and Fletcher's criticisms might seem relatively damning, there are a number of very serious problems with concluding that social network effects do not exist:
(1) Different data: Cohen-Cole and Fletcher use a different dataset than Christakis and Fowler. The studies differ in location (high schools versus homes), population (teenagers versus the general population), time (less than one decade versus three), and so forth. It's entirely possible that contagion effects do not exist for teenagers in high schools, but do exist for members of Framingham, Massachusetts. This would make sense if, for example, social imprinting effects were weaker among teenagers with possibly ephemeral relations in high schools rather than adults with more durable relationships.
(2) Presence vs. absence: The presence of putatively implausible contagion effects for variables such as height and acne does not demonstrate the absence of plausible contagion effects for variables such as obesity and happiness. Another way to think about it: studies showing that smoking causes cancer are just as valid even if, say, frequent use of a fireplace in a house does not cause cancer. Although the mechanisms are broadly similar (inhaling smoke), there may be important differences (the substance that is smoked).
(3) Definition of "implausible": Cohen-Cole and Fletcher assert that network effects for acne, headaches, and height are "implausible." Is this really the case, however? All the data from AddHealth on acne, headaches, and height are self-reported; thus, if I have a friend who complains of headaches then I may very well find it easier to complain of headaches due to changing norms or a driving desire to be more like my friends. Or perhaps my friend has found an effective way to prevent headaches (such as a daily dose of aspirin), and I've adopted my friend's behaviors. Or it's even possible that headaches are associated with stress, and that what is actually spreading through social networks is stress along with headaches. These same points can also be said about acne. What about height, though? Self-reported height could also spread through networks; if my friends are tall, then I'm likely to nudge up or down my actual height. We've all known short guys who have added a few inches and tall women who've subtracted a few. In addition, even assuming that height in AddHealth is accurately measured, there is substantial evidence that height during adolescence is greatly influenced by diet, exercise, and smoking habits. To the extent that any of these flow through networks, then so will height.
(4) Flaws with fixed effects: Cohen-Cole and Fletcher used fixed effects to "control" for time-invariant unobserved differences between schools. Although superficially a good way to adjust for all stable differences between schools, it is well-known that the fixed effects estimator reduces the size of the standard errors (intuitively, because less information is used to estimate the coefficients). This is especially the case if the number of waves is small (as is the case with AddHealth, which consists of only three waves of data, although each school has numerous individuals). More importantly, however, is that with a lot of inefficiency not only are the standard errors larger, but the point estimates may also become small and fragile, resulting in erroneous inferences (in the same way that a biased coefficient would). For this reason I'd recommend Cohen-Cole and Fletcher re-do their analyses using fixed effects vector decomposition, which can improve the efficiency of estimates yet retain some of the advantages of the traditional fixed effects estimator.
(5) Neglected information: Another problem is that Cohen-Cole and Fletcher neglect important ancillary evidence that supports the interpretation of the Framingham data given by Christakis and Fowler; to wit, there is substantial evidence from social psychology that people automatically mimic each other in myriad ways (e.g., emotions and behavior, including other people's tone of voice and body posture). Social mimicry is a highly plausible mechanism that buttresses the network diffusion interpretation of the data.
(6) Relevance of networks: Even if network effects are not causal, such descriptive information is still highly relevant. Most importantly, knowing the social patterning of behaviors and attitudes is extraordinarily useful for maximizing the impact of social interventions. For example, regarding obesity, clinicians could provide information on the benefits of dietary changes to people and then offer incentives to overweight people to recruit others in their social network. In this way, the population in most need of a particular intervention would be efficiently affected.
Understanding social causes is complex, and necessarily requires an accumulation of different kinds of knowledge. This is how epidemiologists discovered that smoking causes cancer: a few observational studies combined with information on the substances in cigarette smoke, not to mention a dash of qualitative evidence from clinicians. There was no single "smoking gun" (pun intended). In the same way, understanding social network effects will not be ruled out by a single study using a fixed effects estimator on a handful of variables gathered from a few thousand teenagers living in late 20th century America. Rather, understanding the degree to which social network effects can be understood as causal will require results from various sources, including both nationally-representative datasets (e.g., AddHealth and Framingham) as well as additional studies on contagion mechanisms (such as the findings on social mimicry from social psychology).
Subscribe to:
Posts (Atom)

