(A.1) Who can hit a softball the best?
Consider these data on the softball hitting differences for both genders in this class over the last ten years (download here). The average distance for females is 92 feet, whereas males hit 172 feet. The averages clearly show males to be superior hitters.
The comparison is a fair one. Both genders in the class represent ordinary students. It is not like the male baseball team tends to enroll in this class but the female softball team does not. Although there are more observations for males than females, that does not systematically inflate the average for males (or deflate the average, for that matter)
Most readers have taken an introductory statistics course, and if so, they might protest that one cannot simply compare the average from two samples. A statistical hypothesis test must be conducted to determine if the average for males is truly higher than the average for females. That is, we must make sure the higher male average is not the result of random chance, and "stochastic" is another name for random. We want to know if males hit further in the data, and if they do, whether it is caused by gender (a deterministic factor) or randomness (the stochastic factor).
I disagree, and claim we can simply look at the averages, dismissing all notions that the discrepancy between the two averages are due to random chance. First, the sample size is relatively large: 443 observations for males and 248 observations for females. Second, the average for males is 79 feet higher than females, and it is hard to believe that in a sample this large that difference could be attributable to randomness. If the male average was only three feet higher than females, then I would not conclude that males definitely hit further, and would say a statistical test is necessary to make this determination. Given the sample size and the large differences in hitting distances between the two genders, I conclude males from this class really do hit further than females, on average, and I feel comfortable making that claim without the use of a statistical test.
(A.2) Check for yourself
There is no need to just assume I'm right though. Download the data at the previous link and test it yourself. You probably recall from your statistics course that you need some sort of t-test to determine if the average for males is really larger than the average for females, but you probably don't remember exactly how to use the test. Nor is there a need to recall. Excel remembers for you. Just go to Excel, and explore the statistical functions (under formulas), looking for the t-test. You will eventually find a function labeled T.TEST , which gives you the p-value of your test. Google "T.TEST in Excel" for detailed information on the formula.
A quick remark on the difference between sample averages and population averages. From the class data we know the average hitting distance for our sample of males and females in the data. What we are interested in, though, is whether the average hitting distance for all males and females between the ages of 18-22 (our population of interst) is different. The sample averages are used to make inferences about the population averages.
Video 1—Watch Bailey Use the T.TEST Function In Excel
The p-value tells you the probability you are wrong if you say the population averages are different (well, not exactly, but that is an intuitive way of explaining it that is close to being accurate, and that is how we will interpret p-values in this class). Researchers generally don't like to say two population averages are different if the chance of them being wrong is greater than 5%, so when the p-value from the T.TEST formula is greater than 0.05 we say the difference between the two averages is not statstically significant (that is just statisticians' way of saying the averages do not seem to be different). When the p-value is 0.05 or less, then we say the population averages really are different, because the chance we are wrong is less than 5%.
Testing the hypothesis that two averages are different:
Two sample averages are used to test whether the population averages are different. Using the T.TEST
function in Excel we get the p-value for this hypothesis test. If the p-value is greater than 0.05 then we
say the difference between the two population averages is not statistically significant. If the p-value is 0.05
or less we say the difference between population averages are statistically significant, knowing the chance of us being
wrong is less than 5%.
Using the T.TEST function for the softball data the p-value is absurdly low, telling us our intuition was correct that males really do hit further than females. The low p-value means we can say this knowing there is basically a 0% chance we are wrong. But we knew this, didn't we? The sample is so large, and the male and female average so different, a hypothesis test is really not needed.
Figure 1—Using T.TEST Function in Excel To See If Males Hit Further Than Females
If difficult to see click here.
When is a sample so large that you can dispense with statistical tests? This is a judgment call. To be on the safe side you can always use a test, but most of you will not remember how to perform such tests, and consultants can be expensive. You have the right to conclude for yourself that the sample size is so large, and what you remember about statistics is so little, that a statistical test is not warranted.
Note on T.TEST function: I use a one-tail test because we are testing whether one is larger than another, not whether they are equal, and I set Type=3 because I assume hitting distances between males and females probably have different variances. In this classes, we always use 0.05 as the threshold for T.Test, meaning if the p-value from the T-Test is greater than 0.05, then we say the averages from the two groups are not statistically significant.
(A.3) When hypothesis tests are needed: the effect of beef commercials on beef demand
There are times when sample sizes are not large, and the differences between two averages so small, that a hypothesis test really is needed to determine if one average differs from another. Consider the example of beef advertising. Does it work? When people see a commercial for beef do they increase the value they place on beef, and consequently purchase more beef? Ideally one would collect data on actual beef purchases and study how they fluctuate before and after beef commercials are aired. The problem is that there are so many things affecting beef demand other than commercials that it is nearly impossible to identify the effect of commercials alone.
Video 2—A 1985 Beef Commercial Starring James Garner
(A.4) A very unusual auction
Some economists have instead looked to the laboratory for an answer. One group of Cornell researchers recruited individuals to attend a research session.(M1) For one group (the Control Group) they simply held an auction for hamburgers to measure they value they place on hamburgers. This auction was specifically designed to elicit the person's true value for the hamburger. If a person submits a bid of $1.00, that means they will buy the hamburger if the price is less than $1.00, but if the price is one penny above $1.00 they will not.
An example is a second-price auction, where the highest bidder wins the burger but the person only pays a price equal to the second-highest bid. This means that if the true value you place on a burger is $3.20 (meaning you will pay up to $3.20 but no more) then you should enter a bid equal to $3.20. This maximizes your chance of winning the auction, and you are guaranteed a good deal if you win. After all, you will always pay less than your maximum willingness-to-pay. Moreover, there is never a time when you regret winning the auction, so you never bid "too high," again, so long as your bid equals the true value you place on the burger.
Back to the Cornell experiment. Another group (the Commercial Group) saw a five minute montage of beef commercial videos before participating in an identical hamburger auction. There were 92 individuals in the Control Group and 44 in the Commercial Group. If advertising has a positive impact on beef demand then the Commercial Group should submit higher bids than the Control Group (assuming the demographic profile of both groups is roughly the same, which it is). The researchers were kind enough to make the data available to this class, and can be downloaded here.
The average of the Commercial Group is indeed higher, as shown in Figure 2. People in the Control Group valued the hamburger at $2.14 while the average value for those who saw the beef commercials is $2.52. Should we conclude from these averages alone that beef commercials enhance beef demand? No, we should be careful about making such claims. The sample size is small, and so we are trying to make bold claims based off only a few individuals. The difference in the averages could be the result of randomness, like a coin-flip, and the sample sizes are not large enough to make randomness irrelevant. This is reflected in the standard deviations of the values, which are not only larger than the differences between the two averages, but are larger than the averages themselves!
Figure 2—Does Beef Advertising Work?
| Value of a hamburger | ||
| Control Group | Commercial Group |
|
| Sample Size | 92 | 44 |
| Average | $2.14 | $2.52 |
| Standard Deviation | $2.27 | $3.04 |
We know the sample average for the Commercial Group is larger than the Control Group, but what about the population average? If this experiment could be conducted with tens of thousands of participants, would the Commercial Group's average still be higher? A hypothesis test is warranted to discern whether the larger value in the Commercial Group was due to the beef commercial or randomness. Using the T.TEST function (with a one-tail test and allowing for different variances in each sample) the result is 0.24. This suggests that if we claim beef commercials enhance beef demand we have a 24% chance of being wrong. This is a chance we cannot take, for we need this chance to be 5% or smaller. As a result, we conclude that there is no statistical difference between the values in the two groups. Advertising is not found to increase the value consumers place on beef.
If the sample sizes were in the thousands and we found the same sample averages as in Figure 2, then the difference in the averages would almost certainly be statistically significant, and one could probably get by without a hypothesis test without receiving much criticism.
(A.6) When hypothesis tests are needed: Mad Cow disease and beef commercials
There are other media exposures that can influence a person's beef demand. When Mad Cow disease started killing people, people who watched the news found out all sorts of repugnant things about the beef industry. Mad Cow disease was spread by feeding sheep carcasses to cows—feeding sheep carcasses to cows! That sounds so unnatural that is sounds unsafe. Most people know that cows are not carnivores, and when they discovered cows were being fed carcasses (even if the carcasses were highly, highly processed) they lost confidence in the judgment of the cattle industry. Animal scientists approved this practice, believing it to have no harmful side-effects, but then the practice was shown to infect beef consumers with a pathogen that turns their brains into sponges. If animal scientists were wrong about this feeding practice, what other dangerous practices might the scientists be approving? You cannot not blame people for responding to the Mad Cow scare by valuing beef less and thus purchasing less beef.
From the beef industry's perspective, they would like to know how news about Mad Cow disease affects beef demand. Like before, it is difficult to study consumers' real purchases of beef, because these purchases are affected by myriad variables, and determining whether rises or falls in purchases are due to news about Mad Cow disease or other variables (e.g., beef price, pork price, season) is almost impossible. Again, agricultural economists have turned to the laboratory for answers.
The same Cornell researchers who studied the effect of beef commercials also included two other groups in their experiment: the Mad Cow Group and the Mad Cow / Commercial Group. The data for all four groups can be downloaded here. As the Excel spreadsheet explains, the Mad Cow group watched a brief NOVA video on Mad Cow disease, after which they participated in the hamburger auction. If information about Mad Cow reduces beef demand then the average value for the Mad Cow group should be lower than the Control Group.
The Mad Cow / Commercial group also watched the NOVA video, but then also watched the beef commercials shown to the Commercial Group. Then, they participated in the hamburger auction. The objective was to see if running beef commercials on television after a Mad Cow scare can restore beef demand, or at least mitigate damage that Mad Cow inflicts. If the average value for this group is higher than the Mad Cow Group, then yes, running commercials after an outbreak of Mad Cow disease might be prudent.
(B.1) When averages can be used without any other statistics: randomized trial experiments
There are many other settings when you can compare two different samples without statistics. Remember that every sample is taken from a population. The shopping preferences for all single mothers can be estimated by studying a sample of all single mothers. It would be unreasonable to ask that every single mother in America must be studied before anything can be said about their shopping habits. So long as the characteristics of the sample closely mimic the characteristics of the population, and that sample is very large, averages of the sample can be used without the use of statistical hypothesis tests. Consider a few examples, all of which refer to "randomized trial" experiments.
(B.1.i) Credit Indemnity, in South Africa, is a large micro-lender, meaning it makes small loans to individuals, and they advertise by mailing letters to tens of thousands of people. Some people respond to the letters by taking out a loan and others don't. Credit Indemnity wanted to know the extent to which the interest-rate advertised in the letters matter, so they conducted an experiment where they mailed 50,000 letters, half of them randomly selected to contain an interest rate of 3.25% and half with a 11.75% rate. It is important to recognize that the samples for each interest rate were virtually identical. They didn't send one interest rate to everyone in the northern part of South Africa and the other rate to those in the southern region, because then they wouldn't be able to determine if the differences in the response rates were due to the interest-rate or the region. Also, the letters sent to all samples were absolutely identical except for the interest rate printed. They both had the same pictures, same paper, same everything.
When a higher pecentage of people responded to the letters with the lower interest-rate, Credit Indemnity assumed the difference to be real and did not conduct a statistical hypothesis test. They knew that even if a statistical test were performed, so long as the percentages differed by only a little it would be deemed "statistically significant." With a sample size of 50,000 (25,000 for one rate and 25,000 for the other) any differences due to random flukes would not affect the average of 25,000 observations.
This firm conducted other similar experiments, finding that certain pictures printed on the letters elicited a greater response. Imagine the value of this discovery. All they had to do is throw a picture on the letter and their profits rise!(A1)
(B.1.ii) The Chinese educational system has a problem in that about 10% of school children have vision problems, yet hardly any wear eyeglasses. Would they perform much better if given free eyeglasses? A randomized trial was performed to answer this question. Researchers took 165 schools in Western China, where 103 schools were randomly selected to receive free eyeglasses and, by random chance, the remaining 62 did not. After one year, the average scores on standardized tests for the schools receiving free eyeglasses increased by a much larger margin than the schools that did not receive free eyeglasses. Because the trail involves over 19,000 students, and those who did and did not get glasses was randomly determined, the researchers knew it was the eyeglasses that caused those 103 schools to improve so much.(G1)
(B.1.iii) Researchers and governments also engage in these large-scale experiments, often to evaluate government programs to help the poor. If the U.S. implements a nationwide poverty-assistance program it will be difficult to determine its impacts, because during the time the program is implemented many other things in the world will be changing. If the program is implemented from 2008-2010, they could not simply look back and compare the poor before 2008 to after 2008, because the Great Recession began in 2008! They would be unable to determine the extent to which the program or the Great Recession caused changes in the status of the poor. Instead, what they should do is provide the program to only half of the poor, and randomly select those recipients. If the sampling is completely random, then the people who receive and don't receive assistance from the program are virtually identical—except for the fact that some receive government help, and that is the point. If the sample size is, say, above 10,000 people, the government knows that whatever differences exist between the two groups in 2010 is the result of the program.
(B.1.iv) That is what Esther Duflo from MIT did to determine whether an anti-poverty program in Bengal would work. Of a number of families applying for help, a portion of them are randomly selected to be given a productive asset—cows, goats, or chickens—and are taught how to care for them. Those random selectees, the study finds, eat more, make more money, and save more money; they also experience better mental health. Note that when we say "eat more" we are comparing them to those who applied for help but due to a random selection did not receive help. Those who do not apply for help are not included. The sample receiving and not receiving help are virtually identical, except for whether they received aid, and so all the differences between the samples can be attributable to the aid.(E1)
(B.1.v) Does medicaid help people? On the one hand, it seems obvious that free health care is beneficial. On the other hand, Medicaid is poorly administered, and the low rates they pay doctors often discourage doctors from accepting Medicaid patients. There is already a safety net in the form of charities, and this "other hand" says that diverting people away from the better run charities detracts from their health.
Then comes Oregon to settle the debate. Oregon announced around 2009 that it would allow 10,000 new slots for its Medicaid program, but they knew many more would apply, so they held a lottery to determine who would win the slots. Indeed, around 90,000 people applied, and by randomly selecting the recipients, researchers could simply look at how the lives of those who received Medicaid changed, compared to those who did not receive Medicaid but applied. There is no self-selection to deal with, no demographic differences, and the sample is large, so simple averages from both groups tell the whole store.
Results clearly show that Medicaid helps people. Those randomly selected to receive aid have better health, better job prospects, and higher incomes. While one can debate whether the costs of funding Medicaid are worth the benefits to its recipients, we can no longer debate whether the benefits exist.(P1)
(B.1.vi) School vouchers have been debated in the U.S., but we now have evidence in favor of the voucher program. New York City established a program in 1996 where families can apply for a voucher for their child to attend a private school of their choice. More applications were received than availability, so the program randomly selected those who would receive the voucher. By comparing the 2666 recipients not to the general population but to people who applied for but didn't receive a voucher, we can determine the failure or success of vouchers. The evidence is clear. Vouchers are good, especially for African Americans, as the vouchers increase their likelihood of entering college by 24 percentage points.(C1)
(C.1) Sex Researcher Alfred C. Kinsey and the stratified sample
No biologist shocked conservative America more than Alfred C. Kinsey, who changed from studying gall wasps to human sexuality when he recognized how little his students knew about sex. This transition occurred in the 1930's, at a time when homosexuality was "a disease" and young people entered marriage with very little knowledge of what to do on their honeymoon. While most of society avoided any discussion of sex, Kinsey set about interviewing thousands of individuals about their sexual habits, and then publishing these statistics in a book. The reader is almost certainly aware of the bigotry that existed against homosexuality before the 1990's. Any gay person who wanted to live a normal life in normal society would keep their sexual predilections secret. Imagine, then, how Americans responded when they read this in the 1948 book Sexual Behavior in the Human Male.
Figure 3—Kinsey's 1948 Claims About Homosexual Tendencies (narrative)(K1)
 narrative.jpg)
Imagine the jaws that dropped when reading this. Could it really be true that 45% of males who mature early have homosexual experience?
I argue it isn't true. To see why, one must understand where the data in Figure 4 (below) were collected. As anyone who saw the movie about Kinsey knows, the data came from intensive interviews with volunteers. The word "volunteer" is stressed, because Kinsey could only acquire data from individuals willing to share, and so long as differences exist between the people willing and not willing to share details of their sexual life, Kinsey's data will be biased, in that the averages in Kinsey's sample do not accurately reflect the averages of the population at-large. They only reflect averages of people who agreed to be in the sample. This is referred to as sample-selection bias, where the researcher's sample is not representative of the population being studied, often because the sample is biased towards one subset of this population. The percent of the sample belonging to this subset is higher or lower than the percent of the subset in the population.
For those who have not watched the movie, see the brief excerpt below to get a feeling for how the data were collected.
Video 3—How Kinsey Interviewed People(K1)
(For browsers other than Internet Explorer)
If you have trouble seeing the video click here.
Figure 4—Kinsey's 1948 Claims About Homosexual Tendencies (data)(K2)
We will now develop a simple model to show why Kinsey's data are biased and how it can conceptually, but not realistically, be corrected. Take the entire population of Americans with at least one year of college and who became an adolescent at the age of 11. These are Americans represented by the sample in Figure 4. For every 100 of these males, what is the average number who had a homosexual experience by the age of 20 (making this average essentially a percentage between 0 and 100%)? That is the question. Suppose further that the true average out of the entire population of these males equals X, but a sample average of this percentage is x-bar. Kinsey's data (according to Figure 4) estimates the population average, X, with the sample average x-bar = 37.6. Is 37.6, a sample average, a good predictor of X, the true average?
X = the true % of all Americans who matured at age 11 and who had a homosexual experience by age 20
37.6 = x-bar, the estimated percent, calculated from a sample of people who volunteered
to share their sexual history
Does X = 37.6?
Is it close?
Let us remember Kinsey only had access to males who would voluntarily talk about their sexual past, and these are only a subset of the population. Hereafter, I call them volunteers. It seems reasonable to assume that males with a greater willingness to talk about their intimate past are more open to sexual taboos and are therefore more likely to have had a homosexual experience. This means that if 37.6 is the average for volunteers, the average for the non-volunteers is less than 37.6. That is, if a male does not wish to talk about sex, he is less likely to have had a homosexual experience.
Denote the true average for volunteers as XV and the true average for non-volunteers as XN. Also, let PV be the percent of these males who are volunteers and let PN be the percent of these males who are non-volunteers. By definition, the value of X (the average for all the males) equals:
[Equation 1] X = (XV)(PV) + (XN)(PN).
XV = true average of X for all who will share their sexual past
XN = true average of X for all who will NOT share their sexual past
PV = the percent of the population of interest who will share their sexual past
PN =the percent who will NOT share their sexual past
True value of X for ALL Americans (including those who will and will NOT share
their sexual past is:(i)
X = XVPV + XNPN
Kinsey was unable to calculate Equation 1 because he could not calculate the value of XN. Instead, he replaced XN with XV, and since we said XN < XV, his estimate of V was overestimated. Do not take this as a criticism of Kinsey, for he cannot make people talk to him. The error he made was in how he promoted his work. Instead of acknowleding this bias as an unavoidable consequence of studying people in a free country, he promoted the numbers in Figure 4 as descriptions of how Americans truly behaved.
Now you see why Kinsey made Americans uncomfortable. The public believed only a small percentage of males had the "disease" of homosexuality, when Kinsey was claiming it was far more prevalent.
Suppose that Kinsey had a mechanism for getting the non-volunteers to talk to him and tell him the truth. Maybe he could pay the non-volunteers a large sum of money to overcome their caution, and used some survey tricks to entice them to be honest. If this were the case, then Kinsey really could estimate XN. This still isn't enough to calculate Equation 1. He would also need an estimate of PN. After that, Kinsey would then have all the variables he needed to calculate Equation 1, and his estimates would no longer be biased.
This is referred to as stratified sampling, where the population is parsed into sub-populations and are assigned their own average, and then the average for the population is calculated as a weighted-average of the sub-population, where each sub-population receive a weight equal to its representation in the population.
As an example, suppose Kinsey could study volunteers and non-volunteers, and estimated XV to equal 37.6 and XN to equal 4. Further, suppose that volunteers represent about 5% of all these males and non-volunteers represent 95%. The average for the entire population is then calculated as the following weighted average.
[Equation 2] Sample average for everyone (using stratified sampling) = (37.6)(0.05) + (4)(0.95) = 5.68.
The population average of 5.68 is far less than Kinsey's estimate of 39.6, because Kinsey's method of sampling from the population was biased from the start.
(C.2) Weren't we talking about large sample sizes?
The student may understandably feel this article has lost its focus, so let's return to the issue of when statistics are and are not needed. In the early stages of Kinsey's research he sought advice on his methodology, and was particularly interested in whether he would be criticized for his statistics, like those in Figure 4. An unfortunate piece of bad advice was passed his way, which tricked him into thinking the bias in the previous section did not exist. Below is an excerpt from Kinsey's biography, which served as the basis for the subsequent movie.
Kinsey was particularly pleased by Pearl's advice on sampling. "When I told him that someone else would have
to handle the mathematics of the material, he told me that when one has such
quantitites of material as I have it needs very little manipulation," beamed
Kinsey. "[Pearl] points out that statistical theory is largely a substitute for
adequate data," he continued. "All this encourages me greatly."
—Jones, James H. Alfred C. Kinsey. Chapter 15.(J1)
It is true that when sample sizes become large, statistics can sometimes be ignored, allowing one to focus on averages as if they were the 100% accurate answer. The problem with Kinsey's research was a bias in his sampling, though, and that cannot be eliminated by larger sample sizes—and Kinsey did have large sample sizes. In regards to estimate the average number of males who had a homosexual experience, Equation 2 is the correct equation, regardless of whether the sample size is 100 or 100,000.
Video 4—The Colbert Report on Surveys and Sampling
(For browsers other than Internet Explorer)
If you have trouble seeing the video click here.
(C.3) The Kinsey Reporter app
The controversy surrounding Kinsey's research didn't just help Americans sort their feelings about morality and sex. The debates about the validity of his statistics helped researchers resolve exactly how they should conduct surveys and interpret the results. The debates about Kinsey's research has largely been settled, and his life has now become an American institution. He more than any other influenced how Americans think about homosexuals and other sexual behaviors that were once considered immoral. It is not surprising, then, that when Indiana University created a free app allowing the public could anonymously report information on their sexual behavior, they named it the Kinsey Reporter.
This app has the potential to overcome some of the obstacles Kinsey faced in his research. If it is true that some portion of the population refused to participate in interviews because they did not want to talk with another person face-to-face about their sexual history, this app allows them to anonymously reveal their sexual history to a smartphone. Perhaps this will make those who were reluctant to participate in this research now share freely.
Figure 5—The Kinsey Reporter app
(D) Making non-representative samples look representative
In my 2013 class we constructed a survey where we asked individuals the question in Figure 6. The respondents were roughly half students and half their parents, and this sample is certainly not representative of how America as a whole would answer. It over-samples the young, relative to the U.S. population, and under-samples adults.
You can download the data here.
Figure 6—Survey Question
The figure below shows the average response from young females, adult females, young males, and adult males. It also shows that roughly half of the sample is comprised of young people less than 25 years of age, when in the population of all Americans that percentage would be far less.
So our survey is not representative of the U.S. population. That doesn't mean we can't use the survey to make inferences about the U.S. population. That is, we can make a non-representative survey behave like a representative survey. All we have to do is calculate a weighted average of the survey responses from each of the four age / gender categories, and weight each average by their percentage of the U.S. population. The figure below presents a hypothetical scenario where young females constitute 15% of all Americans. The following weighted average is then an unbiased estimate of what the average survey response would be, if the sample were actually representative.
[Equation 2] U.S. Average Response = (0.15)(4.23) + (0.35)(4.38) + (0.1)(4.64) + (0.4)(4.63) = 4.49
If you calculate the sample average the result is 4.38, meaning the average response for all those who took the survey is 4.38. However, based on this sample, we anticipate that if all Americans had taken it the average would be 4.49.
What this means is that if we did not use weighted averages to make the data representative of the American public we might underestimate the popularity of low-carb diets.
Figure 7—Data from survey
(E) Sample sizes and polling data
Accompanying every major political election is a barrage of statistics from polling data, usually numbers detailing the percent of people who support a particular candidate. These sample sizes are typically larger than 1,000, and that is considered large, but not so large that statistics can be ignored. All polls have a margin-of-error associated with them, which are reported only when candidates have similar support rates.
Often, if one candidate has a support rate of 52% and the others support rate is 48%, the media will report it as a "dead-heat", because the differences in these support rates could be due to stochastic errors in measurement. If another poll was conducted, identical in every way except the specific people who were randomly recruited, the lead could easily switch to the other person, simply because they happened to call different people.
Most of the time polls are said to have a 3% margin-of-error, which means when the estimated support rate is 52%, it could really as high as 55% or as low as 49%. This interval of 49-55% is referred to as a confidence interval. An example is shown below, where the percent of people who (in May of 2012) thought Obama will better handle "the economy" is 45%, while the number for Romney is 48%. Because these numbers are so similar, the reporter is careful to note the margin-of-error of 3.1 percentage points. What this means is that Romney and Obama are basically tied when it comes to "the economy."
Figure 8—Excerpt from Poll Watch Daily (May 22, 2012)(P2)
Footnotes
(i) This is the proof that the population average, X, equals X = XVPV + XNPN, see this proof.
References
(A1) Ayres, Ian. 2007. SuperCrunchers. Bantam Books: NY, NY.
(C1) Chingos, Matthew M. and Paul E. Peterson. August 23, 2012. "A Generation of School-Voucher Success." Letter to the Editor. The Wall Street Journal. A13.
(E1) The Economist. May 12, 2012. "Hope springs a trap." Free Exchange.
(G1) Glewwe, Paul, Albert Park, and Meng Zhao. December, 2010. "The Impact of Eyeglasses on the Academic Performance of Primary School Students: Evidence from a Randomized Trial in Rural China." Accessed August 6, 2012 at http://federation.ens.fr/ydepot/semin/texte0708/GLE2007IMP.pdf.
(J1) Jones, James H. 1997. Alfred C. Kinsey: a life. Norton Publishers: NY, NY.
(K1) Producers: Francis Ford Coppola and Gail Mutrux. Director:Bill Condon. 2004. Kinsey. Fox Searchlight Pictures, et. al.
(K2) Kinsey, Alfred C., Wardell B. Pomeroy, and Clyde E. Martin. 1948. Sexual Behavior in the Human Male. Indiana University Press: IN.
(M1) Messer, Kent D., Harry M. Kaiser, Collin Payne, and Brian Wansink. 2011. "Can generic advertising alleviate consumer concerns over food scares?" Applied Economics. 43:1535-1549.
(P1) Planet Money. June 15, 2012. "Does Medicaid Actually Help People?" Podcast episode 379. National Public Radio.
(P2) Poll Watch Daily. Excerpt. Accessed May 22, 2017 at http://www.pollwatchdaily.com/.