Vote buying – the exchange of particularistic goods and services for political support (Kitschelt and Wilkinson 2007; Nichter 2014) – is said to dominate the politics of the global south. Yet, it remains an elusive concept for scholars to study due to what is known as social desirability bias (Bradburn et al. 1978; DeMaio 1984). Such bias occurs when survey participants falsely answer a question in order to provide a more socially desirable response. Such bias can lead to either underreporting or overreporting of attitudes or behaviors such as prejudice (Hatchett and Schuman 1975) or voting in elections (Campbell 1960). Engaging in vote buying (either as the buyer or seller) is illegal in most countries (Schaffer 2007) and may be considered to be immoral or detrimental to society (Gonzales-Ocantos et al. 2014); thus, survey questions about its practice are often expected to be plagued by social desirability bias. Researchers may also include a list experiment in their research in order to demonstrate that responses to their questions of interest are not sensitive. For example, Kao and Revkin (Forthcoming) use such an experiment in Mosul, Iraq to show that the estimated sensitivity bias to a direct question about the Islamic State’s governance performance could not be distinguished from zero.
To solve the issue of social desirability bias, social scientists often employ list experiments, a statistical technique that has been in use for more than 3 decades (Miller 1984). List experiments give respondents the opportunity to truthfully report their behavior while remaining anonymous to the interviewer (Kuklinski et al. 1997). Essentially, respondents can “hide” truthful responses to a sensitive item among responses to a list of control questions. Instead of answering a sensitive question directly, the respondent is read a set of binary questions; she counts up and reports how many affirmative responses to this list of questions are applicable to her. Respondents are randomly selected to receive either a list of innocuous items (control group) or the same list including the sensitive item (treatment group). A comparison of the average count of affirmative items among those in the control with those in the treatment group allows the researcher to reveal the true prevalence of the sensitive item. If the difference is not statistically significant, the item can be considered to be non-sensitive or at least to the extent that it would greatly affect average responses from among a sample.
Concerning vote buying specifically, list experiments have been widely employed. Gonzales-Orcantes et al. (2012) find that when asked directly, only 2 percent of Nicaraguan voters admit to being offered gifts or services in exchange for their votes compared to nearly one in four reporting vote buying offers in a list experiment. Likewise, Çarkoğlu and Aytaç (2015)’s list experiment in Turkey finds that 35 percent of the sample received vote buying offers, while only 16 percent reported this behavior in a direct question. Blair, Coppock, and Moore (2020, 1309) consider 19 studies that have employed a list experiment to study vote buying and find that the average reporting error is –8 points with a 95% confidence interval ranging between –13 and –3 points.
Such enthusiastic faith in the list experiment to reveal sensitivity bias – to the extent that they are dubbed the “statistical truth serum” (Glynn 2013) – motivates questions about the drawbacks of the technique and how accurate it really is. One major criticism of the method is that it is inefficient, requiring a high number of respondents to detect effects (Corstange 2009; Blair, Coppock, and Moore 2020). The literature also notes important conditions that list experiments must satisfy in order to be accurate including consideration of “ceiling” and “floor” effects (Blair and Imai 2012; Glynn 2013) and assumptions that inclusion of the sensitive item does not affect responses to the control items (Blair and Imai 2012). In this study we focus on the extent to which they work as expected in practice. Based on decades of survey research demonstrating that complicated survey questions are difficult to answer and particularly so for subgroups of respondents (Krosnick 1991), we have good reason to expect that outcomes from these experiments are more imprecise than is often implied in research that employs them.
Scholars speculate that there is high potential for either surveyors or respondents to botch the implementation of list experiments due to their long and confusing format (e.g., Kramon and Weghorst 2019), but there is little systematic, empirical evidence of just how prevalent problems are in the field. In particular, we expect that list experiments often suffer a major implementation failure during fielding: respondents end up revealing their answers to list items despite being asked not to do so. As far as we know, the extent of list experiment implementation failure has not been studied before. As noted above, we have reasons to believe that respondents often fail to understand list experiments, including the purpose behind them meaning they do not realize the anonymity this experiment affords them. This is a serious concern as it invalidates the main purpose of the experiment. We also suspect that these implementation errors occur more frequently among certain subpopulations. Thus, using list experiments to detect sensitivity of subjects among subpopulations may instead reveal differences in implementation. Finally, since this problem has not been directly studied before, scholars do not fully understand the reasons these errors in list experiments may occur.
Our research seeks to fill these gaps. We interrogate two possible reasons for why implementation errors may occur: 1) violations of anonymity, such that answers to specific list items are revealed, and 2) cognitive overload, leading respondents to be unable or unwilling to carefully consider each item and count up their affirmative answers to all of the list items in order to report them at the end of the question. Both of these implementation errors invalidate the purposes of a list experiment and the latter error is expected to explain why the former occurs.
The study will make a contribution to the field regardless of the outcome. It will either help to validate the reliability of list experiments or demonstrate the extent to which the design is prone to implementation error, particularly among populations in the global south and among uneducated populations in particular. Further, the study outcomes will shed some insight into why these errors, if present, may occur. Doing so, it aims to help researchers employ the technique appropriately and identify solutions for improving the implementation of list experiments in the field.