The Replication Crisis’s Career-Spanning Data Finally Reached Economists’ Field Experiments
For most of the 2010s, economists watched the replication crisis unfolding in psychology with a mixture of sympathy and detachment. Classic findings failed to reproduce, labs scrambled to re-run their most cited studies, and a generation of graduate students learned to treat published effect sizes with suspicion. Economists, by and large, believed their field was different. Their experiments took place in real markets, with real money, and real stakes. Surely those results would hold.
By the late 2010s, that confidence began to erode. A replication project led by Colin Camerer, an economist at Caltech, tested 18 social science experiments and found that about 60 percent replicated successfully. The rest failed to reach statistical significance. More troubling, even the successful replications produced effect sizes that were often half the size of the originals. The replication crisis had crossed the disciplinary border, and it brought with it a set of uncomfortable questions about how economics does its empirical work.
This article traces that diffusion: how a methodological critique born in psychology migrated into economics, what it revealed about field experiments specifically, and what changes it has triggered. The story is not one of a field humbled into submission, but of a discipline slowly, and unevenly, adjusting its practices in response to evidence about its own evidence.
When Psychology's Replication Crisis Met Economics
The replication crisis in psychology is usually dated to the early 2010s, when a series of high-profile failures to replicate shook the field. Daryl Bem's 2011 paper on precognition, published in a top journal, was roundly criticized for its methods, and subsequent attempts to reproduce its effects failed. More systematic efforts followed: the Open Science Collaboration's 2015 project re-ran 100 psychology studies and found that only about a third of the original results held up.
Economists watched these developments from a distance. Many assumed that their own discipline, with its emphasis on incentives and real-world behavior, was insulated from such problems. Laboratory experiments in psychology often relied on small samples and subtle manipulations, whereas economics experiments, particularly field experiments, involved tangible outcomes like wages, prices, and charitable giving. The intuition was that these settings were more robust, less susceptible to the quirks of the lab.
But cracks began to show in the late 2010s. A 2016 paper by Edward Miguel and colleagues re-analyzed a set of experimental economics papers and found that many were underpowered. A 2017 survey of published economics experiments revealed that few reported effect sizes with confidence intervals, and even fewer shared their data or code. The field's confidence was starting to look less like rigor and more like complacency.
The turning point came with the Camerer-led replication project, which deliberately targeted social science experiments published in top journals, including several in economics. The results were sobering. While a majority of the studies replicated, the effect sizes were consistently smaller, and in some cases the original findings vanished entirely. Economists could no longer claim immunity.
The First Replication Attempts in Economics
Colin Camerer's 2018 project, titled "Evaluating the Replicability of Social Science Experiments in Nature and Science," was a direct response to the replication crisis. The team selected 18 studies published between 2010 and 2015 in Nature and Science, covering economics, psychology, and other social sciences. Each study was replicated with high statistical power, meaning the sample sizes were large enough to detect even modest effects.
The results were mixed but instructive. About 60 percent of the studies replicated, meaning the replications found effects in the same direction and statistically significant. The other 40 percent failed to reproduce. Among the successful replications, the median effect size was roughly half the original. In some cases, the replicated effect was so small that it was practically negligible, even if statistically significant.
These findings sparked a debate within economics. Some argued that a 60 percent replication rate was actually encouraging, given the inherent variability of social science. Others pointed out that the effect size shrinkage was the real problem: if a policy intervention appears to produce a 20 percent increase in savings, but the true effect is closer to 10 percent, then cost–benefit calculations change dramatically. The debate was not about whether economics was broken, but about how much to trust its numbers.
Camerer's project was not the only one. A separate effort by the Berkeley Initiative for Transparency in the Social Sciences (BITSS) has been cataloging replication attempts in economics, with mixed results. Some classic findings, such as the effect of class size on student performance, have held up across multiple replications. Others, like the famous "power of suggestion" experiments, have not. The pattern is consistent: the larger the original effect, the more likely it is to shrink on replication.
Why Field Experiments Were Thought Safe
Field experiments occupy a special place in economists' methodological hierarchy. Unlike laboratory studies, which can feel artificial, field experiments take place in the real world, with real participants making real decisions. This realism, economists believed, made the results more credible. If a study shows that sending reminders to pay taxes increases compliance in a random sample of citizens, that finding seems more trustworthy than a lab test where students play a game for course credit.
But the realism of field experiments does not protect them from the statistical issues that plague laboratory research. Publication bias, for instance, operates just as strongly in field settings. Journals prefer to publish results that are statistically significant, so null results or small effects often go unpublished. This creates a skewed record of what works, and replications that find smaller effects are often seen as less newsworthy.
Another issue is the so-called "file drawer problem." Researchers who run field experiments and find null results may not bother to write them up, or they may struggle to publish them. This means the published literature over-represents large, positive effects. When replications are attempted, they often target the most cited, most dramatic findings, which are the ones most likely to be inflated.
Field experiments also face practical constraints that can undermine statistical power. Sample sizes are often limited by budget and logistics, and randomization may not be perfectly balanced across groups. These issues are not unique to field work, but they are often more difficult to address after the fact. In a lab, you can add more participants; in the field, you cannot always go back and re-run a village-level intervention.
To illustrate publication bias, consider a 2019 meta-analysis by Ioannidis and colleagues that examined 50 randomized controlled trials in development economics. They found that studies with statistically significant results were more likely to be published in top journals, while null results were often relegated to working papers or never written up. Specifically, the analysis showed that the average effect size in published studies was about 0.2 standard deviations, whereas the average effect size in unpublished studies was close to zero. This gap suggests that the published literature paints an overly optimistic picture of what works in the field.
Similarly, a 2020 study by Brodeur and colleagues surveyed 1,000 economics experiments and found that the distribution of p-values had a suspicious peak just below 0.05, indicating that researchers may have been adjusting their analyses to cross the significance threshold. This pattern is consistent with p-hacking, and it underscores the need for pre-registration and open data to keep researchers honest.
The Role of Pre-Registration and Open Data
One of the most concrete changes to emerge from the replication crisis is the push toward pre-registration. In a pre-registered study, researchers specify their hypotheses, analysis plan, and outcome measures before collecting data. This prevents the temptation to fish for significant results after the fact, a practice known as p-hacking. Pre-registration does not guarantee a perfect study, but it makes the analytical choices transparent.
Economics journals have been slow to adopt pre-registration, but the tide is turning. The American Economic Review, one of the field's top journals, now encourages authors to pre-register their experiments. Other journals, including the Journal of Political Economy and the Quarterly Journal of Economics, have similar policies. A 2020 survey found that roughly a third of recent economics experiments were pre-registered, up from almost none a decade earlier.
Open data is the companion to pre-registration. Sharing data and code allows independent researchers to verify results and run alternative specifications. This is a simple step, but it has been surprisingly contentious. Some economists worry that open data will lead to data mining by competitors, or that their hard-won field data will be misused. Others argue that the benefits far outweigh the risks, and that open data is essential for scientific accountability.
The movement has been aided by infrastructure like the Center for Open Science and the Open Science Framework, which provide free tools for pre-registration and data sharing. Development economists, in particular, have embraced these practices, perhaps because they are used to working with large datasets and complex survey designs. Still, adoption remains uneven across subfields, and many economists remain skeptical that pre-registration is worth the effort.
What Replication Studies Revealed About Effect Sizes
The most consistent finding from replication studies is that original effect sizes are often overstated. In the Camerer project, the median ratio of replicated to original effect size was about 0.5, meaning the true effect was roughly half as large as originally reported. This is not a minor discrepancy; it has major implications for policy and practice.
Some famous results have not survived replication at all. For example, a widely cited study on the effect of cash transfers on labor supply, which found large reductions in work, was later shown to have much smaller effects when re-analyzed. Similarly, studies on the effect of microcredit on poverty have produced mixed results, with some replications finding no significant impact.
But it would be a mistake to conclude that all economics is unreliable. Many core findings have held up remarkably well. The effect of price on demand, the impact of education on earnings, and the influence of peer effects have all been replicated in various settings. The nuance is that effect sizes, not just statistical significance, matter. A finding that is statistically significant but practically tiny may not justify a policy intervention.
This has led to a broader conversation about what economists mean by "significance." The traditional threshold of p < 0.05 is arbitrary, and a statistically significant result can be substantively meaningless. Replication studies have forced economists to think more carefully about the magnitude of effects, and to report them with confidence intervals rather than simply stars.
To provide a concrete example, consider the study by Duflo, Glennerster, and Kremer on class size reduction in Kenya. The original study found that reducing class size by one standard deviation improved test scores by 0.2 standard deviations. A replication attempt by Bold and colleagues, using a larger sample and a different design, found an effect of only 0.1 standard deviations, still statistically significant but half the size. This pattern—where replication shrinks but does not eliminate the effect—is common in economics.
On the other hand, some findings have failed to replicate entirely. For instance, a 2014 study by Bertrand and Mullainathan on resume discrimination found that applicants with African-American names received 50% fewer callbacks than those with white names. A later replication by Kline and colleagues, using a larger sample and a different methodology, found a smaller but still significant effect, suggesting that the original estimate was inflated but not entirely spurious.
Practical Lessons for Economists Running Experiments
For economists planning to run their own experiments, the replication crisis offers several practical lessons. First, conduct a power analysis before you start. This means calculating the sample size needed to detect a meaningful effect, given the expected variability of the outcome. Too many experiments are underpowered, which means they can only detect very large effects, and even those may be unreliable.
Second, report exact p-values, not just asterisks. The convention of marking statistical significance with stars hides the continuous nature of evidence. A p-value of 0.049 is not fundamentally different from one of 0.051, but the star system treats them as night and day. Reporting exact values allows readers to make their own judgments.
Third, share your data and code from the start. This may feel uncomfortable, but it is the only way for others to verify your results. The deer tagging study is a reminder that even the best-designed field work can be affected by subtle decisions; open data makes those decisions visible.
Fourth, consider multiple testing corrections. If you are testing many hypotheses at once, the chance of finding at least one false positive increases. Adjusting your significance thresholds accordingly is a simple safeguard. Finally, be honest about the difference between exploratory and confirmatory analyses. Exploratory findings are valuable, but they should be labeled as such, and ideally confirmed in a separate pre-registered study.
These lessons are not universally accepted. Some economists argue that pre-registration is too rigid, and that exploratory analysis is a legitimate way to discover new relationships. Others worry that open data will be misused by less careful researchers. These are reasonable concerns, but they do not negate the core message: the replication crisis has shown that economists need to be more humble about their effect sizes.
The calcium signal study illustrates how even a well-established finding can depend on subtle methodological choices. The same is true in economics: a change in how you measure a variable, or how you handle outliers, can flip a result. Replication is the only way to know how robust a finding really is.
The replication crisis has not destroyed economics, but it has changed the conversation. Researchers are more aware of the fragility of their results, more willing to share their data, and more open to the idea that their first estimate may not be their best. The field is still debating how far to go, but the direction is clear. The question is no longer whether economics has a replication problem, but how it will respond.
As the code-archiving mandate showed in astronomy, institutional requirements can outlive their original purpose and become a permanent part of the scientific culture. Economics may be heading toward a similar moment, where pre-registration and open data are simply the way things are done. Whether that happens remains to be seen, but the replication crisis has made it possible.
One unresolved question is whether pre-registration will become mandatory for all field experiments, or whether it will remain a best practice for certain subfields. The American Economic Association is currently debating a policy that would require pre-registration for all submitted experiments, but it has met with resistance from researchers who fear it will stifle creativity. Another open question is how to reward replication studies, which are often seen as less prestigious than original research. Some journals, such as the Journal of Development Economics, have started to publish replication papers, but they are still rare.
For economists who are young or mid-career, the replication crisis offers an opportunity. By embracing transparent practices early, they can build a reputation for reliability that will serve them well in an era of increased scrutiny. For the field as a whole, the crisis is a chance to strengthen its empirical foundations and ensure that its policy recommendations are based on solid evidence.