Design: Identification and Beyond

Design is crucial to ensure causality but also matters beyond identification. Here we discuss both aspects

Date

September 23, 2026

Objective

This session aims to underline the importance of design, in terms of identification but also beyond identifications.

Summary

Design is central to applied economics research, first and foremost through identification: we review the potential outcomes framework, the assignment mechanisms that can eliminate selection bias, and the resulting identification strategies.

But, in particular due to statistical power and exaggeration questions, design matters beyond identification. Low power leads significant estimates from an unbiased estimator to exaggerate the true effect.

Studies also usually pursue multiple goals beyond the average treatment effect (capturing heterogeneity, effects on multiple outcomes, and extrapolating to other populations) which have implications for the choices made at the design stage.

Session Outline

  1. Definition and importance of design
  2. Design and Identification
  3. Design Beyond Identification: Low Power
  4. The multiple goals pursued in studies retroactively motivate design choices
  5. Improving and Assessing Design
  6. Calibrating Simulations

Materials

Open slides in html

Open slides in pdf

Specific resources for this lecture

If you should read only one thing (ok, two)

Chapter 2 of Angrist and Pischke (2009) and Gelman and Carlin (2014)

Identification

Publication Bias in Economics

  • Doucouliagos and Stanley (2013) 60% of research areas in economics feature substantial publication bias (strongest when dominant theory stronger because difficult to defend results that go against it)
  • Brodeur et al. (2016) document publication bias in top econ journals (+ show that it comes more from the author’s side)
  • Vivalt (2019) studies the extent of p-hacking in impact evaluations (but decrease over time for RCTs)
  • Andrews and Kasy (2019): provides a publication bias correction based on the probability of publication conditional on result + method to identify this probability
  • Brodeur, Cook, and Heyes (2020) compare to what extent different causal identification strategies suffer from publication bias. IV (and DiD) suffer more than RCT and RDD
  • Chopra et al. (2024) using an experiment with researchers as subject, show that “studies with a null result are perceived to be less publishable, of lower quality and of lower importance”
  • Brodeur et al. (2023): issues of marginal significance (publication bias) come more from authors’ behavior than from the peer review process
  • Table 2 in Christensen and Miguel (2018) summarizes this literature

Evidence of low statistical power (and exaggeration)

In Economics

  • Ioannidis, Stanley, and Doucouliagos (2017) use meta-analyses to compute the statistical power of the studies “contained” in these meta-analyses
  • Ferraro and Shukla (2020) use the same techniques as Ioannidis et al (2017) to show that there are power issues in environmental economics
  • Ferraro and Shukla (2023) same in agricultural economics
  • DellaVigna and Linos (2022): shows that academic papers studying nudges find effects that are much larger than in large nudge experiment ran by nudges companies. Explain this with low power
  • Black et al. (2022): show the importance of taking power into account and show how to implement power calculations
  • Young (2022) documents a lack of power of IVs in economics (among other things)
  • In a non-directly related context, Roth (2022) underlines that a lack of power of pre-trend tests in event-study designs can lead to bias on the main estimate

In Political Science

  • Arel-Bundock et al. (2024) documents a lack of power in political sciences (median power 10% and only 1 in 10 tests have 80% power to detect the consencessus effects reported in the literature)
  • Lal et al. (2024) documents a lack of power of IVs in political science (among other things)
  • Stommes, Aronow, and Sävje (2023) shows in RD in political sciences are under-powered to detect anything but large effects and lead to exaggeration

Mechanisms behind exaggeration

  • Gelman and Tuerlinckx (2000) and Gelman and Carlin (2014) introduce the concept of Type-M error (exaggeration)
  • Lu, Qiu, and Deng (2019) and van Zwet and Cator (2021) derive mathematical proof of the evolution of exaggeration with effect size and precision of the estimator

Comparison IV and OLS

  • Young (2022) replicate 30 papers from the economics literature (AEA journals). Find that:
    • 75% of the 2SLS 95% CI contain the corresponding OLS point estimates (67.3% of main results)
    • IV estimates often larger (in absolute terms) or opposite sign than OLS: “greater than 0.5 times the absolute value of the OLS point estimate in .73 of headline regressions”
    • 2SLS estimates usually do not provide meaningful information regarding the extent to which OLS are biased
  • Lal et al. (2024) replicate 67 papers from the political science literature. Find that:
    • For 97% of designs studied 2SLS > OLS (34% at least 5 times larger)

References

Andrews, Isaiah, and Maximilian Kasy. 2019. “Identification of and Correction for Publication Bias.” American Economic Review 109 (8): 2766–94. https://doi.org/10.1257/aer.20180310.
Angrist, Joshua D., and Jörn-Steffen Pischke. 2009. Mostly Harmless Econometrics: An Empiricist’s Companion. 1 edition. Princeton: Princeton University Press.
Arel-Bundock, Vincent, Ryan C. Briggs, Hristos Doucouliagos, Marco M. Aviña, and T. D. Stanley. 2024. “Quantitative Political Science Research Is Greatly Underpowered.” The Journal of Politics 88 (1): 36–46. https://doi.org/10.1086/734279.
Black, Bernard, Alex Hollingsworth, Letícia Nunes, and Kosali Simon. 2022. “Simulated Power Analyses for Observational Studies: An Application to the Affordable Care Act Medicaid Expansion.” Journal of Public Economics 213 (September): 104713. https://doi.org/10.1016/j.jpubeco.2022.104713.
Brodeur, Abel, Scott Carrell, David Figlio, and Lester Lusher. 2023. “Unpacking P-hacking and Publication Bias.” American Economic Review 113 (11): 2974–3002. https://doi.org/10.1257/aer.20210795.
Brodeur, Abel, Nikolai Cook, and Anthony Heyes. 2020. “Methods Matter: P-Hacking and Publication Bias in Causal Analysis in Economics.” American Economic Review 110 (11): 3634–60. https://doi.org/10.1257/aer.20190687.
Brodeur, Abel, Mathias Lé, Marc Sangnier, and Yanos Zylberberg. 2016. “Star Wars: The Empirics Strike Back.” American Economic Journal: Applied Economics 8 (1): 1–32. https://doi.org/10.1257/app.20150044.
Chopra, Felix, Ingar Haaland, Christopher Roth, and Andreas Stegmann. 2024. “The Null Result Penalty.” The Economic Journal 134 (657): 193–219. https://doi.org/10.1093/ej/uead060.
Christensen, Garret, and Edward Miguel. 2018. “Transparency, Reproducibility, and the Credibility of Economics Research.” Journal of Economic Literature 56 (3): 920–80. https://doi.org/10.1257/jel.20171350.
DellaVigna, Stefano, and Elizabeth Linos. 2022. “RCTs to Scale: Comprehensive Evidence From Two Nudge Units.” Econometrica 90 (1): 81–116. https://doi.org/10.3982/ECTA18709.
Doucouliagos, Chris, and T.d. Stanley. 2013. “Are All Economic Facts Greatly Exaggerated? Theory Competition and Selectivity.” Journal of Economic Surveys 27 (2): 316–39. https://doi.org/10.1111/j.1467-6419.2011.00706.x.
Ferraro, Paul J., and Pallavi Shukla. 2020. “Feature—Is a Replicability Crisis on the Horizon for Environmental and Resource Economics?” Review of Environmental Economics and Policy 14 (2): 339–51. https://doi.org/10.1093/reep/reaa011.
———. 2023. “Credibility Crisis in Agricultural Economics.” Applied Economic Perspectives and Policy 45 (3): 1275–91. https://doi.org/10.1002/aepp.13323.
Gelman, Andrew, and John Carlin. 2014. “Beyond Power Calculations: Assessing Type S (Sign) and Type M (Magnitude) Errors.” Perspectives on Psychological Science 9 (6): 641–51. https://doi.org/10.1177/1745691614551642.
Gelman, Andrew, and Francis Tuerlinckx. 2000. “Type S Error Rates for Classical and Bayesian Single and Multiple Comparison Procedures.” Computational Statistics 15 (3): 373–90. https://doi.org/10.1007/s001800000040.
Ioannidis, John P. A., T. D. Stanley, and Hristos Doucouliagos. 2017. “The Power of Bias in Economics Research.” The Economic Journal 127 (605): F236–65. https://doi.org/10.1111/ecoj.12461.
Lal, Apoorva, Mackenzie Lockhart, Yiqing Xu, and Ziwen Zu. 2024. “How Much Should We Trust Instrumental Variable Estimates in Political Science? Practical Advice Based on 67 Replicated Studies.” Political Analysis, May, 1–20. https://doi.org/10.1017/pan.2024.2.
Lu, Jiannan, Yixuan Qiu, and Alex Deng. 2019. “A Note on Type S/M Errors in Hypothesis Testing.” British Journal of Mathematical and Statistical Psychology 72 (1): 1–17. https://doi.org/10.1111/bmsp.12132.
Roth, Jonathan. 2022. “Pre-Test with Caution: Event-Study Estimates After Testing for Parallel Trends.” American Economic Review: Insights. https://doi.org/10.1257/aeri.20210236.
Stommes, Drew, P. M. Aronow, and Fredrik Sävje. 2023. “On the Reliability of Published Findings Using the Regression Discontinuity Design in Political Science.” Research & Politics 10 (2). https://doi.org/10.1177/20531680231166457.
van Zwet, Erik, and Eric Cator. 2021. “The Significance Filter, the Winner’s Curse and the Need to Shrink.” Statistica Neerlandica 75 (4): 437–52. https://doi.org/10.1111/stan.12241.
Vivalt, Eva. 2019. “Specification Searching and Significance Inflation Across Time, Methods and Disciplines.” Oxford Bulletin of Economics and Statistics 81 (4): 797–816. https://doi.org/10.1111/obes.12289.
Young, Alwyn. 2022. “Consistency Without Inference: Instrumental Variables in Practical Application.” European Economic Review 147 (August): 104112. https://doi.org/10.1016/j.euroecorev.2022.104112.