Topics in Econometrics - M2 ENS Lyon
2026-09-23
Anyone wants to share theirs?
What is it about?
What did you particularly like about this paper?
Anything else worth sharing with the others?
Design: decisions of data collection and measurement
Modeling: define statistical models
Analysis: estimation and questions of statistical inference
Typically in economics, we are interested in estimating causal effects
In (non-experimental) economics, design presented in this lexicographic order:
Identification
Unbiasedness
Minimum variance
Robustness to misspecification somewhere in the mix
Design includes identification but not only
These steps interact with one another
It describes how to use observational data to recover the effect of interest
It defines a source of identifying variation in the treatment
It specifies the assumptions under which this variation identifies the causal effect
Types of experiments:
A true experiment: treatment is randomly assigned by the researcher
A natural experiment: treatment is plausibly randomly assigned by “nature”
A quasi-experiment: treatment is close-to (but not) randomly assigned
| Individual Treatment Effects (TEs) | \(Y_i^1-Y_i^0, \forall i\) | What we would ideally estimate |
| Average Treatment Effects (ATE) | \(\mathbb{E}[Y_i^1-Y_i^0]\) | What we reasonably want to estimate |
| Average Treatment Effects on the Treated (ATT) | \(\mathbb{E}[Y_i^1-Y_i^0 \vert D_i = 1]\) | What we reasonably want to estimate |
| Conditional Average Treatment Effects (CATE) | \(\mathbb{E}[Y_i^1-Y_i^0 \vert X_i]\) | What we reasonably want to estimate |
| Difference in average observed outcomes | \(\mathbb{E}[Y_i \vert D_i = 1, X_i] - \mathbb{E}[Y_i \vert D_i = 0, X_i]\) | What we can estimate |
Stable unit treatment value assumption (SUTVA):
Each unit has only 2 potential outcomes: \(Y_i^0, Y_i^1\)
Assumes no spillover effects
Assumes no general equilibrium effects
Often not realistic in economics
\(\underbrace{\mathbb{E}[Y_i | D_i = 1, X_i] - \mathbb{E}[Y_i | D_i = 0, X_i]}_{\text{Difference in average observed outcomes}} = \\ \qquad \underbrace{\mathbb{E}[Y_i^1-Y_i^0 \vert D_i = 1, X_i]}_{CATT} + \underbrace{\mathbb{E}[Y_i^0 \vert D_i = 1, X_i] - \mathbb{E}[Y_i^0 \vert D_i = 0, X_i]}_{\text{Selection Bias}}\)
Goal: eliminate this selection bias to be able to say something about the quantity of interest (the CATT)
Selection bias: average difference in \(Y_i^0\) between the treated and untreated
Assumptions regarding the assignment mechanism (identifying assumptions) can help eliminate it
Random assignment (eg experiments)
Treatment independent of potential outcomes \(\Rightarrow\) no selection bias in expectation
It is the Independence Assumption (IA): \((Y_i^0, Y_i^1) \perp D_i\)
Selection on observables
Random assignment conditional on some pre-treatment characteristic \(X\)
It is the Conditional Independence Assumption (CIA): \((Y_0, Y_1) \perp D_i | X_i\)
Compare outcomes within each stratum of \(X_i\)
Selection on unobservables
Goal: identifying causal effects
ie a difference between two potential outcomes
But, we cannot observe the potential outcomes
We only see the differences in observed outcomes
If (C)IA holds, we can estimate an unbiased ATT
But (C)IA rarely holds \(\Rightarrow\) need an identification strategy to elimate selection bias
Randomized experiments (RCT)
Difference-in-differences (DiD), event studies, synthetic control methods (SCM)
Instrumental variables (IV) or regression discontinuity (RD)
Matching estimators:
Can anyone summarize the idea?
What do you get out of it?
What did you think of it?
Randomness comes from treatment assignment
Identification rest on how this treatment was assigned
Credibility comes from knowledge of the assignment process
Other identification “types”:
We want to have an accurate measure of the quantity of interest
\(\to\) research design:
Data collection and measurement choices to allow us to draw credible conclusions
Includes the causal identification strategy, but not only
A great strategy useless if a poor design prevents us detecting the effect
For instance if statistical power is low
Power is a key implication of design choices
Definition:
\[\text{Power} = 1 - \text{rate of Type II error}\]
Power is a function of design: poor designs can lead to low statistical power
We want to be able to detect an effect if there is one (that is large enough to be relevant)
Because costly to run a study for “nothing”
In RCT, typical threshold for power: 80%
In observational settings, why not run a study with say 20% power?
Because low statistical power \(\Rightarrow\) exaggeration
Definition: \[E = \dfrac{\mathbb{E} [ | \hat{\beta} | | \text{ signif } ]}{| \beta_1 |} = \dfrac{\mathbb{E} [ | \hat{\beta} | | \beta_{1}, \sigma, | \hat{\beta} | > z_{\alpha} \sigma ]}{| \beta_1 |}\]
Exaggeration \(\searrow\) with statistical power and thus:
When power is low, significant estimates from an unbiased estimator ALWAYS exaggerate the true effect
There are also less straightforward drivers (we are going to discuss them later today)
Significant results are favored
Evidence of a significance filter in economics
(Rosenthal 1979; Andrews and Kasy 2019; Abadie 2020; Brodeur et al. 2016; Brodeur, Cook, and Heyes 2020)
Low statistical power
Median power in economics: 18%
(Ioannidis, Stanley, and Doucouliagos 2017; Ferraro and Shukla 2020)
Editorial process favors significant results for publication
In a way, that makes sense if a non-significant result reflects a poor research question \(\Rightarrow\) importance of theory
But, might also be that the effect is difficult to capture
File drawer problem: tend to give up projects more when results are non-significant (put them away in a drawer)
Forking paths: we make many choices when implementing a study and they may be more likely to lead to a significant outcome
In economics, nearly 80% of estimates are exaggerated by a factor of 2 (Ioannidis, Stanley, and Doucouliagos 2017)
Not all designs suffer from exaggeration
But exaggeration is likely substantial in many studies
Have a large enough sample size and we’re good?
Not so simple!
Other aspects than sample size affect power (and interact multiplicatively):
Often, goal of an econometrics study: estimate the ATE (Does the treatment work?)
But also, where and when does it work?:
Capture heterogeneity: treatment effect varies across time and individuals
Often consider the effect on multiple outcomes
Extrapolate
They have intertwined implications for how we approach design
Not possible to have high power for everything
Goals can be competing
Can take action at the design stage, acknowledging these multiple goals
Treatment effect rarely homogeneous
The phrase “Average Treatment Effect” implicitly acknowledges this
Variation across individuals, time, space, etc
There are therefore potential confounders:
Need to adjust for such variables
Measure them
An usual approach to account for heterogeneity is to use interactions
To measure interactions, we need 16 times the sample size:
The estimates has twice the s.e. of the main effect
Reasonable to assume that interaction have half the magnitude of the main effect
Thus Signal to Noise Ratio \(\left( SNR = \frac{\text{True effect}}{\text{s.e.}} \right)\) is 4 times smaller for interaction
Thus need \(4^2 = 16\) times the sample size
Issues:
When treatment effect heterogenous (in time or across groups)
Treated units in the control group
Negative weights
The literature addressed it as a analysis problem: proposed alternative estimators
But can see it as non-modeled heterogeneity
Rough approximation of the median number of estimates per paper in top economics journals: 19
Bonferroni correction:
Underlines that need more power \(\Rightarrow\) need to take that into account
External validity
When increase the sample size, often changes the underlying estimand
eg, increasing sample size by increasing the time frame
or the spatial frame
Increasing sample size not always a silver bullet
Without publication bias this issue disappears
Abandoning the 5% significance threshold
Interpretation of CI’s width to embrace uncertainty
Replication of studies with similar designs
Increased sample size
Increased effect size
Focus on units with the largest effect
Increase take-up of the treatment
Decreased inferential uncertainty
More pre-treatment information
Better measurement of outcomes
Weave empirical models with substantive theory
Adjust the research question
Measure intermediate outcomes
Use simple design calculations
Simulations (you now got that hopefully)
Retrodesign calculations
Goal: choose a design that would yield an adequate statistical power
Compute the expected power, in this setting, as a function of design and in particular sample size
Find the necessary sample size
Before implementing the analysis
Common practice in experimental economics, much less in observational settings
Statistical power is a function of true effect size and s.e. of the estimator
Strictly increasing with true effect sizes
Strictly decreasing with s.e. of the estimator
Slightly complex closed form
Need to hypothesize a s.e. and a true effect size
s.e. unknown before the analysis
Basically boil the analysis down to a difference of average outcome between treatment and controls
\[se_{\bar{y_t} - \bar{y_c}} = \sqrt{\dfrac{\sigma_T^2}{n_T} + \dfrac{\sigma_C^2}{n_C}}\]
\(\sigma_T^2\) and \(\sigma_C^2\) variance of the outcome for the treatment and control group respectively (after partialing out controls)
Assuming \(\sigma_T^2 = \sigma_C^2 = \sigma^2\) and for \(p_T = \frac{n_T}{n}\), this simplifies to \(se_{\bar{y_t} - \bar{y_c}} = \frac{\sigma}{\sqrt{n}}\sqrt{\frac{1}{p_T(1-p_T)}}\)
Consider the proportion of affected individuals
Consider a range of effects (make several assumptions)
Derived from the literature
Based on theory
Consider what could be reasonable deviations from these effects
Multiply the fraction of non-zero effect with the hypothesize effects
Once an estimate has been obtained
Ask the question would my design allow me to detect a smaller effect (of magnitude \(m\))?
Need the standard error of your estimate and an hypothetical true effect size ( \(m\) )
One line of r code: retrodesign::retrodesign(m, se)
Run it for a range of values
Design matters:
In terms of identification
And beyond
Even after a significant estimate has been obtained
When power is low, significant estimates from an unbiased estimator are always far from the true effect
Might have important implications for policy making