Topics in Econometrics - M2 ENS Lyon
2025-09-24
Introduction and Fundamentals
Simulations
Design: Identification and Beyond
Controls and Fixed Effects
IV and RDD
Modelling
Analysis
Communication
Controls: what do controls do to our regression?
Fixed effects are extremely common in applied economics
What are they really doing?
More generally, what are we really estimating in a specific model?
What are we comparing to what?
Where does the identifying variation come from?
\[Y = X\beta + W\delta + U\]
\[Y^{\perp W} = X^{\perp W}\tilde{\beta} + U^{\perp W}\]
where \(.^{\perp W}\) denotes each variable where \(W\) has been residualized
ie its projection onto the orthogonal space to W
Obtained using:
eg \(X^{\perp W} = M_W X\)
Fixed effects regression = regression on variables after partialling out the fixed effects
To compute the partialled out version of a regression:
Exercise. Using the data bellow, run two regressions and compare the estimates obtained:
l_murder on l_pris, male, population and incomel_murder and l_pris on one another (partialling out male, population and income)# Graph levels
guns |>
ggplot(aes(x = prisoners, y = murder)) +
geom_point() +
labs(
title = "Relationship between incarceration and murder rates",
subtitle = "Variables in level: need to transform it",
x = "Incarceration rate",
y = "Murder rate"
)
# Graph logs
guns |>
ggplot(aes(x = l_pris, y = l_murder)) +
geom_point() +
geom_smooth(method = "lm") +
labs(
title = "Relationship between incarceration and murder rates",
subtitle = "Log are better suited",
x = "Log of incarceration rate",
y = "Log of murder rate"
)#demeaning and showing that equal to residuals
guns_demean <- guns |>
mutate(
l_murder_res = lm(data = guns, formula = l_murder ~ male + population + income) |>
residuals(),
l_pris_res = lm(data = guns, formula = l_pris ~ male + population + income) |>
residuals()
)
reg <- guns |>
lm(formula = l_murder ~ l_pris + male + population + income) |>
broom::tidy() |>
mutate(reg = "raw", .before = 1)
reg_res <- guns_demean |>
lm(formula = l_murder_res ~ l_pris_res - 1) |>
broom::tidy() |>
mutate(reg = "residualized", .before = 1)
rbind(reg, reg_res) |>
filter(str_starts(term, "l_pris")) |>
kable()| reg | term | estimate | std.error | statistic | p.value |
|---|---|---|---|---|---|
| raw | l_pris | 0.8702664 | 0.0251012 | 34.67037 | 0 |
| residualized | l_pris_res | 0.8702664 | 0.0250583 | 34.72968 | 0 |
In economics, we are typically interested in one coefficient, that of the treatment
The true relationship of interest is the partialled out one
Can bring back a complex linear model to a bivariate regression
What matters is the variation in the treatment after partialling out
Where does the variation in the treatment comes from?
Interpretation of the coefficient of interest: “Comparing …”
The estimate of the treatment coefficient is in fact a weighted average of individual treatment effects \(\tau_i\)
Weight: \(w_{i} = (T_{i} - \mathbb{E}[T_{i} | X_{i}])^{2}\)
The weight represents:
How well the controls explain the treatment status
The conditional variance of the treatment, given \(X_i\)
Actually equivalent to leverage in the residualized regression
Observations whose treatment status is largely explained by covariates therefore contribute little, if at all, to estimation
For FE: if for some groups there is little within variation, these groups do not contribute to identification
Implications for external validity and representativity
May weigh more some characteristics
Implications for statistical power: the effective sample might be much smaller than the nominal sample
If need a CIA, weights are not equal
Figure from Aronow and Samii (2016)
Can distinguish controls between those that are identification-related and those that are not
Identification-related controls are those that are necessary for the CIA to hold
Other controls are here to improve precision of the estimator
Repeated observations over some dimension allow adjusting for all the unobserved characteristics that are constant across that dimension
Transform each variable into its deviation from the group mean
Only keep within variation (discards the between)
Two approaches to do that:
Basically build a counterfactual
Objective: estimate the impact of some treatment at a certain time
Leverages repeated observations, typically panel data
Builds a counterfactual that can be explicit or more implicit (eg TWFE):
Potentially, all units are treated
Assumed counterfactual: group’s past value
Within variation only
Flexible, allows looking at whether effects are dynamic
Difficult to rule out other things changing at the same time
\[Y_{it} = \sum_{t = -K}^{\tau - 2} [\beta_t \mathbb{1}\{t\}] + \beta_{\tau} \mathbb{1}\{\tau\} + \sum_{t = \tau + 1}^{L} [\beta_t \mathbb{1}\{t\}] + e_{it}\]
\[Y_{it} = \beta G_{i}P_t + \lambda_G + \lambda_P + e_{it}\]
Group FEs: compare individuals within the group
Time FEs: compare individuals within a time period
TWFEs:
Average of TEs identified from variation within group and variation within period
\(\neq\) variation within “that group that year” (this would be group-year FEs)
Including FEs changes the estimand: we compare observation within a group or within a time period
#demeaning and showing that equal to residuals
sample_demean <- guns |>
mutate(
l_murder_res = feols(data = guns, fml = l_murder ~ 1 | state) |>
residuals()
) |>
group_by(state) |>
mutate(mean_l_murder = mean(l_murder)) |>
ungroup() |>
mutate(
l_murder_demean = l_murder - mean_l_murder
) |>
select(l_murder_res, l_murder_demean) |>
head(10) | l_murder_res | l_murder_demean |
|---|---|
| 0.2824963 | 0.2824963 |
| 0.2170183 | 0.2170183 |
| 0.2094711 | 0.2094711 |
| 0.2094711 | 0.2094711 |
| 0.1057927 | 0.1057927 |
| -0.0098917 | -0.0098917 |
| -0.1515422 | -0.1515422 |
| -0.1300360 | -0.1300360 |
| -0.0883633 | -0.0883633 |
| -0.0582103 | -0.0582103 |
library(fixest)
#demeaning and showing that equal to residuals
guns_demean <- guns |>
mutate(
l_murder_res = feols(data = guns, fml = l_murder ~ 1 | state) |>
residuals(),
l_pris_res = feols(data = guns, fml = l_pris ~ 1 | state) |>
residuals()
)
reg_fe <- guns |>
fixest::feols(fml = l_murder ~ l_pris | state, cluster = "state") |>
broom::tidy() |>
mutate(reg = "fixed_effects", .before = 1)
reg_res <- guns_demean |>
feols(fml = l_murder_res ~ l_pris_res - 1, cluster = "state") |>
broom::tidy() |>
mutate(reg = "residualized", .before = 1)
rbind(reg_fe, reg_res) |>
kable()| reg | term | estimate | std.error | statistic | p.value |
|---|---|---|---|---|---|
| fixed_effects | l_pris | -0.15834 | 0.0365294 | -4.334587 | 7.05e-05 |
| residualized | l_pris_res | -0.15834 | 0.0365138 | -4.336438 | 7.01e-05 |
Let’s run some R code together to understand how fixed effects work
Let’s stil will use the guns dataset
Let’s consider several regressions, with various sets of fixed effects
Build graphs and interpret the coefficients
When adding FE (or controlling in general), we partial out or absorb some of the variation
We throw out variation
Good if throw out variation that:
Bad if throw out identifying variation, ie variation that allows you to identify the effect of interest
Each unit does not contribute equally: there are many implicit weights
At least two different objects can be weighted:
Outcomes: weighted difference in mean outcomes
Treatment effects: weighted mean of treatment effects
Often important when effects are heterogenous
Also important: which comparisons are used (ie who counts as “control”)
\[\hat{\beta} = \sum_i \hat{w}_i \, Y_i\]
Chattopadhyay and Zubizarreta (2023): characterizes these implied weights of linear regression (FE or not)
Need not be non-negative, nor look like the weights of a randomized experiment
\(\Rightarrow\) possible extrapolation: \(\hat{\beta}\) may not represent any real population
Useful as a diagnostic: inspect the weights before trusting \(\hat{\beta}\) (lmw package)
When treatment staggered, TWFE \(\hat{\beta}\) is weighted average of many pairwise DiDs (Goodman-Bacon 2021)
Each compares a newly treated group to a “control”:
Never-treated units, or
Not-yet-treated units
Issue: some “controls” are already-treated units
\[\hat{\beta}^{TWFE} = \sum_{g,t} w_{gt} \; ATT_{g,t}, \quad \sum_{g,t} w_{gt} = 1, \quad w_{gt} \ \text{can be} < 0\]
\(\Rightarrow\) \(\hat{\beta}^{TWFE}\) can have the wrong sign
A simple simulation easily illustrates this
Especially problematic when effects vary in time or are heterogeneous
New estimators reweight to recover an interpretable, convex average
Can anyone summarize the idea?
What do you get out of it?
What did you think of it?
TWFE \(\hat{\beta}\) = weighted sum of \(ATT\)s, weights can be negative
Arises in many standard staggered-adoption designs
Reviews for binary/discrete/continuous treatments, staggered or not:
Revisits Wolfers (2006) on divorce laws: conclusions change across estimators
So far, we considered very simple simulations, with “naive” distributions
Calibrating can help make simulations more realistic
But simulations will never be truly realistic
Yet can still allow to run some sort of robustness check on the ability of your design to retrieve the effects of interest
Also allows you to think about the DGP, your identification strategy, and so on
Read the literature
Get a sense of typical effect sizes and of relationships between variables
Make assumptions on those relationships. Acknowledge them.
Complexify later if needed. You choose when you stop.
Varying parameters values might change the distribution of some “variables”
eg of the error term
Difficult to work ceteris paribus
Start from an existing data set
Not yours. At least not the subset you are interested in
Try to pick a subset where there is not already a treatment effect
Define a treatment allocation mechanism
Add an artificial treatment effect to the outcome variable in your initial data set, eg
\[Y_i(1) = Y_i(0) + \beta_i T_i\]
There is only one artificial aspect in such simulations: the treatment
We can play on only 2 components:
Who is treated? Treatment allocation
Everyone
Only a subset of the population
How? Treatment effect
Homogenous
Heterogenous but random
Some specific correlation structure
Today we reviewed:
Hopefully you have a better understanding of:
The choice of FE is crucial and affects the estimand
FE can remove a lot of variation:
Regressions (with or without FE) implicitly weight units
With heterogeneous or dynamic effects, plain TWFE can be misleading