Topics in Econometrics - M2 ENS Lyon
2026-09-22
Applied economics aim to produce accurate causal estimates (eg inform public policy)
Whole game in our metrics analyses: approximate the DGP
Objective of the class: discuss practical issues that may prevent us from doing so
Can arise in any of the steps of research: design, modeling and analysis
There are some fundamental hurdles to estimating causal effects
Simulations can help spot and understand these hurdles
There is some chronology in hurdles: each only matters to the extent that the previous one are addressed
We need to have, in that order of concern:
A good research question, grounded in theory
A good design: enough data, well measured, etc
A good identification strategy to avoid some fundamental hurdles (reverse causality, confounders, etc)
An adequate specification that allows us to estimate the quantity we want to estimate
Make reliable inference based on well-constructed standard errors for instance
Reliable communication: make sure the audience receive what we intend to share
What was the idea behind the implementation of simulations last week?
Objective: explore how several parameters affect the estimate of interest
Approach: Generate fake data (we thus know the whole DGP) and run an analysis
Were they useful? If so, in what way?
Understand how various parameters affect the estimate of interest, without deriving the maths
Help to shape intuition and understanding
A simulation is a process in which we:
Generate artificial data (and thus know the DGP)
From scratch (fake data simulation) or
On top of an existing data set (real data simulation)
Run an analysis on this data
Repeat the process many times
Allows us to assess the performance of our analysis:
Can we accurately estimate the true effect of interest?
Are there hurdles to doing so and can we overcome them?
To understand econometric concepts
To design a study, before having the data
To design a study, once having the data
Tests and checks, after running the analysis
As a rhetorical tool
No maths required and allows to consider many general cases easily
Useful to get intuition on how econometric aspects work
Understand general concepts:
Understand conceptual hurdles specific to our context:
What did you learn from the exercise you had to do?
What affects leverage? How does it affect the parameter of interest?
Present the intuition behind leverage and influence
How did you implement your simulation?
Any cool graphs/outputs?
Useful to get started on a concrete reflection about:
The setting
What we want to estimate, exactly
The data needed and its granularity
The identification strategy
As a proof of concept (to apply for grants, data access, etc)
Useful to think about:
Threats to identification and important assumptions
The statistical power of our study (difficult to do without a simulation)
Explore where to best invest resources:
Larger sample
Improved data precision (reduce measurement error)
Does our analysis detects the effect we are interested in, in a pristine setting?
If the analysis faces issues in simulations, it will probably also in an actual setting
What happens to the product of our analysis if the setting is slightly more complex?
What happens if some hypotheses do not hold?
All this can be discussed even after the analysis has been run
Simplify what we are working on to the bare minimum
What is the simplest way of pitching the analysis and the identification strategy?
Can help build simple visualizations
Can be useful to illustrate why a given approach does not work
To argue why we chose a certain approach
In a referee report
Start with a simple DGP:
Simple correlation structure
Our model represents the actual DGP
Does our analysis recover the effect in a rather “pristine” setting?
Then complexify the DGP
Define a DGP and the distribution of variables
Set parameters values
Generate a data set
Estimate the effect in the generated data set
Repeat many times
Compute the measure of interest
Change parameters values
Understand how the measure of interest is affected by a given parameter
eg how does bias evolve with the correlation between \(x_1\) and \(x2\)?
Complexify the DGP
Repeat
Impact of receiving extra lessons on students’ grades
Simulate an experiment (RCT):
\(\forall i \in \{1, .., n\}, \quad Grade_i = \alpha_0 + \beta_0 Treat_i + u_i\)
Which sample size and proportion of treated to have a high probability of detecting the effect?
Simulate many experiments
Compute the proportion of effects detected
Instructions here
A few random take-away points
Points can be thought of as linked to the regression line through a rope: outliers can pull the regression line
Measurement error in \(x\) creates an attenuation bias
Bad controls introduce bias
If we cannot retrieve the effect in a simulation, there are no chances we will recover it in an actual setting
Simulations can be helpful just to make us think about our estimand, the type of data we need, the relationships between the variables, effect sizes, etc
To build simulations, always start simple, complexify later