arXiv · 2211.01221
Propensity score models are better when post-calibrated
Abstract
Theoretical guarantees for causal inference using propensity scores are partly based on the scores behaving like conditional probabilities. However, scores between zero and one, especially when outputted by flexible statistical estimators, do not necessarily behave like probabilities. We perform a simulation study to assess the error in estimating the average treatment effect before and after applying a simple and well-established post-processing method to calibrate the propensity scores. We find that post-calibration reduces the error in effect estimation for expressive uncalibrated statistical estimators, and that this improvement is not mediated by better balancing. The larger the initial lack of calibration, the larger the improvement in effect estimation, with the effect on already-calibrated estimators being very small. Given the improvement in effect estimation and that post-calibration is computationally cheap, we recommend it will be adopted when modelling propensity scores with expressive models.
Explore related subjects
Keep this discovery
Rom Gutman, Ehud Karavani, Yishai Shimoni. 2022-11-02. Propensity score models are better when post-calibrated. https://doi.org/10.1097/ede.0000000000001733
Cite the original work for its findings. Save a collection to share your selection of sources.