SearcharxivSearch

arXiv subjects

Simon Freyaldenhoven

Publications and source records attributed to Simon Freyaldenhoven.

2 recordsLinked to original sources

When Can We Work in Embedding Space? What Text Embeddings Preserve

When do text embeddings work as inputs to empirical analysis? Their use rests on an assumption: that we can trade text for its low-dimensional embedding, and lose little in doing so. I make that assumption precise under a generative model in which documents are mixtures of latent topics. I study two uses---clustering units in embedding space and controlling for high-dimensional text. A cluster of embeddings is a set of documents with similar topic mixtures; controlling for the embedding is equivalent to controlling for the topic mixture, so validity reduces to whether that mixture captures the confounding. In an application to 363 U.S. metropolitan areas, embedding-based clusters of LLM-generated economic descriptions recover interpretable economic archetypes and separate local employment dynamics more sharply than clustering on model residuals, or on a curated set of industry and demographic covariates.

econ.EM

(Visualizing) Plausible Treatment Effect Paths

We consider point estimation and inference for the treatment effect path of a policy. Examples include dynamic treatment effects in microeconomics, impulse response functions in macroeconomics, and event study paths in finance. We present two sets of plausible bounds to quantify and visualize the uncertainty associated with this object. Both plausible bounds are often substantially tighter than traditional confidence intervals, and can provide useful insights even when traditional (uniform) confidence bands appear uninformative. Our bounds can also lead to markedly different conclusions when there is significant correlation in the estimates, reflecting the fact that traditional confidence bands can be ineffective at visualizing the impact of such correlation. Our first set of bounds covers the average (or overall) effect rather than the entire treatment path. Our second set of bounds imposes data-driven smoothness restrictions on the treatment path. Post-selection Inference (Berk et al. [2013]) provides formal coverage guarantees for these bounds. The chosen restrictions also imply novel point estimates that perform well across our simulations.

econ.EM