SearcharxivSearch

arXiv subjects

Lucas D. Konrad

Publications and source records attributed to Lucas D. Konrad.

3 recordsLinked to original sources

Bayesian Indicator-Saturated Regression

Structural break detection has emerged as an important tool for assessing the effects of policies in settings where conventional policy evaluation methods might not be applicable.In this paper, we introduce a unified Bayesian framework for detecting structural breaks with unknown timing and arbitrary sequence in longitudinal data. The proposed setup builds on a indicator-saturated regression design and uses a spike-and-slab prior for selection among indicators. We establish that a non-local prior as the slab component is a necessary condition to provide model selection consistency in this model class. Simulation results show that the method outperforms comparable frequentist approaches, particularly in environments with a high probability of structural breaks. We illustrate the proposed framework by analysing climate policies in the European road transport sector.

econ.EM

Finding Most Influential Sets

Identifying most influential sets (MIS) - size-$k$ subsets whose removal maximally changes a target estimand - is typically infeasible because it requires searching over $\binom{n}{k}$ subsets. For estimands with linear-fractional leave-set-out effects, we show that MIS selection reduces to a one-parameter sequence of top-$k$ problems. Dinkelbach's method yields an algorithm with $\mathcal{O}(n)$ cost per iteration and finite termination. For fixed residualized inputs, the algorithm returns a globally optimal set for the univariate ratio objective, including the oracle-residualized partial linear model. With estimated nuisance functions, uniform denominator and generated-score stability imply approximation to the first-order oracle orthogonal-score objective; exact set recovery follows under a separation condition. Simulations and applications show that the method recovers exact MIS that were previously computationally inaccessible.

stat.ML

Testing Most Influential Sets

Small influential data subsets can dramatically impact model conclusions, with a few data points overturning key findings. While recent work identifies these most influential sets, there is no formal way to tell when maximum influence is excessive rather than expected under natural random sampling variation. We address this gap by developing a principled framework for most influential sets. Focusing on linear least-squares, we derive a convenient exact influence formula and identify the extreme value distributions of maximal influence - the heavy-tailed Fréchet for constant-size sets and heavy-tailed data, and the well-behaved Gumbel for growing sets or light tails. This allows us to conduct rigorous hypothesis tests for excessive influence. We demonstrate through applications across economics, biology, and machine learning benchmarks, resolving contested findings and replacing ad-hoc heuristics with rigorous inference.

stat.ML