SearcharxivSearch

arXiv subjects

Artem Prokhorov

Publications and source records attributed to Artem Prokhorov.

7 recordsLinked to original sources

Change Detection in Probability Flow ODE: Online Testing in Diffusion Latent Spaces

A rapidly growing range of sequential data tasks, such as identifying trend reversals in financial markets, auto-segmenting video and audio recordings, detecting changes in movement direction from motion sensors cannot be fully addressed without detection of distributional shifts in time-ordered data. We consider a sequential change-point detection problem where the conditional density switches at an unknown time, yet neither the pre- nor post-change distribution admits a closed-form. Classical likelihood-ratio statistics are inapplicable in this settings. A conditional diffusion model, trained on pre-change-point data with a frozen context encoder, defines a deterministic bijection via the probability flow ODE. Pre-change observations are mapped onto standard Gaussian latent variables. Post-change observations, processed through the same frozen map, deviate from this reference. We employ the Maximum Mean Discrepancy as the test statistic, derive closed-form expressions for its components under the Gaussian null, and establish its asymptotic distribution as a degenerate U-statistic. Afterwards we apply an online detection procedure of Shiryaev--Roberts to the resulting statistic with exact threshold calibration. The method detects arbitrary distributional shifts, including covariance rotations and higher-order structural breaks, without parametric assumptions on either regime.

cs.LG

The Post Double LASSO for Efficiency Analysis

Big data and machine learning methods have become commonplace across economic milieus. One area that has not seen as much attention to these important topics yet is efficiency analysis. We show how the availability of big (wide) data can actually make detection of inefficiency more challenging. We then show how machine learning methods can be leveraged to adequately estimate the primitives of the frontier itself as well as inefficiency using the `post double LASSO' by deriving Neyman orthogonal moment conditions for this problem. Finally, an application is presented to illustrate key differences of the post-double LASSO compared to other approaches.

econ.EM

Improved Semi-Parametric Bounds for Tail Probability and Expected Loss: Theory and Applications

Many management decisions involve accumulated random realizations for which only the first and second moments of their distribution are available. The sharp Chebyshev-type bound for the tail probability and Scarf bound for the expected loss are widely used in this setting. We revisit the tail behavior of such quantities with a focus on independence. Conventional primal-dual approaches from optimization are ineffective in this setting. Instead, we use probabilistic inequalities to derive new bounds and offer new insights. For non-identical distributions attaining the tail probability bounds, we show that the extreme values are equidistant regardless of the distributional differences. For the bound on the expected loss, we show that the impact of each random variable on the expected sum can be isolated using an extension of the Korkine identity. We illustrate how these new results open up abundant practical applications, including improved pricing of product bundles, more precise option pricing, more efficient insurance design, and better inventory management. For example, we establish a new solution to the optimal bundling problem, yielding a 17% uplift in per-bundle profits, and a new solution to the inventory problem, yielding a 5.6% cost reduction for a model with 20 retailers.

econ.EM

Change-Point Detection in Time Series Using Mixed Integer Programming

We use cutting-edge mixed integer optimization (MIO) methods to develop a framework for detection and estimation of structural breaks in time series regression models. The framework is constructed based on the least squares problem subject to a penalty on the number of breakpoints. We restate the $l_0$-penalized regression problem as a quadratic programming problem with integer- and real-valued arguments and show that MIO is capable of finding provably optimal solutions using a well-known optimization solver. Compared to the popular $l_1$-penalized regression (LASSO) and other classical methods, the MIO framework permits simultaneous estimation of the number and location of structural breaks as well as regression coefficients, while accommodating the option of specifying a given or minimal number of breaks. We derive the asymptotic properties of the estimator and demonstrate its effectiveness through extensive numerical experiments, confirming a more accurate estimation of multiple breaks as compared to popular non-MIO alternatives. Two empirical examples demonstrate usefulness of the framework in applications from business and economic statistics.

econ.EM

Bank Cost Efficiency and Credit Market Structure Under a Volatile Exchange Rate

We study the impact of exchange rate volatility on cost efficiency and market structure in a cross-section of banks that have non-trivial exposures to foreign currency (FX) operations. We use unique data on quarterly revaluations of FX assets and liabilities (Revals) that Russian banks were reporting between 2004 Q1 and 2020 Q2. {\it First}, we document that Revals constitute the largest part of the banks' total costs, 26.5\% on average, with considerable variation across banks. {\it Second}, we find that stochastic estimates of cost efficiency are both severely downward biased -- by 30\% on average -- and generally not rank preserving when Revals are ignored, except for the tails, as our nonparametric copulas reveal. To ensure generalizability to other emerging market economies, we suggest a two-stage approach that does not rely on Revals but is able to shrink the downward bias in cost efficiency estimates by two-thirds. {\it Third}, we show that Revals are triggered by the mismatch in the banks' FX operations, which, in turn, is driven by household FX deposits and the instability of Ruble's exchange rate. {\it Fourth}, we find that the failure to account for Revals leads to the erroneous conclusion that the credit market is inefficient, which is driven by the upper quartile of the banks' distribution by total assets. Revals have considerable negative implications for financial stability which can be attenuated by the cross-border diversification of bank assets.

econ.EM

An early warning system for emerging markets

Financial markets of emerging economies are vulnerable to extreme and cascading information spillovers, surges, sudden stops and reversals. With this in mind, we develop a new online early warning system (EWS) to detect what is referred to as `concept drift' in machine learning, as a `regime shift' in economics and as a `change-point' in statistics. The system explores nonlinearities in financial information flows and remains robust to heavy tails and dependence of extremes. The key component is the use of conditional entropy, which captures shifts in various channels of information transmission, not only in conditional mean or variance. We design a baseline method, and adapt it to a modern high-dimensional setting through the use of random forests and copulas. We show the relevance of each system component to the analysis of emerging markets. The new approach detects significant shifts where conventional methods fail. We explore when this happens using simulations and we provide two illustrations when the methods generate meaningful warnings. The ability to detect changes early helps improve resilience in emerging markets against shocks and provides new economic and financial insights into their operation.

econ.EM

Efficient estimation of parameters in marginals in semiparametric multivariate models

We consider a general multivariate model where univariate marginal distributions are known up to a parameter vector and we are interested in estimating that parameter vector without specifying the joint distribution, except for the marginals. If we assume independence between the marginals and maximize the resulting quasi-likelihood, we obtain a consistent but inefficient QMLE estimator. If we assume a parametric copula (other than independence) we obtain a full MLE, which is efficient but only under a correct copula specification and may be biased if the copula is misspecified. Instead we propose a sieve MLE estimator (SMLE) which improves over QMLE but does not have the drawbacks of full MLE. We model the unknown part of the joint distribution using the Bernstein-Kantorovich polynomial copula and assess the resulting improvement over QMLE and over misspecified FMLE in terms of relative efficiency and robustness. We derive the asymptotic distribution of the new estimator and show that it reaches the relevant semiparametric efficiency bound. Simulations suggest that the sieve MLE can be almost as efficient as FMLE relative to QMLE provided there is enough dependence between the marginals. We demonstrate practical value of the new estimator with several applications. First, we apply SMLE in an insurance context where we build a flexible semi-parametric claim loss model for a scenario where one of the variables is censored. As in simulations, the use of SMLE leads to tighter parameter estimates. Next, we consider financial risk management examples and show how the use of SMLE leads to superior Value-at-Risk predictions. The paper comes with an online archive which contains all codes and datasets.

econ.GN