Searcharxiv⌕ Search

arXiv subjects

Getachew K. Befekadu

Publications and source records attributed to Getachew K. Befekadu.

At least 19 recordsLinked to original sources

Simulation-based parameter estimation via a combination of embedded normalizing flows and implied empirical probabilities under moment restrictions

In this work, we present a simulation-based parameter estimation framework for a model defined by a computational simulation of a physical system. We specifically outline an estimation framework consisting of two closely-integrated steps that facilitate an overall end-to-end parameter estimation scheme. The first step involves utilizing an embedded normalizing flow which is used to transform the unknown complex distribution of the residual information into a simple base distribution corresponding to the transformed residual information. In the second step, an empirical-likelihood estimator, under moment restrictions, is utilized for imposing an indirect constrain on the base distribution, where such an instantiated task reasonably allows us to treat the transformed residual information as random variables arising from discretely distribution population with each transformed data point as a single-cell from a set of finite-cell contingencies. Moreover, we use first-order gradient methods for updating the estimated parameter values of the model defined by the computational simulation and the corresponding parametrized embedded normalizing flow, that call for all gradient-related information by leveraging implicitly differentiations of the empirical-likelihood function, which is constructed from the implied empirical probabilities under moment restrictions. Here, it is worth mentioning that the problem formulation presented in this work, which highlights an information-theoretic interpretation, allows to present a computational framework for algorithmic implementations. Finally, as a-by-product, the inverse of the parametrized embedded normalizing flow, w.r.t. the estimated parameter values, serves as a surrogate model for the computational simulation model, which provides useful information for quantifying model discrepancies and sensitivity analysis.

stat.ME↗

Simulation-based Bayesian inference with ameliorative learned summary statistics -- Part I

This paper, which is Part 1 of a two-part paper series, considers a simulation-based inference with learned summary statistics, in which such a learned summary statistic serves as an empirical-likelihood with ameliorative effects in the Bayesian setting, when the exact likelihood function associated with the observation data and the simulation model is difficult to obtain in a closed form or computationally intractable. In particular, a transformation technique which leverages the Cressie-Read discrepancy criterion under moment restrictions is used for summarizing the learned statistics between the observation data and the simulation outputs, while preserving the statistical power of the inference. Here, such a transformation of data-to-learned summary statistics also allows the simulation outputs to be conditioned on the observation data, so that the inference task can be performed over certain sample sets of the observation data that are considered as an empirical relevance or believed to be particular importance. Moreover, the simulation-based inference framework discussed in this paper can be extended further, and thus handling weakly dependent observation data. Finally, we remark that such an inference framework is suitable for implementation in distributed computing, i.e., computational tasks involving both the data-to-learned summary statistics and the Bayesian inferencing problem can be posed as a unified distributed inference problem that will exploit distributed optimization and MCMC algorithms for supporting large datasets associated with complex simulation models.

stat.ML↗

A brief note on learning problem with global perspectives

This brief note considers the problem of learning with dynamic-optimizing principal-agent setting, in which the agents are allowed to have global perspectives about the learning process, i.e., the ability to view things according to their relative importances or in their true relations based-on some aggregated information shared by the principal. Whereas, the principal, which is exerting an influence on the learning process of the agents in the aggregation, is primarily tasked to solve a high-level optimization problem posed as an empirical-likelihood estimator under conditional moment restrictions model that also accounts information about the agents' predictive performances on out-of-samples as well as a set of private datasets available only to the principal. In particular, we present a coherent mathematical argument which is necessary for characterizing the learning process behind this abstract principal-agent learning framework, although we acknowledge that there are a few conceptual and theoretical issues still need to be addressed.

stat.ML↗

On improving generalization in a class of learning problems with the method of small parameters for weakly-controlled optimal gradient systems

In this paper, we provide a mathematical framework for improving generalization in a class of learning problems which is related to point estimations for modeling of high-dimensional nonlinear functions. In particular, we consider a variational problem for a weakly-controlled gradient system, whose control input enters into the system dynamics as a coefficient to a nonlinear term which is scaled by a small parameter. Here, the optimization problem consists of a cost functional, which is associated with how to gauge the quality of the estimated model parameters at a certain fixed final time w.r.t. the model validating dataset, while the weakly-controlled gradient system, whose the time-evolution is guided by the model training dataset and its perturbed version with small random noise. Using the perturbation theory, we provide results that will allow us to solve a sequence of optimization problems, i.e., a set of decomposed optimization problems, so as to aggregate the corresponding approximate optimal solutions that are reasonably sufficient for improving generalization in such a class of learning problems. Moreover, we also provide an estimate for the rate of convergence for such approximate optimal solutions. Finally, we present some numerical results for a typical case of nonlinear regression problem.

math.OC↗

Further extensions on the successive approximation method for hierarchical optimal control problems and its application to learning

In this paper, further extensions of the result of the paper "A successive approximation method in functional spaces for hierarchical optimal control problems and its application to learning, arXiv:2410.20617 [math.OC], 2024" concerning a class of learning problem of point estimations for modeling of high-dimensional nonlinear functions are given. In particular, we present two viable extensions within the nested algorithm of the successive approximation method for the hierarchical optimal control problem, that provide better convergence property and computationally efficiency, which ultimately leading to an optimal parameter estimate. The first extension is mainly concerned with the convergence property of the steps involving how the two agents, i.e., the "leader" and the "follower," update their admissible control strategies, where we introduce augmented Hamiltonians for both agents and we further reformulate the admissible control updating steps as as sub-problems within the nested algorithm of the hierarchical optimal control problem that essentially provide better convergence property. Whereas the second extension is concerned with the computationally efficiency of the steps involving how the agents update their admissible control strategies, where we introduce intermediate state variable for each agent and we further embed the intermediate states within the optimal control problems of the "leader" and the "follower," respectively, that further lend the admissible control updating steps to be fully efficient time-parallelized within the nested algorithm of the hierarchical optimal control problem.

math.OC↗

A successive approximation method in functional spaces for hierarchical optimal control problems and its application to learning

We consider a class of learning problem of point estimation for modeling high-dimensional nonlinear functions, whose learning dynamics is guided by model training dataset, while the estimated parameter in due course provides an acceptable prediction accuracy on a different model validation dataset. Here, we establish an evidential connection between such a learning problem and a hierarchical optimal control problem that provides a framework how to account appropriately for both generalization and regularization at the optimization stage. In particular, we consider the following two objectives: (i) The first one is a controllability-type problem, i.e., generalization, which consists of guaranteeing the estimated parameter to reach a certain target set at some fixed final time, where such a target set is associated with model validation dataset. (ii) The second one is a regularization-type problem ensuring the estimated parameter trajectory to satisfy some regularization property over a certain finite time interval. First, we partition the control into two control strategies that are compatible with two abstract agents, namely, a leader, which is responsible for the controllability-type problem and that of a follower, which is associated with the regularization-type problem. Using the notion of Stackelberg's optimization, we provide conditions on the existence of admissible optimal controls for such a hierarchical optimal control problem under which the follower is required to respond optimally to the strategy of the leader, so as to achieve the overall objectives that ultimately leading to an optimal parameter estimate. Moreover, we provide a nested algorithm, arranged in a hierarchical structure-based on successive approximation methods, for solving the corresponding optimal control problem. Finally, we present some numerical results for a typical nonlinear regression problem.

math.OC↗

A new perspective on the learning dynamics for a class of learning problems via averaged gradient systems coupled with diffusion-transmutation processes

In the first part of this paper, we consider a family of continuous-time dynamical systems coupled with diffusion-transmutation processes. Under certain conditions, such randomly perturbed dynamical systems can be interpreted as an averaged dynamical system, whose weighting coefficients, that depend on the state trajectory of the underlying averaged system, are assumed to be strictly positive with sum unity. Here, we provide a large deviation result for the corresponding family of processes, i.e., a variational problem formulation modeling the most likely sample path leading to certain noise-induced rare-events. This remarkably allows us to provide a computational algorithm for solving the corresponding variational problem. In the second part of the paper, we use some of the insights from the first part and provide a new perspective on the learning dynamics for a class of learning problems, whose averaged gradient dynamical systems, from continuous-time perspective, are guided by a set of subsampled datasets that are obtained from the original dataset via bootstrapping or other related resampling-based techniques. Finally, we present some numerical results for a typical nonlinear regression problem, where the corresponding averaged gradient system is interpreted as random walks on a graph, whose outgoing edges are uniformly chosen at random.

math.OC↗

Embedding generalization within the learning dynamics: An approach based-on sample path large deviation theory

We consider a typical learning problem of point estimations for modeling of nonlinear functions or dynamical systems in which generalization, i.e., verifying a given learned model, can be embedded as an integral part of the learning process or dynamics. In particular, we consider an empirical risk minimization based learning problem that exploits gradient methods from continuous-time perspective with small random perturbations, which is guided by the training dataset loss. Here, we provide an asymptotic probability estimate in the small noise limit based-on the Freidlin-Wentzell theory of large deviations, when the sample path of the random process corresponding to the randomly perturbed gradient dynamical system hits a certain target set, i.e., a rare event, when the latter is specified by the testing dataset loss landscape. Interestingly, the proposed framework can be viewed as one way of improving generalization and robustness in learning problems that provides new insights leading to optimal point estimates which is guided by training data loss, while, at the same time, the learning dynamics has an access to the testing dataset loss landscape in some form of future achievable or anticipated target goal. Moreover, as a by-product, we establish a connection with optimal control problem, where the target set, i.e., the rare event, is considered as the desired outcome or achievable target goal for a certain optimal control problem, for which we also provide a verification result reinforcing the rationale behind the proposed framework. Finally, we present a computational algorithm that solves the corresponding variational problem leading to an optimal point estimates and, as part of this work, we also present some numerical results for a typical case of nonlinear regression problem.

math.OC↗

On the rare-event simulations of diffusion processes pertaining to a chain of distributed systems with small random perturbations

In this paper, we consider an importance sampling problem for a certain rare-event simulations involving the behavior of a diffusion process pertaining to a chain of distributed systems with random perturbations. We also assume that the distributed system formed by $n$-subsystems -- in which a small random perturbation enters in the first subsystem and then subsequently transmitted to the other subsystems -- satisfies an appropriate Hörmander condition. Here we provide an efficient importance sampling estimator, with an exponential variance decay rate, for the asymptotics of the probabilities of the rare events involving such a diffusion process that also ensures a minimum relative estimation error in the small noise limit. The framework for such an analysis basically relies on the connection between the probability theory of large deviations and the values functions for a family of stochastic control problems associated with the underlying distributed system, where such a connection provides a computational paradigm -- based on an exponentially-tilted biasing distribution -- for constructing efficient importance sampling estimators for the rare-event simulation. Moreover, as a by-product, the framework also allows us to derive a family of Hamilton-Jacobi-Bellman for which we also provide a solvability condition for the corresponding optimal control problem.

math.OC↗

Optimal residence time control for stochastically perturbed prescription opioid epidemic models

In this paper, we consider an optimal control problem for a prescription opioid epidemic model that describes the interaction between the regular prescription or addictive use of opioid drugs, and the process of rehabilitation and that of relapsing into opioid drug use. In particular, our interest is in the situation, where the control appearing linearly in the opioid epidemics is interpreted as the rate at which the susceptible individuals are effectively removed from the population due to an opioid-related intervention policy or when the dynamics of the addicted is strategically influenced due to an accessible addiction treatment facility, while a small perturbing noise enters through the dynamics of the susceptible group in the population compartmental model. To this end, we introduce a mathematical apparatus that minimizes the asymptotic exit-rate with which the solution for such stochastically perturbed prescription opioid epidemics exits from a given bounded open domain. Moreover, under certain assumptions, we also provide an admissible optimal Markov control for the corresponding optimal control problem that optimally effected removal of the susceptible or recovered individuals from the population dynamics.

math.OC↗

Optimal control of diffusion processes pertaining to an opioid epidemic dynamical model with random perturbations

In this paper, we consider the problem of controlling a diffusion process pertaining to an opioid epidemic dynamical model with random perturbation so as to prevent it from leaving a given bounded open domain. Here, we assume that the random perturbation enters only through the dynamics of the susceptible group in the compartmental model of the opioid epidemic dynamics and, as a result of this, the corresponding diffusion is degenerate, for which we further assume that the associated diffusion operator is hypoelliptic. In particular, we minimize the asymptotic exit rate of such a controlled-diffusion process from the given bounded open domain and we derive the Hamilton-Jacobi-Bellman equation for the corresponding optimal control problem, which is closely related to a nonlinear eigenvalue problem. Finally, we also prove a verification theorem that provides a sufficient condition for optimal control.

math.OC↗

A further study on the opioid epidemic dynamical model with random perturbation

In this paper, we consider an opioid epidemic dynamical model with random perturbation that typically describes the interplay between regular prescription use, addictive use, and the process of rehabilitation from addiction and vice-versa. In particular, we provide two-sided bounds on the solution of the transition density function for the Fokker-Planck equation that corresponds to the opioid epidemic dynamical model, when a random perturbation enters only through the dynamics of the susceptible group in the compartmental model. Here, the proof for such bounds basically relies on the interpretation of the solution for the transition density function as the value function of a certain optimal stochastic control problem. Finally, as a possible interesting development in this direction, we also provide an estimate for the attainable exit probability with which the solution for the randomly perturbed opioid epidemic dynamical model exits from a given bounded open domain during a certain time interval. Note that such qualitative information on the first exit-time as well as two-sided bounds on the transition density function are useful for developing effective and fact-informed intervention strategies that primarily aim at curbing opioid epidemics or assisting in interpreting outcome results from opioid-related policies.

math.OC↗

On the asymptotic of exit problems for controlled Markov diffusion processes with random jumps and vanishing diffusion terms

In this paper, we study the asymptotic of exit problem for controlled Markov diffusion processes with random jumps and vanishing diffusion terms, where the random jumps are introduced in order to modify the evolution of the controlled diffusions by switching from one mode of dynamics to another. That is, depending on the state-position and state-transition information, the dynamics of the controlled diffusions randomly switches between the different drift and diffusion terms. Here, we specifically investigate the asymptotic exit problem concerning such controlled Markov diffusion processes in two steps: (i) First, for each controlled diffusion model, we look for an admissible Markov control process that minimizes the principal eigenvalue for the corresponding infinitesimal generator with zero Dirichlet boundary conditions -- where such an admissible control process also forces the controlled diffusion process to remain in a given bounded open domain for a longer duration. (ii) Then, using large deviations theory, we determine the exit place and the type of distribution at the exit time for the controlled Markov diffusion processes coupled with random jumps and vanishing diffusion terms. Moreover, the asymptotic results at the exit time also allow us to determine the limiting behavior of the Dirichlet problem for the corresponding system of elliptic partial differential equations containing a small vanishing parameter.

math.DS↗

On the stochastic decision problems with backward stochastic viability property

In this paper, we consider a stochastic decision problem for a system governed by a stochastic differential equation, in which an optimal decision is made in such a way to minimize a vector-valued accumulated cost over a finite-time horizon that is associated with the solution of a certain multi-dimensional backward stochastic differential equation (BSDE). Here, we also assume that the solution for such a multi-dimensional BSDE {\it almost surely} satisfies a backward stochastic viability property w.r.t. a given closed convex set. Moreover, under suitable conditions, we establish the existence of an optimal solution, in the sense of viscosity solutions, to the associated system of semilinear parabolic PDEs. Finally, we briefly comment on the implication of our results.

math.OC↗

On the hierarchical risk-averse control problems for diffusion processes

In this paper, we consider a risk-averse control problem for diffusion processes, in which there is a partition of the admissible control strategy into two decision-making groups (namely, the {\it leader} and {\it follower}) with different cost functionals and risk-averse satisfactions. Our approach, based on a hierarchical optimization framework, requires that a certain level of risk-averse satisfaction be achieved for the {\it leader} as a priority over that of the {\it follower's} risk-averseness. In particular, we formulate such a risk-averse control problem involving a family of time-consistent dynamic convex risk measures induced by conditional $g$-expectations (i.e., filtration-consistent nonlinear expectations associated with the generators of certain backward stochastic differential equations). Moreover, under suitable conditions, we establish the existence of optimal risk-averse solutions, in the sense of viscosity solutions, for the corresponding risk-averse dynamic programming equations. Finally, we briefly comment on the implication of our results.

math.OC↗

Large deviation principle for dynamical systems coupled with diffusion-transmutation processes

In this paper, we introduce a mathematical apparatus that is relevant for understanding a dynamical system with small random perturbations and coupled with the so-called transmutation process -- where the latter jumps from one mode to another, and thus modifying the dynamics of the system. In particular, we study the exit problem, i.e., an asymptotic estimate for the exit probabilities with which the corresponding processes exit from a given bounded open domain, and then formally prove a large deviation principle for the exit position joint with the type occupation times as the random perturbation vanishes. Moreover, under certain conditions, the exit place and the type of distribution at the exit time are determined and, as a consequence of this, such information also give the limit of the Dirichlet problems for the associated partial differential equation systems with a vanishing small parameter.

math.DS↗

Acceptable risks and related decision problems with multiple risk-averse agents

In this paper, we consider a risk-averse decision problem for controlled-diffusion processes, with dynamic risk measures, in which multiple risk-averse agents choose their decisions in such a way to minimize their individual accumulated risk-costs over a finite-time horizon. In particular, we introduce multi-structure dynamic risk measures induced from conditional $g$-expectations, where the latter are associated with the generator functionals of certain BSDEs that implicitly take into account the risk-cost functionals of the risk-averse agents. Here, we also require that such solutions of the BSDEs to satisfy a stochastic viability property with respect to a given closed convex set. Moreover, using a result similar to that of the Arrow-Barankin-Blackwell theorem, we establish the existence of consistent optimal decisions for the risk-averse agents, when the set of all Pareto optimal solutions, in the sense of viscosity, for the associated dynamic programming equations is dense in the given closed convex set. Finally, we briefly comment on the characteristics of acceptable risks vis-á-vis some uncertain future costs or outcomes, in which results from the dynamic risk analysis constitute part of the information used in the risk-averse decision criteria.

math.OC↗

On the dynamic consistency of hierarchical risk-averse decision problems

In this paper, we consider a risk-averse decision problem for controlled-diffusion processes, with dynamic risk measures, in which there are two risk-averse decision makers (i.e., {\it leader} and {\it follower}) with different risk-averse related responsibilities and information. Moreover, we assume that there are two objectives that these decision makers are expected to achieve. That is, the first objective being of {\it stochastic controllability} type that describes an acceptable risk-exposure set vis-á-vis some uncertain future payoff, and while the {\it second one} is making sure the solution of a certain risk-related system equation has to stay always above a given continuous stochastic process, namely {\it obstacle}. In particular, we introduce multi-structure, time-consistent, dynamic risk measures induced from conditional $g$-expectations, where the latter are associated with the generator functionals of two backward-SDEs that implicitly take into account the above two objectives along with the given continuous obstacle process. Moreover, under certain conditions, we establish the existence of optimal hierarchical risk-averse solutions, in the sense of viscosity solutions, to the associated risk-averse dynamic programming equations that formalize the way in which both the {\it leader} and {\it follower} consistently choose their respective risk-averse decisions. Finally, we remark on the implication of our result in assessing the influence of the {\it leader'}s decisions on the risk-averseness of the {\it follower} in relation to the direction of {\it leader-follower} information flow.

math.OC↗