SearcharxivSearch

arXiv subjects

Asaf Cohen

Publications and source records attributed to Asaf Cohen.

At least 19 recordsLinked to original sources

Quantitative comparison of closed- and open-loop linear-quadratic $N$-player differential games

We compare closed-loop and open-loop Nash equilibria in a finite-horizon stochastic linear-quadratic $N$-player game with decoupled state dynamics and interaction through the state costs. We introduce a block-diagonal reference game and a nearby perturbed game, and study each under both information structures. The closed-loop and open-loop problems lead to different Riccati systems, but for the reference game their induced equilibrium state and control processes coincide exactly. We then prove solvability and stability under weak perturbations of the state costs, yielding quantitative bounds between the corresponding equilibria. In particular, when the perturbation decreases sufficiently fast with the population size, the closed-loop and open-loop equilibria of the perturbed game become asymptotically equivalent. The comparison is carried out directly at the finite-player level, without requiring exchangeability or a mean-field limit.

math.OC

A Principal-Agent Mean-Field Game Model of Insurance with Risk Interdependence

We study an insurance contract-design problem under moral hazard, endogenous participation, and strategic risk interdependence. Because the resulting $N$-agent game suffers from the curse of dimensionality, we approximate the strategic interactions via a heterogeneous mean-field game. We rigorously establish the existence of a lower-level mean-field Nash equilibrium using measurable selection arguments and the Kakutani fixed-point theorem. By proving the $L^1$-Lipschitz continuity of the aggregate participation threshold, we further establish equilibrium uniqueness via a contraction mapping. We then embed this mean-field response into the insurer's upper-level Stackelberg optimization problem. We formulate the objective through general performance envelopes to accommodate potential equilibrium multiplicity, proving the existence of upper-level $\varepsilon$-optimal contracts, and demonstrating the existence of an exact Stackelberg equilibrium under the uniqueness regime. We conclude by extending the model to finite contract menus, providing numerical evidence that multi-contract screening improves the principal's expected payoff in interdependent risk environments.

math.OC

Importance Sampling for Event Discovery via Guesswork

Traditional importance sampling (IS) is designed to estimate rare-event probabilities by minimizing estimator variance. However, many applications prioritize rapid discovery: the generation of a trajectory within a rare set $A_n$. This requires a shift from ensemble-based estimation to a design principle focused on the hitting time $\tau_{A_n} := \inf\{t \ge 1 : Y_t^n \in A_n\}$. We formalize a Quality of Discovery problem as the problem of minimizing the description length (surprisal) of the discovered trajectory under the nominal model $p$. We prove that minimizing this description length is equivalent to minimizing the nominal rank exponent $J_{\mathrm{rank}}(q_n) := \lim_{n\to\infty} \frac{1}{n} \log G_n(Y^n)$, where $G_n(x^n)$ is the guesswork of sequence $x^n$. For i.i.d.\ models and type-defined rare sets $\Gamma$, we show that while classical IS targets the mass-dominating type $Q_{\mathrm{IS}}^* \in \arg\min_{Q \in \Gamma} D(Q\|p)$, discovery optimality is achieved by $Q_{\mathrm{GW}}^* \in \arg\min_{Q \in \Gamma} [H(Q) + D(Q\|p)]$. This framework identifies a fundamental rule: minimizing the guesswork exponent ensures the discovered sequence is the "least surprising" representative of the set relative to the nominal model's search order. We further demonstrate that under budgetary constraints, this exponent serves as a lexicographic tie-breaker when the hitting-time minimizer is not unique. This establishes $H(Q) + D(Q\|p)$ as a natural objective for discovery-based importance sampling, providing a formal bridge between randomized sampling and systematic search.

cs.IT

Conditional Diffusion Under Linear Constraints: Langevin Mixing and Information-Theoretic Guarantees

We study zero-shot conditional sampling with pretrained diffusion models for linear inverse problems, including inpainting and super-resolution. In these problems, the observation determines only part of the unknown signal. The remaining degrees of freedom must be sampled according to the correct conditional data distribution. Existing projection-based samplers enforce measurement consistency by correcting the observed component during reverse diffusion. However, measurement consistency alone does not determine how probability mass should be distributed along the feasible set, and this can lead to biased conditional samples. We analyze this issue through a normal--tangent decomposition of the score function. For Gaussian noising, the observed-direction score is exactly determined by the measurement; only the tangent conditional score is unknown. We prove that the error from replacing this score by the unconditional tangent score is upper bounded by a dimension-free conditional mutual information between observed and unobserved components. This gives an information-theoretic decomposition into initialization and pathwise score-mismatch errors. Motivated by the theory, we propose a projected-Langevin initialization followed by guided reverse denoising, which outperforms a strong projection-based baseline in inpainting and super-resolution experiments.

cs.LG

Operator Learning for Families of Finite-State Mean-Field Games

Finite-state mean-field games (MFGs) arise as limits of large interacting particle systems and are governed by an MFG system, a coupled forward-backward differential equation consisting of a forward Kolmogorov-Fokker-Planck (KFP) equation describing the population distribution and a backward Hamilton-Jacobi-Bellman (HJB) equation defining the value function. Solving MFG systems efficiently is challenging, with the structure of each system depending on an initial distribution of players and the terminal cost of the game. We propose an operator learning framework that solves parametric families of MFGs, enabling generalization without retraining for new initial distributions and terminal costs. We provide theoretical guarantees on the approximation error, parametric complexity, and generalization performance of our method, based on a novel regularity result for an appropriately defined flow map corresponding to an MFG system. We demonstrate empirically that our framework achieves accurate approximation for two representative instances of MFGs: a cybersecurity example and a high-dimensional quadratic model commonly used as a benchmark for numerical methods for MFGs.

math.OC

Thompson Sampling Algorithm for Stochastic Games

We study a stochastic differential game with $N$ competitive players in a linear-quadratic framework with ergodic cost, where $d$-dimensional diffusion processes govern the state dynamics with an unknown common drift (matrix). Assuming a Gaussian prior on the drift, we use filtering techniques to update its posterior estimates. Based on these estimates, we propose a Thompson-sampling-based algorithm with dynamic episode lengths to approximate strategies. We show that the Bayesian regret for each player has an error bound of order $O(\sqrt{T\log(T)})$, where $T$ is the time-horizon, independent of the number of players. This implies that average regret per unit time goes to zero. Finally, we prove that the algorithm results in a Nash equilibrium.

math.OC

On Cost-Aware Designs for Sequential Hypothesis Testing

We introduce Cost-Aware (CA) Sequential Hypothesis Testing (CASHT), in which an active decision-maker selects sensing actions with differing, random costs to identify the true hypothesis under an average-error constraint $\delta$ while minimizing the expected total cost rather than the number of samples. For fixed costs, we prove that the optimal expected total cost scales as $\Theta(\log(1/\delta))$, and is achievable by Multihypothesis Sequential Probability Ratio Test-based procedures. We show that the CA design principle is to maximize the ratio of expected information gain to expected cost under the policy-induced action distribution. Guided by this principle, we adapt two classic policies to the CA setting and establish their asymptotic optimality. We then treat random costs under two revelation models: ex-post, where costs are disclosed only after a sample is obtained, and the cost-error tradeoff coincides with the fixed-cost case, and ex-ante, where costs accrue before acquisition, and the decision maker may cancel an action mid-operation. For the ex-ante model, we characterize when cancellation lowers the total cost and analyze several cost distributions in detail. Simulations confirm our findings that the CA variants consistently reduce total cost relative to their classical counterparts, and when action cancellation helps or hurts.

cs.IT

Iterative Hypothesis Pruning and Distribution-based Early Labeling for Sequential Hypothesis Testing

We consider the framework of Sequential Hypothesis Testing (SHT), in which a decision maker (DM) selects actions that generate samples from known, action-dependent distributions, while the realized distribution is determined by an unknown true hypothesis. To identify this hypothesis, we adopt the elimination perspective and propose three deterministic, adaptive, multi-iteration algorithms with a common structure, termed $\Phi$, $\Phi$-$\Delta$, and $I$. In each iteration, the DM selects an action and repeatedly applies it to collect samples, after which hypotheses inconsistent with the observed data are eliminated. The algorithms differ in the criterion used to terminate each iteration: $\Phi$ continues until one hypothesis dominates all others; $\Phi$-$\Delta$ first clusters hypotheses whose per-action distributions are close in total variation and then proceeds in the spirit of $\Phi$; $I$ continues until one hypothesis can be safely discarded. We analyze our algorithms, establishing: (i) controlled error-rates, (ii) controlled sample complexity, (iii) asymptotic optimality, (iv) computational complexity, and (v) NP-hardness of the optimal action-sequence selection for minimal sample complexity.

cs.IT

Active Sequential Hypothesis Testing with Non-Homogeneous Costs

We study the Non-Homogeneous Sequential Hypothesis Testing (NHSHT), where a single active Decision-Maker (DM) selects actions with heterogeneous positive costs to identify the true hypothesis under an average error constraint \(\delta\), while minimizing expected total cost paid. Under standard arguments, we show that the objective decomposes into the product of the mean number of samples and the mean per-action cost induced by the policy. This leads to a key design principle: one should optimize the ratio of expectations (expected information gain per expected cost) rather than the expectation of per-step information-per-cost ("bit-per-buck"), which can be suboptimal. We adapt the Chernoff scheme to NHSHT, preserving its classical \(\log 1/\delta\) scaling. In simulations, the adapted scheme reduces mean cost by up to 50\% relative to the classic Chernoff policy and by up to 90\% relative to the naive bit-per-buck heuristic.

cs.IT

Secure Best Arm Identification in the Presence of a Copycat

Consider the problem of best arm identification with a security constraint. Specifically, assume a setup of stochastic linear bandits with $K$ arms of dimension $d$. In each arm pull, the player receives a reward that is the sum of the dot product of the arm with an unknown parameter vector and independent noise. The player's goal is to identify the best arm after $T$ arm pulls. Moreover, assume a copycat Chloe is observing the arm pulls. The player wishes to keep Chloe ignorant of the best arm. While a minimax--optimal algorithm identifies the best arm with an $\Omega\left(\frac{T}{\log(d)}\right)$ error exponent, it easily reveals its best-arm estimate to an outside observer, as the best arms are played more frequently. A naive secure algorithm that plays all arms equally results in an $\Omega\left(\frac{T}{d}\right)$ exponent. In this paper, we propose a secure algorithm that plays with \emph{coded arms}. The algorithm does not require any key or cryptographic primitives, yet achieves an $\Omega\left(\frac{T}{\log^2(d)}\right)$ exponent while revealing almost no information on the best arm.

cs.LG

Turnpike properties in linear quadratic Gaussian N-player differential games

We consider the long-time behavior of equilibrium strategies and state trajectories in a linear quadratic $N$-player game with Gaussian initial data. By comparing the finite-horizon game with its ergodic counterpart, we establish exponential convergence estimates between the solutions of the finite-horizon generalized Riccati system and the associated algebraic system arising in the ergodic setting. Building on these results, we prove the convergence of the time-averaged value function and derive a turnpike property for the equilibrium pairs of each player. Importantly, our approach avoids reliance on the mean field game limiting model, allowing for a fully uniform analysis with respect to the number of players $N$. As a result, we further establish a uniform turnpike property for the equilibrium pairs between the finite-horizon and ergodic games with $N$ players. Numerical experiments are also provided to illustrate and support the theoretical results.

math.OC

Uniform-in-Time Convergence Rates to a Nonlinear Markov Chain for Mean-Field Interacting Jump Processes

We consider a system of $N$ particles interacting through their empirical distribution on a finite state space in continuous time. In the formal limit as $N\to\infty$, the system takes the form of a nonlinear (McKean--Vlasov) Markov chain. This paper rigorously establishes this limit. Specifically, under the assumption that the mean field system has a unique, exponentially stable stationary distribution, we show that the weak error between the empirical measures of the $N$-particle system and the law of the mean field system is of order $1/N$ uniformly in time. Our analysis makes use of a master equation for test functions evaluated along the measure flow of the mean field system, and we demonstrate that the solutions of this master equation are sufficiently regular. We then show that exponential stability of the mean field system is implied by exponential stability for solutions of the linearized Kolmogorov equation with a source term. Finally, we show that our results can be applied to the study of mean field games and give a new condition for the existence of a unique stationary distribution for a nonlinear Markov chain.

math.PR

On the light-curves of disk and bulge novae

We examine the light curves of a sample of novae, classifying them into single-peaked and multiple-peaked morphologies. Using accurate distances from Gaia, we determine the spatial distribution of these novae by computing their heights, $Z$, above the Galactic plane. We show that novae exhibiting a single peak in their light curves tend to concentrate near the Galactic plane, while those displaying multiple peaks are more homogeneously distributed, reaching heights up to 1000 pc above the plane. A KS test rejects the null hypothesis that the two distributions originate from the same population at a significance level corresponding to $4.2\sigma$.

astro-ph.GA

Covert Adversarial Actuators in Finite MDPs

We consider a Markov decision process (MDP) in which actions prescribed by the controller are executed by a separate actuator, which may behave adversarially. At each time step, the controller selects and transmits an action to the actuator; however, the actuator may deviate from the intended action to degrade the control reward. Given that the controller observes only the sequence of visited states, we investigate whether the actuator can covertly deviate from the controller's policy to minimize its reward without being detected. We establish conditions for covert adversarial behavior over an infinite time horizon and formulate an optimization problem to determine the optimal adversarial policy under these conditions. Additionally, we derive the asymptotic error exponents for detection in two scenarios: (1) a binary hypothesis testing framework, where the actuator either follows the prescribed policy or a known adversarial strategy, and (2) a composite hypothesis testing framework, where the actuator may employ any stationary policy. For the latter case, we also propose an optimization problem to maximize the adversary's performance.

cs.IT

Multi-Stage Active Sequential Hypothesis Testing with Clustered Hypotheses

We consider the problem where an active Decision-Maker (DM) is tasked to identify the true hypothesis using as few as possible observations while maintaining accuracy. The DM collects observations according to its determined actions and knows the distributions under each hypothesis. We propose a deterministic and adaptive multi-stage hypothesis-elimination strategy where the DM selects an action, applies it repeatedly, and discards hypotheses in light of its obtained observations. The DM selects actions based on maximal separation expressed by the distance between the parameter vectors of each distribution under each hypothesis. Close distributions can be clustered, simplifying the search and significantly reducing the number of required observations. Our algorithms achieve vanishing Average Bayes Risk (ABR) as the error probability approaches zero, i.e., the algorithm is asymptotically optimal. Furthermore, we show that the ABR is bounded when the number of hypotheses grows. Simulations are carried out to evaluate the algorithm's performance compared to another multi-stage hypothesis-elimination algorithm, where an improvement of several orders of magnitude in the mean number of observations required is observed.

cs.IT

PIR Over Wireless Channels: Achieving Privacy With Public Responses

In this paper, we address the problem of Private Information Retrieval (PIR) over a public Additive White Gaussian Noise (AWGN) channel. In such a setup, the server's responses are visible to other servers. Thus, a curious server can listen to the other responses, compromising the user's privacy. Indeed, previous works on PIR over a shared medium assumed the servers cannot instantaneously listen to other responses. To address this gap, we present a novel randomized lattice -- PIR coding scheme that jointly codes for privacy, channel noise, and curious servers which may listen to other responses. We demonstrate that a positive PIR rate is achievable even in cases where the channel to the curious server is stronger than the channel to the user.

cs.IT

Convergence of the Deep Galerkin Method for Finite State Mean Field Control Problems

We establish the convergence of the deep Galerkin method (DGM), a deep learning-based scheme for solving high-dimensional nonlinear PDEs, for Hamilton-Jacobi-Bellman (HJB) equations that arise from the study of mean field control problems (MFCPs). Based on a recent characterization of the value function of the MFCP as the unique viscosity solution of an HJB equation on the simplex, we establish both an existence and convergence result for the DGM. First, we show that the loss functional of the DGM can be made arbitrarily small given that the value function of the MFCP possesses sufficient regularity. Then, we show that if the loss functional of the DGM converges to zero, the corresponding neural network approximators must converge uniformly to the true value function on the simplex. We also provide numerical experiments demonstrating the DGM's ability to generalize to high-dimensional HJB equations.

math.OC

Asymptotic Nash Equilibria of Finite-State Ergodic Markovian Mean Field Games

Mean field games (MFGs) model equilibria in games with a continuum of weakly interacting players as limiting systems of symmetric $n$-player games. We consider the finite-state, infinite-horizon problem with ergodic cost. Assuming Markovian strategies, we first prove that any solution to the MFG system gives rise to a $(C/\sqrt{n})$-Nash equilibrium in the $n$-player game. We follow this result by proving the same is true for the strategy profile derived from the master equation. We conclude the main theoretical portion of the paper by establishing a large deviation principle for empirical measures associated with the asymptotic Nash equilibria. Then, we contrast the asymptotic Nash equilibria using an example. We solve the MFG system directly and numerically solve the ergodic master equation by adapting the deep Galerkin method of Sirignano and Spiliopoulos. We use these results to derive the strategies of the asymptotic Nash equilibria and compare them. Finally, we derive an explicit form for the rate functions in dimension two.

math.OC