SearcharxivSearch

arXiv subjects

Bingyan Han

Publications and source records attributed to Bingyan Han.

17 recordsLinked to original sources

Volterra Generative Models

Score-based diffusion models typically use Brownian perturbations, which provide tractable reverse-time dynamics but impose memoryless noising. We introduce Volterra generative models, a continuous-time score-based framework whose forward process injects path-dependent noise through fractional kernels. To handle the non-Markovian and non-semimartingale dynamics, we construct finite-dimensional Markovian lifts using Gaussian quadrature in both regimes and a hybrid finite-difference exponential approximation in the smooth regime. We prove squared error bounds, derive an augmented linear-Gaussian forward process, and show that the learning can remain data-dimensional by considering residual states and analytic auxiliary Gaussian scores. We also identify covariance and reverse-time degeneracies caused by shared Brownian factors and signed smooth-regime weights. The degeneracy motivates stabilized conditioning and, for stiff larger lifts, a Gaussian-bridge reconstruction sampler. Experiments on MNIST and CIFAR-10 show that persistent fractional perturbations with small Markovian lifts can improve score-based generation on MNIST and provide a promising extension to natural images, while the bridge sampler provides a stability mechanism for larger lifts.

cs.LG

Continuous-time Online Learning via Mean-Field Neural Networks: Regret Analysis in Diffusion Environments

We study continuous-time online learning where data are generated by a diffusion process with unknown coefficients. The learner employs a two-layer neural network, continuously updating its parameters in a non-anticipative manner. The mean-field limit of the learning dynamics corresponds to a stochastic Wasserstein gradient flow adapted to the data filtration. We establish regret bounds for both the mean-field limit and finite-particle system. Our analysis leverages the logarithmic Sobolev inequality, Polyak-Lojasiewicz condition, Malliavin calculus, and uniform-in-time propagation of chaos. Under displacement convexity, we obtain a constant static regret bound. In the general non-convex setting, we derive explicit linear regret bounds characterizing the effects of data variation, entropic exploration, and quadratic regularization. Finally, our simulations demonstrate the outperformance of the online approach and the impact of network width and regularization parameters.

cs.LG

Goal-based portfolio selection with fixed transaction costs

We study a goal-based portfolio selection problem in which an investor aims to meet multiple financial goals, each with a specific deadline and target amount. Trading the stock incurs a strictly positive transaction cost. Using the stochastic Perron's method, we show that the value function is the unique viscosity solution to a system of quasi-variational inequalities. The existence of an optimal trading strategy and goal funding scheme is established. Numerical results reveal complex optimal trading regions and show that the optimal investment strategy differs substantially from the V-shaped strategy observed in the frictionless case.

math.OC

Goal-based portfolio selection with mental accounting

We present a continuous-time portfolio selection framework that reflects goal-based investment principles and mental accounting behavior. In this framework, an investor with multiple investment goals constructs separate portfolios, each corresponding to a specific goal, with penalties imposed on fund transfers between these goals, referred to as mental costs. By applying the stochastic Perron's method, we demonstrate that the value function is the unique constrained viscosity solution of a Hamilton-Jacobi-Bellman equation system. Numerical analysis reveals several key features: the free boundaries exhibit complex shapes with bulges and notches; the optimal strategy for one portfolio depends on the wealth level of another; investors must diversify both among stocks and across portfolios; and they may postpone reallocating surplus from an important goal to a less important one until the former's deadline approaches.

q-fin.PM

The McCormick martingale optimal transport

Martingale optimal transport (MOT) often yields broad price bounds for options, constraining their practical applicability. In this study, we extend MOT by incorporating causality constraints among assets, inspired by the nonanticipativity condition of stochastic processes. This, however, introduces a computationally challenging bilinear program. To tackle this issue, we propose McCormick relaxations to ease the bicausal formulation and refer to it as McCormick MOT. The primal attainment and strong duality of McCormick MOT are established under standard assumptions. Empirically, we apply McCormick MOT to basket and digital options. With natural bounds on probability masses, the average price reduction for basket options is approximately 1.08% to 3.90%. When tighter probability bounds are available, the reduction increases to 12.26%, compared to the classic MOT, which also incorporates tighter bounds. For most dates considered, there are basket options with suitable payoffs, where the price reduction exceeds 10.00%. For digital options, McCormick MOT results in an average price reduction of over 20.00%, with the best case exceeding 99.00%.

q-fin.MF

Existence of Markov equilibrium control in discrete time

For time-inconsistent stochastic controls in discrete time and finite horizon, an open problem in Bj\"ork and Murgoci (Finance Stoch, 2014) is the existence of an equilibrium control. A nonrandomized Borel measurable Markov equilibrium policy exists if the objective is inf-compact in every time step. We provide a sufficient condition for the inf-compactness and thus existence, with costs that are lower semicontinuous (l.s.c.) and bounded from below and transition kernels that are continuous in controls under given states. The control spaces need not to be compact.

math.OC

Robust Time-inconsistent Linear-Quadratic Stochastic Controls: A Stochastic Differential Game Approach

This paper studies robust time-inconsistent (TIC) linear-quadratic stochastic control problems, formulated by stochastic differential games. By a spike variation approach, we derive sufficient conditions for achieving the Nash equilibrium, which corresponds to a time-consistent (TC) robust policy, under mild technical assumptions. To illustrate our framework, we consider two scenarios of robust mean-variance analysis, namely with state- and control-dependent ambiguity aversion. We find numerically that with time inconsistency haunting the dynamic optimal controls, the ambiguity aversion enhances the effective risk aversion faster than the linear, implying that the ambiguity in the TIC cases is more impactful than that under the TC counterparts, e.g., expected utility maximization problems.

math.OC

Fitted value iteration methods for bicausal optimal transport

We develop a fitted value iteration (FVI) method to compute bicausal optimal transport (OT) where couplings have an adapted structure. Based on the dynamic programming formulation, FVI adopts a function class to approximate the value functions in bicausal OT. Under the concentrability condition and approximate completeness assumption, we prove the sample complexity using (local) Rademacher complexity. Furthermore, we demonstrate that multilayer neural networks with appropriate structures satisfy the crucial assumptions required in sample complexity proofs. Numerical experiments reveal that FVI outperforms linear programming and adapted Sinkhorn methods in scalability as the time horizon increases, while still maintaining acceptable accuracy.

stat.ML

Distributionally robust Kalman filtering with volatility uncertainty

This work presents a distributionally robust Kalman filter to address uncertainties in noise covariance matrices and predicted covariance estimates. We adopt a distributionally robust formulation using bicausal optimal transport to characterize a set of plausible alternative models. The optimization problem is transformed into a convex nonlinear semi-definite programming problem and solved using the trust-region interior point method with the aid of $LDL^\top$ decomposition. The empirical outperformance is demonstrated through target tracking and pairs trading.

math.OC

Equilibrium transport with time-inconsistent costs

Given two probability measures on sequential data, we investigate the transport problem with time-inconsistent preferences in a discrete-time setting. Motivating examples are nonlinear objectives, state-dependent costs, and regularized optimal transport with general $f$-divergence. Under the bicausal constraint, we introduce the concept of equilibrium transport. Existence is proved in the semi-discrete Markovian case and the continuous non-Markovian case with strict quasiconvexity, while uniqueness also holds in the second case. We apply our framework to study mean-variance dynamic matching, nonlinear or state-dependent objectives with Gaussian data, and mismatches in job markets. Numerical results indicate a positive relationship between mismatches and state dependence.

math.OC

Can maker-taker fees prevent algorithmic cooperation in market making?

In a semi-realistic market simulator, independent reinforcement learning algorithms may facilitate market makers to maintain wide spreads even without communication. This unexpected outcome challenges the current antitrust law framework. We study the effectiveness of maker-taker fee models in preventing cooperation via algorithms. After modeling market making as a repeated general-sum game, we experimentally show that the relation between net transaction costs and maker rebates is not necessarily monotone. Besides an upper bound on taker fees, we may also need a lower bound on maker rebates to destabilize the cooperation. We also consider the taker-maker model and the effects of mid-price volatility, inventory risk, and the number of agents.

q-fin.TR

Cooperation between Independent Market Makers

With the digitalization of the financial market, dealers are increasingly handling market-making activities by algorithms. Recent antitrust literature raises concerns on collusion caused by artificial intelligence. This paper studies the possibility of cooperation between market makers via independent Q-learning. Market making with inventory risk is formulated as a repeated general-sum game. Under a stag-hunt type payoff, we find that market makers can learn cooperative strategies without communication. In general, high spreads can have the largest probability even when the lowest spread is the unique Nash equilibrium. Moreover, introducing more agents into the game does not necessarily eliminate the presence of supra-competitive spreads.

q-fin.TR

Distributionally robust risk evaluation with a causality constraint and structural information

This work studies the distributionally robust evaluation of expected values over temporal data. A set of alternative measures is characterized by the causal optimal transport. We prove the strong duality and recast the causality constraint as minimization over an infinite-dimensional test function space. We approximate test functions by neural networks and prove the sample complexity with Rademacher complexity. An example is given to validate the feasibility of technical assumptions. Moreover, when structural information is available to further restrict the ambiguity set, we prove the dual formulation and provide efficient optimization methods. Our framework outperforms the classic counterparts in the distributionally robust portfolio selection problem. The connection with the naive strategy is also investigated numerically.

q-fin.MF

Algorithmic pricing with independent learners and relative experience replay

In an infinitely repeated general-sum pricing game, independent reinforcement learners may exhibit collusive behavior without any communication, raising concerns about algorithmic collusion. To better understand the learning dynamics, we incorporate agents' relative performance (RP) among competitors using experience replay (ER) techniques. Experimental results indicate that RP considerations play a critical role in long-run outcomes. Agents that are averse to underperformance converge to the Bertrand-Nash equilibrium, while those more tolerant of underperformance tend to charge supra-competitive prices. This finding also helps mitigate the overfitting issue in independent Q-learning. Additionally, the impact of relative ER varies with the number of agents and the choice of algorithms.

econ.GN

Time-inconsistency with rough volatility

In this paper, we consider equilibrium strategies under Volterra processes and time-inconsistent preferences embracing mean-variance portfolio selection (MVP). Using a functional It\^o calculus approach, we overcome the non-Markovian and non-semimartingale difficulty in Volterra processes. The equilibrium strategy is then characterized by an extended path-dependent Hamilton-Jacobi-Bellman equation system under a game-theoretic framework. A verification theorem is provided. We derive explicit solutions to three problems, including MVP with constant risk aversion, MVP for log returns, and a mean-variance objective with a linear controlled Volterra process. We also thoroughly examine the effect of volatility roughness on equilibrium strategies. Numerical experiments demonstrate that trading rules with rough volatility outperform the classic counterparts.

q-fin.MF

Merton's portfolio problem under Volterra Heston model

This paper investigates Merton's portfolio problem in a rough stochastic environment described by Volterra Heston model. The model has a non-Markovian and non-semimartingale structure. By considering an auxiliary random process, we solve the portfolio optimization problem with the martingale optimality principle. Optimal strategies for power and exponential utilities are derived in semi-closed form solutions depending on the respective Riccati-Volterra equations. We numerically examine the relationship between investment demand and volatility roughness.

q-fin.PM

Mean-variance portfolio selection under Volterra Heston model

Motivated by empirical evidence for rough volatility models, this paper investigates continuous-time mean-variance (MV) portfolio selection under the Volterra Heston model. Due to the non-Markovian and non-semimartingale nature of the model, classic stochastic optimal control frameworks are not directly applicable to the associated optimization problem. By constructing an auxiliary stochastic process, we obtain the optimal investment strategy, which depends on the solution to a Riccati-Volterra equation. The MV efficient frontier is shown to maintain a quadratic curve. Numerical studies show that both roughness and volatility of volatility materially affect the optimal strategy.

q-fin.PM