SearcharxivSearch

arXiv subjects

Xunyu Zhou

Publications and source records attributed to Xunyu Zhou.

6 recordsLinked to original sources

UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models

The abundance of casually captured monocular videos and images on social media provides a valuable source for immersive content creation, where generating novel views from such sparse observations can greatly enhance user experiences. However, producing photorealistic and geometrically consistent views with precise camera control remains challenging when input coverage is extremely limited. Reconstruction-based approaches such as NeRF and 3D Gaussian Splatting (3DGS) deteriorate severely under sparse inputs and fail to explicitly handle occlusions. Generative methods ease data requirements but still struggle with large-baseline view synthesis due to inaccurate or implicit geometric guidance. To overcome these limitations, we introduce UniWorld-View, a unified framework for controllable large-baseline novel view synthesis from monocular inputs. UniWorld-View integrates explicit 3D guidance with generative diffusion modeling to enable precise camera control and geometrically consistent view generation. The geometric guidance is obtained through an occlusion-aware point cloud rendering strategy that resolves visibility ambiguities and provides accurate priors for diffusion-based synthesis. By coupling this rendering strategy with powerful video diffusion backbones, UniWorld-View achieves high-fidelity novel view generation even under extreme camera motions and wide-baseline changes, and can further provide multi-view videos for downstream dynamic 3DGS reconstruction. Experiments on the WorldScore benchmark and zero-shot NVS benchmarks demonstrate the effectiveness of UniWorld-View in controllability, geometric consistency, and visual fidelity.

cs.CV

ART for Diffusion Sampling: A Reinforcement Learning Approach to Timestep Schedule

We consider time discretization for score-based diffusion models to generate samples from a learned reverse-time dynamic on a finite grid. Uniform and hand-crafted grids can be suboptimal given a budget on the number of time steps. We introduce Adaptive Reparameterized Time (ART), which controls the clock speed of a reparameterized time variable to redistribute computation along the sampling trajectory while preserving the terminal time, with the objective of minimizing the aggregate Euler discretization error. We derive a randomized companion ART-RL that recasts ART as a continuous-time reinforcement learning problem with Gaussian policies, and prove a two-directional bridge between the two: the deterministic ART optimum lifts to an optimal Gaussian policy, and conversely any optimal Gaussian policy must recover the ART control through its mean. This bridge turns continuous-time actor--critic learning into a principled, rather than heuristic, route to the deterministic timestep optimum. Within the official EDM pipeline, ART-RL improves FID on CIFAR--10 across a wide range of budgets; after one-time offline training, the distilled deterministic schedule transfers without retraining to AFHQv2, FFHQ, and ImageNet at no extra inference cost.

cs.LG

Exploration versus exploitation in reinforcement learning: a stochastic control approach

We consider reinforcement learning (RL) in continuous time and study the problem of achieving the best trade-off between exploration of a black box environment and exploitation of current knowledge. We propose an entropy-regularized reward function involving the differential entropy of the distributions of actions, and motivate and devise an exploratory formulation for the feature dynamics that captures repetitive learning under exploration. The resulting optimization problem is a revitalization of the classical relaxed stochastic control. We carry out a complete analysis of the problem in the linear--quadratic (LQ) setting and deduce that the optimal feedback control distribution for balancing exploitation and exploration is Gaussian. This in turn interprets and justifies the widely adopted Gaussian exploration in RL, beyond its simplicity for sampling. Moreover, the exploitation and exploration are captured, respectively and mutual-exclusively, by the mean and variance of the Gaussian distribution. We also find that a more random environment contains more learning opportunities in the sense that less exploration is needed. We characterize the cost of exploration, which, for the LQ case, is shown to be proportional to the entropy regularization weight and inversely proportional to the discount rate. Finally, as the weight of exploration decays to zero, we prove the convergence of the solution of the entropy-regularized LQ problem to the one of the classical LQ problem.

math.OC

Two explicit Skorokhod embeddings for simple symmetric random walk

Motivated by problems in behavioural finance, we provide two explicit constructions of a randomized stopping time which embeds a given centered distribution $μ$ on integers into a simple symmetric random walk in a uniformly integrable manner. Our first construction has a simple Markovian structure: at each step, we stop if an independent coin with a state-dependent bias returns tails. Our second construction is a discrete analogue of the celebrated Azéma-Yor solution and requires independent coin tosses only when excursions away from maximum breach predefined levels. Further, this construction maximizes the distribution of the stopped running maximum among all uniformly integrable embeddings of $μ$.

math.PR

A stochastic maximum principle via Malliavin calculus

This paper considers a controlled Itô-Lévy process where the information available to the controller is possibly less than the overall information. All the system coefficients and the objective performance functional are allowed to be random, possibly non-Markovian. Malliavin calculus is employed to derive a maximum principle for the optimal control of such a system where the adjoint process is explicitly expressed.

math.OC

Behavioral Portfolio Selection in Continuous Time

This paper formulates and studies a general continuous-time behavioral portfolio selection model under Kahneman and Tversky's (cumulative) prospect theory, featuring S-shaped utility (value) functions and probability distortions. Unlike the conventional expected utility maximization model, such a behavioral model could be easily mis-formulated (a.k.a. ill-posed) if its different components do not coordinate well with each other. Certain classes of an ill-posed model are identified. A systematic approach, which is fundamentally different from the ones employed for the utility model, is developed to solve a well-posed model, assuming a complete market and general Itô processes for asset prices. The optimal terminal wealth positions, derived in fairly explicit forms, possess surprisingly simple structure reminiscent of a gambling policy betting on a good state of the world while accepting a fixed, known loss in case of a bad one. An example with a two-piece CRRA utility is presented to illustrate the general results obtained, and is solved completely for all admissible parameters. The effect of the behavioral criterion on the risky allocations is finally discussed.

q-fin.PM