SearcharxivSearch

arXiv subjects

Haoyuan Sun

Publications and source records attributed to Haoyuan Sun.

At least 19 recordsLinked to original sources

On the extension of K\"ahler currents on compact complex manifolds

Let $(X,\omega)$ be a compact K\"ahler manifold and let $V\subset X$ be a closed complex submanifold. Coman-Guedj-Zeriahi proposed the problem: is every $\omega|_V$-plurisubharmonic function on $V$ the restriction of an $\omega$-plurisubharmonic function on $X$? In this paper, we solve this problem affirmatively, even for a compact Hermitian manifold.

math.CV

Capacity Stability of Complex Monge-Amp\`ere Equations with Moving Prescribed Singularities

For complex Monge-Amp\`ere equations with moving big cohomology classes and prescribed model singularities of positive Monge-Amp\`ere mass, we prove that, under total variation convergence of the right-hand side non-pluripolar positive Radon measures, convergence of the prescribed model potentials in Monge-Amp\`ere capacity is equivalent to convergence in capacity of the associated normalized solutions. We further prove that the ceiling operator coincides with the singularity envelope for potentials associated to a big $(1,1)$-class, regardless of their Monge-Amp\`ere mass, thereby resolving a conjecture of Darvas-Di Nezza-Lu. Consequently, the singularity envelope is idempotent without the positivity assumption on the mass.

math.CV

Xiaomi-GUI-0 Technical Report

Graphical user interface (GUI) agents build on vision-language models to complete user tasks end-to-end in real applications through interface actions such as tapping, swiping, text entry, and navigation. However, existing GUI agents are trained and evaluated largely on offline trajectories, simulated environments, and standardized benchmarks. These differ substantially from real applications in interface layout, interaction logic, and abnormal-state distribution, and cannot faithfully characterize execution stability in real-world use, where account states, permission dialogs, payment authentication, and risk control continually reshape the state distribution and open a persistent gap between benchmark scores and real usability. To close this gap, we propose Xiaomi-GUI-0, a native multimodal GUI agent for real mobile environments, trained and evaluated within a real-device closed loop. At its core is a real-device-dominant hybrid infrastructure, where physical devices are the primary execution environment and sandboxes provide auxiliary support, so that data collection, training, rollout, and evaluation share an execution distribution close to real deployment. We construct multi-source training data spanning high-frequency head tasks, high-generalization data for long-tail intents, and capability-enhancement data for reflection and memory, and introduce an error-driven data flywheel that turns failure trajectories into corrected actions, reflective explanations, and recovery demonstrations. The model is trained through a progressive three-stage pipeline of supervised fine-tuning, step-level reinforcement learning, and agentic reinforcement learning. Evaluated on public benchmarks and our in-house RealMobile, Xiaomi-GUI-0 achieves 72.0% success on RealMobile and 78.9% on AndroidWorld, while substantially improving execution stability and abnormal-state recognition in real-world tasks.

cs.AI

The Cauchy-Dirichlet Problem for Complex Hessian Flows: From A Priori Estimates to Pluripotential Theory

We study the Cauchy--Dirichlet problem for parabolic complex Hessian equations on Hermitian manifolds and on bounded strictly m-pseudoconvex domains. In the smooth setting, we prove global existence and uniqueness of classical solutions under the presence of an admissible parabolic subsolution, by establishing a priori estimates up to the parabolic boundary. The estimates combine parabolic boundary techniques for complex Hessian equations with interior second order estimates and a blow-up argument. We then develop a general pluripotential framework for degenerate right-hand sides with L^p densities, p>n/m, and bounded Cauchy--Dirichlet data. Since the usual automorphism and Walsh-type arguments do not directly apply in a variable Hermitian background, we use approximation by smooth data, balayage, parabolic Perron envelopes, and a continuous obstacle approximation based on Harvey--Lawson--Plis subequation theory. The resulting solution is continuous for positive time, locally uniformly Lipschitz and semi-concave in time, and continuous up to the initial slice when the initial datum is continuous. We also prove a parabolic comparison principle via time regularization, Riemann sum approximations, and mixed Hessian inequalities.

math.AP

Convergence of the Chern-Ricci flow on complex minimal surfaces of general type

We prove uniform diameter estimates, volume non-collapsing estimates and Gromov-Hausdorff convergence for the normalized Chern-Ricci flow on smooth complex minimal surfaces of general type, starting from an arbitrary Hermitian metric. This removes the local Kahler assumption near the null locus used in our previous work and confirms the Tosatti-Weinkove conjecture in complex dimension two. The main analytic ingredients are a surface torsion estimate, a uniform total variation bound for Delta |G|, a Green-weighted L^2 estimate for the torsion, and a linear iteration of real Poisson equations, which together give the required Green function estimates.

math.DG

Power Reinforcement Post-Training of Text-to-Image Models with Super-Linear Advantage Shaping

Recently, post-training methods based on reinforcement learning, with a particular focus on Group Relative Policy Optimization (GRPO), have emerged as the robust paradigm for further advancement of text-to-image (T2I) models. However, these methods are often prone to reward hacking, wherein models exploit biases in imperfect reward functions rather than yielding genuine performance gains. In this work, we identify that normalization could lead to miscalibration and directly removing the prompt-level standard deviation term yields an optimal policy ascent direction that is linear in the advantage but still limits the separation of genuine signals from noise. To mitigate the above issues, we propose Super-Linear Advantage Shaping (SLAS) by revisiting the functional update from an information geometry perspective. By extending the Fisher-Rao information metric with advantage-dependent weighting, SLAS introduces a non-linear geometric structure that reshapes the local policy space. This design relaxes constraints along high-advantage directions to amplify informative updates, while tightening those in low-advantage regions to suppress illusory gradients. In addition, batch-level normalization is applied to stabilize training under varying reward scales. Extensive evaluations demonstrate that SLAS consistently surpasses the DanceGRPO baseline across multiple backbones and benchmarks. In particular, it yields faster training dynamics, improved out-of-domain performance on GenEval and UniGenBench++, and enhanced robustness to model scaling, while mitigating reward hacking and preserving semantic and compositional fidelity in generations.

cs.CV

Gromov-Hausdorff limits of the Chern-Ricci flow on smooth Hermitian minimal models of general type

We establish uniform diameter estimates and volume non-collapsing estimates for the Chern-Ricci flow on smooth Hermitian minimal models of general type, assuming the initial metric is K\"ahler in a neighborhood of the null locus of the canonical bundle. This yields subsequential Gromov-Hausdorff convergence, partially resolving a conjecture of Tosatti and Weinkove. When the underlying manifold is K\"ahler, we further prove the uniqueness of the limit space. Analytically, we overcome the difficulties posed by non-K\"ahler torsion in the Green's formula by exploiting our local K\"ahler assumption, successfully adapting recent estimates of K\"ahler Green's function to the Hermitian setting. To prove the uniqueness of the limit, we introduce Perelman's reduced length to the Chern-Ricci flow. By establishing a uniform Chern scalar curvature bound and an almost monotonicity formula for the reduced volume, we deduce an almost-avoidance principle for the singular set, allowing us to effectively compare the flow distance with the canonical limit distance.

math.DG

Online Learning for Supervisory Switching Control

We study supervisory switching control for partially-observed linear dynamical systems. The objective is to identify and deploy a suitable controller for the unknown system by periodically selecting among a collection of $N$ candidate controllers, some of which may destabilize the underlying system. While classical estimator-based supervisory control guarantees asymptotic stability, it lacks quantitative finite-time performance bounds. Conversely, current non-asymptotic methods in both online learning and system identification require restrictive assumptions that are incompatible in a control setting, such as system stability, which preclude testing potentially unstable controllers. To bridge this gap, we propose a novel, non-asymptotic analysis of supervisory control that adapts multi-armed bandit algorithms to a control-theoretic setting. The proposed data-driven algorithm evaluates candidate controllers via scoring criteria that leverage system observability to isolate the effects of state history, enabling both detection of destabilizing controllers and accurate system identification. We present two algorithmic variants with dimension-free, finite-time guarantees, where each identifies the matching controller in $O(N \log^2 N)$ steps, while simultaneously achieving finite $L_2$-gain with respect to system disturbances.

math.OC

When to Lock Attention: Training-Free KV Control in Video Diffusion

Maintaining background consistency while enhancing foreground quality remains a core challenge in video editing. Injecting full-image information often leads to background artifacts, whereas rigid background locking severely constrains the model's capacity for foreground generation. To address this issue, we propose KV-Lock, a training-free framework tailored for DiT-based video diffusion models. Our core insight is that the hallucination metric (variance of denoising prediction) directly quantifies generation diversity, which is inherently linked to the classifier-free guidance (CFG) scale. Building upon this, KV-Lock leverages diffusion hallucination detection to dynamically schedule two key components: the fusion ratio between cached background key-values (KVs) and newly generated KVs, and the CFG scale. When hallucination risk is detected, KV-Lock strengthens background KV locking and simultaneously amplifies conditional guidance for foreground generation, thereby mitigating artifacts and improving generation fidelity. As a training-free, plug-and-play module, KV-Lock can be easily integrated into any pre-trained DiT-based models. Extensive experiments validate that our method outperforms existing approaches in improved foreground quality with high background fidelity across various video editing tasks.

cs.CV

Let the Agent Search: Autonomous Exploration Beats Rigid Workflows in Temporal Question Answering

Temporal Knowledge Graph Question Answering (TKGQA) is challenging because it requires multi-hop reasoning under complex temporal constraints. Recent LLM-based approaches have improved semantic modeling for this task, but many still rely on fixed reasoning workflows or costly post-training, which can limit adaptability and make error recovery difficult. We show that enabling an off-the-shelf Large Language Model (LLM) to determine its next action is already effective in a zero-shot setting. Based on this insight, we propose AT2QA, an Autonomous and Training-free Agent for TKG Question Answering. AT2QA empowers the LLM to iteratively interact with the TKG via a generic search tool, inherently enabling autonomous exploration and dynamic self-correction during reasoning. To further elicit the LLM's potential for complex temporal reasoning, we introduce a training-free experience mining mechanism that distills a compact few-shot demonstration library from successful self-generated trajectories. AT2QA also yields a transparent audit trail for every prediction. Experiments on three challenging benchmarks -- MultiTQ, Timeline-CronQuestion, and Timeline-ICEWS-Actor -- show that AT2QA achieves new state-of-the-art performance, surpassing the strongest baselines by 10.7, 4.9, and 11.2 absolute points, respectively. Our code is available at https://github.com/AT2QA-Official-Code/AT2QA-Official-Code

cs.CL

CoLoGen: Progressive Learning of Concept-Localization Duality for Unified Image Generation

Unified conditional image generation remains difficult because different tasks depend on fundamentally different internal representations. Some require conceptual understanding for semantic synthesis, while others rely on localization cues for spatial precision. Forcing these heterogeneous tasks to share a single representation leads to concept-localization representational conflict. To address this issue, we propose CoLoGen, a unified diffusion framework that progressively learns and reconciles this concept-localization duality. CoLoGen uses a staged curriculum that first builds core conceptual and localization abilities, then adapts them to diverse visual conditions, and finally refines their synergy for complex instruction-driven tasks. Central to this process is the Progressive Representation Weaving (PRW) module, which dynamically routes features to specialized experts and stably integrates their outputs across stages. Experiments on editing, controllable generation, and customized generation show that CoLoGen achieves competitive or superior performance, offering a principled representational perspective for unified image generation.

cs.CV

Mixed Hessian inequalities on Hermitian manifolds and applications

Let $(X,\omega)$ be a compact Hermitian manifold of complex dimension $n$. In this paper we establish a Ko\l odziej-Nguyen type weak convergence theorem of complex Hessian operators. Utilizing this result, we prove a general mixed Hessian inequality with respect to a background Hermitian metric, covering both local and global case. As an application, we prove the existence of bounded solutions of complex Hessian equations where the right-hand side measure is well dominated by capacities.

math.DG

Degenerate Complex Hessian type equations on compact Hermitian manifolds and Applications

The aim of this paper is to further develop the theory of the degenerate complex Hessian equations on compact Hermitian manifolds. Building upon the generalization of the Bedford-Taylor pluripotential theory to complex Hessian equations by Ko\l odziej-Nguyen, we solve these equations in the $(\omega, m)$-positive cone, $(\omega, m)$-big classes and in nef classes, where $\omega$ is a reference Hermitian metric. These results are also new in the K\"ahler case. Moreover, we adapt our techniques to solve complex Monge-Amp\`ere equations in nef classes with mild singularities. The solutions we obtain, in the compact K\"ahler case, coincide with those for the complex Monge-Amp\`ere equations in the sense of the non-pluripolar product introduced by Boucksom-Eyssidieux-Guedj-Zeriahi. One of the key ingredients in the proof is the adaption, to the Hermitian setting, of a new a priori $L^\infty$-estimate established by Guo-Phong-Tong and Guo-Phong-Tong-Wang.

math.CV

ViSS-R1: Self-Supervised Reinforcement Video Reasoning

Complex video reasoning remains a significant challenge for Multimodal Large Language Models (MLLMs), as current R1-based methodologies often prioritize text-centric reasoning derived from text-based and image-based developments. In video tasks, such strategies frequently underutilize rich visual information, leading to potential shortcut learning and increased susceptibility to hallucination. To foster a more robust, visual-centric video understanding, we start by introducing a novel self-supervised reinforcement learning GRPO algorithm (Pretext-GRPO) within the standard R1 pipeline, in which positive rewards are assigned for correctly solving pretext tasks on transformed visual inputs, which makes the model to non-trivially process the visual information. Building on the effectiveness of Pretext-GRPO, we further propose the ViSS-R1 framework, which streamlines and integrates pretext-task-based self-supervised learning directly into the MLLM's R1 post-training paradigm. Instead of relying solely on sparse visual cues, our framework compels models to reason about transformed visual input by simultaneously processing both pretext questions (concerning transformations) and true user queries. This necessitates identifying the applied transformation and reconstructing the original video to formulate accurate final answers. Comprehensive evaluations on six widely-used video reasoning and understanding benchmarks demonstrate the effectiveness and superiority of our Pretext-GRPO and ViSS-R1 for complex video reasoning. Our codes and models will be publicly available.

cs.CV

Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization

Recent Progress in post-training flow matching for text-to-image (T2I) generation with Group Relative Policy Optimization (GRPO) has demonstrated strong potential. However, it is hindered by a critical limitation: inaccurate advantage attribution. In this work, we argue that aggregating consecutive steps into a coherent 'chunk' and shifting the policy optimization paradigm from GRPO's step level to the chunk level can effectively mitigate the negative impact of this issue. Building on this insight, we propose Group Chunking Policy Optimization (GCPO), the first chunk-level reinforcement learning approach for post-training flow matching. Extensive experiments demonstrate that GCPO achieves superior performance on both standard T2I benchmarks and preference alignment, with up to 43% relative gains over GRPO, highlighting the promise of chunk-level policy optimization. The code is available on https://github.com/xingzhejun/GCPO.

cs.CV

Reinforcement Learning Meets Masked Generative Models: Mask-GRPO for Text-to-Image Generation

Reinforcement learning (RL) has garnered increasing attention in text-to-image (T2I) generation. However, most existing RL approaches are tailored to either diffusion models or autoregressive models, overlooking an important alternative: masked generative models. In this work, we propose Mask-GRPO, the first method to incorporate Group Relative Policy Optimization (GRPO)-based RL into this overlooked paradigm. Our core insight is to redefine the transition probability, which is different from current approaches, and formulate the unmasking process as a multi-step decision-making problem. To further enhance our method, we explore several useful strategies, including removing the KL constraint, applying the reduction strategy, and filtering out low-quality samples. Using Mask-GRPO, we improve a base model, Show-o, with substantial improvements on standard T2I benchmarks and preference alignment, outperforming existing state-of-the-art approaches. The code is available on https://github.com/xingzhejun/Mask-GRPO

cs.CV

Distribution Preference Optimization: A Fine-grained Perspective for LLM Unlearning

As Large Language Models (LLMs) demonstrate remarkable capabilities learned from vast corpora, concerns regarding data privacy and safety are receiving increasing attention. LLM unlearning, which aims to remove the influence of specific data while preserving overall model utility, is becoming an important research area. One of the mainstream unlearning classes is optimization-based methods, which achieve forgetting directly through fine-tuning, exemplified by Negative Preference Optimization (NPO). However, NPO's effectiveness is limited by its inherent lack of explicit positive preference signals. Attempts to introduce such signals by constructing preferred responses often necessitate domain-specific knowledge or well-designed prompts, fundamentally restricting their generalizability. In this paper, we shift the focus to the distribution-level, directly targeting the next-token probability distribution instead of entire responses, and derive a novel unlearning algorithm termed \textbf{Di}stribution \textbf{P}reference \textbf{O}ptimization (DiPO). We show that the requisite preference distribution pairs for DiPO, which are distributions over the model's output tokens, can be constructed by selectively amplifying or suppressing the model's high-confidence output logits, thereby effectively overcoming NPO's limitations. We theoretically prove the consistency of DiPO's loss function with the desired unlearning direction. Extensive experiments demonstrate that DiPO achieves a strong trade-off between model utility and forget quality. Notably, DiPO attains the highest forget quality on the TOFU benchmark, and maintains leading scalability and sustainability in utility preservation on the MUSE benchmark.

cs.LG

The Hidden Link Between RLHF and Contrastive Learning

Alignment of large language models (LLMs) with human values has recently garnered significant attention, with prominent examples including the canonical yet costly Reinforcement Learning from Human Feedback (RLHF) and the simple Direct Preference Optimization (DPO). In this work, we demonstrate that both RLHF and DPO can be interpreted from the perspective of mutual information (MI) maximization, uncovering a profound connection to contrastive learning. Within this framework, both RLHF and DPO can be interpreted as methods that performing contrastive learning based on the positive and negative samples derived from base model, leveraging the Donsker-Varadhan (DV) lower bound on MI (equivalently, the MINE estimator). Such paradigm further illuminates why RLHF may not intrinsically incentivize reasoning capacities in LLMs beyond what is already present in the base model. Building on the perspective, we replace the DV/MINE bound with the Jensen-Shannon (JS) MI estimator and propose the Mutual Information Optimization (MIO). Comprehensive theoretical analysis and extensive empirical evaluations demonstrate that MIO mitigates the late-stage decline in chosen-likelihood observed in DPO, achieving competitive or superior performance across various challenging reasoning and mathematical benchmarks.

cs.LG